Skip to main content

Setting Up a PDF Dataset

Walk through Nexadata's Process PDF flow: table detection, rotation, keep or discard decisions, and building the Dataset.

PDF Datasets let you bring documents into Nexadata as structured, repeatable inputs to your Workflows. Invoices, statements, and billing reports carry real data, but they carry it in printed tables built for a human reader rather than rows and columns a system can consume. The Process PDF flow closes that gap: Nexadata reads the document, finds the tables inside it, and hands you the controls to decide what each one is and how it should be shaped.

This article walks through that flow end to end, starting at the Connect Data step and following through to a finished Dataset. The running example is a legal invoice, chosen because it is awkward in all the ways real documents are: a header block with its labels running down the side, a line-item table spanning several pages, and a summary table on the last page.

You do this once per document format, not once per document. Everything described below is setup. The decisions you make while working through your first invoice, statement, or report are saved as a template, and every later document in that same format is processed against it without repeating any of this. The format has to stay the same; the data is expected to change. Next quarter's report has new numbers in the same tables, and that is exactly the case the template is built for.

The example is a legal invoice. The capability is not about legal work. PDFs arrive in every shape and size across every function, and nothing in this flow is specific to law firms or billing. If the numbers you need for planning and analytics are locked inside a PDF, this is how you get them out: supplier and freight invoices, bank and brokerage statements, utility bills, remittance advices, shipping manifests, lab and inspection results, insurance schedules, fund and board reports, regulatory filings. Length is not a constraint either. A 600-page private equity quarterly report is as legitimate an input as a two-page invoice. The document changes. The flow does not.

What You Are Actually Doing

It helps to know the shape of the flow before you start clicking, because you will see two different progress bars along the way. They look alike and are easy to mistake for each other.

The outer wizard is Create New Dataset. Selecting PDF as your Data Format gives it three steps instead of two:

Connect Data → Process PDF → Define Columns

The middle step is not a single screen. Submitting Connect Data opens a dedicated Process PDF workspace that takes over the window and runs a progress bar of its own, with three stages:

Processing → Table Selection → Dataset Builder

To leave that workspace, use the back link in its top left. It reads Dataset, sits beside the name of your Dataset, and returns you to the outer wizard.

Because both progress bars number their stages 1 to 3, it is worth being precise about which is which. Step 2 of the wizard is Process PDF; stage 2 of Process PDF is Table Selection. The rest of this article names which one it means.

The three inner stages do the following:

  1. Processing: you start extraction, and Nexadata reads the document and detects the tables inside it.

  2. Table Selection: you review each detected table, rename it, rotate the ones whose headers run down the side instead of across the top, and decide which to keep. This is where your judgment matters most.

  3. Dataset Builder: you nominate one table as the main table, then layer on transformations that join and enrich it.

A clean, single-table document moves through Table Selection in a few clicks. A dense invoice with a header block, a multi-page line-item table, and a summary table is where the controls below earn their keep.

Read the whole flow as a build rather than a chore. Every choice you make in these three stages is recorded in the template, so you are not processing one document, you are teaching Nexadata how to read that kind of document. See Datasets and Templates below for what that produces and how it gets reused.

Step 1: Connect Your Data

Create your Dataset from SetupDatasetsCreate New Dataset. The Connect Data step asks for the following:

  • Name. The label you and your team will use to find this Dataset. It must be unique across your organization, and it is the name that appears throughout the Process PDF workspace, so make it descriptive.

  • Data Connection. The storage location holding your document. PDF Datasets work with Nexadata's file-based connections.

  • Data Format. Select PDF. This is the choice that adds the Process PDF step to the wizard, so the progress bar at the top goes from two steps to three the moment you pick it.

  • Allow Any File Type. Off by default. Nexadata normally checks that the file extension matches the format you selected. Enable this only when your source system produces files with no extension or an unexpected one, since it turns off the check that catches format mismatches.

  • Choose file for your Dataset. The document you picked appears under Selected file. Use Choose different file to browse your connection and swap it for another.

Because this first pass defines the template, pick a representative document rather than an unusual one. A month with the full set of tables present will produce a template that holds up better than a short or atypical one.

Click Submit. You go straight into the Process PDF workspace, and the wizard's progress bar is replaced by the three-stage one described above.

The Create New Dataset screen with PDF selected as the Data Format. Note the three-step progress bar: selecting PDF adds the Process PDF step.

For a field-by-field explanation of the Connect Data step and how the other formats differ, see Supported Data Formats in Nexadata.

Step 2: Process the Document

Extraction does not start on its own. You arrive on the Processing stage, which names your Dataset, offers to "extract tables from it to begin reviewing and curating the results", and gives you a single Process PDF button. Nothing happens until you click it, so if you are looking at a screen that seems idle, this is the button you want.

Extraction waits for you. Click Process PDF to begin. The stepper along the top shows the three stages of the Process PDF step.

Once started, Nexadata reads the document, identifies every table it contains, and stitches tables that continue across page breaks back into one. A line-item table running over three pages arrives as a single continuous table rather than three fragments, so you do not have to reassemble it yourself.

Nexadata extracting tables from the document. Larger or more complex files take longer.

How Large a Document Can Be

This scales much further than a few pages. A 600-page private equity quarterly report is a realistic input, with its schedule of investments running for dozens of pages, capital account statements, management fee and expense breakdowns, and performance tables scattered through the commentary. Nexadata reads the whole document, stitches every table that spans a page break, and presents the complete set for review.

The point of doing that is what happens next. A document nobody could reasonably key in by hand becomes a Dataset your Workflows treat like any other source, so quarter-over-quarter fund performance, fee drag, and capital account movement become things you can chart and model rather than things you read.

Long documents take longer to process and produce a longer review queue, but the work is the same work described below: name what matters, discard what does not, and let the template carry those decisions into next quarter's report.

When extraction finishes, you move on to Table Selection.

Step 3: Review Each Detected Table

Table Selection is a review queue. The left rail lists every table Nexadata found, with a count at the top. Expect that count to be higher than the number of tables you actually care about: a document you think of as having three tables can easily detect ten, because footers, address blocks, and signature panels all read as tables to an extraction engine. Seeing 10 tables detected is normal, not a sign something has gone wrong.

Each entry in the rail shows the page it came from and its shape, such as p.1 and 3 rows × 2 cols, which is often enough to recognize a table before opening it. Filter chips above the list (All, Pending, Approved, Rejected) narrow the rail to what still needs attention, and a counter along the bottom of the screen tracks how far you have got, reading 0 / 10 reviewed before you begin. On a long document that count runs into the hundreds, and the Pending filter becomes the main way to work through it without losing your place.

Selecting a table opens it in the center of the screen, where you will find:

  • An editable Name field, covered in the next section.

  • A confidence score for the extraction, shown beside the table position as something like 90.83% confidence. This is Nexadata telling you how cleanly it read that table. A low score is a prompt to compare the extracted values against the source document before you keep it.

  • A position indicator reading TABLE 1 OF 10 · PAGE 1, and a footer restating what was pulled, such as "Extracted from page 1 · 3 rows · 2 columns".

  • The Column Editor panel on the right, which shows the column definitions for the selected table. Not every table has them, and the panel will tell you when there are none available.

Read each table for three things: what it should be called, whether it is oriented correctly, and whether you need it at all. This is the stage that most defines the template, so it is worth the attention. Every judgment you make here is one you will not have to make again next month.

The Table Selection stage. The left rail lists every detected table, the center shows the one you have open, and the Column Editor sits on the right.

Rename Each Table

Tables arrive with generic names: Table 1, Table 2, and so on. Those names are all you will have to work with later, in the Dataset Builder, where you can no longer see what each table contains. Renaming now, while the contents are in front of you, is the difference between choosing a main table confidently and guessing.

To rename a table, highlight the text in the Name field above the table and type the new name. There is no edit mode and no separate save.

Name each table for what it holds rather than where it came from, so "Billing Details" or "Timekeeper Rates" rather than "Table 1" or "page 6 table".

Highlight the text in the Name field and type over it. The new name is what you will see in the Dataset Builder.

Rotate Tables Whose Headers Run Down the Side

Not every table in a document is laid out as columns across the top. Invoice header blocks are the common case: the field names run down the left edge with their values beside them, so they read as rows rather than as columns.

The example document contains exactly this. One table extracts as three rows by two columns, carrying Invoice Date, Invoice No., and Matter No. down the left with their values on the right. Loaded as-is, that gives you a two-column table of labels and values, which is not what you want.

Click the Rotate button at the bottom left of the table view and the orientation flips: the labels become column headers and the values become a single row of data. You can watch the shape change in the footer beneath the table, which goes from three rows by two columns to two rows by three columns.

Choose Which Values to Project

Rotating does one more thing, and it is easy to miss. A Header Metadata list appears beneath the rotated table, listing each field with a checkbox and the value it holds, under the prompt "Select values to project into your dataset columns".

This is where you decide which of those fields actually become columns on your Dataset. Ticking Invoice Date and Invoice No. carries both onto every row, so each billing line records which invoice it came from and when. Leaving Matter No. unticked keeps it out of the Dataset entirely.

The choice is per field, so take what downstream reporting needs and leave the rest. A rotated table with nothing ticked projects nothing, which is the most common reason invoice-level fields go missing from a finished Dataset.

Before and after rotating. In the first panel the field names run down the side. In the second they have become column headers, and the Header Metadata list underneath is where you tick the values to project into your dataset columns.

Rule of thumb: if reading the table left to right gives you a field name followed by its value, it needs rotating. If reading left to right gives you several different fields belonging to one record, it is already correct.

Keep or Discard Each Table

Documents contain tables you do not want. Page footers, signature blocks, remittance instructions, and formatting artifacts all get detected alongside the data you actually came for.

Two buttons at the top of the table view carry the decision: Not relevant drops the table from the Dataset, Keep table carries it forward. Every detected table needs one or the other, so a document that detected ten tables needs ten decisions, and the reviewed counter along the bottom tracks how many you have made.

Your decisions show up immediately in the left rail. A kept table gets a green check beside it, a discarded one gets a red cross, and anything still undecided keeps a neutral marker. Those icons are what the Pending, Approved, and Rejected chips filter on, so you can hide what you have already handled and work down to an empty Pending list.

The rail reflects the rest of your work on each table too. A renamed table shows its new name, and a rotated one shows its new shape, so Invoice Details reads 2 rows × 3 cols once flipped rather than the 3 rows × 2 cols it arrived as. If a change did not take, the rail is the quickest place to notice.

Be deliberate rather than permissive here. Every table you keep is one you will have to account for in the Dataset Builder, and discarding the noise now makes the next stage considerably simpler. On a long report that is most of the job: a few hundred detected tables can reduce to a handful worth carrying forward.

The Not relevant and Keep table buttons, and the status icons your decisions leave behind in the left rail.

A worked example, staying with the legal invoice. It yields three tables worth keeping:

  1. Invoice Details, the header block carrying Invoice Date, Invoice No., and Matter No. Renamed, rotated, its invoice-level values projected into columns, then kept.

  2. Billing Details, the line-item table of dated entries by timekeeper. Spans multiple pages, already stitched, kept as the core data.

  3. Timekeeper Rates, the summary table listing each timekeeper's hourly rate. Kept so it can be joined in later.

Everything else the document threw off gets marked Not relevant.

That three-part shape is worth recognizing because it recurs far beyond legal billing. A supplier invoice has a header block, line items, and a tax summary. A brokerage statement has an account block, transactions, and a holdings summary. A utility bill has a service block, usage detail, and a rate schedule. A fund report has a fund identification block, a schedule of investments, and a performance summary. Different documents, same decisions: identify the reference block, find the table that carries the detail, and keep whatever summary you can join back to it.

Save Tables

When every table has been named, oriented, and decided on, click Save Tables at the bottom right, beside the reviewed counter. Your decisions are committed, and you move on to the Dataset Builder.

Step 4: Choose the Main Table

The Dataset Builder opens on a short setup form. Give the build a name, unique across your organization, add an optional description, and then make the decision this screen exists for: select the main table.

The main table is the spine of your Dataset. It determines the grain, meaning what one row represents. Every other table you kept becomes something that augments it rather than something that stands alongside it.

In the legal invoice example, Billing Details is the main table, because one row should equal one billable entry. The values you projected from the rotated Invoice Details block attach to every row, and the Timekeeper Rates table is available to join. Choose a summary table as your main table by mistake, and you get a Dataset at the wrong grain, so it is worth pausing on this one.

This is the screen where the names you set in Table Selection pay off. The dropdown lists your tables by name and nothing else.

Selecting the main table in the Dataset Builder. Every other table you kept will augment this one.

When the main table is set, click Start Transforming.

Step 5: Build the Dataset with Transformations

The Dataset Builder shows your Transformations as a stack on the left and a live preview of the resulting data on the right. Each transformation is a step in a chain: the preview reflects everything applied so far, so you can watch the Dataset take shape as you go.

Click any transformation to open it and edit its configuration. The breadcrumb at the top of the screen tracks where you are in the chain, so you can move between steps without losing your place.

The Dataset Builder with two transformations applied. The preview on the right reflects every step in the chain.

Two transformations do most of the work on a document like this one:

  • A join that brings a supplemental table into the main table. Joining Timekeeper Rates onto Billing Details by timekeeper name attaches the correct hourly rate to every billing line.

  • A calculation that derives a value the document never printed. With hours on each line and rate now joined in, calculating the billable amount per entry gives you a figure the original invoice only ever showed as a total.

This is the step where a PDF stops being a document and starts being data. The source file showed totals; the Dataset gives you the arithmetic behind them, per line, ready to aggregate any way you need. The same pattern applies whatever the document is, whether that means costing a freight invoice against a rate table, extending a usage figure by a tariff, checking a statement's transactions against its stated closing balance, or netting a fund's fees against its gross return.

Add whatever further transformations the data needs, including dropping columns you no longer want to carry. These transformations are part of the template as well, so the same joins and calculations run automatically against every later document of this format.

Step 6: Finish the Dataset

When the preview matches what you expect, continue to Define Columns, the final step of the outer wizard, where you confirm column names and data types exactly as you would for a Tabular or Spreadsheet Dataset. Save, and the Dataset is available to add to your Workflows.

Note: The name you give the Dataset must be unique across your organization before it can be saved.

That is the end of the setup. From here on, this document format is configured, and the next section explains what you actually have.

Datasets and Templates

Two things come out of this flow, and it is worth being clear about which is which.

Each PDF you process generates a Dataset. That is the structured output: the rows and columns produced from the tables you kept, shaped by the orientation choices and transformations you applied.

The Dataset is stored with the template. The template holds the configuration rather than the data, which is the record of which tables mattered, how they were oriented, which one was the main table, and what transformations were applied on top. Because the Dataset is stored with it, the template is both a reusable configuration and the place you go to find what has been processed through it.

The next document does not start over. When the following month's invoice, statement, or report arrives in the same layout, you process it against the template you already built. The same tables are recognized, named, oriented, and joined the way you decided the first time, and you get a new Dataset out of it without repeating any of the review. Only the document changes; the configuration stays put.

What the template depends on is the layout, not the values. New numbers in the same tables are exactly what it is built for, so a fresh quarter, a new billing period, or a different account of the same kind all run straight through. What a template cannot absorb is a structural change: your vendor redesigns the invoice, a report gains or drops a schedule, or a table moves its columns around. When that happens, set up the new format once, the same way you set up this one.

The practical consequence is that the effort you spend curating tables is spent once per format, not once per document. That matters most where the documents are largest. Curating a 600-page quarterly report is real work the first time, and close to free every quarter after, which is what turns a recurring pile of PDFs into a dependable input to your planning and analytics.

If the Results Are Not What You Expected

A few situations come up often enough to name:

  • Nothing is happening on the Processing stage. Extraction does not start automatically. Click the Process PDF button.

  • Far more tables were detected than the document appears to contain. This is expected. Footers, address blocks, and signature panels all read as tables. Work through the list and mark the ones you do not need as Not relevant.

  • A table came through as labels and values in two columns. It needed rotating. Return to Table Selection, open that table, and click Rotate at the bottom left.

  • Document-level fields are missing from the finished Dataset. Rotating a header block is only half the job. Check the Header Metadata list beneath the rotated table and tick the values you want projected into your dataset columns.

  • The extracted values do not match the document. Check the confidence score on that table. A low score means Nexadata had difficulty reading it, and the values are worth verifying against the source before you keep the table.

  • You are not sure which tables you have already handled. Filter the rail to Pending. What remains is what still needs a decision.

  • Your row counts look wrong. Check the main table selection. A Dataset built on a summary table carries one row per group rather than one row per record.

  • You cannot tell your tables apart in the Dataset Builder. Return to Table Selection and rename them there, while you can still see their contents.

Did this answer your question?