How to Extract Structured Data with OpenAI in a Jinba Workflow

How to Extract Structured Data with OpenAI in a Jinba Workflow

Summary

  • OPENAI_INVOKE calls the OpenAI API from a Jinba workflow and returns JSON for downstream steps.
  • Store the OpenAI key in Jinba's credentials system, and configure the step with config (version, prompt) and input.text.
  • Use explicit JSON instructions, schema keys, and few-shot examples so the model returns consistent structures.
  • Validate outputs with Jinba Modules Checker V2 before sending data to Snowflake or Google Sheets.
  • Jinba Flow lets you build this extraction workflow on-premise and connect it to downstream systems in days.

Structured data extraction from unstructured text is the step most data pipelines underestimate. Template-based approaches fail when document layouts change. OCR produces text, not structure. The extraction step itself becomes the bottleneck.

The Jinba OPENAI_INVOKE tool addresses this directly. It calls the OpenAI API from within a Jinba workflow, returns a JSON object, and passes that object to downstream steps such as a Snowflake insert or a Google Sheets write. This guide covers the configuration required for this integration, including credentials, YAML shape, prompt design, downstream wiring, and deterministic validation afterward.

What structured data extraction actually involves

Structured extraction converts free-form text into a defined JSON schema. The extracted object is then programmatically accessible for analysis, reporting, or ingestion into a data warehouse.

The pattern applies across industries:

  • Financial services: Pulling loan applicant details from bank statements for underwriting
  • Healthcare: Extracting patient fields from clinical notes to reduce manual transcription errors
  • E-commerce: Parsing product reviews into sentiment scores and feature mentions for product development

In each case, the source text varies in layout and phrasing. A model-based approach handles that variation without requiring a separate template per document type. That is the core advantage of openai workflow automation over rule-based parsers for this class of problem.

Step 1: Obtain and store your OpenAI API key

An API key authenticates every request OPENAI_INVOKE makes to OpenAI. To generate one:

  1. Create an account at platform.openai.com
  2. Verify your email address
  3. Log in, click your profile icon, and select View API keys
  4. Click Create new secret key
  5. Copy and save the key immediately. OpenAI does not display it again after this screen

Once generated, store the key in Jinba's credentials system, not in workflow YAML or source repositories. Jinba resolves stored credentials at runtime, keeping the key out of version control and execution logs.

To control spend, set a usage limit inside your OpenAI account before connecting anything to production. OpenAI's pricing varies by model, and gpt-4 costs significantly more per token than gpt-3.5-turbo. Match the model to the complexity of the extraction task.

Step 2: Configure OPENAI_INVOKE in your workflow

The OPENAI_INVOKE tool takes two top-level blocks: config and input. Both are singular, not plural.

config holds the model settings:

  • version: the OpenAI model identifier, such as gpt-3.5-turbo or gpt-4
  • prompt: the instruction the model receives, including the extraction schema definition

input holds the data the step processes:

  • text: the unstructured content to extract from

A complete step is configured as follows:

- tool: OPENAI_INVOKE
config:
version: "gpt-3.5-turbo"
prompt: "Extract the full name, email address, and company from the following text.
Return the result as a JSON object with the keys 'name', 'email', and 'company'."
input:
text: "Contact person: Jane Doe from ExampleCorp. Reach her at jane.doe@examplecorp.com for inquiries."

The version field determines which model processes the request. For extraction tasks involving short, consistently structured documents, gpt-3.5-turbo performs well and keeps token costs low. For longer documents with complex or ambiguous layouts, gpt-4 produces more reliable output.

Step 3: Write prompts that produce consistent JSON

The prompt is the contract between the workflow and the model. Vague prompts produce inconsistent output structures, which break downstream steps that depend on specific key names.

Three practices produce reliable JSON from OPENAI_INVOKE:

Instruct explicitly for JSON. State in the prompt that the output must be a JSON object; the format is not left implicit.

Define the schema in the prompt. Name every key the downstream step will reference. Include the expected data type where it matters, for example specifying that monthly_rent should be a number, not a string.

Use few-shot examples. Including one worked example inside the prompt significantly reduces format deviation, particularly for documents with irregular phrasing. The model treats the example as the target structure.

An extraction prompt for lease agreements using these three practices:

You are an assistant that extracts key information from lease agreements into JSON.

Example:
Text: 'This lease is between Landlord John Smith and Tenant Sarah Johnson for the
property at 123 Main St, starting June 1, 2024, for $2000/month.'
JSON: {"landlord_name": "John Smith", "tenant_name": "Sarah Johnson",
"property_address": "123 Main St", "start_date": "2024-06-01", "monthly_rent": 2000}

Now extract from the following:
{lease_document_text}

This structure anchors the model to an exact output shape before it sees the target document.

Step 4: Pass the JSON result to downstream steps

Jinba resolves step outputs via result references using the pattern {{steps.<step_id>.result}}. After OPENAI_INVOKE completes, the parsed JSON object is available to every subsequent step in the workflow.

To insert the extracted fields into Snowflake:

- tool: SNOWFLAKE_QUERY
config:
credential: snowflake_prod
input:
query: >
INSERT INTO contacts (name, email, company)
VALUES (
'{{steps.OPENAI_INVOKE.result.name}}',
'{{steps.OPENAI_INVOKE.result.email}}',
'{{steps.OPENAI_INVOKE.result.company}}'
);

The same reference pattern applies when writing to Google Sheets. Each field from the extracted JSON maps to its corresponding column using {{steps.OPENAI_INVOKE.result.<key>}}.

Downstream steps receive the output of the extraction through Jinba's result reference system, not through polling or external status checks. A step that references a prior result executes only after that result is available, giving the workflow implicit dependency management.

Validating after extraction: Modules Checker V2

OPENAI_INVOKE is probabilistic. Given the same input on successive runs, the model output usually conforms to the prompt schema, but not on every run. Fields can be missing, key names can drift, and numeric values can be returned as strings. Testing a prompt in development and deploying it without a validation layer is the most common point of failure in production extraction workflows.

Jinba Modules Checker V2 is the deterministic complement to OPENAI_INVOKE. It validates the extracted JSON against a defined rule set before the result reaches any downstream system. Required fields, data types, and value constraints are enforced as hard rules, not sampled against a probability.

The two-step pattern for reliable extraction:

  1. OPENAI_INVOKE performs the flexible, layout-agnostic extraction and returns a JSON object
  2. Jinba Modules Checker V2 validates that object, halting the workflow if any rule fails before a malformed record reaches Snowflake or Sheets

This separation preserves the role of each tool. The LLM handles variance and ambiguity in the source text. The checker enforces the schema contract the downstream system requires. Neither tool replaces the other.

Governance: audit logging, key security, and private models

Jinba's enterprise workflow orchestration captures a complete audit log for each workflow execution, including inputs, outputs, and decisions at every step. For extraction workflows processing financial records, contracts, or personally identifiable information, that log is the evidence trail for compliance reviews.

API key security follows from the credential store setup in Step 1. Keys stored in Jinba's credentials system are resolved at runtime and are never written into YAML, logs, or repositories. Rotation is handled in the credential store without changes to the workflow definition.

For data subject to GDPR or other privacy regulations, teams processing sensitive text evaluate whether a private or self-hosted model is appropriate. Sending PII to a third-party API endpoint can fall outside the data processing terms acceptable to a regulator or a customer contract. Jinba supports credential-based routing, so the version field can point to a private endpoint rather than the public OpenAI API without changing the rest of the workflow configuration.

What to build next

A working extraction workflow in Jinba requires four things:

  • An OpenAI API key stored in the Jinba credentials system
  • An OPENAI_INVOKE step with a schema-defining prompt and a version match to the task complexity
  • A validation step using Jinba Modules Checker V2 to enforce the output contract
  • Downstream steps referencing {{steps.OPENAI_INVOKE.result.<key>}} to route the validated data to Snowflake, Google Sheets, or any other target

The combination of probabilistic extraction and deterministic validation makes this architecture production-safe rather than prototype-grade. Start with a single document type, verify the schema contract holds across a sample of real inputs, and extend from there.

Review the Jinba tools documentation for the full list of available tool identifiers and configuration options before building the workflow.

Frequently Asked Questions

What is Jinba's OPENAI_INVOKE tool used for?

OPENAI_INVOKE is a Jinba workflow tool that calls the OpenAI API to extract structured data from unstructured text and returns a JSON object. It is designed for tasks like pulling fields from documents, emails, or forms so the data can be passed to downstream systems such as Snowflake or Google Sheets.

How do I configure OPENAI_INVOKE in a Jinba workflow?

The tool uses two top-level blocks: config and input. Under config, set version to the OpenAI model (e.g., gpt-3.5-turbo or gpt-4) and prompt to the extraction instructions. Under input, provide the text field with the unstructured content. A complete step looks like:

- tool: OPENAI_INVOKE
config:
version: "gpt-3.5-turbo"
prompt: "Extract the full name, email address, and company from the following text. Return the result as a JSON object with the keys 'name', 'email', and 'company'."
input:
text: "Contact person: Jane Doe from ExampleCorp. Reach her at jane.doe@examplecorp.com for inquiries."

How do I get an OpenAI API key for Jinba?

Create an account at platform.openai.com, verify your email, go to View API keys, and click Create new secret key. Copy and save the key immediately because OpenAI does not show it again. Then store it in Jinba's credentials system, and never in workflow YAML or source code. Jinba resolves the key at runtime to keep it secure.

How do I write prompts that produce consistent JSON?

Use three practices: explicitly instruct the model to output JSON, define the exact schema with key names and data types, and include a few-shot example that shows the desired output shape. For example, show a sample text and its corresponding JSON object before the actual input. This anchors the model to the expected format.

How do I pass OPENAI_INVOKE output to Snowflake or Google Sheets?

Use result references in the form {{steps.<step_id>.result.<key>}}. After the extraction step completes, each field from the JSON object is accessible. For Snowflake, embed the references in a SNOWFLAKE_QUERY step's SQL. For Google Sheets, map each field to its column using the same reference pattern. Jinba automatically manages dependencies so downstream steps wait for the result.

How do I validate the JSON returned by OPENAI_INVOKE?

Use Jinba Modules Checker V2, a deterministic validation tool that checks the extracted JSON against rules for required fields, data types, and value constraints. Place it immediately after OPENAI_INVOKE in the workflow. If validation fails, the workflow halts before malformed data reaches downstream systems. This two-step pattern, probabilistic extraction followed by deterministic validation, is essential for production reliability.

Can I use a private or self-hosted model with OPENAI_INVOKE?

Yes. Jinba supports credential-based routing, so the version field can point to a private endpoint instead of the public OpenAI API. This supports processing sensitive data under GDPR or other privacy regulations. The rest of the workflow configuration stays the same, only the model endpoint changes.

What are the costs of using OPENAI_INVOKE with OpenAI?

Costs depend on the model and token usage. gpt-3.5-turbo is cheaper per token than gpt-4, so for short, consistently structured documents, gpt-3.5-turbo is sufficient. For complex, lengthy documents, gpt-4 is required, but at higher cost. Set a usage limit in the OpenAI account before connecting to production, and monitor spending via OpenAI's pricing page.

Build your way.

The AI layer for your entire organization.

Get Started