> For the complete documentation index, see [llms.txt](https://docs.diaflow.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.diaflow.io/agent-builder/agent-builder/document-understanding.md).

# Document understanding

## Quick Start

What you can do:

\- Analyze uploaded documents with AI

\- Extract text, labels, and data from documents

\- Get descriptions and summaries of visual content

\- Compare documents and identify differences

\- Answer questions about document content

<br>

**Quick Start Steps:**

1\. Navigate to Chat: Go to the AI Chat interface

2\. Upload an Image:

&#x20;  \- Click the \`+\` icon

&#x20;  \- Select "Upload from device" or "Diaflow Drive” or you can simply drag and drop your files or copy and paste your images onto the chat frame.

Supported File Formats (Format support may vary by model):

* Images: .jpg, .jpeg, .png, .webp
* Documents: .docx, .txt, .pdf

3\. Ask Your Question: Type a prompt asking about the image (e.g., "Describe this image" or "Extract the text from this receipt")

4\. Send: Click Send and wait for AI analysis

5\. Review Results: AI provides structured analysis, descriptions, or extracted data

<figure><img src="/files/0pDJEObIOOQk3eZwUyLK" alt=""><figcaption></figcaption></figure>

**Expected Result**: AI analyzes the documents and provides detailed insights, descriptions, or extracted information based on your prompt.

<figure><img src="/files/NFwAvfRoaWfWbAXRp5wj" alt=""><figcaption></figcaption></figure>

## Common Uses

### 1. Image Description and Analysis

Get detailed descriptions of images for:

\- Accessibility (alt-text generation)

\- Content moderation

\- Visual content understanding

<br>

**Example Prompts:**

\- "Describe this image in detail"

\- "What objects are visible in this image?"

\- "Analyze the composition and color scheme"

<figure><img src="/files/RkcWPl3FztTLKZkYJePf" alt=""><figcaption></figcaption></figure>

**Output:**

<figure><img src="/files/dc0YNeOX4tP4GH8oSVOy" alt=""><figcaption></figcaption></figure>

### 2. Data Extraction and Structuring

Extract structured data from files:

\- Tables and charts

\- Forms and surveys

\- Product information

\- Contact details

<br>

**Example Prompts:**

\- "Extract the table data from this image and format as CSV"

\- "Identify all products and their prices in this catalog file"

\- "Convert this form into structured JSON data"<br>

<figure><img src="/files/nse5ZQzUVcAePOm85nRP" alt=""><figcaption></figcaption></figure>

\ <br>

### 3. Content Analysis

Analyze visual content for:

\- Brand compliance

\- Design feedback

\- Accessibility issues

\- Quality assessment

<br>

**Example Prompts:**

\- "Check this UI design for accessibility issues."

\- "Analyze the color contrast in this image."

\- "Provide design feedback on this mockup."

<figure><img src="/files/KjbWzan4yfGDDRsE91V8" alt=""><figcaption></figcaption></figure>

\
\ <br>

### 4. Object Detection and Classification

Identify and classify objects in images:

\- Product identification

\- Scene understanding

\- Object counting

\- Category classification

<br>

**Example Prompts:**

\- "Count the number of people in this image"

\- "Identify all products in this store shelf image"

\- "What type of vehicle is shown in this image?"
