Skip to main content

ragfs_extract/
lib.rs

1//! # ragfs-extract
2//!
3//! Content extraction from various file formats for the RAGFS indexing pipeline.
4//!
5//! This crate provides the extraction layer that reads files and produces
6//! [`ExtractedContent`](ragfs_core::ExtractedContent) for downstream chunking and embedding.
7//!
8//! ## Supported Formats
9//!
10//! | Extractor | Formats | Features |
11//! |-----------|---------|----------|
12//! | [`TextExtractor`] | `.txt`, `.md`, `.rs`, `.py`, `.js`, `.ts`, `.go`, `.java`, `.json`, `.yaml`, `.toml`, `.xml`, `.html`, `.css`, and 30+ more | UTF-8 text extraction |
13//! | [`PdfExtractor`] | `.pdf` | Text extraction + embedded images (JPEG, PNG, JPEG2000) |
14//! | [`OfficeExtractor`] | `.docx`, `.xlsx`, `.pptx`, `.odt` | ZIP+XML visible text (not binary `.doc`) |
15//! | [`ImageExtractor`] | `.png`, `.jpg`, `.gif`, `.webp`, `.bmp` | Metadata extraction, optional vision captioning |
16//!
17//! ## Usage
18//!
19//! ```rust,ignore
20//! use ragfs_extract::{ExtractorRegistry, TextExtractor, PdfExtractor, ImageExtractor, OfficeExtractor};
21//! use std::path::Path;
22//!
23//! // Create a registry with all extractors
24//! let mut registry = ExtractorRegistry::new();
25//! registry.register("text", TextExtractor);
26//! registry.register("pdf", PdfExtractor::new());
27//! registry.register("image", ImageExtractor::new(None));
28//! registry.register("office", OfficeExtractor::new());
29//!
30//! // Extract content from a file
31//! let content = registry.extract(Path::new("document.pdf"), "application/pdf").await?;
32//! println!("Extracted {} bytes", content.text.len());
33//! ```
34//!
35//! ## PDF Image Extraction
36//!
37//! The [`PdfExtractor`] can extract embedded images from PDF documents:
38//!
39//! - **Supported formats**: JPEG (`DCTDecode`), PNG (`FlateDecode`), JPEG2000 (`JPXDecode`)
40//! - **Color spaces**: RGB, Grayscale, CMYK (auto-converted to RGB)
41//! - **Limits**: 100 images max, 50MB total, 50px minimum dimension
42//!
43//! ## Vision Captioning
44//!
45//! The [`ImageExtractor`] supports optional vision-based captioning via the
46//! [`ImageCaptioner`] trait. A [`PlaceholderCaptioner`] is provided as a no-op default.
47//!
48//! ## Components
49//!
50//! | Type | Description |
51//! |------|-------------|
52//! | [`ExtractorRegistry`] | Routes files to appropriate extractors by MIME type |
53//! | [`TextExtractor`] | UTF-8 text/code/markup by extension (not Office parsers) |
54//! | [`PdfExtractor`] | PDF text and image extraction |
55//! | [`OfficeExtractor`] | OOXML/ODT text (docx/xlsx/pptx/odt) |
56//! | [`ImageExtractor`] | Image metadata and optional captioning |
57//! | [`ImageCaptioner`] | Trait for vision model integration |
58//! | [`PlaceholderCaptioner`] | No-op captioner implementation |
59
60pub mod image;
61pub mod office;
62pub mod pdf;
63pub mod registry;
64pub mod text;
65pub mod vision;
66
67pub use image::ImageExtractor;
68pub use office::OfficeExtractor;
69pub use pdf::PdfExtractor;
70pub use registry::ExtractorRegistry;
71pub use text::TextExtractor;
72#[cfg(feature = "vision")]
73pub use vision::BlipCaptioner;
74pub use vision::{CaptionConfig, CaptionError, ImageCaptioner, PlaceholderCaptioner};