Data Marketplace

Use verified, licensed data with confidence. You can download right away or check the data through inquiry.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Pre-training DataAudio

    Large-Scale Maghrebi Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Maghrebi Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

  • Pre-training DataImage

    Multilingual Menu OCR Image Dataset

    A multilingual OCR image dataset for menu images in seven languages, including Korean, English, Chinese, and Japanese, comprising original images, translated-text images, and inpainted images with professionally reviewed annotations.

  • Pre-training DataImage

    English Handwritten Document OCR Image Dataset

    An English handwritten document OCR image dataset with bounding-box annotations for multilingual document understanding, OCR model training, and layout-aware text recognition.

  • Pre-training DataImage

    English Manufacturing Document OCR and Parsing Image Dataset

    An English OCR and document parsing sample dataset for manufacturing documents, built with bounding-box JSON annotations for PoC, Document AI, and model evaluation use cases.

  • Pre-training DataImage

    Japanese Handwritten Document OCR Image Dataset

    A Japanese handwritten document OCR image dataset with bounding-box annotations for multilingual document understanding, OCR model training, and layout-aware text recognition.

  • Alignment DataText

    Korean Document Summarization Text Dataset

    A Korean document summarization text dataset built by converting Korean news articles into markdown-based outputs such as key trends, summary reports, and document summaries.