DataSunrise Achieves AWS Data & Analytics Competency. Learn more →

Part ofPosture and Discovery

Sensitive Data Discovery

Classification Across Databases, Files, and Storage

Find regulated and business-sensitive data across databases, files, storage, and collaboration platforms with classification, OCR, and scheduled scans.

  • Database and Content Scanning
  • Information Types and OCR
  • Searchable Findings

What this product covers

Sensitive Data Discovery turns scattered records and content into an inventory that teams can search, review, and report on.

Scheduled tasks inspect database metadata and selected values or extract content from files and objects. Information Types use patterns, dictionaries, context checks, NLP, ML/ONNX scoring, and OCR to classify each match.

Every finding points to the exact database object or content path and records how the match was classified.

Run a Repeatable Discovery Workflow

Move from a connected database or content repository to reusable findings in four clear steps.

  1. Connect and set the scope

    Add credentials, verify access, and choose the databases, schemas, paths, buckets, shares, files, or other content to inspect.

  2. Configure classification

    Choose Information Types, security standards, custom definitions, OCR, and extraction options.

  3. Run and review

    Start the task manually or on a schedule, then review matched locations, classifications, statistics, and processing errors.

  4. Report and reuse

    Export CSV or PDF reports and add selected findings to Object Groups for later security policies.

Each run keeps its scope, configuration, and results together, so teams can compare scans over time.

See the workflow in DataSunrise

Product screens show how policy, evidence, and protection appear in practice.

Explore Sensitive Data Discovery

1 / 4

Inspect Sensitive Data Across Your Environment

SourceWhat DataSunrise inspectsHow it connects
Databases and analyticsSchemas, tables, columns, metadata, and selected value samplesDatabase connection and Periodic Data Discovery task
Object and cloud storageObjects, files, documents, images, and archives in Amazon S3, Azure storage, and Google Cloud StorageStorage connector
Local and network storageFiles reached through local paths, SMB2/3, FTP/FTPS, or NFSFile-storage connector and discovery worker
Microsoft OneDrive and SharePointFiles, document libraries, metadata, and extracted contentMicrosoft 365 authentication and Cloud File Discovery

Classify With Layered Detection Methods

Information Types define what DataSunrise should find and how it should validate a match. Teams can use built-in security standards or create definitions for their own identifiers.

Detection methodHow it helps
Patterns and regular expressionsFind structured account, identity, contact, and payment values
Names, metadata, and data typesUse column, field, file, or object context to narrow matches
Dictionaries and lexiconsMatch approved value sets, terms, and language-specific vocabularies
NLP and context validationEvaluate surrounding content and proximity or semantic conditions
Lua validationApply custom validation logic
ML/ONNX-assisted scoringAdd AI Score evidence to selected classification workflows
OCRExtract text from images, image-based PDFs, and embedded document images

Control Scan Depth and Execution

Database tasks can use top, random, all-row, incremental, and randomized strategies. File and object tasks can process complete content or a configured byte range, apply source tags and path filters, and distribute work across multiple DataSunrise servers.

Amazon S3 also offers S3 Inventory mode for manifest-driven discovery.

Teams can start with a narrow sample for a fast inventory, then increase scan depth for broader coverage. Processing reports make encrypted, damaged, inaccessible, and unreadable content visible for follow-up.

Turn Findings Into Security Work

Discovery gives teams a concrete starting point for cleanup, compliance evidence, and protection. Instead of reviewing entire systems by hand, they can focus on the database objects and content paths where sensitive information was found.

Object Groups let teams reuse a selected set of findings in audit, masking, firewall, or test-data work. Teams choose and configure the policy that fits the data.

File and Object Storage Discovery explains file formats and connectors. Active DSPM adds cloud inventory and posture. Data Protection and Enforcement covers controls for live traffic and protected copies.

FAQ

Frequently Asked Questions

Does discovery change source data?

No. Discovery reads metadata or content and records findings. Static Data Masking creates a separate de-identified target.

How do teams turn a discovery result into protection?

Findings can populate Object Groups and give teams the exact data context needed to configure masking, audit, or firewall policies.

Can DataSunrise inspect images and image-based documents?

Yes. DataSunrise combines native OCR with Amazon Textract for S3 to inspect images, image-based PDFs, and embedded document images.

How does DataSunrise scan different data environments?

Databases use sampling, storage uses connectors and file extraction, and images use OCR. The results arrive in the same discovery workflow.

See how DataSunrise works with your technology stack

View Integration Examples