Files
mve-micro-docs/COMO_PROBAR.md
Ernesto Herrera fcc516c9b3 feat: Agregar soporte OCR con Tesseract para PDFs escaneados
- Integrar Tesseract OCR para leer PDFs escaneados automáticamente
- Detectar automáticamente si el PDF tiene texto o requiere OCR
- Agregar servicio ocr_service.py con funciones de OCR
- Actualizar Dockerfile con tesseract-ocr, tesseract-ocr-spa y poppler-utils
- Agregar variables de configuración OCR (OCR_ENABLED, OCR_LANGUAGE, OCR_DPI, OCR_TIMEOUT)
- Crear endpoint de debug para ver texto extraído (/api/v1/debug/extract-text)
- Agregar scripts de instalación y prueba (install_ocr.ps1, test_ocr.py, debug_pdf.ps1)
- Documentación completa (OCR_SETUP.md, DOCKER_OCR.md, COMO_PROBAR.md)
- Actualizar docker-compose.yml con variables de entorno OCR
- Modificar pdf_text.py para usar OCR cuando sea necesario
- Actualizar requirements.txt con pytesseract, Pillow, pdf2image
2026-03-04 08:21:41 -07:00

7.5 KiB

🧪 Guía de Prueba - OCR Funcionando

Estado del Sistema

Tu sistema está completamente configurado con OCR:

  • Docker Compose ejecutándose
  • Tesseract 5.5.0 instalado en contenedores
  • Idioma español (spa) disponible
  • Variables OCR configuradas en .env
  • API corriendo en http://localhost:9876

🚀 Cómo Probar

1. Obtener un Token de Autenticación

# Opción A: Usando curl (si lo tienes instalado)
curl -X POST http://localhost:9876/api/auth/login `
  -H "Content-Type: application/x-www-form-urlencoded" `
  -d "username=admin&password=12345"

# Opción B: Usando Invoke-RestMethod (PowerShell nativo)
$body = @{
    username = "admin"
    password = "12345"
}
$response = Invoke-RestMethod -Uri "http://localhost:9876/api/auth/login" `
    -Method Post `
    -ContentType "application/x-www-form-urlencoded" `
    -Body $body

# Ver el token
$response.access_token

# Guardar el token en una variable
$token = $response.access_token
Write-Host "Token obtenido: $token"

2. Probar con un PDF

Opción A: Endpoint Síncrono (respuesta inmediata)

# Reemplaza "C:\ruta\al\documento.pdf" con tu PDF
$pdfPath = "C:\ruta\al\documento.pdf"

# Subir y procesar
curl -X POST "http://localhost:9876/api/v1/incrementables/parse" `
  -H "Authorization: Bearer $token" `
  -F "file=@$pdfPath"

Opción B: Endpoint Asíncrono (recomendado para PDFs grandes o escaneados)

# 1. Iniciar procesamiento
$pdfPath = "C:\ruta\al\documento.pdf"
$taskResponse = curl -X POST "http://localhost:9876/api/v1/incrementables/parse-async" `
  -H "Authorization: Bearer $token" `
  -F "file=@$pdfPath"

# 2. Extraer el task_id de la respuesta
# Busca algo como: "task_id":"abc-123-def"

# 3. Consultar el estado de la tarea
$taskId = "TU_TASK_ID_AQUI"
curl -X GET "http://localhost:9876/api/v1/incrementables/task/$taskId" `
  -H "Authorization: Bearer $token"

3. Ver Logs en Tiempo Real

Para ver qué está pasando cuando procesas un PDF:

# Ver logs de la API
docker-compose logs -f mve-incrementables-parser

# Ver logs del worker (procesamiento asíncrono)
docker-compose logs -f celery-worker

# Ver todos los logs
docker-compose logs -f

4. Identificar si se Usó OCR

En la respuesta del API, busca el campo extraction_method:

{
  "status": "success",
  "data": {
    "extraction_info": {
      "method": "ocr",  // ← "ocr" = PDF escaneado, "pymupdf"/"pdfplumber" = PDF con texto
      "anchors_found": [...],
      "warnings": [...]
    }
  }
}

También verás en los logs:

  • "PDF appears to be scanned, attempting OCR extraction" → Detectó PDF escaneado
  • "Starting OCR extraction" → Iniciando OCR
  • "OCR extraction completed: X characters" → OCR terminado

📋 Script Completo de Prueba

Copia y pega esto en PowerShell:

# 1. Obtener token
Write-Host "=== Obteniendo token ===" -ForegroundColor Cyan
$body = @{
    username = "admin"
    password = "12345"
}
$loginResponse = Invoke-RestMethod -Uri "http://localhost:9876/api/auth/login" `
    -Method Post `
    -ContentType "application/x-www-form-urlencoded" `
    -Body $body

$token = $loginResponse.access_token
Write-Host "Token obtenido: $token" -ForegroundColor Green
Write-Host ""

# 2. Verificar que tienes un PDF
$pdfPath = Read-Host "Ingresa la ruta completa al PDF (ej: C:\docs\documento.pdf)"

if (-not (Test-Path $pdfPath)) {
    Write-Host "Error: El archivo no existe" -ForegroundColor Red
    exit
}

Write-Host "Archivo encontrado: $pdfPath" -ForegroundColor Green
Write-Host ""

# 3. Procesar PDF (asíncrono)
Write-Host "=== Procesando PDF ===" -ForegroundColor Cyan
$headers = @{
    "Authorization" = "Bearer $token"
}

# Crear el form data
$fileName = Split-Path $pdfPath -Leaf
$fileContent = [System.IO.File]::ReadAllBytes($pdfPath)
$boundary = [System.Guid]::NewGuid().ToString()
$LF = "`r`n"

$bodyLines = (
    "--$boundary",
    "Content-Disposition: form-data; name=`"file`"; filename=`"$fileName`"",
    "Content-Type: application/pdf$LF",
    [System.Text.Encoding]::GetEncoding("iso-8859-1").GetString($fileContent),
    "--$boundary--$LF"
) -join $LF

try {
    $response = Invoke-RestMethod -Uri "http://localhost:9876/api/v1/incrementables/parse-async" `
        -Method Post `
        -Headers $headers `
        -ContentType "multipart/form-data; boundary=$boundary" `
        -Body $bodyLines

    Write-Host "Tarea iniciada!" -ForegroundColor Green
    Write-Host "Task ID: $($response.task_id)" -ForegroundColor Cyan
    Write-Host ""
    
    # 4. Consultar estado
    $taskId = $response.task_id
    Write-Host "=== Consultando estado ===" -ForegroundColor Cyan
    
    $maxAttempts = 30
    $attempt = 0
    
    do {
        Start-Sleep -Seconds 2
        $attempt++
        
        $statusResponse = Invoke-RestMethod -Uri "http://localhost:9876/api/v1/incrementables/task/$taskId" `
            -Method Get `
            -Headers $headers
        
        Write-Host "Intento $attempt - Estado: $($statusResponse.status)" -ForegroundColor Yellow
        
        if ($statusResponse.status -eq "completed") {
            Write-Host ""
            Write-Host "=== RESULTADO ===" -ForegroundColor Green
            Write-Host ($statusResponse | ConvertTo-Json -Depth 10)
            break
        } elseif ($statusResponse.status -eq "failed") {
            Write-Host ""
            Write-Host "=== ERROR ===" -ForegroundColor Red
            Write-Host ($statusResponse | ConvertTo-Json -Depth 10)
            break
        }
        
    } while ($attempt -lt $maxAttempts)
    
    if ($attempt -eq $maxAttempts) {
        Write-Host "Tiempo de espera agotado" -ForegroundColor Red
    }
    
} catch {
    Write-Host "Error: $_" -ForegroundColor Red
}

🔍 Qué Buscar en los Logs

Cuando procesas un PDF, verás algo como:

PDF con Texto Normal (sin OCR):

INFO: Text extracted: 2500 characters, 3 pages, method: pymupdf
INFO: Extracted 2500 chars from 3 pages using pymupdf

PDF Escaneado (con OCR):

INFO: PDF appears to be scanned (only 5 characters found)
INFO: Starting OCR extraction with language=spa, dpi=300
INFO: Converted PDF to 3 images
INFO: Page 1: Extracted 1200 characters
INFO: Page 2: Extracted 1500 characters
INFO: Page 3: Extracted 800 characters
INFO: OCR extraction completed: 3500 characters from 3 pages
INFO: Text extracted: 3500 characters, 3 pages, method: ocr

🐛 Solución de Problemas

"401 Unauthorized"

  • El token expiró o es incorrecto
  • Vuelve a obtener un token con el paso 1

"Failed to extract text"

  • El PDF puede estar dañado o cifrado
  • Revisa los logs: docker-compose logs -f

OCR muy lento

  • Normal para PDFs escaneados (5-30 segundos por página)
  • Usa el endpoint asíncrono (/parse-async)
  • Reduce OCR_DPI=200 en .env para mayor velocidad

"No incrementables found"

  • El PDF no tiene la sección de incrementables
  • Revisa el campo warnings en la respuesta
  • El OCR puede no haber detectado el texto correctamente

📊 Ver Configuración Actual

# Ver variables de entorno en el contenedor
docker exec mve-incrementables-parser env | Select-String OCR

# Ver configuración OCR
docker exec mve-incrementables-parser tesseract --version
docker exec mve-incrementables-parser tesseract --list-langs

🎯 Siguientes Pasos

  1. Prueba con un PDF que tenga texto seleccionable primero
  2. Luego prueba con un PDF escaneado (imagen)
  3. Compara los tiempos de respuesta
  4. Revisa el campo extraction_method en la respuesta

¡Listo para probar! 🚀