Data Wrangling

Parsing, classificação, mídia e enriquecimento

Author

Marlon L.

Published

May 7, 2026

---------------------------------------------------------------------------
ModuleNotFoundError                       Traceback (most recent call last)
Cell In[1], line 11
      7 # Carrega .env e adiciona src ao path
      8 load_dotenv()
      9 sys.path.insert(0, os.getenv('PROJECT_ROOT') + '/src')
     10 
---> 11 from config import PROJECT_ROOT, PATHS
     12 from wrangling import run_wrangling_pipeline, WRANGLING_STEPS
     13 
     14 print(f"📁 Projeto: {PROJECT_ROOT}")

ModuleNotFoundError: No module named 'config'

Data Wrangling

Com o arquivo limpo gerado no Data Cleaning, esta etapa transforma o TXT em um DataFrame estruturado, vincula arquivos de mídia, integra transcrições e exporta os dados enriquecidos.

Objetivo

Transformar o arquivo TXT limpo em dados estruturados e enriquecidos:

  • ✅ Parsing: TXT → DataFrame
  • ✅ Classificação: Identificar tipos de mensagem
  • ✅ Mídia: Vincular arquivos físicos
  • ✅ Transcrição: Integrar transcrições existentes
  • ✅ Enriquecimento: Substituir mídias por transcrições
  • ✅ Exportação: CSV + TXTs de corpus

Configuração do Pipeline

---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[2], line 27
     23 # 📁 CONFIGURAÇÃO DE PATHS
     24 # =============================================================================
     25 
     26 # Arquivo de entrada (saída do cleaning)
---> 27 INPUT_FILE = PATHS['interim'] / 'raw-data_cln7.txt'
     28 
     29 # Diretório de saída
     30 OUTPUT_DIR = PATHS['processed']

NameError: name 'PATHS' is not defined

Execução do Pipeline

---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[3], line 2
      1 # Executa o pipeline
----> 2 result = run_wrangling_pipeline(
      3     order=PIPELINE_ORDER,
      4     input_file=INPUT_FILE,
      5     output_dir=OUTPUT_DIR,

NameError: name 'run_wrangling_pipeline' is not defined

Pipeline de Transformação

---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[4], line 2
      1 for i, step_id in enumerate(PIPELINE_ORDER):
----> 2     step = WRANGLING_STEPS[step_id]
      3     step_num = i + 1
      4 
      5     print(f'''<details style="margin: 0.5em 0; list-style: none; border: 1px solid #d4d4d4; border-radius: 5px; padding: 10px;">

NameError: name 'WRANGLING_STEPS' is not defined

Estatísticas do Processamento

📊 Estatísticas do Pipeline:
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[5], line 3
      1 print("📊 Estatísticas do Pipeline:\n")
      2 
----> 3 if 'total_linhas_txt' in stats:
      4     print(f"   • Linhas no TXT: {stats['total_linhas_txt']:,}")
      5 if 'total_mensagens' in stats:
      6     print(f"   • Mensagens parseadas: {stats['total_mensagens']:,}")

NameError: name 'stats' is not defined

Distribuição de Tipos de Mensagem

---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[6], line 1
----> 1 if 'tipo_mensagem' in df.columns:
      2     print("📊 Distribuição por tipo de mensagem:\n")
      3 
      4     type_counts = df['tipo_mensagem'].value_counts()

NameError: name 'df' is not defined
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[7], line 1
----> 1 if 'tipo_mensagem' in df.columns and 'remetente' in df.columns:
      2     print("\n📊 Distribuição por tipo e remetente:")
      3     pd.crosstab(df['tipo_mensagem'], df['remetente'], margins=True)

NameError: name 'df' is not defined

Arquivos Gerados

📁 Arquivos gerados:
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[8], line 3
      1 print("📁 Arquivos gerados:\n")
      2 
----> 3 for name, path in outputs.items():
      4     size = os.path.getsize(path) if os.path.exists(path) else 0
      5     size_mb = size / (1024 * 1024)
      6     print(f"   • {name:<20} {path.name:<30} {size_mb:>6.2f} MB")

NameError: name 'outputs' is not defined

Tip🎯 Resultado Final
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[9], line 1
----> 1 print(f"📊 DataFrame final: {len(df):,} mensagens × {len(df.columns)} colunas")
      2 print(f"📋 Colunas: {', '.join(df.columns[:10])}...")
      3 print(f"\n📅 Período: {df['timestamp'].min()} até {df['timestamp'].max()}")
      4 

NameError: name 'df' is not defined

Preview dos Dados

---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[10], line 3
      1 # Seleciona colunas para preview
      2 preview_cols = ['timestamp', 'remetente', 'tipo_mensagem', 'conteudo']
----> 3 if 'conteudo_enriquecido' in df.columns:
      4     preview_cols.append('conteudo_enriquecido')
      5 
      6 print("🔍 Primeiras mensagens:\n")

NameError: name 'df' is not defined
---------------------------------------------------------------------------
NameError                                 Traceback (most recent call last)
Cell In[11], line 1
----> 1 if 'tem_transcricao' in df.columns:
      2     df_trans = df[df['tem_transcricao'] == True]
      3 
      4     if len(df_trans) > 0:

NameError: name 'df' is not defined

Exportação Manual (Opcional)

from wrangling import export_corpus_files, export_to_csv

# Re-exportar só texto (sem tags de mídia)
df_text = df[~df['tipo_mensagem'].str.contains('omitted|attached')]
export_corpus_files(df_text, OUTPUT_DIR / 'text_only', use_enriched=False)

# Exportar CSV com colunas específicas
cols = ['timestamp', 'remetente', 'conteudo', 'tipo_mensagem']
export_to_csv(df, OUTPUT_DIR / 'messages_minimal.csv', columns=cols)

Próximos Passos

Com os dados estruturados e enriquecidos, seguimos para:

  1. Feature Engineering — Criação de variáveis derivadas
  2. Análises Descritivas — Estatísticas e visualizações