--------------------------------------------------------------------------- ModuleNotFoundError Traceback (most recent call last) Cell In[1], line 11 7 # Carrega .env e adiciona src ao path 8 load_dotenv() 9 sys.path.insert(0, os.getenv('PROJECT_ROOT') + '/src') 10 ---> 11 from config import PROJECT_ROOT, PATHS 12 from wrangling import run_wrangling_pipeline, WRANGLING_STEPS 13 14 print(f"📁 Projeto: {PROJECT_ROOT}") ModuleNotFoundError: No module named 'config'
Data Wrangling
Parsing, classificação, mídia e enriquecimento
Data Wrangling
Com o arquivo limpo gerado no Data Cleaning, esta etapa transforma o TXT em um DataFrame estruturado, vincula arquivos de mídia, integra transcrições e exporta os dados enriquecidos.
Objetivo
Transformar o arquivo TXT limpo em dados estruturados e enriquecidos:
- ✅ Parsing: TXT → DataFrame
- ✅ Classificação: Identificar tipos de mensagem
- ✅ Mídia: Vincular arquivos físicos
- ✅ Transcrição: Integrar transcrições existentes
- ✅ Enriquecimento: Substituir mídias por transcrições
- ✅ Exportação: CSV + TXTs de corpus
Configuração do Pipeline
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[2], line 27 23 # 📁 CONFIGURAÇÃO DE PATHS 24 # ============================================================================= 25 26 # Arquivo de entrada (saída do cleaning) ---> 27 INPUT_FILE = PATHS['interim'] / 'raw-data_cln7.txt' 28 29 # Diretório de saída 30 OUTPUT_DIR = PATHS['processed'] NameError: name 'PATHS' is not defined
Execução do Pipeline
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[3], line 2 1 # Executa o pipeline ----> 2 result = run_wrangling_pipeline( 3 order=PIPELINE_ORDER, 4 input_file=INPUT_FILE, 5 output_dir=OUTPUT_DIR, NameError: name 'run_wrangling_pipeline' is not defined
Pipeline de Transformação
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[4], line 2 1 for i, step_id in enumerate(PIPELINE_ORDER): ----> 2 step = WRANGLING_STEPS[step_id] 3 step_num = i + 1 4 5 print(f'''<details style="margin: 0.5em 0; list-style: none; border: 1px solid #d4d4d4; border-radius: 5px; padding: 10px;"> NameError: name 'WRANGLING_STEPS' is not defined
Estatísticas do Processamento
📊 Estatísticas do Pipeline:
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[5], line 3 1 print("📊 Estatísticas do Pipeline:\n") 2 ----> 3 if 'total_linhas_txt' in stats: 4 print(f" • Linhas no TXT: {stats['total_linhas_txt']:,}") 5 if 'total_mensagens' in stats: 6 print(f" • Mensagens parseadas: {stats['total_mensagens']:,}") NameError: name 'stats' is not defined
Distribuição de Tipos de Mensagem
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[6], line 1 ----> 1 if 'tipo_mensagem' in df.columns: 2 print("📊 Distribuição por tipo de mensagem:\n") 3 4 type_counts = df['tipo_mensagem'].value_counts() NameError: name 'df' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[7], line 1 ----> 1 if 'tipo_mensagem' in df.columns and 'remetente' in df.columns: 2 print("\n📊 Distribuição por tipo e remetente:") 3 pd.crosstab(df['tipo_mensagem'], df['remetente'], margins=True) NameError: name 'df' is not defined
Arquivos Gerados
📁 Arquivos gerados:
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[8], line 3 1 print("📁 Arquivos gerados:\n") 2 ----> 3 for name, path in outputs.items(): 4 size = os.path.getsize(path) if os.path.exists(path) else 0 5 size_mb = size / (1024 * 1024) 6 print(f" • {name:<20} {path.name:<30} {size_mb:>6.2f} MB") NameError: name 'outputs' is not defined
Preview dos Dados
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[10], line 3 1 # Seleciona colunas para preview 2 preview_cols = ['timestamp', 'remetente', 'tipo_mensagem', 'conteudo'] ----> 3 if 'conteudo_enriquecido' in df.columns: 4 preview_cols.append('conteudo_enriquecido') 5 6 print("🔍 Primeiras mensagens:\n") NameError: name 'df' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[11], line 1 ----> 1 if 'tem_transcricao' in df.columns: 2 df_trans = df[df['tem_transcricao'] == True] 3 4 if len(df_trans) > 0: NameError: name 'df' is not defined
Exportação Manual (Opcional)
Note📋 Como re-exportar com configurações diferentes
from wrangling import export_corpus_files, export_to_csv
# Re-exportar só texto (sem tags de mídia)
df_text = df[~df['tipo_mensagem'].str.contains('omitted|attached')]
export_corpus_files(df_text, OUTPUT_DIR / 'text_only', use_enriched=False)
# Exportar CSV com colunas específicas
cols = ['timestamp', 'remetente', 'conteudo', 'tipo_mensagem']
export_to_csv(df, OUTPUT_DIR / 'messages_minimal.csv', columns=cols)Próximos Passos
Com os dados estruturados e enriquecidos, seguimos para:
- Feature Engineering — Criação de variáveis derivadas
- Análises Descritivas — Estatísticas e visualizações