--------------------------------------------------------------------------- ModuleNotFoundError Traceback (most recent call last) Cell In[1], line 12 8 # Carrega .env e adiciona src ao path 9 load_dotenv() 10 sys.path.insert(0, os.getenv('PROJECT_ROOT') + '/src') 11 ---> 12 from config import PROJECT_ROOT, PATHS 13 from utils import ( 14 audit_transformation, 15 audit_pipeline, ModuleNotFoundError: No module named 'config'
Data Cleaning
Limpeza e normalização do arquivo bruto
Data Cleaning
Com base nas descobertas documentadas no Data Profiling, esta etapa implementa as transformações de limpeza necessárias para preparar o arquivo bruto para parsing e análise.
Objetivo
Transformar o arquivo bruto exportado do WhatsApp em um arquivo limpo e otimizado, removendo:
- ✅ Caracteres invisíveis (U+200E)
- ✅ Timestamps vazios (múltiplas mídias)
- ✅ Linhas vazias
- ✅ Espaços em branco excessivos
- ✅ Nomes de participantes (anonimização)
- ✅ Delimitadores redundantes de timestamp
- ✅ Indentação desnecessária em linhas de continuação
Pipeline de Transformação
▸ Etapa 1: Remoção do caractere U+200E
O caractere U+200E (Left-to-Right Mark) é um caractere de formatação invisível usado para controle de direção de texto. O WhatsApp o insere sistematicamente durante a exportação, mas não carrega informação semântica útil para nossa análise.
Impacto identificado no profiling:
- ~48.700 ocorrências
- ~243 KB do arquivo
- Presente em ~50% das linhas
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[2], line 18 14 15 return total_occurrences 16 17 # Define arquivos ---> 18 input_1 = RAW_FILE 19 output_1 = INTERIM_DIR / 'raw-data_cln1.txt' 20 21 # Executa transformação NameError: name 'RAW_FILE' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[3], line 1 ----> 1 audit_1 = audit_transformation(input_1, output_1, "Remoção U+200E") NameError: name 'audit_transformation' is not defined
▸ Etapa 2: Remoção de timestamps vazios
Quando múltiplas mídias são enviadas simultaneamente, o WhatsApp registra uma linha com timestamp e remetente vazios, seguida das mídias. Exemplo:
[24/08/25, 16:38:36] Lê 🖤:
[24/08/25, 16:38:37] Lê 🖤: <attached: foto1.jpg>
[24/08/25, 16:38:38] Lê 🖤: <attached: foto2.jpg>
A primeira linha é redundante e será removida.
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[4], line 38 34 35 return len(lines_to_skip) 36 37 # Define arquivos ---> 38 input_2 = output_1 39 output_2 = INTERIM_DIR / 'raw-data_cln2.txt' 40 41 # Executa transformação NameError: name 'output_1' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[5], line 1 ----> 1 audit_2 = audit_transformation(input_2, output_2, "Remoção timestamps vazios") NameError: name 'audit_transformation' is not defined
▸ Etapa 3: Remoção de linhas vazias
Linhas completamente vazias (apenas \n) aparecem entre mensagens e não carregam informação. Serão removidas preservando a formatação interna das mensagens multilinha.
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[6], line 13 9 10 return len(lines) - len(cleaned_lines) 11 12 # Define arquivos ---> 13 input_3 = output_2 14 output_3 = INTERIM_DIR / 'raw-data_cln3.txt' 15 16 # Executa transformação NameError: name 'output_2' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[7], line 1 ----> 1 audit_3 = audit_transformation(input_3, output_3, "Remoção linhas vazias") NameError: name 'audit_transformation' is not defined
▸ Etapa 4: Normalização de espaços internos
Espaços excessivos, tabs e trailing whitespace no meio do conteúdo são normalizados:
- Múltiplos espaços consecutivos → espaço único
- Tabs → espaço único
- Espaços no final das linhas → removidos
- Preserva: Indentação inicial (será tratada na Etapa 7)
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[8], line 32 28 29 return original - cleaned 30 31 # Define arquivos ---> 32 input_4 = output_3 33 output_4 = INTERIM_DIR / 'raw-data_cln4.txt' 34 35 # Executa transformação NameError: name 'output_3' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[9], line 1 ----> 1 audit_4 = audit_transformation(input_4, output_4, "Normalização espaços") NameError: name 'audit_transformation' is not defined
▸ Etapa 5: Anonimização de participantes
Os nomes dos participantes são substituídos por identificadores genéricos para preservar privacidade:
| Original | Anônimo |
|---|---|
| Marlon | P1 |
| Lê 🖤 | P2 |
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[10], line 28 24 25 return replacements 26 27 # Define arquivos ---> 28 input_5 = output_4 29 output_5 = INTERIM_DIR / 'raw-data_cln5.txt' 30 31 # Executa transformação NameError: name 'output_4' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[11], line 1 ----> 1 audit_5 = audit_transformation(input_5, output_5, "Anonimização") NameError: name 'audit_transformation' is not defined
▸ Etapa 6: Otimização do formato de timestamp
Remove delimitadores redundantes para facilitar o parsing:
De: [28/11/24, 19:30:05] P1: Mensagem
Para: 28/11/24 19:30:05 P1: Mensagem
Economiza 4 caracteres por mensagem ([, ], ,, espaço extra).
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[12], line 27 23 24 return count 25 26 # Define arquivos ---> 27 input_6 = output_5 28 output_6 = INTERIM_DIR / 'raw-data_cln6.txt' 29 30 # Executa transformação NameError: name 'output_5' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[13], line 1 ----> 1 audit_6 = audit_transformation(input_6, output_6, "Otimização timestamps") NameError: name 'audit_transformation' is not defined
▸ Etapa 7: Normalização de indentação
Mensagens multilinha frequentemente contêm espaços de indentação no início das linhas de continuação. Exemplo:
28/11/24 19:30:05 P2: *🪙 Colombia 🪙*
R$ 22 o grama.
25g. $20 ▪️ 50g. $18
Esses espaços iniciais são visuais e podem ser removidos, preservando o conteúdo.
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[14], line 29 25 26 return spaces_removed 27 28 # Define arquivos ---> 29 input_7 = output_6 30 output_7 = INTERIM_DIR / 'raw-data_cln7.txt' 31 32 # Executa transformação NameError: name 'output_6' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[15], line 1 ----> 1 audit_7 = audit_transformation(input_7, output_7, "Normalização indentação") NameError: name 'audit_transformation' is not defined
Auditoria Final do Pipeline
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[17], line 1 ----> 1 format_audit_table(df_audit, include_chars=True).style.hide(axis='index') NameError: name 'format_audit_table' is not defined
--------------------------------------------------------------------------- NameError Traceback (most recent call last) Cell In[18], line 1 ----> 1 audits = [audit_1, audit_2, audit_3, audit_4, audit_5, audit_6, audit_7] 2 format_summary_table(audits).style.hide(axis='index') NameError: name 'audit_1' is not defined
Arquivo de Saída
O arquivo raw-data_cln7.txt agora possui:
- ✅ Formato otimizado:
DD/MM/YY HH:MM:SS Remetente: Conteúdo - ✅ Participantes anonimizados (P1, P2)
- ✅ Sem caracteres invisíveis
- ✅ Sem linhas/timestamps vazios
- ✅ Espaços e indentação normalizados
# Padrão para identificar início de mensagem
message_pattern = r'^(\d{2}/\d{2}/\d{2}) (\d{2}:\d{2}:\d{2}) (.+?): (.*)$'
# Grupos de captura:
# 1: Data (DD/MM/YY)
# 2: Hora (HH:MM:SS)
# 3: Remetente (P1 ou P2)
# 4: Conteúdo da mensagemPróximos Passos
Com o arquivo limpo, seguimos para:
- Data Wrangling — Parsing, agregação de multilinha, vinculação de mídia
- Feature Engineering — Criação de variáveis derivadas