A method and system for optimizing the generation of cosmetic adverse reaction analysis reports

By obtaining cosmetics production batch data, performing timestamp conflict detection and data completion, and constructing a causal network and spatiotemporal knowledge graph, the problem of incomplete batch traceability information in the cosmetics supply chain is solved, and the accuracy of cosmetics adverse reaction analysis and decision-making efficiency are improved.

CN120298011BActive Publication Date: 2025-09-05南昌市检验检测中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510772893.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-05
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The complexity of the cosmetics production supply chain leads to incomplete batch traceability information. The existing report generation system is unable to effectively build a reliable causal chain between adverse reaction events and the actual responsible batches, affecting the credibility of analytical conclusions and the accuracy of regulatory decisions.

Method used

By obtaining raw material procurement, production and processing, and warehousing logistics data of cosmetics production batches, timestamp conflict detection and data completion are performed, causal networks and spatiotemporal knowledge graphs are constructed, potential related batches are identified, and an analysis report with labeled batch association probability weights is generated.

Benefits of technology

It has significantly improved the accuracy and credibility of cosmetic adverse reaction analysis reports, achieved multi-level precise positioning of risk batches and a basis for decision-making at all stages, and improved the timeliness and pertinence of risk intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298011B_ABST
    Figure CN120298011B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for optimizing the generation of cosmetic adverse reaction analysis reports, specifically relating to the technical field of cosmetic safety supervision, and is used to solve the problem of inaccurate adverse reaction attribution caused by dispersed supply chain data and inconsistent batch labels in the prior art. By integrating full-process data of raw material procurement, production and processing, warehousing and logistics, and adverse reaction events, a batch labeling system with a unified index is constructed, and cross-link data deviations are repaired based on timestamp conflict detection and conditional generative adversarial networks to generate complete batch data after elimination. Further, a causal network is constructed through an attribute-constrained variational autoencoder to extract potential risk factors, and a spatiotemporal knowledge graph is constructed by combining supply chain disruption events and user location trajectories to identify abnormally associated batches. Finally, causal weights and spatiotemporal probabilities are integrated to generate a batch risk quantification report, realizing a closed-loop analysis from data conflict resolution to accurate attribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cosmetics safety supervision, and more specifically, to a method and system for optimizing the generation of cosmetics adverse reaction analysis reports. Background Art

[0002] In the field of cosmetics safety supervision, adverse reaction analysis reports can serve as a basis for identifying product risks, tracing problem batches, and formulating recall strategies. Current report generation methods rely on product batch traceability information to establish the correlation between adverse reactions and specific production batches, and locate potential quality defects or ingredient safety hazards by integrating supply chain data, consumer feedback, and test results. However, cosmetics production involves multi-level supply chain collaboration, and the collection and transmission of batch data in links such as raw material procurement, production and processing, warehousing and logistics rely on manual records or heterogeneous systems, which makes it easy for batch marking information to be recorded with deviations or update lags during cross-link circulation.

[0003] In the existing technology, due to the complexity of the supply chain and the imperfect data transmission mechanism, the integrity and accuracy of cosmetic batch traceability information are difficult to guarantee. When there are errors, confusion or delays in batch labeling, the report generation system cannot effectively establish a reliable causal chain between adverse reaction events and the actual responsible batches, resulting in inaccurate identification of the root causes of risks and a decrease in the credibility of the analysis conclusions, which in turn leads to a chain of problems such as insufficient basis for regulatory decision-making, expanded product recall scope or lagging risk control measures, seriously affecting the efficiency and accuracy of cosmetics safety management. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a method and system for optimizing the generation of a cosmetic adverse reaction analysis report to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A method for optimizing the generation of a cosmetic adverse reaction analysis report, comprising:

[0007] S1. Obtain raw material purchase batch marks, production and processing batch records, warehousing and logistics batch codes, and adverse reaction events for cosmetic production batches;

[0008] S2. Perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes to identify conflicting data for the same batch at different stages;

[0009] S3. For the batch fields missing in the conflicting data, the conditional generative adversarial network is used to generate data to complete the missing fields based on the known fields, and the discriminator confidence is combined to generate the resolved batch labeled data.

[0010] S4. Construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and filter out unexplained isolated event data;

[0011] S5. For isolated event data, a spatiotemporal knowledge graph of batch regional diffusion and migration is constructed based on supply chain disruption event records and user location time series trajectories to identify potentially related batches.

[0012] S6. Fusion of potentially associated batches and resolved batch labeling data to generate an adverse reaction analysis report with batch association probability weights.

[0013] In a preferred embodiment, obtaining raw material purchase batch marks, production and processing batch records, warehousing and logistics batch codes, and adverse reaction events of cosmetic production batches includes:

[0014] Extract the raw material purchase batch mark provided by the supplier from the raw material purchase order, and associate the raw material purchase batch mark with the corresponding raw material quality inspection report for storage;

[0015] Export the production timestamp, process parameters and operator identification of each batch from the production execution system as production and processing batch records;

[0016] Call the logistics management interface to extract the batch delivery time, transportation route and receipt status to generate the warehouse logistics batch code;

[0017] Collect symptom description texts, occurrence time and usage scenario images submitted by consumers, and convert the symptom description texts into standardized medical codes to form adverse reaction events.

[0018] In a preferred embodiment, timestamp conflict detection is performed on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes to identify conflicting data of the same batch at different stages, including:

[0019] Extract the purchase timestamp in the raw material purchase batch mark and align it with the production start timestamp in the production and processing batch record for verification;

[0020] Compare the production end timestamp in the production and processing batch record of the same batch with the delivery timestamp in the warehousing and logistics batch code in chronological order;

[0021] Based on the preset time window threshold, determine whether the interval between the purchase timestamp and the production start timestamp exceeds the allowable range, detect whether the production end timestamp is later than the delivery timestamp, and identify the reverse order or overlap of the timestamp sequences in different links of the same batch;

[0022] The same batch whose interval between the purchase timestamp and the production start timestamp exceeds the allowable range, whose production end timestamp is later than the delivery timestamp, or whose batch has reverse order or overlap is marked as conflicting data and stored in the conflict dataset.

[0023] In a preferred embodiment, for batch fields missing in conflicting data, a conditional generative adversarial network is used to generate data to complete the missing fields based on known fields, and the resolved batch labeled data is generated in combination with the discriminator confidence screening, including:

[0024] Extract known fields from conflicting data as input features for conditional generative adversarial networks;

[0025] Generate missing field completion data based on known fields through the generator of the conditional generative adversarial network;

[0026] The discriminator of the conditional generative adversarial network is used to score the rationality of the completed data, and the completed data with a score higher than the preset confidence threshold is selected;

[0027] Merge the filtered completed data with the known fields in the original conflicting data to generate resolved batch labeled data;

[0028] Perform format check on the digested batch tag data to ensure that the value range and logical relationship of the completed fields are consistent with those of the known fields.

[0029] In a preferred embodiment, known fields include batch number, timestamp and associated raw material name; missing fields include logistics receipt status or process parameters.

[0030] In a preferred embodiment, a causal network of the digested batch labeled data and adverse reaction events is constructed, and potential causal factors are extracted and unexplained isolated event data are screened through an attribute-constrained variational autoencoder, including:

[0031] The batch number, raw material name, process parameters and logistics receipt status in the digested batch tag data are used as batch node attributes in the causal network;

[0032] Map the standardized medical codes in adverse reaction events to symptom nodes in the causal network and associate them with the corresponding batch nodes through batch numbers;

[0033] Encode batch node attributes to generate potential causal factors through attribute-constrained variational autoencoders;

[0034] Calculate the association weights between batch nodes and symptom nodes based on potential causal factors to generate edge relationships in the causal network;

[0035] Symptom nodes and their associated batch nodes whose correlation weights are lower than a preset causal threshold are filtered and marked as unexplained isolated event data.

[0036] In a preferred embodiment, for isolated event data, a spatiotemporal knowledge graph of batch regional diffusion and migration is constructed based on supply chain disruption event records and user location time series trajectories to identify potentially related batches, including:

[0037] Extract batch numbers and their associated symptom codes and occurrence timestamps from isolated event data;

[0038] Match the disruption type, disruption occurrence time, and affected geographical area corresponding to the batch number in the isolated event data from the supply chain disruption event records;

[0039] Convert the geographic location information in the user's location time series trajectory into GeoHash code;

[0040] Based on the impact area of ​​the interruption event and the user's use of GeoHash encoding to build a regional diffusion node of the spatiotemporal knowledge graph, the attributes of the regional diffusion node include batch number, interruption type and geographical scope;

[0041] Generate the migration path edge relationship of the spatiotemporal knowledge graph based on the time overlap between the interruption occurrence time and the user's usage timestamp;

[0042] Screen out abnormal areas in the regional diffusion nodes that are not covered by the logistics path but have a user GeoHash coding concentration higher than the preset concentration threshold, and mark them as potential related batches;

[0043] The potential associated batches are associated with the regional diffusion nodes and migration path edge relationships of the spatiotemporal knowledge graph and stored in the batch association analysis library.

[0044] In a preferred embodiment, time overlap is determined as the interval between the interruption occurrence time and the usage timestamp being smaller than a preset time window.

[0045] In a preferred embodiment, the potentially associated batches and the resolved batch marker data are integrated to generate an adverse reaction analysis report with batch association probability weights, including:

[0046] Extract batch numbers from potentially related batches and their associated exception areas and interruption types;

[0047] Merge the batch number, raw material name, process parameters and logistics receipt status fields in the batch mark data after digestion;

[0048] The batch association probability weight is calculated based on the abnormal area concentration of potential associated batches and the number of days of time overlap. The association probability weight is the product of the concentration ratio and the time overlap coefficient;

[0049] Adjust the batch association probability weight based on the association weight between the batch node and the symptom node in the causal network to generate a comprehensive association probability score;

[0050] Sort the batch numbers from high to low according to the comprehensive association probability scores, and generate an analysis report with the batch association probability weights marked.

[0051] In another aspect, the present invention provides a cosmetic adverse reaction analysis report generation optimization system, comprising:

[0052] Data acquisition module, used to obtain raw material procurement batch marks, production and processing batch records, warehousing and logistics batch codes and adverse reaction events of cosmetics production batches;

[0053] The conflict detection module is used to perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes, and identify conflicting data of the same batch at different links;

[0054] The data completion module is used to complete the missing batch fields in the conflicting data by generating missing field data based on known fields through a conditional generative adversarial network, and combining the discriminator confidence screening to generate the resolved batch labeled data;

[0055] The causal analysis module is used to construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and screen out unexplained isolated event data;

[0056] The spatiotemporal analysis module is used to construct a spatiotemporal knowledge graph of batch regional diffusion and migration based on isolated event data and supply chain disruption event records and user location time series trajectories to identify potentially related batches.

[0057] The report generation module is used to integrate the potentially associated batches and the resolved batch label data to generate an adverse reaction analysis report with the batch association probability weights.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. Through multi-dimensional data integration and dynamic association modeling, the accuracy and credibility of cosmetic adverse reaction analysis reports have been significantly improved. Based on the full-chain data collection of raw material procurement, production and processing, warehousing and logistics, and adverse reaction events, a batch labeling system with a unified index is constructed, which solves the problem of traceability information fragmentation caused by traditional methods due to data dispersion or format heterogeneity. Through the data completion mechanism of timestamp conflict detection and conditional generative adversarial networks, batch record deviations across links are automatically identified and repaired, ensuring the temporal consistency and field integrity of batch data, providing a reliable data foundation for subsequent analysis. At the same time, the attribute-constrained variational autoencoder incorporates prior knowledge such as the chemical stability and toxicity risk of raw materials during the data completion process, preventing the generated data from deviating from actual production constraints, further enhancing the authenticity and usability of the data.

[0060] 2. Through dual correlation analysis of causal networks and spatiotemporal knowledge graphs, multi-level precise positioning of risky batches is achieved; causal network modeling dynamically associates adverse reaction symptoms with batch components and process parameters, and quantifies the contribution weight of component risks to symptoms; spatiotemporal knowledge graphs combine supply chain disruption events with user location trajectories to capture abnormal regional associations in unexpected circulation scenarios and effectively identify potential risk spread outside the logistics path; finally, a comprehensive risk assessment report is generated by integrating causal weights and spatiotemporal correlation probabilities, providing a decision-making basis covering the entire production, circulation, and use links, greatly improving the timeliness and pertinence of risk intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A flowchart of a cosmetics adverse reaction analysis report generation optimization method according to the present invention;

[0062] Figure 2 The present invention is a structural diagram of a cosmetics adverse reaction analysis report generation optimization system. DETAILED DESCRIPTION

[0063] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0064] Example 1: Figure 1 The present invention provides a method for optimizing the generation of a cosmetic adverse reaction analysis report, comprising:

[0065] S1. Obtain raw material purchase batch marks, production and processing batch records, warehousing and logistics batch codes, and adverse reaction events for cosmetic production batches;

[0066] S2. Perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes to identify conflicting data for the same batch at different stages;

[0067] S3. For the batch fields missing in the conflicting data, the conditional generative adversarial network is used to generate data to complete the missing fields based on the known fields, and the discriminator confidence is combined to generate the resolved batch labeled data.

[0068] S4. Construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and filter out unexplained isolated event data;

[0069] S5. For isolated event data, a spatiotemporal knowledge graph of batch regional diffusion and migration is constructed based on supply chain disruption event records and user location time series trajectories to identify potentially related batches.

[0070] S6. Fusion of potentially associated batches and resolved batch labeling data to generate an adverse reaction analysis report with batch association probability weights.

[0071] S1. Obtain the raw material purchase batch mark, production and processing batch record, warehousing and logistics batch code and adverse reaction events of the cosmetics production batch. The specific implementation is as follows:

[0072] In the cosmetics production batch data acquisition step, the raw material purchase batch tag is obtained by parsing the electronic raw material purchase order document provided by the supplier. The electronic raw material purchase order document records the supplier name, raw material name, raw material purchase batch tag, and purchase date fields. The raw material purchase batch tag is generated by the supplier according to a preset rule: "supplier abbreviation - raw material category - year and month - serial number." For example, the first batch of "glycerin" raw material purchased in October 2023 from supplier "ABC Company" is marked as "ABC-GLY-202310-001." After extracting the raw material purchase batch tag, it is associated with the quality inspection report for the same batch of raw materials for storage. The quality inspection report contains the raw material test items, test results, and test date. The storage method is to use the raw material purchase batch tag as the primary key in the relational database, and associate it with the file storage path and test result summary fields of the quality inspection report.

[0073] Production and processing batch records are obtained by connecting to the cosmetics manufacturer's production execution system database. Each production batch record in the production execution system database includes the production batch number, production start timestamp, production end timestamp, process parameter set, and operator ID. The production start and end timestamps have an accuracy of seconds and are formatted as "2023-11-05 08:30:45." The process parameter set is dynamically configured based on the product type. For example, the process parameters for facial masks include the mixing temperature range, stirring speed level, and filling pressure threshold. The mixing temperature range is recorded as "45°C to 50°C," the stirring speed level is recorded as "Level 3," and the filling pressure threshold is recorded as "0.5MPa." The operator ID is consistent with the unique employee identifier in the company's human resources system and is used to trace operation records throughout the production process.

[0074] Warehouse logistics batch codes are obtained by calling the application programming interface (API) of the enterprise logistics management system. The logistics management system interface accepts the production batch number as an input parameter and returns the batch's delivery time, a list of transportation route nodes, and the receipt status. The transportation route node list is recorded in chronological order: "warehouse delivery → regional distribution center → destination city distribution station → terminal store." Each node includes an arrival timestamp and a departure timestamp. The receipt status field is an enumeration value, including "delivered," "in transit," and "received." The "received" status requires the name of the recipient and the receipt timestamp.

[0075] Adverse reaction event data is collected through the consumer feedback module on cosmetic companies' official websites and mobile applications. When submitting an adverse reaction event, consumers are required to provide a symptom description, the time of occurrence, and upload an image of the use scenario. Symptom descriptions are free text input, such as "redness and itching on the face after use." This is converted into standardized medical codes using a natural language processing engine. This conversion involves keyword matching against a medical dictionary using the "Preferred Term" coding system of the Medical Dictionary for Reproductive Diseases (MedDRA). For example, "erythema" is mapped to "10017456" and "itching" is mapped to "10022130." The time of occurrence is entered by the consumer selecting a date control in the format of "2023-11-05." The use scenario image is a photograph of the skin taken by the consumer, stored in JPEG format, and linked to the corresponding event record.

[0076] Raw material procurement batch tags, production and processing batch records, warehousing and logistics batch codes, and adverse event data are stored in a unified index relationship, with the production batch number as the index key. The production batch number is generated according to the "brand code - production year and month - serial number" rule. For example, the 25th batch of brand "XYZ" produced in November 2023 is numbered "XYZ-202311-025." All data fields undergo format validation before storage, including batch number length, timestamp format, and enumeration value range validation. Data that fails this validation triggers an exception alarm and is recorded in the error log table.

[0077] The specific process for storing quality inspection reports is as follows: After extracting the raw material purchase batch tag, the corresponding quality inspection report PDF document is downloaded from the enterprise quality management system based on the supplier name and raw material purchase batch tag in the purchase order. The test result table in the PDF document is parsed, and the key test item values ​​are extracted and stored in database fields. For example, the parsed storage fields for the glycerin raw material quality inspection report include "Glycerin Purity (%)," "Heavy Metal Content (ppm)," "Microbiological Test Results," and "Name of the Testing Agency."

[0078] The production execution system database is connected via JDBC protocol to access read-only views provided by the system. These views have preconfigured access permissions for production batch-related fields. The logistics management system interface is synchronized every six hours to avoid overloading the logistics system. When converting medical codes for adverse event data submitted by consumers, if a keyword match fails, the data is forwarded for manual review. Reviewers manually assign codes based on symptom descriptions and record the reasons for the corrections.

[0079] All data acquisition steps are performed on a data collection server. For example, the server operating system is Linux CentOS 7.6, the database is MySQL 8.0, and the network communication protocols are HTTPS and SFTP. Sensitive data involved in the collection process, such as supplier names, employee IDs, and consumer identity information, is anonymized by hashing the raw data using the SHA-256 algorithm to generate irreversible identifiers before storage.

[0080] S2. Perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes to identify conflicting data of the same batch at different stages. The specific implementation is as follows:

[0081] During the timestamp conflict detection step for cosmetics production batch data, the purchase timestamp in the raw material purchase batch tag is aligned with the production start timestamp in the production and processing batch record. The purchase timestamp is extracted from the electronic raw material purchase order document and is formatted as "2023-11-05 08:30:45." The production start timestamp is obtained from the production execution system database and is formatted in the same format as the purchase timestamp. Alignment verification is performed by chronologically arranging the purchase timestamps and production start timestamps corresponding to the same production batch number. If the purchase timestamp is later than the production start timestamp, a timing conflict is determined. For example, if the purchase timestamp for a batch of raw materials is "2023-11-05 10:00:00" and the production start timestamp is "2023-11-05 09:30:00," the purchase timestamp is later than the production start timestamp, triggering a conflict flag.

[0082] For the same batch, the production end timestamp in the batch record is compared with the delivery timestamp in the warehouse logistics batch code in chronological order. The production end timestamp is recorded by the production execution system, and the delivery timestamp is obtained from the logistics management system interface. The comparison rule is that the production end timestamp must be earlier than or equal to the delivery timestamp. If the production end timestamp is "2023-11-06 17:00:00" and the delivery timestamp is "2023-11-06 16:30:00", it is considered a logical conflict. Timestamps are accurate to the second and are formatted in a unified "YYYY-MM-DD HH:MM:SS" format.

[0083] The preset time window threshold determines whether the interval between the purchase timestamp and the production start timestamp exceeds the allowable range. The time window threshold is set according to industry standards. For example, the maximum allowable interval between the purchase and production of cosmetic raw materials is 72 hours. If the purchase timestamp is "2023-11-01 08:00:00" and the production start timestamp is "2023-11-05 09:00:00", the interval is 98 hours, and if it exceeds the threshold, a conflict is flagged. The threshold is based on raw material stability data and the company's historical production cycle statistics. For example, the threshold for perishable raw materials is set at 48 hours, and for standard raw materials it is set at 72 hours.

[0084] When checking whether the production end timestamp is later than the shipment timestamp, if the production end timestamp is "2023-11-06 18:00:00" and the shipment timestamp is "2023-11-06 17:30:00", a logical conflict is identified. The conflict detection logic is based on the strict sequential nature of timestamps, ensuring that the production phase can only enter the logistics phase after completion.

[0085] Identify reverse or overlapping timestamp sequences within the same batch at different stages. The timestamp sequence is arranged in the order of stages: purchase timestamp → production start timestamp → production end timestamp → delivery timestamp. Reversal detection is performed by checking whether the timestamp of the subsequent stage is earlier than the timestamp of the previous stage. For example, if the production start timestamp is "2023-11-05 09:00:00" and the purchase timestamp is "2023-11-05 10:00:00," the sequence is reversed. Overlap detection is performed by checking whether the timestamp intervals of adjacent stages intersect. For example, if the production start timestamp is "2023-11-05 09:00:00" to "2023-11-05 17:00:00" and the delivery timestamp is "2023-11-05 16:30:00," the time intervals are overlapped.

[0086] Batches with an excessive interval between the purchase timestamp and the production start timestamp, a production end timestamp later than the delivery timestamp, or timestamps with reversed or overlapping timestamps are marked as conflicting data and stored in a conflict dataset. The conflict dataset's storage structure includes fields for the conflicting batch number, conflict type code, and conflicting timestamp details. The conflict type code is an enumeration value, with "1" representing an excessive interval, "2" representing a logical reverse order, and "3" representing a time overlap. For example, batch "XYZ-202311-025" is marked with conflict type code "2" because its production end timestamp is later than its delivery timestamp. The conflicting timestamp details are recorded as "Production end timestamp: 2023-11-06 18:00:00; Delivery timestamp: 2023-11-06 17:30:00."

[0087] The time window threshold is configured based on the system configuration file, which is stored in the server's " / etc / timestamp_rules" directory in JSON format. For example, the threshold configuration might be "{"material_type": "Conventional Raw Materials", "max_interval": 72}." Extreme input handling includes filtering for invalid timestamps (such as "0000-00-00 00:00:00"). The filtering rule discards unparseable date formats and logs them in the error log.

[0088] The conflict detection step is performed on the data processing server, running, for example, Linux CentOS 7.6, and relying on Python 3.8 and the Pandas 1.3.3 library for timestamp parsing and comparison. Data preprocessing includes timestamp format standardization and null-filling, using a forward-filling strategy to the most recent valid timestamp. Post-processing includes persistent storage of conflicting datasets and anomaly notifications, which are sent to quality management personnel via the company's email system.

[0089] S3. For the missing batch fields in the conflicting data, the conditional generative adversarial network is used to generate data to complete the missing fields based on the known fields. The discriminator confidence filter is combined to generate the resolved batch labeled data. The specific implementation is as follows:

[0090] During the conflict resolution step for cosmetics production batch data, missing batch fields in the conflicting data are complemented by a conditional generative adversarial network (GAN) based on known fields. Known fields are extracted from the conflicting dataset and include batch number, timestamp, and associated raw material name. The batch number format follows the "brand code - production year and month - serial number" format defined in S1, for example, "XYZ-202311-025." The timestamp format is consistent with the procurement timestamp in S1, "YYYY-MM-DD HH:MM:SS." The raw material name directly references the standardized name from the raw material procurement batch tag in S1. The input features of the known fields are converted into numeric vectors through data preprocessing. This preprocessing method includes splitting the batch number into three separate fields: brand code, production year and month, and serial number, and encoding each as an integer. The timestamp is converted to a numeric value in Unix timestamp format. The raw material name is mapped to a predefined raw material category code table (for example, glycerin is mapped to the code "GLY-001").

[0091] The generator of a conditional generative adversarial network generates data to complete missing fields based on known fields. The generator's network structure consists of an input layer that receives preprocessed known field vectors. The number of input layer nodes is equal to the encoding dimension of the known field. For example, if the known field is encoded as a 128-dimensional vector, the number of input layer nodes is set to 128. The hidden layer uses a fully connected neural network structure with the Reluctant Unit (ReLU) activation function. The number of output layer nodes matches the dimension of the missing field. For example, if the missing field is the logistics receipt status (enumeration value), the number of output layer nodes is set to 3, and the Softmax activation function is used to generate a probability distribution. The generator is trained with historical complete batch data. The Adam optimizer is used for training, with a learning rate of 0.001 and a cross-entropy loss function. Training is terminated when the validation set accuracy fluctuates by less than 1% for five consecutive epochs.

[0092] The discriminator of the conditional generative adversarial network scores the plausibility of the completed data. The discriminator's input layer receives the concatenation vector of the completed data and the known fields output by the generator. For example, if the known fields are 128-dimensional and the completed fields are 3-dimensional, the number of input layer nodes is set to 131. The hidden layer structure is the same as that of the generator, and the output layer uses a sigmoid function to generate a confidence score between 0 and 1. The preset confidence threshold is set by statistically analyzing the score distribution of manually reviewed completed data in historical data and taking the 95th percentile of this score distribution as the threshold. For example, if 95% of valid completed data in historical data have a score above 0.85, the threshold is set to 0.85. The discriminator is trained on the historical completed data and the manually annotated plausibility labels (0 / 1). The learning rate is set to 0.0001 to reduce the risk of overfitting, and the batch size is fixed at 32.

[0093] The filtered completed data is merged with the known fields in the original conflicting data to generate resolved batch tag data. This merging method inserts the completed fields into the corresponding positions in the original conflicting data by field name. For example, the completed logistics receipt status is inserted into the "Receipt Status" field of the warehouse logistics batch code, and the process parameters are inserted into the "Process Parameters" field of the production batch record. The merged data undergoes format verification, using the same verification rules as those in S1: the batch number must be 12 characters long (e.g., "XYZ-202311-025"), the timestamp format must match the regular expression "^\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}$", and the process parameter value range must conform to the safety threshold defined in the production execution system (e.g., the mixing temperature range is 45°C to 50°C). If the verification finds that the completed process parameter "Mixing Temperature" is 60°C, which exceeds the safety threshold defined in S1, the verification is considered failed and an exception alarm is triggered.

[0094] The verified digested batch tag data is stored in a standardized batch database. The table structure of the standardized batch database is consistent with the storage specifications of S1, including batch number, raw material name, procurement timestamp, production timestamp, logistics receipt status, and process parameter fields. A batch write strategy is adopted for storage, and a primary key conflict check is performed before each batch of data is written: if the batch number already exists, the original record is updated with the completion field; if it does not exist, a new record is inserted. The database access interface is connected via the JDBC protocol, the maximum number of concurrent connections in the database connection pool is set to 50, and the connection timeout is set to 30 seconds.

[0095] Conditional generative adversarial network training requires GPU-accelerated servers, for example, NVIDIA Tesla V100 graphics cards with 32GB of video memory, CUDA version 11.2, and the TensorFlow 2.6 deep learning framework. Training data must be standardized: numerical fields (such as Unix timestamps) are normalized to the range 0 to 1 using min-max scaling, and categorical fields (such as ingredient category codes) are converted to one-hot encoding. During training, a model snapshot is saved every 1000 iterations. Training logs record the loss function value and validation set accuracy. The log files are stored in the path / training_logs / cgan.

[0096] The extreme input handling mechanism includes filtering out invalid values ​​in the generator's output. For example, if the generator predicts a delivery status of "in transit" but leaves the delivery time blank, the system automatically replaces it with the default value of "shipped" and logs the exception. The exception handling log records the original data output by the generator, the replaced data, and the processing time. The log files are segmented and stored by date and are retained for 30 days.

[0097] S4. Construct a causal network of the digested batch labeled data and adverse reaction events, extract potential causal factors and filter unexplained isolated event data through attribute-constrained variational autoencoders. The specific implementation is as follows:

[0098] During the causal network construction step for cosmetic adverse reaction analysis, the digested batch tag data is associated with adverse reaction events through causal network modeling. The batch number, raw material name, process parameters, and logistics receipt status in the digested batch tag data serve as batch node attributes in the causal network. Batch numbers strictly follow the "brand code - production year and month - serial number" format defined in S1. For example, "XYZ-202311-025" represents the 25th batch of brand XYZ produced in November 2023. Raw material names reference the standardized name library for raw material procurement batch tags in S1, such as "GLY-001" for glycerin and "PG-002" for propylene glycol. Process parameters are inherited from the completed data in S3. For example, the mixing temperature field is stored as "45°C" and the filling pressure field is stored as "0.5 MPa." Logistics receipt status is extracted from the warehouse logistics batch code in S3 and includes three enumerated values: "shipped," "in transit," and "received," encoded in the format of "0," "1," and "2."

[0099] Standardized medical codes for adverse events were extracted from S1's consumer feedback data using the "Preferred Term" coding system of the Medical Dictionary for Reproductive Diseases (MedDRA). For example, "erythema" corresponds to code "10017456," and "pruritus" corresponds to code "10022130." Symptom nodes are linked to batch nodes via batch numbers. This association is achieved by establishing an edge relationship between batch nodes and symptom nodes with the same batch number in the causal network. For example, the batch node corresponding to batch number "XYZ-202311-025" is linked to symptom nodes "10017456" and "10022130," indicating that this batch of product may cause erythema and pruritus.

[0100] An attribute-constrained variational autoencoder encodes batch node attributes to generate latent causal factors. The encoder's input layer receives a batch node attribute vector, which includes the category code for the raw material name (e.g., glycerol is encoded as "GLY-001"), the numerical value of the process parameter (e.g., a mixing temperature of 45°C converted to a floating-point number 45.0), and a one-hot encoding of the logistics receipt status (e.g., "received" is encoded as [0,0,1]). Attribute constraints are implemented using the chemical stability threshold and toxicity risk level of the raw materials. The chemical stability threshold is extracted from the S1 quality inspection report, for example, the glycerol purity threshold is set at ≥99%. The toxicity risk level is obtained from industry toxicology databases (e.g., the ECHA database), for example, propylene glycol has a skin irritation rating of Category B (medium risk). The encoder's hidden layer uses a fully connected neural network architecture with 128 hidden layer nodes and a ReLU activation function. The output layer generates a 64-dimensional latent causal factor vector. The encoder was trained on historical batch node attribute data. The Adam optimizer was used for training, with a learning rate of 0.0005 and a batch size of 64. The loss function was the sum of the mean squared error and the attribute constraint penalty. The attribute constraint penalty was calculated as follows: if the raw material purity fell below the quality inspection report threshold or the toxicity level exceeded the safety limit, the loss function value was increased, and the penalty coefficient was set to 0.1.

[0101] The association weight between batch nodes and symptom nodes is calculated based on potential causal factors. This calculation involves two steps: first, the cosine similarity between the batch node's potential causal factor vector and the symptom node's standardized medical code vector is calculated using the dot product of the two vectors divided by the product of their moduli. The cosine similarity is then scaled to the range of 0 to 1 using min-max normalization, serving as the association weight. For example, if the potential causal factor vector for the batch node "XYZ-202311-025" is [0.2, 0.5, ..., 0.7] and the code vector for the symptom node "10017456" is [0.1, 0.6, ..., 0.8], the cosine similarity calculated is 0.75, resulting in a normalized association weight of 0.75. The preset causal threshold is set based on the distribution of verified causal relationship weights in historical data. Specifically, this is calculated by taking the 90th percentile of all historical association weights. For example, if 90% of the valid association weights in the historical data are ≥ 0.6, the threshold is set to 0.6.

[0102] Symptom nodes and their associated batch nodes whose association weights fall below the preset causal threshold are screened and marked as unexplained isolated event data. For example, if the association weight between the symptom node "10017456" and the batch node "XYZ-202311-025" is 0.45, which is below the threshold of 0.6, it is marked as an isolated event. Details of the association weights between the unexplained isolated event data and the causal network are stored in the isolated event analysis library. The storage structure of the isolated event analysis library includes batch number, symptom code, association weight, and timestamp fields. The format of each field is consistent with the S1 definition: batch number is a 12-character string (such as "XYZ-202311-025"), symptom code is an 8-digit number (such as "10017456"), association weight is a floating point number (such as 0.45), and timestamp format is "YYYY-MM-DD HH:MM:SS" (such as "2023-11-05 08:30:45").

[0103] Training of attribute-constrained variational autoencoders requires a GPU-accelerated server. For example, the server configuration is an NVIDIA Tesla V100 graphics card with 32GB of video memory, CUDA version 11.2, and the deep learning framework PyTorch 1.10. Training data must be pre-normalized: numerical fields (such as mixing temperature) are normalized to the range 0 to 1 using min-max scaling, for example, the temperature range of 45°C to 50°C is mapped to 0.0 to 1.0; categorical fields (such as the raw material name encoding) are converted to a one-hot encoding. During training, a model snapshot is saved every 1000 iterations. Training logs record the loss function value and validation set accuracy. The log file is stored in the path " / training_logs / causal_vae" and is formatted as JSON, containing the fields "epoch," "train_loss," and "val_accuracy."

[0104] The extreme input handling mechanism includes filtering and correcting illegal batch nodes. For example, if a batch node's process parameter, "mixing temperature 200°C," exceeds the safety threshold of "45°C to 50°C" defined in S1, the node is automatically discarded and an exception log is recorded. The exception log includes illegal data details (such as batch number, field name, and illegal value), the handling method (such as "discard"), and a timestamp. The log is stored in the path " / var / log / causal_network_exceptions" and is segmented by date (such as "exceptions_20231105.log"). The log is retained for 30 days.

[0105] S4 constructs a causal network between digested batch labeled data and adverse reaction events, and uses attribute-constrained variational autoencoders to extract latent causal factors. This effectively addresses the problem of inaccurate identification of risky batches due to broken causal chains in cosmetic adverse reaction analysis. Compared to existing technologies that rely solely on timestamp matching or statistical association, S4 accurately models the complex causal relationships between ingredients and adverse reactions by integrating batch attributes (such as raw material chemical stability and toxicity risk level) with deep associations of symptom codes. The introduction of attribute constraints ensures that the generation of latent factors is consistent with domain knowledge and avoids data-driven biases. The screening mechanism of association weights can identify abnormal events missed by traditional methods (such as unrecorded supply chain issues or ingredient interactions), thereby improving the comprehensiveness of the analysis, achieving dynamic and accurate batch risk attribution, and significantly improving report credibility and decision-making efficiency.

[0106] S5. For isolated event data, we build a spatiotemporal knowledge graph of batch regional diffusion and migration based on supply chain disruption event records and user location time series trajectories to identify potentially related batches. The specific implementation is as follows:

[0107] During the spatiotemporal knowledge graph construction step for cosmetic adverse reaction analysis, isolated event data is integrated with supply chain disruption records and user location trajectories to identify potentially associated batches. The batch number, symptom code, and occurrence timestamp in the isolated event data are extracted from the isolated event analysis library in S4. The batch number is identical to the digested batch tag data generated in S3. For example, the digestion data corresponding to the batch number "XYZ-202311-025" includes the process parameter "mixing temperature 45°C" and the logistics receipt status "received." Symptom codes use the terminology of the Medical Dictionary for Reproductive Diseases (MedDRA). For example, the symptom code "10017456" corresponds to "erythema." The occurrence timestamp format is consistent with that defined in S1, "YYYY-MM-DD HH:MM:SS," for example, "2023-11-05 08:30:45."

[0108] Supply chain disruption event records are extracted from S1's logistics management interface data, including the disruption type, disruption time, and affected geographical area. The disruption type is an enumeration, including "transportation delay," "warehouse accident," or "raw material shortage." For example, a disruption type of "transportation delay" indicates that the logistics process's transportation time exceeded the plan. The disruption time format is "YYYY-MM-DD HH:MM:SS," for example, "2023-11-05 10:00:00." The affected geographical area is represented by a GeoHash code. GeoHash codes are generated by converting longitude and latitude coordinates into a string encoding based on a geographic grid. The precision parameter is set to 6, corresponding to a geographic grid of approximately 1.2 km x 1.2 km. For example, the affected area for a transportation delay event is the GeoHash code "wx4g0" to "wx4g3." The matching logic is to filter the interruption events recorded in the logistics management interface for the same batch based on the batch number in the isolated event. For example, the interruption event associated with the batch number "XYZ-202311-025" is "Transportation Delay: 2023-11-05 10:00:00; Affected Region: wx4g0-wx4g3".

[0109] Geographic location information from the user's location time series trajectory was extracted from S1's adverse reaction events, including the usage timestamp and longitude and latitude coordinates. This geographic location information was converted to a GeoHash code using the GeoHash algorithm, with the precision parameter set to 6, corresponding to a geographic grid of approximately 1.2 km x 1.2 km. For example, the longitude and latitude coordinates "39.9042°N, 116.4074°E" were converted to the GeoHash code "wx4g0." The timestamp format used was consistent with the occurrence timestamp, for example, "2023-11-08 15:30:00," which was used for subsequent time overlap determination.

[0110] Based on the impact area of ​​the disruption event and the GeoHash codes used by users, a regional diffusion node is constructed for the spatiotemporal knowledge graph. Attributes of a regional diffusion node include batch number, disruption type, and geographic scope. For example, the geographic scope of the node "XYZ-202311-025_Transportation Delay" is the GeoHash code "wx4g0-wx4g3." The matching logic for the geographic scope is: if the GeoHash code used by a user falls within the impact area of ​​the disruption event, the user's location is associated with the corresponding regional diffusion node. For example, if a user with the GeoHash code "wx4g1" falls within the range "wx4g0-wx4g3," the user is associated with that node.

[0111] Based on the temporal overlap between the interruption occurrence time and the user usage timestamp, migration path edge relationships are generated in the spatiotemporal knowledge graph. Temporal overlap is determined by calculating the number of days between the interruption occurrence time and the user usage timestamp. If the interval is less than a preset time window, a migration path edge relationship is generated. The preset time window is based on the average duration of supply chain disruptions in historical statistical data. For example, by analyzing 1,000 historical interruption events, the distribution of time intervals from the interruption occurrence to user usage is calculated, and the median of the distribution is used as the preset time window (e.g., 7 days). If the interruption occurrence time is "2023-11-05 10:00:00" and the user usage timestamp is "2023-11-08 15:30:00", the interval is 3 days. If it is less than 7 days, it is considered a temporal overlap.

[0112] Screen out anomalous areas within the regional diffusion node where the logistics route is not covered but the user GeoHash code concentration exceeds the preset concentration threshold. Exclusion of the logistics route is determined by checking whether the GeoHash code used by the user is outside the transportation route of the warehouse logistics batch code supplemented by S3. The transportation route is extracted from the logistics receipt status data in S3. For example, the transportation route record is "wx4g0→wx4g1→wx4g2", but the user uses the code "wx4g4". Concentration is calculated by counting the proportion of user usage within the same GeoHash code grid to the total number of times. The preset concentration threshold is set to twice the standard deviation of the historical mean concentration of normal areas. For example, by analyzing historical data for areas unaffected by the disruption, the average concentration is calculated to be 5% with a standard deviation of 2%. The threshold is set to 9% (5% + 2 × 2%). If the concentration of a region is 12%, it is marked as an anomalous area.

[0113] Potentially associated batches are associated with the regional diffusion nodes and migration path edges of the spatiotemporal knowledge graph and stored in the batch association analysis library. The storage structure includes the batch number, the GeoHash code of the abnormal area, the associated interruption event, and time overlap details. For example, a storage entry might be "Batch number: XYZ-202311-025; abnormal area: wx4g4; interruption event: transportation delay; number of days of time overlap: 3 days." The database table structure is compatible with S1's standardized batch database, using the MySQL 8.0 storage engine. The index key is a combination of the batch number and the GeoHash code. A batch write strategy is used for storage, and a primary key conflict check is performed before each batch of data is written. If the batch number and GeoHash code combination key already exists, the existing record is updated; if not, a new record is inserted.

[0114] Among them, the batch number is consistent with the batch marking data after digestion, the supply chain interruption event record is inherited from the logistics management interface of S1, and the geographic location information is extracted from the adverse reaction event and includes the usage timestamp.

[0115] The construction of the spatiotemporal knowledge graph relies on a spatiotemporal data processing server. For example, the server operating system is Ubuntu 20.04, relying on Python 3.8 and the GeoHash 0.9.5 library for geocoding conversion, and the network visualization tool Gephi 0.9.2. The extreme input processing mechanism includes filtering for illegal GeoHash codes (such as those that do not meet the 6-character length requirement or contain illegal characters). The filtering rule discards unparseable codes and records them in the exception log. The exception log includes details of the illegal data (such as the original latitude and longitude, error code), the handling method (such as "discard"), and a timestamp. The log file format is CSV, the storage path is " / var / log / geohash_exceptions", and the retention period is 30 days.

[0116] S5 constructs a spatiotemporal knowledge graph by integrating supply chain disruption events with user location time series trajectories, addressing the potential missed detection of related batches in traditional cosmetics adverse reaction analysis due to unclear geographical diffusion paths or broken temporal correlations. Compared to existing technologies that rely solely on a single data source (such as logistics records or user feedback), S5 dynamically correlates the geographical impact of disruption events with user location clustering through a spatiotemporal knowledge graph, identifying areas not covered by logistics routes but with unusual user concentrations, thereby capturing indirectly related batches not recorded in supply chain data (such as cross-regional transfers or unexpected circulation). The attribute-constrained spatiotemporal overlap determination and aggregation threshold mechanism effectively filter out noise interference, ensure the reliability of the correlation results, and achieve precise batch traceability driven by multi-dimensional data fusion, significantly improving the coverage of cosmetics safety supervision and risk warning capabilities.

[0117] S6. Fusion of potentially associated batches and resolved batch label data to generate an adverse reaction analysis report with batch association probability weights. The specific implementation is as follows:

[0118] During the report generation step for cosmetic adverse reaction analysis, the potentially associated batches are fused with the digested batch labeling data to generate an adverse reaction analysis report annotated with batch association probability weights. The batch numbers, along with their associated anomaly regions and interruption types, are extracted from the potentially associated batches generated by S5. The batch numbers are identical to the digested batch labeling data generated by S3. For example, the digestion data corresponding to the batch number "XYZ-202311-025" includes the process parameter "mixing temperature 45°C" and the logistics receipt status "received." The anomaly regions are GeoHash-coded regions labeled in the spatiotemporal knowledge graph in S5, such as "wx4g4." The interruption types are inherited from the interruption event records in S5, such as "shipping delay."

[0119] Merge the batch number, raw material name, process parameters, and logistics receipt status fields from the digested batch tag data. The digested batch tag data is extracted from S3's standardized batch database, and the field format is exactly the same as that defined in S1: the batch number is a 12-character string (e.g., "XYZ-202311-025"), the raw material name uses the S1 standardized name library (e.g., "GLY-001" for glycerol), the process parameter value ranges comply with the safety thresholds defined in S1 (e.g., the mixing temperature range is 45°C to 50°C), and the logistics receipt status is an enumeration value of "shipped," "in transit," or "received."

[0120] The batch association probability weight is calculated based on the concentration of abnormal regions and the number of days of temporal overlap for potentially associated batches. The concentration of abnormal regions is extracted from the spatiotemporal knowledge graph in S5 and represents the percentage of user usage within the GeoHash-coded region compared to the total number of batch usages. For example, if a region has 120 user usages out of a total of 1,000, the concentration is 12%. The number of days of temporal overlap is the number of days between the supply chain disruption time calculated in S5 and the user usage timestamp. For example, if the disruption time is "2023-11-05 10:00:00" and the user usage timestamp is "2023-11-08 15:30:00," the interval is 3 days. The temporal overlap coefficient is calculated based on the ratio of the time interval to the preset time window in S5. The preset time window is set to 7 days based on historical disruption event analysis. The temporal overlap coefficient is calculated as the time interval divided by the time window, for example, 3 / 7 ≈ 0.43. The batch association probability weight is the product of the clustering ratio and the time overlap coefficient, for example, 12%×0.43≈5.16%.

[0121] The batch association probability weight is adjusted based on the association weights between the batch and symptom nodes in the S4 causal network. The S4 causal network association weight represents the strength of the causal relationship between the batch and the symptom. For example, the association weight between batch "XYZ-202311-025" and the symptom "erythema" is 0.75. The adjustment is performed by multiplying the batch association probability weight by the causal network association weight to generate a comprehensive association probability score, for example, 5.16% × 0.75 ≈ 3.87%. The comprehensive association probability score reflects both the temporal and spatial correlations between the batch and the adverse reaction, as well as the causal strength.

[0122] Sort batch numbers by comprehensive association probability score from high to low, and generate an analysis report with batch association probability weights. The sorting rule is descending order of comprehensive score. For example, if batch "XYZ-202311-026" has a score of 12% and batch "XYZ-202311-025" has a score of 3.87%, the former will be ranked higher than the latter. Analysis report fields include batch number, raw material name, process parameters, logistics receipt status, association probability weight, and comprehensive association probability score. For example, a report entry might be "Batch number: XYZ-202311-025; Raw material name: Glycerin; Process parameters: Mixing temperature: 45°C; Association probability weight: 5.16%; Comprehensive score: 3.87%."

[0123] The analysis report with the batch association probability weights is stored in a standardized report database. The table structure of the report database is compatible with the standardized batch database of S1, and the field types and lengths are exactly the same: the batch number is a 12-character fixed-length string, the raw material name is a 32-character string, the process parameter is a floating point number (such as 45.0), the logistics receipt status is a 3-character enumeration value, and the association probability weight and comprehensive score are floating point numbers in percentage format (such as 5.16). A batch write strategy is adopted for storage, and a primary key conflict check is performed before each batch of data is written: if the batch number already exists, the original record is overwritten with the new data; if it does not exist, a new record is inserted. The database access interface is connected through the JDBC protocol, the maximum concurrency of the database connection pool is set to 50, the connection timeout is set to 30 seconds, and the transaction isolation level is "repeatable read".

[0124] The calculation of batch association probability weights relies on the data processing server, for example, running Ubuntu 20.04, and relying on Python 3.8 and the Pandas 1.3.3 library for numerical computation. Extreme input handling includes filtering for invalid batch numbers or abnormal clustering. For example, if the batch number length does not exceed 12 characters (e.g., "XYZ-202311-025" is valid, "XYZ-202311-02" is invalid) or the clustering is negative, the data is discarded and an exception log is recorded. The exception log includes details of the invalid data (e.g., original batch number, error value), the handling method (e.g., "discard"), and a timestamp. The log file format is CSV, stored in the directory " / var / log / report_exceptions", and is retained for 30 days.

[0125] This embodiment overcomes the limitations of traditional single-data-source analysis by integrating multi-dimensional data fusion and dynamic correlation modeling throughout the overall cosmetics adverse reaction analysis process. Existing technologies typically process production batch data and adverse reaction events independently, or rely on static statistical methods to establish simple correlations, making it difficult to capture the complex causal relationships between supply chain disruptions, user geographic distribution, and ingredient risks. This embodiment dynamically correlates raw material procurement, production and processing, logistics and warehousing, and user feedback data through temporal conflict detection and causal network construction, forming a closed-loop analysis chain. A variational autoencoder based on attribute constraints incorporates raw material chemical stability and toxicity risk thresholds during the data completion phase, ensuring that generated data conforms to actual production constraints rather than relying solely on statistical distributions. A spatiotemporal knowledge graph integrates supply chain disruption events with user location temporal trajectories to identify anomalous clusters outside the logistics path, addressing the blind spots of traditional methods for unexpected distribution scenarios. Ultimately, by dynamically superimposing causal network weights and spatiotemporal correlation probabilities, a comprehensive quantitative assessment of batch risk is generated, achieving a complete closed-loop process from data conflict resolution to precise attribution. This embodiment overcomes the problems of broken causal chains and blurred regional diffusion paths in traditional analysis through a domain knowledge-driven data fusion mechanism and multi-level association modeling.

[0126] Example 2: Figure 2 A schematic diagram of a cosmetics adverse reaction analysis report generation optimization system of the present invention is provided. The cosmetics adverse reaction analysis report generation optimization system comprises:

[0127] Data acquisition module, used to obtain raw material procurement batch marks, production and processing batch records, warehousing and logistics batch codes and adverse reaction events of cosmetics production batches;

[0128] The conflict detection module is used to perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes, and identify conflicting data of the same batch at different links;

[0129] The data completion module is used to complete the missing batch fields in the conflicting data by generating missing field data based on known fields through a conditional generative adversarial network, and combining the discriminator confidence screening to generate the resolved batch labeled data;

[0130] The causal analysis module is used to construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and screen out unexplained isolated event data;

[0131] The spatiotemporal analysis module is used to construct a spatiotemporal knowledge graph of batch regional diffusion and migration based on isolated event data and supply chain disruption event records and user location time series trajectories to identify potentially related batches.

[0132] The report generation module is used to integrate the potentially associated batches and the resolved batch label data to generate an adverse reaction analysis report with the batch association probability weights.

[0133] It should be noted that the present invention can be deployed on the device itself to implement embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0134] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to actual conditions.

[0135] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0136] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0137] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0138] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0139] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0140] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0141] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0142] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for optimizing the generation of a cosmetic adverse reaction analysis report, characterized in that: include: S1. Obtain raw material purchase batch marks, production and processing batch records, warehousing and logistics batch codes, and adverse reaction events for cosmetic production batches; S2. Perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes to identify conflicting data for the same batch at different stages, including: Extract the purchase timestamp in the raw material purchase batch mark and align it with the production start timestamp in the production and processing batch record for verification; Compare the production end timestamp in the production and processing batch record of the same batch with the delivery timestamp in the warehousing and logistics batch code in chronological order; Based on the preset time window threshold, determine whether the interval between the purchase timestamp and the production start timestamp exceeds the allowable range, detect whether the production end timestamp is later than the delivery timestamp, and identify the reverse order or overlap of the timestamp sequences in different links of the same batch; The same batch whose purchase timestamp and production start timestamp interval exceeds the allowable range, whose production end timestamp is later than the delivery timestamp, or whose batches have reverse order or overlap are marked as conflicting data and stored in the conflict dataset; S3. For the batch fields missing in the conflicting data, the conditional generative adversarial network is used to generate data to complete the missing fields based on the known fields, and the discriminator confidence is combined to generate the resolved batch labeled data. S4. Construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and screen unexplained isolated event data, including: The batch number, raw material name, process parameters and logistics receipt status in the digested batch tag data are used as batch node attributes in the causal network; Map the standardized medical codes in adverse reaction events to symptom nodes in the causal network and associate them with the corresponding batch nodes through batch numbers; Encode batch node attributes to generate potential causal factors through attribute-constrained variational autoencoders; Calculate the association weights between batch nodes and symptom nodes based on potential causal factors to generate edge relationships in the causal network; Screening symptom nodes and their associated batch nodes whose correlation weights are lower than a preset causal threshold and marking them as unexplained isolated event data; S5. For isolated event data, a spatiotemporal knowledge graph of batch regional diffusion and migration is constructed based on supply chain disruption event records and user location time series trajectories to identify potentially related batches. S6. Fusion of potentially associated batches and resolved batch labeling data to generate an adverse reaction analysis report with batch association probability weights.

2. The method for optimizing the generation of a cosmetic adverse reaction analysis report according to claim 1, characterized in that: Obtain raw material purchase batch marks, production and processing batch records, warehousing and logistics batch codes, and adverse reaction events for cosmetic production batches, including: Extract the raw material purchase batch mark provided by the supplier from the raw material purchase order, and associate the raw material purchase batch mark with the corresponding raw material quality inspection report for storage; Export the production timestamp, process parameters and operator identification of each batch from the production execution system as production and processing batch records; Call the logistics management interface to extract the batch delivery time, transportation route and receipt status to generate the warehouse logistics batch code; Collect symptom description texts, occurrence time and usage scenario images submitted by consumers, and convert the symptom description texts into standardized medical codes to form adverse reaction events.

3. The method for optimizing the generation of a cosmetic adverse reaction analysis report according to claim 1, characterized in that: For batch fields missing in conflicting data, a conditional generative adversarial network is used to generate data to complete the missing fields based on known fields. This is combined with discriminator confidence filtering to generate resolved batch labeled data, including: Extract known fields from conflicting data as input features for conditional generative adversarial networks; Generate missing field completion data based on known fields through the generator of the conditional generative adversarial network; The discriminator of the conditional generative adversarial network is used to score the rationality of the completed data, and the completed data with a score higher than the preset confidence threshold is selected; Merge the filtered completed data with the known fields in the original conflicting data to generate resolved batch labeled data; Perform format check on the digested batch tag data to ensure that the value range and logical relationship of the completed fields are consistent with those of the known fields.

4. A cosmetic adverse reaction analysis report generation optimization method according to claim 3, characterized in that: Known fields include batch number, timestamp, and associated raw material name; missing fields include logistics receipt status or process parameters.

5. The method for optimizing the generation of a cosmetic adverse reaction analysis report according to claim 1, characterized in that: For isolated event data, we build a spatiotemporal knowledge graph of batch regional diffusion and migration based on supply chain disruption event records and user location time series trajectories to identify potentially related batches, including: Extract batch numbers and their associated symptom codes and occurrence timestamps from isolated event data; Match the disruption type, disruption occurrence time, and affected geographical area corresponding to the batch number in the isolated event data from the supply chain disruption event records; Convert the geographic location information in the user's location time series trajectory into GeoHash code; Based on the impact area of ​​the interruption event and the user's use of GeoHash encoding to build a regional diffusion node of the spatiotemporal knowledge graph, the attributes of the regional diffusion node include batch number, interruption type and geographical scope; Generate the migration path edge relationship of the spatiotemporal knowledge graph based on the time overlap between the interruption occurrence time and the user's usage timestamp; Screen out abnormal areas in the regional diffusion nodes that are not covered by the logistics path but have a user GeoHash coding concentration higher than the preset concentration threshold, and mark them as potential related batches; The potential associated batches are associated with the regional diffusion nodes and migration path edge relationships of the spatiotemporal knowledge graph and stored in the batch association analysis library.

6. A cosmetic adverse reaction analysis report generation optimization method according to claim 5, characterized in that: Time overlap is determined when the interval between the interruption occurrence time and the usage timestamp is less than the preset time window.

7. The method for optimizing the generation of a cosmetic adverse reaction analysis report according to claim 1, characterized in that: Fusion of potentially associated batches and resolved batch label data generates an adverse reaction analysis report with batch association probability weights, including: Extract batch numbers from potentially related batches and their associated exception areas and interruption types; Merge the batch number, raw material name, process parameters and logistics receipt status fields in the batch mark data after digestion; The batch association probability weight is calculated based on the abnormal area concentration of potential associated batches and the number of days of time overlap. The association probability weight is the product of the concentration ratio and the time overlap coefficient; Adjust the batch association probability weight based on the association weight between the batch node and the symptom node in the causal network to generate a comprehensive association probability score; Sort the batch numbers from high to low according to the comprehensive association probability scores, and generate an analysis report with the batch association probability weights marked.

8. A cosmetic adverse reaction analysis report generation optimization system, used to implement the cosmetic adverse reaction analysis report generation optimization method according to any one of claims 1 to 7, characterized in that: include: Data acquisition module, used to obtain raw material procurement batch marks, production and processing batch records, warehousing and logistics batch codes and adverse reaction events of cosmetics production batches; The conflict detection module is used to perform timestamp conflict detection on raw material procurement batch marks, production and processing batch records, and warehousing and logistics batch codes, and identify conflicting data of the same batch at different links; The data completion module is used to complete the missing batch fields in the conflicting data by generating missing field data based on known fields through a conditional generative adversarial network, and combining the discriminator confidence screening to generate the resolved batch labeled data; The causal analysis module is used to construct a causal network between the digested batch labeled data and adverse reaction events, extract potential causal factors through attribute-constrained variational autoencoders, and screen out unexplained isolated event data; The spatiotemporal analysis module is used to construct a spatiotemporal knowledge graph of batch regional diffusion and migration based on isolated event data and supply chain disruption event records and user location time series trajectories to identify potentially related batches. The report generation module is used to integrate the potentially associated batches and the resolved batch label data to generate an adverse reaction analysis report with the batch association probability weights.

Citation Information

Patent Citations

  • Food full-period tracing method based on food safety

    CN118379075A

  • Beverage production whole process tracing method

    CN119180661A