Cosmetic adverse reaction user report information automatic acceptance and feature recognition method

By identifying the timing marker words and implicit operation nodes in the cosmetic user report, a state transfer chart is constructed, and causal weight is calculated based on the topological structure entropy, the problem of missing causal correlation characteristics in the existing technology is solved, and the accuracy of monitoring of adverse reactions in cosmetics is improved.

CN120376027AActive Publication Date: 2025-07-25南昌市检验检测中心
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510873170.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The prior art cannot automatically reconstruct the implicit timing event chain between cosmetic use and symptom attacks, resulting in the lack of causal correlation characteristics, affecting the accuracy of monitoring of adverse cosmetic reactions.

Method used

By obtaining free text in user reports, identifying timing markers, inversely pushing implicit operation nodes, building a state transfer chart of the symptom state change node and the product usage operation node, generating causal weight values based on topological structure entropy, and outputting product-symptom causal correlation characteristics.

Benefits of technology

It significantly improves the accuracy of attribution of adverse reactions in cosmetics, accurately recognizes real cause and effect and time coincidence, adapts to the action mechanisms of different products, and solves the problems of fuzzy expression and logical faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120376027A_ABST
    Figure CN120376027A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic acceptance and feature recognition method for cosmetic adverse reaction user report information, particularly relates to the technical field of adverse reaction monitoring, and is used for solving the problem that causal associated features are missing due to the fact that an implicit time sequence event chain cannot be automatically reconstructed from a user free text in the prior art. The method comprises the following steps: acquiring a user report free text, identifying a time sequence tag word and positioning an operation chain fracture position; inversely deducing recessive operation nodes based on time density characteristics of symptom description; constructing a state transition diagram containing product use operation nodes and symptom state change nodes according to the time sequence tag words and the implicit operation nodes; generating a causal weight value based on the product of the node degree centrality distribution entropy value in the graph and the average shortest path depth from the symptom to the operation node; outputting a product-symptom causal association feature set through dynamic threshold comparison; conversion from fragmented description to structured event chain is realized, and limitation of traditional text analysis on fuzzy expression and logic fault is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of adverse reaction monitoring, and more specifically, to a method for automatically accepting and feature-recognizing user-reported information on cosmetic adverse reactions. Background Art

[0002] Monitoring of cosmetic adverse reactions is a key link in ensuring product safety. At present, enterprises and regulatory agencies generally use online systems to accept adverse reaction reports submitted by users, and automatically extract structured features such as symptoms and product names in the reports through natural language processing technology. Usually, users are required to fill in standardized forms, and at the same time, they are allowed to supplement descriptions in the free text column. For explicit information in the text (such as symptom keywords like rash and itching), the prior art can achieve basic feature extraction through entity recognition models.

[0003] However, the safety assessment of cosmetics also includes confirming the causal relationship between product use and symptom onset, which highly depends on the temporal logic chain (such as symptom remission after discontinuation and recurrence after reuse). Since users generally lack medical training, they often omit key temporal descriptions in the free text or only imply the causal logic with vague expressions (such as "it got better and then recurred later"). The prior art can only identify explicit entities and simple time markers in the text and cannot automatically reconstruct implicit temporal event chains, resulting in the lack of causal attribution ability in the feature set output by the system, making the automated risk assessment results deviate from the actual situation and seriously affecting the monitoring efficiency. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for automatically accepting and feature-recognizing user-reported information on cosmetic adverse reactions to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: A method for automatically accepting and feature-recognizing user-reported information on cosmetic adverse reactions, including: S1. Obtain a user report containing free text descriptions; S2. Identify temporal marker words in the free text description through natural language processing technology; S3. When there is an operation chain break in the temporal marker words, reverse-infer implicit operation nodes according to the time density characteristics of symptom descriptions in the free text description; S4. Construct a state transition diagram of symptom state change nodes and product use operation nodes according to the order and association relationship of the temporal marker words and the implicit operation nodes; S5. Generate the causal weight value between the product usage operation and the symptom change based on the topological structure entropy of the state transition diagram; wherein, the topological structure entropy is the product of the degree centrality distribution entropy value of all nodes in the state transition diagram and the average shortest path depth from the symptom state change node to the product usage operation node; S6. Compare the causal weight value with a preset threshold. When the causal weight value is greater than or equal to the preset threshold, output a feature set including the product-symptom causal association feature.

[0006] Further, obtain a user report including free text descriptions, including: Receive an adverse reaction report submitted by the user through an online system. The adverse reaction report includes standardized form data filled in by the user and supplementary description content in the free text column; Extract the supplementary description content in the free text column as the free text description, which includes records of cosmetic usage operation events and records of skin symptom change events.

[0007] Further, the record of cosmetic usage operation events refers to the cosmetic contact behavior actively performed by the user, and the record of skin symptom change events refers to the description of the biological phenomenon of the change in skin state.

[0008] Further, identify the temporal marker words in the free text description through natural language processing technology, including: Perform word segmentation on the free text description to generate a word sequence, and match the temporal marker words included in the word sequence based on a predefined temporal marker word library; Wherein the temporal marker word library includes four types of association markers: marker words indicating the product discontinuation operation, marker words indicating the symptom remission state, marker words indicating the product reuse operation, and marker words indicating the symptom recurrence state; Establish a position association mapping between the matched temporal marker words and the corresponding cosmetic usage operation events record or skin symptom change events record in the free text description.

[0009] Further, when there is an operation chain break in the temporal marker words, infer the implicit operation node according to the time density feature of the symptom description in the free text description, including: Detect the symptom description text corresponding to the skin symptom change events record in the free text description; Extract the occurrence frequency of the symptom suddenness description words and the occurrence frequency of the symptom progression description words in the symptom description text as the time density feature; When the occurrence frequency of the symptom suddenness description words is greater than the occurrence frequency of the symptom progression description words, generate the product start usage operation in the cosmetic usage operation events record as the implicit operation node; When the frequency of occurrence of progressive symptom descriptors is greater than or equal to the frequency of occurrence of sudden symptom descriptors, generate the continuous product usage operation in the cosmetic usage operation event record as a hidden operation node.

[0010] Further, the sudden symptom descriptors include "suddenly", "immediately", and "abruptly", and the progressive symptom descriptors include "gradually", "slowly", and "step by step".

[0011] Further, according to the order and correlation relationship of the time sequence markers and the hidden operation nodes, construct a state transition diagram of symptom state change nodes and product usage operation nodes, including: Extract the cosmetic usage operation event record corresponding to the time sequence marker as an explicit product usage operation node; Extract the skin symptom change event record corresponding to the time sequence marker as a symptom state change node; Merge the hidden operation node and the explicit product usage operation node to form a complete set of product usage operation nodes; Arrange the complete set of product usage operation nodes and symptom state change nodes in the order of the time when the event records appear in the free text description; Establish a directed connection edge between adjacent nodes to form a state transition diagram from symptom state change nodes to product usage operation nodes, where the product usage operation node points to the symptom state change node directly caused by it.

[0012] Further, generate a causal weight value between product usage and symptom changes based on the topological structure entropy of the state transition diagram, including: Calculate the degree centrality value of all nodes in the state transition diagram. The degree centrality value represents the number of direct connections between a node and other nodes; Form a degree centrality distribution according to the degree centrality values of all nodes, and calculate the entropy value of the degree centrality distribution as the degree centrality distribution entropy value; Calculate the shortest path depth from each symptom state change node to each product usage operation node in the state transition diagram. The shortest path depth refers to the minimum number of edges connecting two nodes; Take the arithmetic mean of the shortest path depths from all symptom state change nodes to product usage operation nodes as the average shortest path depth; Multiply the degree centrality distribution entropy value by the average shortest path depth to obtain the topological structure entropy value, which is used as the causal weight value between product usage and symptom changes.

[0013] Further, compare the causal weight value with a preset threshold. When the causal weight value is greater than or equal to the preset threshold, output a feature set including product-symptom causal association features, including: Compare the causal weight value with a preset causal determination threshold numerically; When the causal weight value is greater than or equal to the causal determination threshold, extract the cosmetic name information corresponding to the product usage operation node in the state transition diagram; Extract the skin symptom description information corresponding to the symptom state change node in the state transition diagram; Generate a product-symptom causal association feature including the cosmetic name information, the skin symptom description information, and the causal weight value; Add the product-symptom causal association feature to the feature set as the output content.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly improve the accuracy of attributing cosmetic adverse reactions through a multi-level event chain reconstruction mechanism; different from traditional text classification that only identifies explicit entities, first accurately capture the discontinuous temporal logic in free text through context-sensitive temporal marking; then reverse-infer the missing operation nodes based on the symptom time density feature to fill the break in the operation chain in the user description; finally, form a complete state transition diagram including explicit and implicit nodes to achieve the dynamic topological association between operations and symptoms; this transformation mechanism from fragmented descriptions to structured event chains solves the defect that the prior art is powerless against fuzzy expressions and logical breaks.

[0015] 2. Break through the limitations of traditional statistical correlation through graph topology quantization technology; the topological structure entropy algorithm combines the network centrality distribution and the path depth feature: reveal the global influence of key operation nodes and quantify the causal conduction efficiency, and the product of the two forms a weight index with causal directionality; combined with the dynamic threshold adjustment based on product categories, the finally output product-symptom causal association feature set has three major advantages: one is to accurately distinguish real causality from temporal coincidence; the second is to identify cross-period implicit associations; the third is to adapt to different product action mechanisms. Description of the Drawings

[0016] Figure 1 It is a flowchart of the method for automatically accepting and feature-recognizing user reports of cosmetic adverse reactions of the present invention; Figure 2 It is a flowchart of reverse-inferring implicit operation nodes of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment: Figure 1The present invention provides an automatic acceptance and feature recognition method for user reports of adverse reactions of cosmetics, including: S1. Obtain user reports containing free-text descriptions; S2. Identify temporal marker words in the free-text descriptions through natural language processing techniques; S3. When there is an operation chain break in the temporal marker words, reverse-infer hidden operation nodes based on the time density characteristics of symptom descriptions in the free-text descriptions; S4. Construct a state transition diagram of symptom state change nodes and product usage operation nodes according to the order and association relationship of the temporal marker words and the hidden operation nodes; S5. Generate a causal weight value between product usage operations and symptom changes based on the topological structure entropy of the state transition diagram. The topological structure entropy is the product of the degree centrality distribution entropy value of all nodes in the state transition diagram and the average shortest path depth from the symptom state change nodes to the product usage operation nodes; S6. Compare the causal weight value with a preset threshold. When the causal weight value is greater than or equal to the preset threshold, output a feature set containing product-symptom causal association features.

[0019] S1. Obtain user reports containing free-text descriptions, and the specific implementation is as follows: When a user submits an adverse reaction report through an online system, the online system provides a standardized form containing fixed fields for the user to fill in. The fixed fields include user basic information fields, cosmetic name fields, usage date fields, and symptom type selection fields; a free-text column input box is set at the end of the standardized form, and the free-text column input box allows the user to input supplementary description content of any length; receive the complete adverse reaction report data packet submitted by the user, and the data packet contains the standardized form data stored in a structured manner and the free-text column supplementary description content string stored independently.

[0020] Separate the free text column supplementary description content string from the data packet as the original free text description. The original free text description is a continuous character sequence containing natural language sentence segments. The semantic parsing unit scans the verb phrase structures in the original free text description. When the verb phrase contains a cosmetic operation action directly acting on human skin tissue, extract the verb phrase and its object component to form a cosmetic usage operation event record. Cosmetic operation actions include smearing operations, spraying operations, applying operations, and removing operations. Among them, the smearing operation refers to the act of applying cosmetics to the skin surface through external force; the spraying operation refers to the act of making cosmetics contact the skin through an atomizing device; the applying operation refers to the act of making cosmetics continuously contact the skin through a carrier material; the removing operation refers to the act of using a cleaning substance to terminate the contact between cosmetics and the skin. At the same time, scan the physiological state description structure in the original free text description. When the description structure contains visible trait changes or subjective perception changes of the skin tissue, extract the description structure to form a skin symptom change event record. Visible trait changes include skin color change phenomena, skin texture change phenomena, and skin integrity change phenomena. Subjective perception changes include abnormal temperature perception phenomena, abnormal tactile perception phenomena, and abnormal pain perception phenomena. Among them, the skin color change phenomenon is manifested as erythema phenomenon or pigmentation phenomenon; the skin texture change phenomenon is manifested as roughness phenomenon or swelling phenomenon; the skin integrity change phenomenon is manifested as desquamation phenomenon or blister phenomenon.

[0021] The cosmetic usage operation event record and the skin symptom change event record together constitute the final free text description data set, and this data set is stored in chronological order of event occurrence based on the timestamp. For example, when the user inputs text containing "smear XX essence", this text is recognized as a cosmetic usage operation event record; when the user inputs text containing "facial patchy erythema", this text is recognized as a skin symptom change event record.

[0022] In the data extraction process, the specific parsing rules for the cosmetic usage operation event record include: when the cosmetic name and the usage action verb are directly collocated in the text, it is determined as a valid operation event; when the cosmetic name is missing but the action verb belongs to the preset cosmetic operation verb library, it is marked as a product information missing operation event and triggers the subsequent supplementary query process.

[0023] The specific parsing rules for the skin symptom change event record include: when the standard skin symptom terms appear in the text, directly match and record; when non-standard descriptions appear, map them to standard terms through the symptom synonym conversion table. For example, map the non-standard description "develop small bumps" to the standard term "papule", and map the non-standard description "peel off" to the standard term "desquamation".

[0024] All event records are associated with position markers in the original text. Position markers record the start character index and end character index in the string of the event description in the free text column. Time logic breaks are detected through event record integrity verification. When the time interval between adjacent event records exceeds the preset threshold, a time logic break marker is added. The event record storage format uses a JSON data structure, which contains an event identifier field, an event type field, an event description text field, a timestamp field, an original text location field, and an associated cosmetic name field.

[0025] The preset resources involved in the execution of this step include: the cosmetics operation verb library contains dozens of basic verbs and their inflected forms, the skin symptom term library contains the core symptom words defined by the international common skin disease classification standard, and the symptom synonym conversion table includes the mapping relationship between consumers' commonly used non-standard description phrases and standard terms. All resource libraries are loaded into the memory hash table through the configuration file to achieve fast matching. The hash table key is the original text entry and the value is the standardized output result. The sliding window scanning mechanism is used to process the string of the supplementary description content in the free text column. The window size is set according to the longest matching principle of Chinese word segmentation. When the text in the window matches the cosmetics operation verb and the skin symptom term at the same time, it is divided into independent event records according to the order of the event occurrence. The timestamp extraction uses regular expressions to match the date and time pattern. When the text does not clearly specify the time information, the relative time series is assigned by default according to the order of the statement. The final output free text description data set is written into the distributed file system, and the storage path is recorded in the metadata table of the main database.

[0026] To verify the accuracy of data extraction, a test set was used for verification. The extraction error of the cosmetics use operation event record mainly comes from the fact that the dialect verbs are not covered by the operation verb library. This error is solved by expanding the resource library; the extraction error of the skin symptom change event record mainly comes from the fact that the atypical symptom description is not included in the synonym conversion table. This error is solved by expanding the mapping table. The time sequence accuracy of the event record verifies the effectiveness of the timestamp extraction mechanism. The integrity of data storage is verified by the verification mechanism.

[0027] S2. Identify the time sequence marker words in the free text description by natural language processing technology, which is specifically implemented as follows: Obtain a free text description data set as input data, which includes records of cosmetic usage operation events and records of skin symptom change events. Perform Chinese word segmentation on the original text strings in the free text description data set. The Chinese word segmentation is implemented by combining the bidirectional maximum matching algorithm with a probability model. The specific execution process includes: scanning the original text string by character sequence, performing forward maximum matching segmentation and reverse maximum matching segmentation through a pre-constructed dictionary, comparing the number of single characters in the two segmentation results, and selecting the segmentation scheme with fewer single characters as the word segmentation result; calling the probability model to calculate the optimal segmentation path for the segmentation ambiguity positions, and the state transition probability of the probability model is obtained through training with annotated corpora.

[0028] The word segmentation process outputs a sequence of words, which is composed of word units arranged in the order of the original text. Each word unit contains the word text content, the starting position index of the word, and the word length information. For example, when the original text contains "After stopping using the lotion, the erythema subsided", this text is segmented into the word sequence [stop using, lotion, after, erythema, subside], where the character position index range of "stop using" is from 0 to 2.

[0029] Load a predefined time series marker word library to perform matching operations. This library contains four types of associated marker word sets: The product discontinuation operation marker word set contains verbs indicating the termination of cosmetic use, such as words like stop using, discard, no longer use, etc.; The symptom remission status marker word set contains verbs or adjectives indicating symptom alleviation, such as words like subside, alleviate, relieve, etc.; The product reuse operation marker word set contains verb phrases indicating restarting use, such as phrases like reuse, start using again, use again, etc.; The symptom recurrence status marker word set contains verb phrases indicating the reappearance of symptoms, such as phrases like recurrence, reappear, occur again, etc.

[0030] The word library is stored using a hash table data structure. The hash table key is the word text, and the value is the marker type. Traverse each word unit in the word sequence, query whether the text content of the word unit exists in the hash table of the time series marker word library. When it exists, record the time series marker word information that matches successfully, which includes the marker word text content, the marker type, and the position index in the word sequence.

[0031] Establish a position association mapping between the time series marker words and the event records. This process includes: using the position interval of the time series marker word in the original text, which is determined by the starting position index and the ending position index of the word unit; querying in the free text description data set for event records that overlap with this position interval, and the position information of the event records comes from the original text position field recorded in the data set.

[0032] When the position interval of the temporal marker word overlaps with the text position interval of the cosmetic usage operation event record, an association mapping between the temporal marker word and the cosmetic usage operation event record is established; when the position interval of the temporal marker word overlaps with the text position interval of the skin symptom change event record, an association mapping between the temporal marker word and the skin symptom change event record is established. The mapping relationship is stored as a triple data structure, including the temporal marker word identifier, the event record identifier, and the association type identifier. For example, if the position interval of the temporal marker word "stop using" from 0 to 2 overlaps with the position interval of the cosmetic usage operation event record "stop using lotion" from 0 to 4, an association mapping between the two is established.

[0033] When dealing with polysemous words, context analysis rules are adopted: when a word unit matches multiple marker types, analyze the semantic features within the front and back word windows. If a cosmetic name appears within the window, it is preferentially determined as a product operation type marker; if a symptom description appears, it is preferentially determined as a symptom status type marker. A distance constraint mechanism is set up, and when the position distance between the temporal marker word and the nearest event record exceeds the threshold, a processing flow is triggered. For example, this threshold can be set to a 50-character distance. After establishing the mapping relationship, conflict detection is performed. When a temporal marker word is mapped to multiple event records, the event record with the highest text overlap degree is selected. The text overlap degree is calculated as the number of overlapping characters divided by the total number of characters of the temporal marker word.

[0034] The thesaurus construction is achieved by analyzing historical data: collect a set of historical free text description data, and annotate the temporal relationship words in it; use the word frequency statistics algorithm to screen out feature words, and select high-frequency words as the basic thesaurus; add similar words through semantic analysis to form a complete thesaurus. The thesaurus update is set to be incrementally updated regularly, and the unregistered words in the new user reports are processed and added to the library after being processed. A pre-screening mechanism is established during the matching execution process to optimize the processing efficiency. The pre-screening mechanism quickly judges the text of the word unit, and only when the pre-screening mechanism returns a possibility does an exact query execute.

[0035] The test set is used to verify the recognition effect of the temporal marker words. The test set contains the marked positions and types of the temporal marker words. The comparison between the word segmentation processing results and the marked results shows that the processing effect meets the requirements. The matching effect of the temporal marker words is evaluated by the proportion of unregistered words, and the number of newly added marker words is regularly counted to optimize the thesaurus. The accuracy of the position association mapping is verified by calculating the overlap degree, and when the text expression is standardized, the mapping effect meets the requirements. The mapping conflict handling mechanism is verified through cases, and the conflict cases are resolved after being processed. The processing efficiency is verified through tests, and it supports processing user report data.

[0036] Figure 2 The flowchart for reverse-inferring the implicit operation nodes of the present invention is given. In S3, when there is an operation chain break in the temporal marker words, the implicit operation nodes are reverse-inferred according to the time density characteristics of the symptom description in the free text description. The specific implementation is as follows: When an operation chain break marker is detected, the operation chain break marker is derived from the result of the integrity check of event records in the previous step, and this marker indicates that there is a time logic gap between adjacent cosmetic use operation event records. Extract all skin symptom change event records from the free text description data set, and for each skin symptom change event record, obtain its corresponding symptom description text field, where the symptom description text field is a natural language description fragment recording the change of skin state. For example, when the skin symptom change event record contains the text content "sudden redness, swelling and itching on the face", extract this text as the symptom description text being processed currently. Perform preprocessing operations on the symptom description text, and the preprocessing includes removing meaningless auxiliary words and punctuation marks. The list of meaningless auxiliary words includes common function words such as "de", "le", "ne", etc. After preprocessing, the symptom description text is converted into a continuous sequence of valid words.

[0037] Load the predefined descriptor set to perform feature extraction. The descriptor set is divided into two categories: the symptom suddenness descriptor set contains adverbs or adjectives indicating the rapid occurrence of symptoms, such as the words "suddenly", "immediately", "abruptly", "right away", "all of a sudden"; the symptom gradualness descriptor set contains adverbs or adjectives indicating the slow development of symptoms, such as the words "gradually", "slowly", "step by step", "little by little", "day by day". Count the number of occurrences of suddenness descriptors and gradualness descriptors in the symptom description text. The number of occurrences is obtained by traversing the sequence of valid words and precisely matching with the descriptor set. Calculate the occurrence frequency of suddenness descriptors and the occurrence frequency of gradualness descriptors. The frequency value is the number of occurrences of a specific type of descriptor divided by the total number of valid words in the symptom description text. The total number of valid words refers to the number of remaining words after removing meaningless auxiliary words. For example, after preprocessing the symptom description text "red spots suddenly appear and slowly spread", the sequence of valid words obtained is ["red spots", "suddenly", "appear", "and", "slowly", "spread"], the total number of valid words is 6, the suddenness descriptor "suddenly" appears 1 time, and the gradualness descriptor "slowly" appears 1 time. Then the occurrence frequency of the suddenness descriptor is one-sixth, and the occurrence frequency of the gradualness descriptor is one-sixth.

[0038] The time density feature is defined as a binary group containing the frequency value of the sudden descriptor and the frequency value of the progressive descriptor. Compare the size relationship of the two frequency values: when the frequency of the sudden descriptor is strictly greater than the frequency of the progressive descriptor, a new cosmetic use operation event record is generated as an implicit operation node, and the event type field of the event record is set to "product start use operation"; when the frequency of the progressive descriptor is greater than or equal to the frequency of the sudden descriptor, a new cosmetic use operation event record is generated as an implicit operation node, and the event type field of the event record is set to "product continuous use operation". For example, when the calculated frequency of the sudden descriptor is 0.25 and the frequency of the progressive descriptor is 0.1, because 0.25 is greater than 0.1, the generation of an implicit operation node with the event type of "product start use operation" is triggered.

[0039] The attribute setting of the implicit operation node follows the following rules: the cosmetic name attribute is inherited from the explicit product use operation node that is closest in time order. If there is no explicit node before or after, the attribute is marked as "unknown product"; the timestamp attribute is calculated based on the time when the associated symptoms appear. For the "product start use operation" type node, its timestamp is set to the timestamp of the associated skin symptom change event record shifted forward by a fixed time. The fixed time is set to an empirical value such as 24 hours based on the cosmetic sensitization incubation period; for the "product continuous use operation" type node, its timestamp is set to the midpoint value of the timestamps of two adjacent explicit product use operation nodes. The generated implicit operation node is added to the cosmetic use operation event record collection, and the "implicit derivation" flag is set in the metadata. For example, when the associated symptom appears at 9:00 a.m. on May 10, 2023 and the "product start use operation" node is generated, the timestamp of the node is set to 9:00 a.m. on May 9, 2023.

[0040] The specific conditions for detecting the breakage of the operation chain are as follows: In the free text description data set, if the time interval difference between two adjacent cosmetic usage operation event records exceeds the maximum reasonable usage interval corresponding to the product category, it is marked that there is a breakage in the operation chain at that position. The maximum reasonable usage interval is set according to the product usage characteristics. For example, it is set to 7 days for skin care products and 3 days for makeup products. The symptom description text frequency calculation uses the relative frequency algorithm. The numerator is the actual number of occurrences of a specific type of descriptor in the text, and the denominator is the total number of valid words after preprocessing the symptom description text. The handling of boundary cases includes the following rules: When no sudden or progressive descriptors appear in the symptom description text, a hidden node with the event type of "continuous product usage operation" is generated by default; when the occurrence frequency of the sudden descriptor is equal to the occurrence frequency of the progressive descriptor, a secondary analysis mechanism is activated. The secondary analysis mechanism calculates the weighted value of the intensifying adverbs in the symptom description text. The intensifying adverbs include words such as "very", "extremely", "especially", etc. Each occurrence of an intensifying adverb is counted as 1 point. When the weighted value exceeds the threshold, a "product start usage operation" node is generated; otherwise, a "continuous product usage operation" node is generated.

[0041] To verify the reliability of the hidden node derivation mechanism, multiple groups of test cases were designed for verification. In Case 1, the symptom description text is "Suddenly, the face became red, swollen, and painful after using the facial mask". After processing, it was detected that "suddenly" is a sudden descriptor and there is no progressive descriptor. The occurrence frequency of the sudden descriptor is one-fifth, triggering the generation of a "product start usage operation" node. After manual review, this derivation conforms to the sudden characteristics of cosmetic allergies. In Case 2, the symptom description text is "After continuous use for a week, the skin gradually became dry and flaky". It was detected that "gradually" is a progressive descriptor and there is no sudden descriptor. The occurrence frequency of the progressive descriptor is one-sixth, triggering the generation of a "continuous product usage operation" node. The review result conforms to the cumulative irritation response pattern. In Case 3, the symptom description text is "The itching recurs", and no descriptors appear, triggering the default rule to generate a "continuous product usage operation" node. The timestamp calculation logic was verified through historical case backtracking. For example, a certain case records "Suddenly developed a rash on May 10th", generating a "started using on May 9th" node. After subsequent confirmation by the user, the actual usage date was the evening of May 9th, verifying that setting the 24-hour forward push is reasonable. In the frequency calculation stability test, when the length of the symptom description text exceeds 20 words, the coefficient of variation of the calculation result is less than 5%, meeting the data processing accuracy requirements. The handling of boundary cases was verified through hundreds of special samples, and the secondary analysis mechanism effectively distinguished degree difference scenarios such as "very sudden" and "slightly gradual". The integrated test of the entire process shows that after the addition of hidden nodes, the elimination rate of the operation chain breakage mark meets the requirements, and the integrity of the state transition diagram construction is guaranteed.

[0042] S4. According to the order and association relationship of the timing marker words and the implicit operation nodes, construct a state transition diagram of the symptom state change nodes and the product usage operation nodes. The specific implementation is as follows: Obtain the position association mapping data of the timing marker words and the implicit operation node data. The position association mapping data of the timing marker words is derived from the association relationship data between the timing marker words and the event records generated in step S2. This data includes the timing marker word identifier, the event record identifier, and the association type identifier. The implicit operation node data is derived from the implicit derivation nodes in the cosmetic usage operation event records generated in step S3. This node includes an event type field, a timestamp field, and an implicit derivation identifier. Screen all event records associated with the timing marker words from the set of cosmetic usage operation event records as explicit product usage operation nodes. The attributes of the explicit product usage operation nodes include a cosmetic name attribute, an operation type attribute, and a timestamp attribute. The cosmetic name attribute is directly inherited from the associated cosmetic name field of the event record. The operation type attribute is set according to the type of the timing marker word. For example, when the timing marker word is a product deactivation operation marker word, the operation type attribute is set to "deactivation operation". At the same time, extract all event records associated with the timing marker words from the set of skin symptom change event records as symptom state change nodes. The attributes of the symptom state change nodes include a symptom description text attribute and an occurrence timestamp attribute. The occurrence timestamp attribute is taken from the timestamp field of the event record. For example, when the position association mapping data includes the timing marker word "deactivation" associated with the cosmetic usage operation event record "deactivate lotion", an explicit product usage operation node "deactivate lotion" is generated; when the timing marker word "subsided" is associated with the skin symptom change event record "the erythema subsided", a symptom state change node "the erythema subsided" is generated.

[0043] Merge the implicit operation nodes and the explicit product usage operation nodes to form a complete set of product usage operation nodes. During the merging process, retain the original attribute data of all nodes, add an additional node type attribute to the implicit operation nodes and assign it the value "implicit", and assign the value "explicit" to the node type attribute of the explicit product usage operation nodes. The complete set of product usage operation nodes and the set of symptom state change nodes together constitute all the node elements of the state transition diagram. The node data structure is stored in a unified format. Each node contains five attribute fields: the node identifier field uses a globally unique code, and the coding rule is the user number plus the timestamp hash value; the node type field takes the values "product operation" or "symptom state"; the timestamp field records the event occurrence time, accurate to the minute level of precision; the associated entity name field stores the cosmetic name or the symptom standard term; the description text field saves the original event description text.

[0044] Perform a chronological sorting operation on all nodes in the set of operation nodes and the set of symptom state change nodes for the complete product. The sorting is based on the numerical values of the node timestamp fields in ascending order, and the timestamp values are stored in Unix timestamp format. When the timestamp values of different nodes are the same, a secondary sorting is performed according to the original character position index of the node corresponding event record in the free text description, and the node with a smaller character position index is ranked first. The sorting algorithm is implemented using the quicksort algorithm, which completes the sorting through recursive partitioning operations, with a time complexity of O(n log n), meeting the real-time requirements of the system. After sorting, an ordered node sequence is generated, which completely reflects the chronological relationship of event occurrences. For example, when there are three nodes: an explicit product usage operation node P1 (timestamp 1654041600), a symptom state change node S1 (timestamp 1654128000), and an implicit operation node P2 (timestamp 1654214400), the sorted node sequence is [P1, S1, P2].

[0045] Establish directed connection edges between adjacent nodes in the ordered node sequence. The connection rules are determined according to the type combination of adjacent nodes: when the predecessor node in the adjacent nodes is a product usage operation node and the successor node is a symptom state change node, a directed edge is established from the predecessor node to the successor node; when both adjacent nodes are product usage operation nodes, a directed edge is established from the previous operation node to the next operation node; when both adjacent nodes are symptom state change nodes, a directed edge is established from the previous symptom node to the next symptom node. Each directed edge is set with a weight attribute, and the initial weight value is uniformly set to 1. For example, when the ordered node sequence is product usage operation node P1 → symptom state change node S1 → product usage operation node P2 → symptom state change node S2, the established connection edges include: a directed edge from P1 to S1, a directed edge from S1 to P2, and a directed edge from P2 to S2.

[0046] The storage structure of the state transition graph is implemented using an adjacency list data structure. The adjacency list uses the node identifier as the primary key, and the corresponding value is the list of successor nodes directly connected to the node. The list item contains the successor node identifier and the connection edge weight value. Calculate the influence strength index for each product usage operation node, which is obtained by calculating the ratio of the number of outgoing edges pointing to symptom state change nodes to the total number of outgoing edges of the node. The visual rendering of the state transition graph uses a timeline layout algorithm, which converts the timestamp into the abscissa value, the vertical coordinate of the product usage operation node is set at 50 pixels above the timeline, the vertical coordinate of the symptom state change node is set at 50 pixels below the timeline, and the directed edge is rendered as a Bezier curve with an arrow.

[0047] The special scenario processing mechanism includes the following rules: When there is no adjacent product usage operation node for the symptom status change node in the sequence, search forward for the nearest product usage operation node to establish a cross-node connection edge. The maximum number of cross-nodes is set to 3 position intervals, and this value is set according to the statistical analysis of event density. When different event timestamps overlap, sort by node type priority, and the product usage operation node has a higher priority than the symptom status change node. When a single product usage operation node is associated with multiple symptom status change nodes, establish a one-to-many connection relationship. For example, the operation node "using a facial mask" can point to the symptom nodes "facial flushing" and "itching" at the same time. When the symptom status change node describes the subsiding or recurrence state, establish a reverse connection edge from this symptom node to the subsequent operation node and set the edge type attribute to "status feedback".

[0048] Verify through typical test cases: In Case 1, the user report contains the description "using Essence A on May 1, facial swelling on May 3, stopping using this product on May 5, and the swelling subsiding on May 7". The constructed ordered node sequence is: the explicit product usage operation node "using Essence A" (timestamp May 1) → the symptom status change node "facial swelling" (May 3) → the explicit product usage operation node "stopping using Essence A" (May 5) → the symptom status change node "swelling subsiding" (May 7). The established connection edges include "using Essence A" pointing to "facial swelling", "facial swelling" pointing to "stopping using Essence A", and "stopping using Essence A" pointing to "swelling subsiding". In Case 2, when processing the scenario of "sudden rash on May 10 (no previous operation record)", generate an implicit operation node "unknown product started to be used" (timestamp May 9) through implicit node derivation, forming a sequence: implicit operation node → the symptom status change node "sudden rash", and establish a directed edge from the operation node to the symptom node. In Case 3, when processing the multi-symptom scenario of "burning sensation and erythema appear after using sunscreen", when the symptom node timestamps are the same, generate a sequence: the explicit product usage operation node "using sunscreen" → the symptom status change node "burning sensation" → the symptom status change node "erythema", and establish connection edges of "using sunscreen" pointing to "burning sensation" and "burning sensation" pointing to "erythema".

[0049] The adjacency list data structure supports efficient query operations. For example, the query response time is less than 10 milliseconds at a scale of 1000 nodes; the timeline layout accurately reflects the sequence of events, and the coordinate conversion error rate is less than 1%; the special scenario processing rules effectively solve 98.5% of the boundary cases. The state transition graph data output includes three parts: a node set, an edge set, and metadata. The metadata records the mapping relationship between the original event identifier and the graph node for subsequent causal analysis processes to call.

[0050] S5. Generate the causal weight value between the product usage operation and the symptom change based on the topological structure entropy of the state transition diagram. The specific implementation is as follows: Obtain the state transition diagram data structure constructed in the previous step as the input data. This state transition diagram data structure includes node set data, edge set data, and an adjacency list storage structure. The node set data includes all product usage operation nodes and symptom state change nodes. Each node contains five attributes: The node identifier attribute stores the globally unique identification code in string format; the node type attribute takes values of "product operation" or "symptom state"; the timestamp attribute records the event occurrence time and is stored as a Unix timestamp value; the associated entity name attribute saves the cosmetic name or the symptom standard terminology text; the description text attribute retains the original event description content. The edge set data includes directed connection edge information. Each edge includes a start node identifier attribute, an end node identifier attribute, and a weight attribute value. The initial value of the weight attribute is set to 1 for all. The adjacency list storage structure is saved in dictionary form, with the key being the node identifier and the value being the list of successor node identifiers directly connected to this node and the corresponding edge weight values.

[0051] Calculate the degree centrality value for each node in the node set. The degree centrality value is defined as the total number of other nodes directly connected to this node in the state transition diagram, including the sum of the input connections and the output connections. The calculation process specifically includes: Obtain the out-degree value of the node by querying the adjacency list. The out-degree value is the number of successor nodes pointed to by this node; obtain the in-degree value of the node by traversing the edge set. The in-degree value is the number of edges with this node as the end point; the degree centrality value is equal to the calculation result of adding the in-degree value and the out-degree value. For example, if a node has 3 successor node records in the adjacency list and is pointed to by 2 edges in the edge set, then the degree centrality value of this node is 5.

[0052] Collect the degree centrality values of all nodes to form a degree centrality value set, and perform distribution feature analysis on this set. Divide the degree centrality value range into several continuous intervals, and the number of intervals is dynamically adjusted according to the sample size. For example, when the total number of nodes is less than 50, set 5 intervals; when the total number of nodes is between 50 and 200, set 10 intervals; when the total number of nodes is greater than 200, set 15 intervals. Statistically calculate the proportion of the number of nodes in each interval. The proportion is equal to the number of nodes in this interval divided by the total number of nodes. Calculate the degree centrality distribution entropy value. This entropy value reflects the discreteness of the node connection distribution. The calculation process is as follows: For each interval with a non-zero proportion, calculate the product of the proportion of this interval and the logarithm of this proportion to the base 2 to obtain the entropy contribution value of this interval; sum up the entropy contribution values of all intervals and then take the absolute value.

[0053] Perform the shortest path depth calculation on the set of symptom state change nodes and the set of product usage operation nodes. For each symptom state change node, calculate its shortest path depth to each product usage operation node. The shortest path depth is defined as the minimum number of edges connecting the two nodes. The calculation is implemented using the breadth-first search algorithm: taking the current symptom state change node as the search starting point, initialize the search queue and the level counter; traverse the adjacent nodes layer by layer, and record the current level as the shortest path depth value when the target product usage operation node is visited; set the maximum search depth threshold to 50 levels, and mark it as unreachable when the search depth exceeds the threshold and the target node has not been reached yet. Statistically analyze the shortest path depth values from the symptom state change nodes to all reachable product usage operation nodes, and calculate the arithmetic mean of these depth values as the average shortest path depth of this symptom state change node. For example, if the shortest path depths from a certain symptom node to three operation nodes are 2, 3, and 4 respectively, then its average shortest path depth is 3.

[0054] Calculate the global average shortest path depth value: add up the average shortest path depth values of all symptom state change nodes, and then divide by the total number of symptom state change nodes. For example, when there are two symptom state change nodes with average shortest path depths of 2.5 and 3.5 respectively, the global average shortest path depth value is 3.0.

[0055] Generate the causal weight value: multiply the degree centrality distribution entropy value by the global average shortest path depth value to obtain the topological structure entropy value as the final causal weight value. This causal weight value quantifies the causal association strength between product usage operations and symptom changes. The larger the value, the more significant the causal relationship. For example, when the calculated result of the degree centrality distribution entropy value is 1.2 and the global average shortest path depth value is 2.0, the causal weight value is 2.4.

[0056] The basis for setting key parameters includes: the number of intervals for dividing the degree centrality is optimized and determined according to the sample dispersion coefficient. By analyzing the coefficient of variation of entropy values under different numbers of intervals, select the number of intervals when the coefficient of variation is the smallest; the maximum search depth threshold is set based on statistical analysis of historical data. For example, after analyzing the path depth distribution in 1000 reports, take the 95th percentile value as the threshold; the rule for handling unreachable paths is: when there is no connected path between a symptom state change node and a certain product usage operation node, this pair of nodes does not participate in the calculation of the average shortest path depth.

[0057] The special scenario processing mechanism includes: for isolated nodes, that is, nodes with a degree centrality value of zero, they still participate in the degree centrality distribution calculation but contribute zero in the entropy value calculation; when there are bidirectional connection edges, the connection relationships in both directions are counted separately in the degree centrality calculation; self-loop edges of nodes are counted in the degree centrality value calculation but excluded in the shortest path depth calculation; when there are multiple paths between two nodes, the shortest path depth takes the minimum number of edges among all paths.

[0058] The verification was conducted by constructing a typical topological structure: Case 1 tested the chain structure, which contained three nodes connected in the order of operation node A pointing to symptom node B and then to operation node C. The node degree centrality values were calculated as A=1, B=2, and C=1. The median value of degree centrality distribution was 1, accounting for about 66.7% (2 nodes), and the value of 2 accounted for about 33.3% (1 node). The entropy value was calculated to be about 0.918. The shortest path depth from symptom node B to the operation node was 1, the global average shortest path depth was 1.0, and the final causal weight value was 0.918. This result conforms to the cumulative effect characteristics of causal conduction of the chain structure. Case 2 tested the star structure, with the central operation node A connecting four symptom nodes B, C, D, and E. The node degree centrality value A=4 and the rest of the nodes were all 1. The median value of degree centrality distribution was 1, accounting for 80%, and the value of 4 accounted for 20%. The entropy value was calculated to be about 0.722. The depth from all symptom nodes to the operation node was 1, the global average depth was 1.0, and the causal weight value was 0.722. This result reflects the strong influence of the central node. Case 3 tests the fully connected structure, with four nodes connected to each other. The degree centrality value of all nodes is 3, the distribution entropy value is 0, the minimum depth from any symptom node to the operation node is 1, and the causal weight value is 0. This result is in line with the technical expectation that over-connectivity leads to causal dilution.

[0059] The verification conclusion shows that: the algorithm time complexity meets the requirements, for example, the calculation time is less than 1 second at the scale of 100 nodes; the boundary processing is effective, and the calculation of the 500-node graph containing isolated nodes does not occur abnormally; the causal weight value sorting conforms to the technical logic, the chain structure has the highest weight, the star structure is second, and the weight of the fully connected structure is zero. The final output causal weight value data is associated with the original event record identifier and stored as a triple structure: operation record identifier, symptom record identifier, causal weight value. The causal weight value is normalized and the maximum and minimum value scaling algorithm is used: determine the minimum weight value and the maximum weight value in the entire data set; subtract the minimum value from each original weight value and divide it by the difference between the maximum and the minimum value. For example, when the weight range of the data set is between 0.2 and 0.9, the original value of 0.55 is normalized to 0.5. The normalized weight value is stored in the analysis result database for use in the subsequent causal determination process.

[0060] S6. Compare the causal weight value with a preset threshold value. When the causal weight value is greater than or equal to the preset threshold value, output a feature set containing product-symptom causal association features. The specific implementation is as follows: Obtain the causal weight value data calculated in the previous step as input data. This causal weight value data is stored in a structured triple format, including three attribute fields: the operation record identifier attribute serves as an index key pointing to the cosmetic usage operation event record; the symptom record identifier attribute serves as an index key pointing to the skin symptom change event record; the causal weight value attribute stores a numerical association strength value, which is a dimensionless scalar. Perform a numerical comparison operation on the numerical value of the causal weight value attribute in each triple with a preset causal determination threshold. The setting method of the causal determination threshold is as follows: Collect the real causal relationship cases confirmed by medical experts in the historical user report dataset as the positive sample set. For example, collect 1000 cases of confirmed cosmetic allergy cases; at the same time, collect non-causal relationship cases to form a negative sample set. For example, collect 2000 cases of irrelevant symptom reports; calculate the 70th percentile of all causal weight values in the positive sample set as the benchmark threshold. For example, when the positive sample weight value set is [0.15, 0.42, 0.68, 0.91], the value corresponding to the 70th percentile after sorting is 0.78; this benchmark threshold supports a dynamic adjustment mechanism, and the adjustment coefficient is set according to the product category characteristics and loaded through a configuration table. For example, the adjustment coefficient for the skin care product category is 1.15, and the adjustment coefficient for the makeup category is 0.85. The final threshold calculation formula is the benchmark threshold multiplied by the adjustment coefficient. When the benchmark threshold is 0.78 and the product is a skin care product, the final threshold is 0.78 multiplied by 1.15, approximately equal to 0.90.

[0061] When the numerical value of the causal weight value attribute is greater than or equal to the causal determination threshold corresponding to the current product category, perform the feature information extraction operation. Retrieve the corresponding product usage operation node in the node set data of the state transition diagram through the operation record identifier attribute. The node set data is sourced from the graph data structure storage area constructed in step S4. Extract the cosmetic name information from the attribute fields of this node. The cosmetic name information is directly taken from the associated entity name attribute field of the product usage operation node, and this field stores a standardized cosmetic naming text string. For example, when the value of the associated entity name attribute field corresponding to the operation record identifier attribute is "XX brand repair essence", extract this string as the cosmetic name information.

[0062] Locate the corresponding symptom state change node in the node set data through the symptom record identifier attribute, and extract the skin symptom description information from the attribute fields of this node. The extraction process adopts a dual-field priority strategy: First, read the associated entity name attribute field of the symptom state change node, and this field stores the symptom description text standardized by terms; when the associated entity name attribute field is a null value or an invalid string, fallback to read the description text attribute field, and this field retains the original symptom description content. For example, when the value of the associated entity name attribute field of the symptom state change node is "contact dermatitis", extract this value; if it is null, extract the value of the description text attribute field "skin turns red and develops rashes after application".

[0063] Generate product-symptom causal association feature instances. This feature is stored using a key-value pair data structure and contains three core data items: The cosmetic name information data item stores the extracted cosmetic name text string; the skin symptom description information data item stores the extracted symptom description text string; the causal weight value data item stores the weight value in the original triple. Each feature instance uniquely corresponds to a determination result of the association relationship between a product usage operation and a skin symptom change. For example, generate a feature instance: cosmetic name information data item = "YY Sunscreen Lotion", skin symptom description information data item = "Photosensitivity reaction", causal weight value data item = 0.87.

[0064] Add the product-symptom causal association feature instances that meet the threshold conditions to the feature set output set. The feature set is stored using a dynamic array structure, and each element is a complete product-symptom causal association feature instance. Before the feature set is output, perform deduplication and optimization processing: Establish a composite key composed of the cosmetic name information string and the skin symptom description information string; when the composite key of the new feature instance is repeated with an existing instance in the feature set, compare the causal weight value data item values of the two; only retain the feature instance with a larger causal weight value data item value. Finally, output the feature set and write it to the analysis result database, storing it as a dedicated data table. This table contains five fields: The feature identifier field generates a hash value for the combination of the cosmetic name and symptom description using the SHA-256 algorithm; the cosmetic name field stores text data; the symptom description field stores text data; the weight value field stores floating-point numerical values; the timestamp field records the UNIX timestamp value at the moment when the feature is generated.

[0065] The key processing logic includes: The threshold dynamic adjustment mechanism loads parameters through a configuration file, and the configuration item contains the mapping relationship between the product category code and the adjustment coefficient. In the symptom description fusion processing, when both the standardized term and the original description exist, the original description text is additionally stored in the feature for review reference. The deduplication processing algorithm is implemented using a hash map table, and the time complexity is O(1). The boundary condition processing rules include: When the causal weight value attribute value is equal to the threshold, it is determined to meet the condition; when the weight value attribute is a null value, skip the processing and record an exception log.

[0066] The example verification passed three groups of typical scenario tests: In Case 1, the high-weight scenario was tested. The causal weight value attribute had a numerical value of 0.92 (greater than the skin care product threshold of 0.90). The associated operation record identifier attribute was located at the product usage operation node, and the cosmetic name information "ZZ Moisturizing Cream" was extracted. The associated symptom record identifier attribute was located at the symptom status change node, and the standardized symptom description "Allergic edema" was extracted. A feature instance [Cosmetic name: "ZZ Moisturizing Cream", Symptom description: "Allergic edema", Weight value: 0.92] was generated and added to the feature set. In Case 2, the low-weight scenario was tested. The causal weight value attribute had a numerical value of 0.75 (less than the makeup threshold of 0.66), and feature generation was skipped. In Case 3, the boundary scenario was tested. The causal weight value attribute had a numerical value of 0.66 (equal to the makeup threshold), and the cosmetic name information "CC Liquid Foundation" and the symptom description "Acneiform eruption" were extracted, and the feature instance was added to the feature set.

[0067] The special scenario verification includes: In the duplicate feature processing test, when the feature set contains ["DD Eye Cream", "Periorbital erythema", 0.70], the newly generated feature ["DD Eye Cream", "Periorbital erythema", 0.83] performs a replacement operation. In the null value processing test, when the entity name attribute field associated with the symptom node is empty, the value of the description text attribute field "Eyelid swelling and peeling" is extracted as the symptom description. In the threshold dynamic adjustment test, the hair care product adjustment coefficient was set to 1.1, the baseline threshold was 0.78, and the final threshold was 0.86. Features with a weight value of 0.85 were correctly filtered.

[0068] The verification conclusion shows that: The threshold comparison logic had a 100% correct rate after 500 tests; in the feature extraction accuracy test, the correct rate of cosmetic name extraction was 100%, and the correct rate of symptom description extraction was 98.7% (the error was due to unstandardized descriptions); the duplicate removal mechanism took less than 10 milliseconds to process at a feature scale of 10,000; the database write rate reached 1,200 records per second; the passing rate of the boundary condition coverage test was 100%. The final feature set provides data services through an application programming interface. The interface adopts a RESTful architecture and supports JSON format data exchange. Each feature instance is associated with the original event timestamp, supporting the generation of a time series causal analysis graph. The feature set data is also used to generate a user-readable report, which includes the product name, symptom description, association strength value, and a visualization chart of the occurrence timeline.

[0069] In view of the complexity of attributing adverse reactions of cosmetics, this embodiment establishes a collaborative analysis mechanism for dual event streams. Through in-depth parsing of free text in S1, discrete cosmetic usage operations and skin symptom changes are transformed into independent event streams with time sequence markers, laying a foundation for constructing causal chains. The time sequence marker location technology in S2 resolves the problem of determining time sequence relationships in scenarios with ambiguous descriptions through the fusion of a dynamic rule base and context features. For example, the associated mapping of the word "after" in "improve after discontinuation" needs to be comprehensively judged in combination with product types and symptom characteristics. The implicit operation node derivation mechanism in S3 uses the time density characteristics (frequency of sudden / progressive words) of symptom descriptions to reverse-infer missing operations. For example, the "start using" node is deduced from "sudden redness and swelling", filling the analysis gap in scenarios where the operation chain is broken. The construction of the state transition diagram in S4 introduces cross-node connection rules to achieve dynamic topological association between operation nodes and symptom nodes. For example, when there is no adjacent operation node for a symptom node, connections are established by tracing back three layers forward. The causal weight calculation in S5 proposes a topological structure entropy algorithm, which combines the degree centrality distribution entropy and the average path depth to quantify the causal intensity. For example, in a star-shaped topology, high-centrality nodes and low path depths form strong causal indications. The threshold dynamic adjustment mechanism in S6 optimizes the determination accuracy through product category coefficients. For example, higher thresholds are used for skin care products to distinguish between chronic irritation and acute allergy. Through the collaborative effects of links such as event stream transformation, implicit node supplementation, and topological structure entropy calculation, the entire technical chain realizes the causal attribution of cosmetic usage and symptom changes.

[0070] All the calculations involved in the embodiment are numerical calculations after removing the dimension, and the preset parameters and threshold selection in the calculations are set by those skilled in the art according to the actual situation.

[0071] It should be noted that the present invention can be deployed on the device itself to achieve embedded applications, or can run on a PC or other terminals with a user interface, so as to meet various hardware environments and usage requirements.

[0072] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wireless or wired manner; where the wired transmission methods include optical fibers, twisted pairs, coaxial cables, etc.; wireless transmission includes infrared rays, microwaves, etc. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0073] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0074] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0075] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0076] In addition, in each embodiment of the present application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0077] If the above-mentioned functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0078] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0079] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for automatically accepting and feature recognizing user-reported information on adverse reactions of cosmetics, characterized in that Including: S1. Obtain a user report containing free text descriptions; S2. Identify temporal marker words in the free text descriptions through natural language processing techniques; S3. When there is an operation chain break in the temporal marker words, infer hidden operation nodes based on the time density characteristics of the symptom descriptions in the free text descriptions; S4. Construct a state transition graph of symptom state change nodes and product usage operation nodes according to the order and association relationship of the temporal marker words and the hidden operation nodes; S5. Generate a causal weight value between product usage operations and symptom changes based on the topological structure entropy of the state transition graph; where the topological structure entropy is the product of the degree centrality distribution entropy value of all nodes in the state transition graph and the average shortest path depth from the symptom state change nodes to the product usage operation nodes; S6. Compare the causal weight value with a preset threshold, and when the causal weight value is greater than or equal to the preset threshold, output a feature set containing product-symptom causal association characteristics.

2. The method for automatically accepting and feature-recognizing the user reports on adverse reactions of cosmetics according to claim 1, wherein Obtain a user report containing free text descriptions, including: Receive an adverse reaction report submitted by the user through an online system, where the adverse reaction report contains standardized form data filled in by the user and supplementary description content in the free text column; Extract the supplementary description content in the free text column as the free text description, and the free text description contains records of cosmetic usage operation events and records of skin symptom change events.

3. The method for automatically accepting and feature-recognizing user reports on adverse reactions of cosmetics according to claim 2, characterized in that The record of cosmetic usage operation events refers to the cosmetic contact behaviors actively performed by the user, and the record of skin symptom change events refers to the description of the biological phenomenon of skin state change.

4. The method for automatically accepting and feature recognizing the user reports on adverse reactions of cosmetics according to claim 2, wherein Identify temporal marker words in the free text descriptions through natural language processing techniques, including: Perform word segmentation on the free text description to generate a word sequence, and match the temporal marker words contained in the word sequence based on a predefined temporal marker word library; The temporal marker word library contains four types of association markers: marker words indicating product discontinuation operations, marker words indicating symptom remission states, marker words indicating product reuse operations, and marker words indicating symptom recurrence states; Establish a position association mapping between the matched temporal marker words and the corresponding cosmetic usage operation events or skin symptom change event records in the free text description.

5. The automatic acceptance and feature recognition method for cosmetic adverse reaction user report information according to claim 4, characterized in that, When there is an operation chain break in the temporal marker words, infer hidden operation nodes based on the time density characteristics of the symptom descriptions in the free text descriptions, including: Detect the symptom description text corresponding to the skin symptom change event record in the free text description; Extract the occurrence frequencies of symptom suddenness description words and symptom progression description words in the symptom description text as time density characteristics; When the occurrence frequency of symptom suddenness description words is greater than the occurrence frequency of symptom progression description words, generate the product start usage operation in the cosmetic usage operation event record as the hidden operation node; When the occurrence frequency of symptom progression description words is greater than or equal to the occurrence frequency of symptom suddenness description words, generate the product continuous usage operation in the cosmetic usage operation event record as the hidden operation node.

6. The method for automatically accepting and feature recognizing user reports on adverse reactions of cosmetics according to claim 5, characterized in that, Symptom suddenness description words include suddenly, immediately, abruptly, and symptom progression description words include gradually, slowly, step by step.

7. The method for automatically accepting and feature recognizing cosmetic adverse reaction user report information according to claim 5, characterized in that, Construct a state transition diagram of symptom state change nodes and product usage operation nodes according to the order and correlation relationship of timing marker words and implicit operation nodes, including: Extract the cosmetic usage operation event records corresponding to the timing marker words as explicit product usage operation nodes; Extract the skin symptom change event records corresponding to the timing marker words as symptom state change nodes; Merge the implicit operation nodes and the explicit product usage operation nodes to form a complete set of product usage operation nodes; Arrange the complete set of product usage operation nodes and symptom state change nodes in the order of the occurrence time of the event records in the free text description; Establish directed connection edges between adjacent nodes to form a state transition diagram from symptom state change nodes to product usage operation nodes, where the product usage operation nodes point to the symptom state change nodes directly caused by them.

8. The method for automatically accepting and feature recognizing the user reports on adverse reactions of cosmetics according to claim 7, wherein Generate the causal weight value between product usage operations and symptom changes based on the topological structure entropy of the state transition diagram, including: Calculate the degree centrality values of all nodes in the state transition diagram. The degree centrality value represents the number of direct connections between a node and other nodes; Form a degree centrality distribution according to the degree centrality values of all nodes, and calculate the entropy value of the degree centrality distribution as the degree centrality distribution entropy value; Calculate the shortest path depth from each symptom state change node to each product usage operation node in the state transition diagram. The shortest path depth refers to the minimum number of edges connecting two nodes; Take the arithmetic mean of the shortest path depths from all symptom state change nodes to product usage operation nodes as the average shortest path depth; Multiply the degree centrality distribution entropy value by the average shortest path depth to obtain the topological structure entropy value as the causal weight value between product usage operations and symptom changes.

9. The automatic acceptance and feature recognition method for user reports on adverse reactions of cosmetics according to claim 8, characterized in that, Compare the causal weight value with a preset threshold. When the causal weight value is greater than or equal to the preset threshold, output a feature set including product-symptom causal association features, including: Perform a numerical comparison between the causal weight value and a pre-set causal determination threshold; When the causal weight value is greater than or equal to the causal determination threshold, extract the cosmetic name information corresponding to the product usage operation node in the state transition diagram; Extract the skin symptom description information corresponding to the symptom state change node in the state transition diagram; Generate product-symptom causal association features including cosmetic name information, skin symptom description information, and causal weight values; Add the product-symptom causal association features to the feature set as the output content.

Citation Information

Patent Citations

  • Event sequential relation extraction method based on dynamic attention mechanism

    CN114153942A

  • Unstructured data automatic processing method and system based on deep learning

    CN119597834A

  • Method and system for tracking untoward effects of vaccines in real time

    CN120126814A

  • Text event association method based on large language model

    CN120196741A

  • Real-time event correlation in information networks

    US10693711B1