Language data processing method and system, terminal equipment and storage medium

By integrating cross-modal text features and historical dialogue data in the dialogue system, a cross-round context reference chain is constructed, and multiple rounds of dialogue directed graphs are used for intention evolution reasoning, the problem of traditional systems understanding user intentions in multiple rounds of dialogue is solved, and a more accurate and personalized dialogue understanding is achieved.

CN120144768AInactive Publication Date: 2025-06-13SHENZHEN WRITER INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510164900.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In multiple rounds of dialogue, traditional dialogue recommendation systems are difficult to accurately understand the user's true intentions, especially when the user's expression is vague or the semantics are unclear, referential errors and semantic understanding bias are prone to occur.

Method used

By extracting cross-modal text features in audio/video and weighted splicing with user dialogue text features, comprehensive dialogue language feature data is generated; combining historical round dialogue text data for pronoun detection and component analysis, cross-round context reference chains are constructed; using multiple rounds of dialogue directed graphs for intent evolution reasoning, and tracking user interaction behavior for negative feedback intention adjustments.

Benefits of technology

It improves the accuracy and personalization of dialogue understanding, can more accurately identify the referential objects in cross-round dialogue, capture the dynamic changes in user intentions, and provide in-depth understanding and personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144768A_ABST
    Figure CN120144768A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language data processing, in particular to a language data processing method and system, terminal equipment and a storage medium. The method comprises the following steps: acquiring current round dialogue data of a user; performing dialogue language text processing according to the current round dialogue data of the user to obtain current round dialogue language feature data; carrying out anonym detection on the current round dialogue language feature data, and carrying out dynamic historical anchor point screening to obtain cross-round screening anchor point list data; performing multi-round dialogue intention evolution reasoning according to cross-round screening anchor point list data to obtain an interaction adjustment intention representation strategy; and according to the interaction adjustment intention representation strategy, carrying out current round dialogue language key referring word determination so as to realize cross-round dialogue type language referring word identification. According to the method, the cross-round dialogue type language anonyms are accurately recognized through multi-round dialogue situation and multi-modal clue language data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language data processing, and in particular, to a method, a system, a terminal device, and a storage medium for language data processing. Background Art

[0002] A conversational recommendation system understands the needs of users through language interaction and provides personalized product or service recommendations. However, in the context of multi-turn conversations, especially in complex scenarios that include colloquial expressions, indefinite references, and multimodal information, the language data in multi-turn conversation scenarios is characterized by interactivity and dynamics. The user's expression is usually not complete in one go, but becomes gradually clear through multi-turn conversations. In this process, the user's language expression often has uncertainties, such as using vague pronouns (such as "it", "that") or semantically unclear phrases (such as "recommend a better one for me"). In multi-turn conversations, the user's expression often does not follow strict written language rules, has a loose grammatical structure, an irregular word order, and may contain a large number of ellipses, repetitions, modal particles, and ambiguous expressions. In many actual application scenarios, the user's intention is conveyed not only through language expression, but also through non-verbal information such as audio, video, or facial expressions. However, the language processing methods in traditional conversational recommendation systems usually rely on static semantic analysis models, and cannot make full use of context information to accurately understand the user's true intention in multi-turn conversations, making it difficult to accurately determine which object the "it" refers to among the previous ones, and prone to reference errors, resulting in semantic understanding deviations. Summary of the Invention

[0003] Based on this, the present invention provides a method, a system, a terminal device, and a storage medium for language data processing to solve at least one of the above technical problems.

[0004] To achieve the above object, a method for language data processing includes the following steps:

[0005] Step S1: Extract the current-round conversation of the user to obtain the user's current-round conversation data; perform dialogue language text processing on the user's current-round conversation data to generate user dialogue text feature data; perform audio / video language cross-modal text extraction on the user's current-round conversation data to obtain cross-modal conversation text feature data; perform weighted splicing on the cross-modal conversation text feature data and the user dialogue text feature data to obtain the current-round conversation language feature data;

[0006] Step S2: Detect the pronouns in the language feature data of the current round of conversation, and perform an analysis of the referent component to obtain the current candidate referential anchor data; obtain the text data of the historical round of conversation; perform dynamic historical anchor screening on the current candidate referential anchor data through the text data of the historical round of conversation to obtain the cross-round screening anchor list data; perform pronoun-referent linking processing according to the cross-round screening anchor list data to generate cross-round context referential chain data;

[0007] Step S3: Perform multi-round conversation directed graph fusion according to the cross-round context referential chain data to construct a user multi-round conversation directed graph; use the user multi-round conversation directed graph to perform multi-round conversation intention evolution reasoning to obtain preliminary conversation language intention vector data; track the subsequent interaction behavior of the user, and perform negative feedback intention adjustment on the preliminary conversation language intention vector data to obtain an interactive adjustment intention representation strategy;

[0008] Step S4: Determine the key pronouns in the language of the current round of conversation according to the interactive adjustment intention representation strategy to achieve cross-round conversational language pronoun recognition.

[0009] By extracting cross-modal text features from audio / video and performing weighted splicing with user dialogue text features, the present invention can capture the user's intentions and emotions more comprehensively, thereby improving the accuracy of dialogue understanding. Especially in cases where the user's expression is ambiguous or the semantics are unclear, the supplementation of cross-modal information can provide additional clues to help the system more accurately understand the user's true needs. For example, when the user says "Recommend a good one for me", it is difficult to determine the specific meaning of "good" relying solely on text information. However, if combined with the user's voice intonation or expression, the system can determine which type of "good" the user is more inclined to, such as "high cost performance" or "powerful function". By detecting pronouns and analyzing constituent elements in the user's current-round dialogue, and dynamically screening historical anchor points in combination with historical-round dialogue text data to construct a cross-round context reference chain, the pronoun can be accurately linked to its referent, thereby eliminating reference ambiguity and improving the accuracy of dialogue understanding. For example, when the user mentions "I want to buy a mobile phone" in the first round and says "It should have good photography" in the second round, the system can accurately refer "It" to "mobile phone" rather than other objects. Constructing a multi-round dialogue into a directed graph, where the nodes represent the key entities and concepts in the dialogue and the edges represent the relationships between them. By analyzing this graph, the system can capture the dynamic changes and evolution trajectories of the user's intentions, thereby more accurately predicting the user's future needs. For example, the user initially only expressed vague needs for a mobile phone, but in subsequent conversations, will gradually clarify specific requirements for aspects such as screen size and camera performance. By analyzing the directed graph of the user's multi-round dialogue, the system can capture these subtle changes and timely adjust the recommendation strategy. By tracking the user's subsequent interaction behaviors and performing negative feedback intention adjustment on the initial dialogue language intention vector, the system can effectively correct misunderstandings of the user's intentions and make it continuously learn and adapt to the user's personalized expression methods. For example, when the system recommends a product and the user indicates dislike, the system will adjust the intention expression strategy according to the user's feedback and avoid similar products in subsequent recommendations. Therefore, a language data processing method of the present invention extracts and processes the user's current-round dialogue, fuses dialogue text features and cross-modal text features to generate comprehensive dialogue language feature data; then, detects pronouns and analyzes the referent constituent elements, dynamically screens anchor points in combination with historical dialogue text to construct a cross-round context reference chain; then, based on the reference chain data, fuses the multi-round dialogue directed graph, infers intention evolution and tracks the user's interaction behaviors for negative feedback adjustment; finally, determines the key pronouns according to the adjusted intention strategy, accurately identifies the referents in cross-round dialogues, and provides the dialogue system with the ability of in-depth understanding and personalized services.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Obtain user ID data;

[0012] Step S12: Extract the current round of conversation for the user based on the user ID data, and perform session time series processing to obtain the user's current round of conversation data;

[0013] Step S13: Perform dialogue language text processing according to the user's current round of conversation data to generate user dialogue text feature data;

[0014] Step S14: Perform audio / video language cross-modal text extraction on the user's current round of conversation data through a preset audio / video language extraction network interface to obtain cross-modal dialogue text feature data;

[0015] Step S15: Use the user dialogue text feature data to evaluate the correlation of the cross-modal dialogue text feature data, and generate cross-modal text correlation score data;

[0016] Step S16: Based on the cross-modal text correlation score data, perform weighted splicing on the cross-modal dialogue text feature data and the user dialogue text feature data to obtain the current round of dialogue language feature data.

[0017] The present invention can distinguish different users through user IDs, avoid confusion between different user data, and thus more accurately capture the personalized needs of each user. At the same time, session time series processing can retain the context information of the conversation, avoid information loss or disorder, and thus more completely restore the user's expression intention. Cross-modal text extraction is performed through a preset audio / video language extraction network interface, and the user dialogue text feature data is used to evaluate the correlation of the cross-modal dialogue text feature data, effectively improving the utilization efficiency and accuracy of cross-modal information. Weighted splicing based on the cross-modal text correlation score data can more effectively fuse text features and cross-modal features, and construct a more comprehensive user expression representation. Compared with simple splicing, weighted splicing can dynamically adjust its weight according to the correlation of different modal information, thus more effectively using cross-modal information and avoiding interference or dilution of cross-modal information on text information. For example, if the user's speech tone and facial expression are consistent with the meaning expressed in the text, the system will increase the weight of cross-modal information; on the contrary, if the speech tone and facial expression are contradictory to the meaning expressed in the text, the system will reduce the weight of cross-modal information or even ignore it, so as to ensure the accuracy and reliability of the final dialogue language feature data.

[0018] Preferably, step S13 includes the following steps:

[0019] Step S131: Perform colloquial text detection on the user's current round of conversation data to generate user colloquial dialogue text data;

[0020] Step S132: Perform dialogue text word segmentation on the user's colloquial dialogue text data, and remove stop words to obtain optimized colloquial dialogue text data;

[0021] Step S133: Perform context-aware part-of-speech tagging on the optimized colloquial dialogue text data to obtain context part-of-speech tagging data;

[0022] Step S134: Perform multi-level nested named entity recognition on the optimized colloquial dialogue text data through the context part-of-speech tagging data to obtain dialogue nested entity data;

[0023] Step S135: Perform semantic role recognition based on the dialogue nested entity data, and perform dialogue text feature processing on the optimized colloquial dialogue text data to generate user dialogue text feature data.

[0024] The present invention performs colloquial text detection on the user's current-round dialogue data, and can identify and process common phenomena in spoken language such as ellipsis, repetition, and modal particles. This provides a clearer input for subsequent text processing steps and avoids interference from these colloquial expressions in word segmentation, part-of-speech tagging and other links. For example, when the user says "Well, I think, um, buy a, uh, mobile phone", the colloquial text detection can filter out these modal particles and repeated words to obtain a more concise text "I want to buy a mobile phone". The system performs dialogue text word segmentation and stop word removal, further streamlining the text content and retaining key information. Stop words are usually some high-frequency but less semantically contributing words, such as "de", "shi", "le", etc. Removing these stop words can reduce the redundancy of text data, improve the efficiency of subsequent processing steps, and highlight key entities and concepts. In spoken language, the same word may have different meanings in different contexts. Context-aware part-of-speech tagging can accurately judge the specific meaning of "da" in the current context based on context information, thus avoiding ambiguity and improving the accuracy of semantic understanding. Through semantic role recognition, the semantic relationship between different components in a sentence can be further analyzed, such as who is the executor of the action and who is the recipient of the action, etc. This provides richer semantic information for subsequent intention understanding. For example, in the sentence "I want to buy a mobile phone", "I" is the executor of "buy", and "mobile phone" is the object of "buy". Through semantic role recognition, the system can more accurately understand that the user's intention is to buy a mobile phone.

[0025] Preferably, step S2 includes the following steps:

[0026] Step S21: Perform pronoun detection on the current-round dialogue language feature data to generate current-round dialogue pronoun data;

[0027] Step S22: Perform referent component analysis based on the current-round dialogue pronoun data to obtain current candidate referent anchor data;

[0028] Step S23: Dynamically retrieve the historical round conversation records based on the user ID data to obtain the historical round conversation text data;

[0029] Step S24: Dynamically screen the current candidate reference anchor data through the historical round conversation text data to obtain the cross-round screening anchor list data;

[0030] Step S25: Use the cross-round screening anchor list data to perform multi-level anchor similarity matching on the current candidate reference anchor data to generate cross-round anchor matching data;

[0031] Step S26: Perform reference target resolution based on the cross-round anchor matching data and perform reference relationship reasoning to obtain the current dialogue language reference relationship graph;

[0032] Step S27: Perform pronoun-referent linking processing based on the current dialogue language reference relationship graph and perform anchor context association to generate cross-round context reference chain data.

[0033] The present invention dynamically retrieves the historical round conversation records according to the user ID, and dynamically screens the current candidate reference anchors by using the historical information, effectively utilizing the context information and improving the accuracy of anchor selection. By introducing the historical dialogue information, it is possible to more accurately judge the true referent of the pronoun. Especially when the information in the current round is insufficient, the historical information can provide important supplementary clues. The use of multi-level anchor similarity matching, rather than simple keyword matching, can more accurately judge the referent. The multi-level matching takes into account various factors such as the semantics, context, and entity type of the words, and can better handle complex reference relationships. For example, "it" can refer to an entity mentioned earlier, or an event or concept mentioned earlier. The multi-level matching can more accurately judge the referential meaning of "it" in a specific context. The solution performs reference target resolution and reference relationship reasoning, constructs the current dialogue language reference relationship graph, and finally generates cross-round context reference chain data. This enables the system to not only identify the referent of the pronoun, but also understand the relationship between them, thus more comprehensively understanding the user's intention. For example, when the user says "I like its color and its feel", the system can not only identify that the two "its" refer to the same object, but also understand that what the user likes are the two attributes of "color" and "feel" of the object.

[0034] Preferably, step S24 includes the following steps:

[0035] Step S241: Mine the historical language dialogue objects according to the historical round conversation text data to obtain the historical language dialogue object data;

[0036] Step S242: Perform historical anchor time series encoding on the historical language dialogue object data to generate historical time series anchor data;

[0037] Step S243: Perform anchor time decay based on the historical time series anchor data to obtain time-decayed anchor data; perform anchor type classification based on the historical time series anchor data to obtain historical classified anchor data;

[0038] Step S244: Extract referential components from the current candidate referential anchor data to obtain current referential component feature data;

[0039] Step S245: Perform cross-round anchor weighted correlation evaluation on the current referential component feature data through the time-decayed anchor data and the historical classified anchor data, and calculate the anchor activation degree to obtain historical anchor activation degree data;

[0040] Step S246: Dynamically screen the historical time series anchor data using the historical anchor activation degree data based on a preset activation degree threshold to obtain a cross-round screened anchor list data.

[0041] The present invention mines historical language dialogue objects from the historical round dialogue text data and performs historical anchor time series encoding, which not only retains the key information of the historical dialogue but also records the time sequence of this information. This provides a necessary basis for subsequent anchor time decay and activation degree calculation. Performing time decay and type classification on the historical time series anchor data further improves the accuracy of anchor selection. The time decay mechanism can reduce the weight of the earlier-appearing anchors because the user's current intention is related to the objects mentioned recently. By extracting referential components from the current candidate referential anchor data and combining the time-decayed anchor data and the historical classified anchor data for cross-round anchor weighted correlation evaluation and anchor activation degree calculation, more accurate anchor screening is achieved. The weighted correlation evaluation not only considers the semantic similarity between the current pronoun and the historical anchor but also factors such as time decay and type matching, thus more comprehensively evaluating the correlation between the historical anchor and the current pronoun. Dynamically screening the historical time series anchor data based on a preset activation degree threshold obtains the final cross-round screened anchor list data, reducing the computational amount of subsequent pronoun resolution, improving efficiency, and ensuring the accuracy of the finally selected anchors.

[0042] Preferably, step S3 includes the following steps:

[0043] Step S31: Perform multi-granularity semantic relationship analysis based on the cross-round context referential chain data to generate dialogue multi-granularity semantic relationship edge data;

[0044] Step S32: Construct multi-modal nodes based on historical temporal anchor data and current candidate reference anchor data to generate multi-modal node data for dialogue language;

[0045] Step S33: Construct dynamic reference relationship edges for cross-turn context reference chain data, and perform multi-turn dialogue directed graph fusion based on multi-granularity semantic relationship edge data and multi-modal node data for dialogue language to construct a user multi-turn dialogue directed graph;

[0046] Step S34: Perform graph embedding representation encoding on the user multi-turn dialogue directed graph to generate an encoded multi-turn dialogue directed graph;

[0047] Step S35: Use a preset graph neural network model to identify the current language intention nodes in the encoded multi-turn dialogue directed graph and perform multi-turn dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data;

[0048] Step S36: Track the subsequent interaction behaviors of the user, and perform negative feedback intention adjustment on the preliminary dialogue language intention vector data to obtain an interactive adjustment intention representation strategy.

[0049] The structured representation of the user's multi-turn dialogue in the present invention has nodes representing key entities and concepts, and edges representing the relationships between them. The multi-granularity semantic relationship analysis can capture semantic associations at different levels, such as attribute relationships and action relationships between entities, so as to more comprehensively depict the user's intention. The construction of multi-modal nodes incorporates text and cross-modal information, enabling the graph to more comprehensively reflect the user's expression. Using a preset graph neural network model to identify the current language intention nodes in the encoded graph and perform multi-turn dialogue intention evolution reasoning can more accurately capture the dynamic changes of the user's intention. The graph neural network model can effectively learn the features of the nodes and edges in the graph and perform global reasoning, so as to more accurately identify the user's current intention and predict their future needs. By tracking the user's subsequent interaction behaviors and performing negative feedback intention adjustment on the preliminary dialogue language intention vector, interactive intention understanding and optimization are achieved. The user's feedback information, such as clicks, purchases, evaluations, etc., can directly reflect the user's true intention. The system adjusts the previous intention understanding according to the user's feedback information, thereby continuously learning and optimizing its own understanding ability to provide a recommendation service that better meets the user's needs.

[0050] Preferably, step S36 includes the following steps:

[0051] Step S361: Track the interaction behaviors of the user to obtain user interaction behavior data;

[0052] Step S362: Analyze the negative feedback based on the user interaction behavior data to obtain negative feedback information data;

[0053] Step S363: Locate the negative object based on the negative feedback information data, and perform negative information analysis to respectively obtain negative keywords and negative object data;

[0054] Step S364: Based on the negative keywords and negative object data, perform multi-level propagation of negative feedback through the user multi-round dialogue directed graph to generate feedback propagation influence node data;

[0055] Step S365: Perform self-adaptive correction of the negative feedback weight on the preliminary dialogue language intention vector data through the feedback propagation influence node data to obtain an interactive adjustment intention representation strategy.

[0056] The present invention tracks the user interaction behavior and analyzes the negative feedback to obtain the specific reasons and objects that the user is dissatisfied with. This provides richer information than simple "like" or "dislike" feedback, which helps the system to more accurately understand the user's needs and preferences. Through negative object location and negative information analysis, negative keywords and negative object data are extracted. This provides the target and direction for subsequent negative feedback propagation. For example, when the user says "I don't like the color of this mobile phone", the system can extract the negative keyword "don't like" and the negative object "color". Based on the negative keywords and negative object data, multi-level propagation of negative feedback is performed on the user multi-round dialogue directed graph, and the propagated influence nodes are used for the correction of the preliminary dialogue language intention vector. This propagation mechanism can transmit the user's negative feedback to relevant nodes, thereby more comprehensively adjusting the intention representation. Self-adaptive correction of the negative feedback weight on the preliminary dialogue language intention vector is performed through the feedback propagation influence node data, achieving a more refined intention adjustment. The self-adaptive correction mechanism can dynamically adjust the weight according to the influence degree of different nodes, thereby more effectively correcting the system's misunderstanding of the user's intention.

[0057] Preferably, the present invention further provides a language data processing system that executes the above-mentioned language data processing method. The language data processing system includes:

[0058] A language feature extraction module, configured to extract the user's current round of dialogue to obtain the user's current round of dialogue data; perform dialogue language text processing according to the user's current round of dialogue data to generate user dialogue text feature data; perform audio / video language cross-modal text extraction on the user's current round of dialogue data to obtain cross-modal dialogue text feature data; perform weighted splicing on the cross-modal dialogue text feature data and the user dialogue text feature data to obtain the current round of dialogue language feature data;

[0059] The anaphora chain dynamic parsing module is used to detect anaphora in the language feature data of the current round of conversation, and perform anaphora object component analysis to obtain the current candidate anaphora anchor data; obtain the text data of the historical round of conversation; perform dynamic historical anchor screening on the current candidate anaphora anchor data through the text data of the historical round of conversation to obtain the cross-round screening anchor list data; perform anaphora-anaphora object link processing according to the cross-round screening anchor list data to generate cross-round context anaphora chain data;

[0060] The multi-round conversation intention reasoning module is used to perform multi-round conversation directed graph fusion according to the cross-round context anaphora chain data to construct a user multi-round conversation directed graph; use the user multi-round conversation directed graph to perform multi-round conversation intention evolution reasoning to obtain preliminary conversation language intention vector data; track the subsequent interaction behavior of the user, and perform negative feedback intention adjustment on the preliminary conversation language intention vector data to obtain an interaction adjustment intention representation strategy;

[0061] The cross-round key anaphora recognition module is used to determine the key anaphora in the language of the current round of conversation according to the interaction adjustment intention representation strategy to realize the recognition of anaphora in cross-round conversational language.

[0062] Preferably, the present invention further provides a terminal device, and the terminal device includes:

[0063] A processor;

[0064] A memory for storing instructions executable by the processor;

[0065] Wherein, the processor is configured to implement the language data processing method described in any one of the above.

[0066] Preferably, the present invention further provides a computer-readable storage medium, storing a computer program, and when the computer program is executed, it implements the language data processing method described in any one of the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a schematic flow chart of the steps of the language data processing method of the present invention;

[0068] Figure 2 is Figure 1 a detailed implementation step flow chart of step S2 in

[0069] Figure 3 is Figure 1 a detailed implementation step flow chart of step S3 in

[0070] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0072] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0073] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0074] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a language data processing method, including the following steps:

[0075] Step S1: Extract the current round of conversation of the user to obtain the user's current round of conversation data; perform conversation language text processing on the user's current round of conversation data to generate user conversation text feature data; perform audio / video language cross-modal text extraction on the user's current round of conversation data to obtain cross-modal conversation text feature data; perform weighted splicing on the cross-modal conversation text feature data and the user conversation text feature data to obtain the current round of conversation language feature data;

[0076] Step S2: Detect the pronouns in the current round of conversation language feature data and perform an analysis of the referent components to obtain the current candidate reference anchor data; obtain the historical round of conversation text data; perform dynamic historical anchor screening on the current candidate reference anchor data through the historical round of conversation text data to obtain the cross-round screening anchor list data; perform pronoun-referent link processing according to the cross-round screening anchor list data to generate cross-round context reference chain data;

[0077] Step S3: Perform multi-turn dialogue directed graph fusion based on the cross-turn context reference chain data to construct a user multi-turn dialogue directed graph; use the user multi-turn dialogue directed graph to perform multi-turn dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data; track the subsequent interaction behavior of the user, and perform negative feedback intention adjustment on the preliminary dialogue language intention vector data to obtain an interaction adjustment intention representation strategy;

[0078] Step S4: Determine the key pronouns of the current turn dialogue language according to the interaction adjustment intention representation strategy to realize cross-turn dialogue language pronoun recognition.

[0079] In the embodiment of the present invention, the language data processing method includes the following steps:

[0080] Step S1: Extract the current turn dialogue of the user to obtain the user's current turn dialogue data; perform dialogue language text processing according to the user's current turn dialogue data to generate user dialogue text feature data; perform audio / video language cross-modal text extraction on the user's current turn dialogue data to obtain cross-modal dialogue text feature data; perform weighted splicing on the cross-modal dialogue text feature data and the user dialogue text feature data to obtain the current turn dialogue language feature data;

[0081] In the embodiment of the present invention, assume that the user inputs voice in the search box of an e-commerce platform: "I want red. The color of the one I saw before was too dark." The system first converts the voice into text: "I want red. The color of the one I saw before was too dark." Subsequently, the system will process this text. For example, it will use a word segmentation tool to break it down into: "I / want / red / , / before / see / the / that / color / too / dark". Then, the system will apply word embedding techniques (such as Word2Vec or BERT) to convert these words into vector representations to form user dialogue text feature data. At the same time, if the user uploads a picture or video, the system will use a pre-trained cross-modal model to extract the visual information therein, such as extracting the main color, shape and other features of the item in the picture, and also convert these features into vector representations to form cross-modal dialogue text feature data. Finally, the system will perform weighted splicing on the text feature vector and the cross-modal feature vector according to the preset weights to obtain the current turn dialogue language feature data. For example, set the text feature weight to 0.7 and the cross-modal feature weight to 0.3, and then splice the two vectors together.

[0082] Step S2: Detect pronouns in the language feature data of the current round of conversation, and perform an analysis of the components of the referent objects to obtain the current candidate referential anchor data; obtain the text data of the historical round of conversation; perform dynamic historical anchor point screening on the current candidate referential anchor data through the text data of the historical round of conversation to obtain the cross-round screening anchor point list data; perform pronoun-referent object linking processing according to the cross-round screening anchor point list data to generate cross-round context referential chain data;

[0083] In the embodiment of the present invention, the system first detects pronouns in the language feature data of the current round of conversation obtained in step S1, and identifies that "that one" is a pronoun. Then, the system analyzes the context information near the pronoun, such as "the color is too dark", and infers that "that one" refers to a certain commodity, and takes "commodity" as the candidate referential anchor data. Next, the system retrieves the historical conversation record according to the user ID. Suppose the historical record contains the user saying: "I took a fancy to a product of type A, and its function is very powerful." and "The product of type B is also good, and its appearance is very fashionable." The system extracts the entities in the historical conversation text data, such as "the product of type A" and "the product of type B", as potential referent objects. Then, the system compares the current candidate referential anchor "commodity" with the entities in the historical conversation, such as calculating the semantic similarity between them. If the similarity between "the product of type A" and "commodity" is relatively high, the system links "that one" to "the product of type A" to form cross-round context referential chain data: {"that one": "the product of type A"}.

[0084] Step S3: Perform multi-round conversation directed graph fusion according to the cross-round context referential chain data to construct a user multi-round conversation directed graph; use the user multi-round conversation directed graph to perform multi-round conversation intention evolution reasoning to obtain preliminary conversation language intention vector data; track the subsequent interaction behavior of the user, and perform negative feedback intention adjustment on the preliminary conversation language intention vector data to obtain an interactive adjustment intention representation strategy;

[0085] In the embodiment of the present invention, the system first constructs a directed graph of the user's multi-turn conversation based on the cross-turn context anaphora chain data generated in step S2. The nodes in the graph represent the entities or concepts mentioned in the conversation, such as "products of type A", "red", and "very powerful in function". The nodes are connected by edges, and the edges represent the relationships between them. For example, there is a "referential" relationship between "that one" and "products of type A", and an "attribute" relationship between "products of type A" and "very powerful in function". Then, the system uses a graph neural network model to analyze this graph to infer the user's intention. For example, based on the user's current conversation "I want red ones" and the historical conversation "I took a fancy to a product of type A", the system can determine that the user wants to buy red products of type A. The system represents this intention as a vector, such as [1 (red), 0 (blue), 1 (type A), 0 (type B)], forming the preliminary conversation language intention vector data. Next, the system tracks the user's subsequent interaction behavior, such as the user clicks on the link of a certain red product of type A. The system adjusts the preliminary intention vector according to the user's behavior, such as increasing the weights of "red" and "type A", forming an interactive adjustment intention representation strategy.

[0086] Step S4: Determine the key anaphora in the current turn of the conversation language according to the interactive adjustment intention representation strategy to achieve cross-turn conversational language anaphora recognition.

[0087] In the embodiment of the present invention, according to the interactive adjustment intention representation strategy obtained in step S3, the system can determine the key anaphora in the current turn of the conversation. For example, if the system finds that the weights of "red" and "type A" are relatively high, it can be judged that "that one" refers to the product of type A and the color attribute is very important. Therefore, the system can identify "that one" as the key anaphora because it directly affects the user's choice of the product. In this way, the system realizes cross-turn conversational language anaphora recognition and can more accurately understand the user's needs, such as recommending red products of type A to the user.

[0088] Preferably, step S1 includes the following steps:

[0089] Step S11: Obtain user ID data;

[0090] Step S12: Extract the user's current turn of the conversation based on the user ID data and perform session time series processing to obtain the user's current turn of the conversation data;

[0091] Step S13: Perform conversation language text processing according to the user's current turn of the conversation data to generate user conversation text feature data;

[0092] Step S14: Extract cross-modal text of the audio / video language from the user's current round of dialogue data through a preset audio / video language extraction network interface to obtain cross-modal dialogue text feature data;

[0093] Step S15: Use the user's dialogue text feature data to evaluate the correlation of the cross-modal dialogue text feature data and generate cross-modal text correlation score data;

[0094] Step S16: Based on the cross-modal text correlation score data, perform weighted splicing on the cross-modal dialogue text feature data and the user's dialogue text feature data to obtain the current round of dialogue language feature data.

[0095] In the embodiments of the present invention, the system first needs to obtain the user's ID data. The user ID can be a unique identifier assigned by the system, such as a string of numbers or letters "user12345". This ID is used to distinguish different users and associate the user's historical conversation data and other relevant information. In the scenario of an e-commerce platform, after the user logs in to the account, the system can obtain their user ID. After obtaining the user ID, the system stores it in the context information of the current session. After obtaining the user ID, the system extracts the user's current turn of conversation data based on the user ID. Suppose the user "user12345" enters in the chat box of the e-commerce platform: "I want to buy one with a faster running speed. The one I looked at before had a bit of a slow running speed." The system will record this sentence and mark it as the current turn of conversation data of the user "user12345". At the same time, the system will record the order and timestamp of this conversation. The system processes the extracted text "I want to buy one with a faster running speed. The one I looked at before had a bit of a slow running speed." For example, using a word segmentation tool to split the sentence into a sequence of words: "I / want / to / buy / one / with / a / faster / running / speed. / The / one / I / looked / at / before / had / a / bit / of / a / slow / running / speed. / "; using a part-of-speech tagging tool to tag the part of speech of each word: "I(pronoun) / want(verb) / to(boundary) / buy(verb) / one(pronoun) / with(preposition) / a(article) / faster(adjective) / running(verb) / speed(noun) / .(punctuation) / The(article) / one(pronoun) / I(pronoun) / looked(verb) / at(preposition) / before(adverb) / had(verb) / a(article) / bit(adverb) / of(preposition) / a(article) / slow(adjective) / running(verb) / speed(noun) / ."; using a named entity recognition tool to identify the entities in the text, such as "running speed". Finally, the system converts this processed information into a vector representation. For example, using models such as Word2Vec or BERT to convert the text into a numerical vector as the user's conversation text feature data. Suppose the user uploads a video in addition to entering text, describing the slow running speed of the previously purchased item. The system will analyze the video content through a preset video language extraction network interface (for example, calling the video content analysis API provided by a cloud service provider) to extract the text information therein. For example, extracting the text converted from the voice in the video: "This program opens so slowly!", and the text description that appears in the video: "Processor: 1.6GHz". The system also converts this extracted cross-modal text information into a vector representation as the cross-modal conversation text feature data. The system calculates the correlation between the user's conversation text feature data and the cross-modal conversation text feature data. For example, calculating the cosine similarity between the text vector "I want to buy one with a faster running speed" and the video-extracted text vector "This program opens so slowly!" to obtain a correlation score, such as 0.8. This score represents the degree of correlation between the text and the video content. The system conducts a correlation assessment on all the extracted cross-modal text information to generate cross-modal text correlation score data.The system will perform weighted splicing on text features and cross-modal features according to the relevance score. For example, if the relevance score between the text and the video is high (e.g., 0.8), a higher weight, such as 0.4, will be assigned to the video features; if the relevance score is low (e.g., 0.2), a lower weight, such as 0.1, will be assigned to the video features. The weight of the text features will be adjusted accordingly. Assuming the weight of the text features is 0.6 and the weight of the video features is 0.4, the system will splice the two vectors together to form the final dialogue language feature data for the current round.

[0096] Preferably, step S13 includes the following steps:

[0097] Step S131: Perform colloquial text detection on the user's current-round dialogue data to generate user colloquial dialogue text data;

[0098] Step S132: Perform dialogue text word segmentation on the user colloquial dialogue text data and remove stop words to obtain optimized colloquial dialogue text data;

[0099] Step S133: Perform context-aware part-of-speech tagging on the optimized colloquial dialogue text data to obtain context part-of-speech tagging data;

[0100] Step S134: Perform multi-level nested named entity recognition on the optimized colloquial dialogue text data through the context part-of-speech tagging data to obtain dialogue nested entity data;

[0101] Step S135: Perform semantic role recognition based on the dialogue nested entity data and perform dialogue text feature processing on the optimized colloquial dialogue text data to generate user dialogue text feature data.

[0102] In the embodiments of the present invention, assume that the user's current round of dialogue data is: "I want to buy one with a faster running speed. The one I looked at before has a bit slow running speed." The system will use a colloquial text detection model to identify the colloquial expressions in the text, such as "faster" and "a bit slow". The detection model can be rule-based, for example, identifying some common colloquial words and sentence patterns; or it can be statistic-based, for example, training a classifier to determine whether the text belongs to the colloquial style. The model will mark the identified colloquial expressions. For example, "faster" is marked as having a high degree of colloquialism, and "running speed" is marked as having a low degree of colloquialism. The final output of the user's colloquial dialogue text data includes the original text and the colloquial marking information. The system will perform word segmentation on the user's colloquial dialogue text data. For example, using the jieba word segmentation tool to split the sentence into a sequence of words: "I / want / to / buy / one / with / a / running / speed / faster / ., / before / looked / at / that / one / with / a / running / speed / a / bit / slow / ." Then, the system will remove stop words, such as words that contribute less to semantic analysis like "I", "want", "one", "of", ",", "a", "bit", etc. The stop word list can be preset or adjusted according to the specific application scenario. The optimized colloquial dialogue text data obtained after removing stop words is: "running / speed / faster / before / looked / at / that / one / running / speed / slow". The system will perform part-of-speech tagging on the optimized colloquial dialogue text data. To improve the accuracy of tagging, the system will use a context-aware part-of-speech tagging model. For example, when processing "running / speed / faster", the model will consider the context information, such as "buy one with a faster running speed", so as to tag "running" as a verb, "speed" as a noun, and "faster" as an adjective. If "running" is processed alone, it may be tagged as a noun or a verb, but in combination with the context, its part of speech can be more accurately determined. The final obtained context part-of-speech tagging data is: "running(v) / speed(n) / faster(a) / before(f) / looked(v) / that(r) / one(m) / running(v) / speed(n) / slow(a)". Here, v represents a verb, n represents a noun, a represents an adjective, f represents a locative word, r represents a pronoun, and m represents a classifier. Use the context part-of-speech tagging data for multi-level nested named entity recognition. For example, in the phrase "running speed faster", "running speed" can be recognized as an entity, and "faster" is an attribute modifying "running speed". The system will identify this nested relationship and construct it into a nested structure, such as {"entity": "running speed", "attribute": "faster"}. The final obtained dialogue nested entity data includes all the identified entities and their attributes, relationships, etc. Perform semantic role recognition on the optimized colloquial dialogue text data. For example, in "the one I looked at before has a slow running speed", "that one" is the patient of the action "looked at", and "has a slow running speed" is an attribute of "that one". The system will extract this semantic role information.Then, the system converts all the processed information, including the word segmentation results, part-of-speech tagging results, named entity recognition results, semantic role recognition results, and colloquialism marking information, etc., into vector representations. For example, it uses the BERT model to convert the text into numerical vectors as the final user dialogue text feature data.

[0103] Preferably, step S2 includes the following steps:

[0104] Step S21: Detect pronouns in the language feature data of the current round of dialogue to generate the pronoun data of the current round of dialogue;

[0105] Step S22: Analyze the referent components according to the pronoun data of the current round of dialogue to obtain the current candidate referent anchor data;

[0106] Step S23: Dynamically retrieve the historical round of dialogue record according to the user ID data to obtain the historical round of dialogue text data;

[0107] Step S24: Dynamically screen the current candidate referent anchor data through the historical round of dialogue text data to obtain the cross-round screening anchor list data;

[0108] Step S25: Use the cross-round screening anchor list data to perform multi-level anchor similarity matching on the current candidate referent anchor data to generate cross-round anchor matching data;

[0109] Step S26: Resolve the referent target according to the cross-round anchor matching data and perform referent relationship reasoning to obtain the current dialogue language referent relationship graph;

[0110] Step S27: Perform pronoun-referent object linking processing according to the current dialogue language referent relationship graph and perform anchor context association to generate cross-round context referent chain data.

[0111] As an example of the present invention, referring to Figure 2 shown, it is Figure 1 the detailed implementation step flow diagram of step S2 in

[0112] Step S21: Detect pronouns in the language feature data of the current round of dialogue to generate the pronoun data of the current round of dialogue;

[0113] In an embodiment of the present invention, it is assumed that the current round of dialogue language feature data generated in step S1 includes the text information: "I want to buy one with a faster running speed. The one I saw before has a bit slow running speed." The system will use a pronoun detection model to analyze this text. The pronoun detection model can be rule-based, for example, predefined a list of pronouns in advance and then match these pronouns in the text; it can also be statistic-based, for example, train a classifier to determine whether a word is a pronoun. In this example, the model will identify that "that one" is a pronoun. The output of the model is the current round of dialogue pronoun data, which includes the identified pronoun and its position information in the text, such as {"pronoun": "that one", "position": [17, 19]}.

[0114] Step S22: Perform anaphoric object component analysis based on the current round of dialogue pronoun data to obtain the current candidate anaphoric anchor data;

[0115] In an embodiment of the present invention, after identifying the pronoun "that one", the system will perform anaphoric object component analysis to try to determine the type or attribute of the object referred to by "that one". For example, the system can analyze the modifiers before and after the pronoun, such as "I saw before" and "has a bit slow running speed". Combining this information, the system can determine that "that one" refers to a certain electronic product, or more specifically, an electronic product with a not-fast-enough running speed. The system will use this information as the candidate anaphoric anchor data, such as {"type": "electronic product", "attribute": "slow running speed"}.

[0116] Step S23: Dynamically retrieve the historical round of dialogue record according to the user ID data to obtain the historical round of dialogue text data;

[0117] In an embodiment of the present invention, the system will dynamically retrieve the historical dialogue record of the user according to the user ID "user12345". The historical dialogue records are stored in the database, and each record includes information such as user ID, dialogue time, and dialogue content. The system will retrieve the previous dialogue records according to the user ID and the current dialogue time. For example, the system can retrieve the dialogue records of the most recent week, or retrieve the historical records related to the current dialogue topic. Assume the retrieved historical dialogue records are as follows: Record 1: User: "I am considering buying a thin and light laptop." Record 2: User: "I have my eye on a tablet with a very high configuration, but the price is a bit expensive." Record 3: User: "I prefer gaming laptops, with strong performance." The system will use this historical dialogue text data as the input for the subsequent steps, for screening the candidate anaphoric anchors and finally determining the specific object referred to by the pronoun "that one". The retrieval process is dynamic because the historical records of each dialogue are different, and the system needs to dynamically retrieve the relevant historical information according to the current dialogue context.

[0118] Step S24: Dynamically filter the current candidate anaphoric anchor data through the historical round dialogue text data to obtain cross-round filtered anchor list data;

[0119] In the embodiment of the present invention, the system analyzes the historical dialogue text data and extracts the entities and their attributes therein. For example, "thin and light laptop" is extracted from record 1, "tablet computer with high configuration" and "a bit expensive in price" are extracted from record 2, and "gaming laptop" and "powerful performance" are extracted from record 3. Then, the system filters these historical entities according to the current candidate anaphoric anchor data. Since the current candidate anaphoric anchors are "electronic products" and "slow running speed", the system will retain all entities belonging to electronic products, namely "thin and light laptop", "tablet computer with high configuration", and "gaming laptop". These filtered entities and their related information form the cross-round filtered anchor list data.

[0120] Step S25: Use the cross-round filtered anchor list data to perform multi-level anchor similarity matching on the current candidate anaphoric anchor data to generate cross-round anchor matching data;

[0121] In the embodiment of the present invention, the system calculates the similarity between the current candidate anaphoric anchor and each entity in the cross-round filtered anchor list. The similarity calculation can be based on various methods, such as word vector similarity, semantic similarity, etc. In this example, the system calculates the similarity between "slow running speed" and "thin and light", "high configuration", and "powerful performance". Since there is a certain semantic association between "slow running speed" and "powerful performance" (slow running speed means insufficiently powerful performance), the similarity score between them will be higher than that of other entities. The system stores these similarity scores in the cross-round anchor matching data, for example: {"that one": [{"anchor": "thin and light laptop", "similarity": 0.2}, {"anchor": "tablet computer with high configuration", "similarity": 0.3}, {"anchor": "gaming laptop", "similarity": 0.6}]}.

[0122] Step S26: Perform anaphoric target resolution according to the cross-round anchor matching data and perform anaphoric relationship reasoning to obtain the current dialogue language anaphoric relationship graph;

[0123] In the embodiments of the present invention, the system will select the most likely referential target according to the similarity score in the cross-turn anchor matching data. In this example, the similarity score of "gaming laptop" is the highest, so the system will use it as the referential target of "that one". Then, the system will perform referential relationship reasoning. For example, since the current conversation mentions "a bit faster running speed", and the "gaming laptop" mentioned in the historical conversation usually has a relatively high running speed, the system can infer that the user may want to buy a product with a running speed faster than the gaming laptop seen before. The system represents these referential relationships and reasoning results in the form of a graph, forming the current conversation language referential relationship graph. For example, the graph contains two nodes: "that one" and "gaming laptop", and there is an edge between them indicating the referential relationship.

[0124] Step S27: Perform the processing of linking the pronouns to the referential objects according to the current conversation language referential relationship graph, and perform the association of the anchor context to generate cross-turn context referential chain data.

[0125] In the embodiments of the present invention, the system will link the pronoun "that one" to its referential object "gaming laptop" according to the current conversation language referential relationship graph. At the same time, the system will associate the referential object "gaming laptop" with its context information in the historical conversation, such as "I quite like gaming laptops, they have strong performance". This information can help the system better understand the user's intention. The finally generated cross-turn context referential chain data includes the pronouns, the referential objects, and the relevant context information, for example: {"that one": {"referential object": "gaming laptop", "context": "I quite like gaming laptops, they have strong performance"}}.

[0126] Preferably, step S24 includes the following steps:

[0127] Step S241: Mine the historical language conversation objects according to the historical turn conversation text data to obtain the historical language conversation object data;

[0128] Step S242: Perform historical anchor time series encoding on the historical language conversation object data to generate historical time series anchor data;

[0129] Step S243: Perform anchor time decay according to the historical time series anchor data to obtain time decay anchor data; perform anchor type classification according to the historical time series anchor data to obtain historical classified anchor data;

[0130] Step S244: Extract the referential components from the current candidate referential anchor data to obtain the current referential component feature data;

[0131] Step S245: Perform cross-round anchor weighted correlation evaluation on the current anaphoric component feature data through time-decayed anchor data and historical classification anchor data, and calculate the anchor activation degree to obtain historical anchor activation degree data;

[0132] Step S246: Dynamically screen historical temporal anchors using the historical anchor activation degree data based on a preset activation degree threshold to obtain cross-round screened anchor list data.

[0133] In the embodiments of the present invention, language object mining is performed on each historical record. For example, named entity recognition and relation extraction techniques are used to extract entities, attributes, and relations therein. For example, "laptop" and its attribute "light and thin" are extracted from record 1; "tablet computer" and its attributes "high configuration" and "expensive price" are extracted from record 2; "gaming laptop" and its attribute "powerful performance" are extracted from record 3. The extracted information constitutes historical language dialogue object data, such as: [{"timestamp": T1, "entity": "laptop", "attribute": "light and thin"}, {"timestamp": T2, "entity": "tablet computer", "attribute": "high configuration", "price": "expensive"}, {"timestamp": T3, "entity": "gaming laptop", "attribute": "powerful performance"}]. Temporal encoding is performed on the historical language dialogue object data to record the chronological order of the appearance of each anchor point. For example, using timestamps as encoding, the historical language dialogue object data is converted into historical temporal anchor point data, such as: [{"timestamp": T1, "anchor": "light and thin laptop"}, {"timestamp": T2, "anchor": "tablet computer with high configuration"}, {"timestamp": T3, "anchor": "gaming laptop"}]. Temporal decay processing is performed on the historical temporal anchor point data according to the timestamps. For example, the farther an anchor point is from the current time, the lower its weight. Assuming the current time is T4, the weight of T1 is the lowest and the weight of T3 is the highest. At the same time, the system will classify the historical temporal anchor point data according to the entity type. For example, "laptop", "tablet computer", and "gaming laptop" are all classified as "electronic products". Time-decayed anchor point data and historical classified anchor point data are obtained. Assuming that the current candidate referential anchor point data obtained in step S22 is {"type": "electronic product", "attribute": "slow running speed"}. The system will perform referential component extraction on this data, extract the key information and convert it into a feature vector representation. For example, "running speed" and "slow" are extracted as key components. The system can use word embedding techniques, such as Word2Vec or BERT, to convert these components into vector representations. In addition, information such as part of speech and dependency relations can also be considered to enrich the feature representation. The finally obtained current referential component feature data is a set containing multiple vectors, and each vector represents a key component. Cross-round anchor weighted correlation evaluation is performed on the current referential component feature data using the time-decayed anchor point data and the historical classified anchor point data. The system will calculate the similarity between the current referential component feature vector and each historical anchor point feature vector. Similarity calculation can use methods such as cosine similarity and Euclidean distance. For example, calculate the similarity between the feature vector of "slow running speed" and the feature vectors of "light and thin laptop", "tablet computer with high configuration", and "gaming laptop".Meanwhile, the system will weight the similarity scores by combining time decay weights. The closer an anchor point is to the current time, the higher its weight and the greater its contribution to the final activation degree. Based on the weighted similarity scores, the system will calculate the activation degree of each historical anchor point. The activation degree represents the degree of relevance between the anchor point and the current anaphor. Multiple methods can be used for activation degree calculation, such as the softmax function. The finally obtained historical anchor point activation degree data is a list containing each historical anchor point and its corresponding activation degree. The historical anchor points are filtered according to a preset activation degree threshold. For example, if the threshold is set to 0.5, only the anchor points with an activation degree greater than or equal to 0.5 will be retained. In this example, the activation degree of "gaming laptop" is 0.7, which is greater than 0.5 and will be retained; while the activation degrees of "thin and light laptop" and "tablet with high configuration" are 0.2 and 0.3 respectively, which are less than 0.5 and will be filtered out. The finally obtained cross-round filtered anchor point list data contains historical anchor points highly relevant to the current anaphor.

[0134] Preferably, step S3 includes the following steps:

[0135] Step S31: Perform multi-granularity semantic relationship analysis based on the cross-round context anaphor chain data to generate dialogue multi-granularity semantic relationship edge data;

[0136] Step S32: Perform multi-modal node construction based on the historical time-series anchor point data and the current candidate anaphor anchor point data to generate dialogue language multi-modal node data;

[0137] Step S33: Construct dynamic anaphor relationship edges for the cross-round context anaphor chain data, and perform multi-round dialogue directed graph fusion based on the dialogue multi-granularity semantic relationship edge data and the dialogue language multi-modal node data to construct a user multi-round dialogue directed graph;

[0138] Step S34: Perform graph embedding representation encoding on the user multi-round dialogue directed graph to generate an encoded multi-round dialogue directed graph;

[0139] Step S35: Use a preset graph neural network model to identify the current language intention nodes in the encoded multi-round dialogue directed graph, and perform multi-round dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data;

[0140] Step S36: Track the user's subsequent interaction behaviors, and perform negative feedback intention adjustment on the preliminary dialogue language intention vector data to obtain an interactive adjustment intention representation strategy.

[0141] As an example of the present invention, refer to Figure 3 shown in Figure 1 is a detailed implementation step flow diagram of step S3 in

[0142] Step S31: Perform multi-granularity semantic relationship analysis based on cross-turn context reference chain data to generate dialogue multi-granularity semantic relationship edge data;

[0143] In the embodiment of the present invention, assume that the cross-turn context reference chain data generated in step S2 is: {\"that one\": {\"referent\": \"gaming laptop\", \"context\": \"I prefer gaming laptops, which have strong performance\"}}. The system will perform multi-granularity semantic relationship analysis on the reference chain data. Multi-granularity means analyzing at different granularities such as the word level, phrase level, and sentence level. For example, at the word level, the system can analyze the reference relationship between \"that one\" and \"gaming laptop\"; at the phrase level, the system can analyze the contrast relationship between \"a little faster running speed\" and \"strong performance\"; at the sentence level, the system can analyze the sequential relationship between the current turn dialogue \"I want to buy one with a little faster running speed, the one I saw before was a bit slow\" and the historical turn dialogue \"I prefer gaming laptops, which have strong performance\". The system will extract these semantic relationships at different granularities and mark them with appropriate labels, such as \"reference\", \"contrast\", \"sequential\", etc. These marked semantic relationships constitute the dialogue multi-granularity semantic relationship edge data, for example: [{\"source\": \"that one\", \"target\": \"gaming laptop\", \"relationship\": \"reference\"}, {\"source\": \"a little faster running speed\", \"target\": \"strong performance\", \"relationship\": \"contrast\"}].

[0144] Step S32: Perform multi-modal node construction based on historical time-series anchor data and current candidate reference anchor data to generate dialogue language multi-modal node data;

[0145] In the embodiment of the present invention, assume that the historical time-series anchor data includes \"thin and light laptop\", \"tablet with high configuration\", \"gaming laptop\", and the current candidate reference anchor data is {\"type\": \"electronic product\", \"attribute\": \"slow running speed\"}. The system will construct these data into multi-modal nodes. Each node can include text information, image information, audio information, etc. For example, the \"gaming laptop\" node can include text descriptions such as \"strong performance\" and \"can play large games\", and can also include pictures or videos of gaming laptops. The current candidate reference anchor data can be constructed into a \"user requirement\" node, including the text information \"a little faster running speed\". The system will store these node data in a structured form to constitute the dialogue language multi-modal node data, for example: [{\"node ID\": \"1\", \"text\": \"thin and light laptop\", \"picture\":...}, {\"node ID\": \"2\", \"text\": \"tablet with high configuration\", \"picture\":...}, {\"node ID\": \"3\", \"text\": \"gaming laptop\", \"picture\":...}, {\"node ID\": \"4\", \"text\": \"a little faster running speed\"}].

[0146] Step S33: Construct dynamic reference relation edges for cross-turn context reference chain data, and perform multi-turn dialogue directed graph fusion based on dialogue multi-granularity semantic relation edge data and dialogue language multi-modal node data to construct a user multi-turn dialogue directed graph;

[0147] In the embodiment of the present invention, the system constructs reference relation edges according to cross-turn context reference chain data. For example, according to {“that one”: {“referent”: “gaming laptop”, “context”: “I prefer gaming laptops as they have strong performance”}}, the system will establish a directed edge between the “user requirement” node (“faster running speed”) and the “gaming laptop” node. The type of the edge is “reference”, and the direction of the edge is from “that one” to “gaming laptop”. Then, the system fuses the dialogue multi-granularity semantic relation edge data and the dialogue language multi-modal node data together to construct a user multi-turn dialogue directed graph. The nodes in the graph represent entities or concepts mentioned in the dialogue, and the edges represent the semantic relations between the nodes. For example, the graph contains nodes such as “thin and light laptop”, “tablet with high configuration”, “gaming laptop”, “faster running speed”, etc., and edges of types such as “reference”, “contrast”, “sequence”, etc.

[0148] Step S34: Perform graph embedding representation encoding on the user multi-turn dialogue directed graph to generate an encoded multi-turn dialogue directed graph;

[0149] In the embodiment of the present invention, the system performs graph embedding representation encoding on the user multi-turn dialogue directed graph constructed in step S33. Graph embedding is a technique for converting graph-structured data into low-dimensional vector representations. Common graph embedding methods include DeepWalk, Node2Vec, GraphSAGE, etc. For example, using the GraphSAGE model, the system can learn the vector representation of each node, such that nodes that are closer in the graph are also closer in the vector space. The vector representation of a node can fuse multi-modal information such as the text information and image information of the node. The type and direction of the edges are also encoded into the vector representation of the nodes. For example, the vector representation of the “gaming laptop” node will contain the text information of “gaming laptop”, attribute information such as “strong performance”, and the “reference” relation information with the “user requirement” node. The finally generated encoded multi-turn dialogue directed graph is a graph-structured data containing the vector representations of the nodes.

[0150] Step S35: Use a preset graph neural network model to identify the current language intention nodes in the encoded multi-turn dialogue directed graph, and perform multi-turn dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data;

[0151] In the embodiments of the present invention, the system uses a preset graph neural network model to identify the current language intention nodes of the encoded multi-round dialogue directed graph. The graph neural network model can learn the complex relationships between the nodes in the graph and identify the nodes most relevant to the user's current intention. For example, in the scenario of e-commerce product recommendation, the system can train a graph neural network model to identify the products that the user is interested in. The input of the model is the encoded multi-round dialogue directed graph, and the output is the intention score of each node. The higher the score of a node, the more relevant it is to the user's current intention. After identifying the current language intention nodes, the system performs multi-round dialogue intention evolution reasoning. The system analyzes the changing trend of the user's intention in the multi-round dialogue. For example, the user may initially only be interested in "thin and light" laptops, and later mention "running speed". Then the system can infer that the user's current intention is to buy a "thin and light laptop with fast running speed". The system converts the inferred intention into a vector representation to obtain the preliminary dialogue language intention vector data. For example, this vector can be represented as [0.7 (thin and light), 0.8 (fast running speed), 0.2 (price),...], where each dimension represents a product attribute, and the value represents the degree of importance the user attaches to this attribute.

[0152] Step S36: Track the subsequent interaction behavior of the user, and perform negative feedback intention adjustment on the preliminary dialogue language intention vector data to obtain an interaction adjustment intention representation strategy.

[0153] In the embodiments of the present invention, the system tracks the subsequent interaction behavior of the user. For example, on an e-commerce platform, the system records behaviors such as which products the user clicks on, which pages the user browses, and which products the user adds to the shopping cart. These interaction behaviors can serve as feedback signals for the user's intention. For example, if the user clicks on multiple "gaming laptops" products, then the attributes related to "gaming laptops" in the preliminary dialogue language intention vector, such as "powerful performance", can be strengthened. If the user's interaction behavior is inconsistent with the preliminary dialogue language intention vector, negative feedback intention adjustment is required. For example, if the preliminary dialogue language intention vector indicates that the user wants to buy a "thin and light laptop", but the user clicks on multiple "gaming laptops" products, then the system needs to adjust the intention vector, reduce the weight of "thin and light", and increase the weight of "powerful performance". In this way, the system can continuously learn and correct the user's intention, and finally obtain an interaction adjustment intention representation strategy, which can guide subsequent conversations and recommendations.

[0154] Preferably, step S36 includes the following steps:

[0155] Step S361: Track the interaction behavior of the user to obtain user interaction behavior data;

[0156] Step S362: Parse the negative feedback based on the user interaction behavior data to obtain the negative feedback information data;

[0157] Step S363: Locate the negative object according to the negative feedback information data and conduct negative information analysis to obtain the negative keyword and the negative object data respectively;

[0158] Step S364: Based on the negative keyword and the negative object data, perform multi-level propagation of negative feedback through the user multi-round dialogue directed graph to generate the feedback propagation influence node data;

[0159] Step S365: Perform negative feedback weight adaptive correction on the preliminary dialogue language intention vector data through the feedback propagation influence node data to obtain the interactive adjustment intention representation strategy.

[0160] The system of the present invention tracks various interaction behaviors of users and records these behaviors to form user interaction behavior data. In an e-commerce scenario, these interaction behaviors can include: clicking on a product, viewing the product details page, adding the product to the shopping cart, submitting an order, evaluating the product, chatting with a customer service robot, etc. Each interaction behavior will be recorded with relevant information such as a timestamp and a product ID. The system analyzes the user interaction behavior data to identify negative feedback information. Negative feedback refers to the behavior where a user expresses dissatisfaction with a certain product or attribute. For example, the user leaves after quickly viewing the details page of a certain product, or the user clearly expresses disinterest in a certain product during a chat with the customer service robot. The system identifies these negative feedback behaviors based on predefined rules or machine learning models. For example, if the user leaves shortly after clicking on a certain product, the system may judge it as negative feedback. The identified negative feedback information forms negative feedback information data. The system locates the negative object based on the negative feedback information data and analyzes the negative information. For example, the negative object is the product "item456" (thin and light laptop). The system further analyzes the negative information, extracts negative keywords and negative object data. Negative keywords can be "dislike", "dissatisfied", etc., or specific product attributes, such as "too expensive", "too heavy", etc. Negative object data can be a product ID, a product category, etc. For example: {"negative keywords": ["thin and light"], "negative object": "laptop computer"}. The system spreads the negative keywords and negative object data in the user multi-round dialogue directed graph. For example, if the user expresses negativity towards "thin and light laptop computers", the system will find the node "thin and light laptop computers" in the graph and spread the negative information to its adjacent nodes. For example, if there are two nodes "thin and light" and "laptop computer" in the graph and they are connected to the node "thin and light laptop computers", then the negative information will also be spread to these two nodes. Multi-level propagation means that the negative information will be propagated in multiple layers along the edges in the graph, and the influence range gradually expands. During the propagation process, the system records the affected nodes and generates feedback propagation influence node data. Weight correction is performed on the preliminary dialogue language intention vector data obtained in step S35. The correction method can be to apply the weight change of the affected nodes to the corresponding dimensions in the intention vector. For example, if the weight of "red" in the preliminary intention vector is 0.8 and the influence of negative feedback propagation is -0.8, the adjusted weight is 0 = 0.8 + (-0.8). After weight correction, the final interactive adjustment intention representation strategy is obtained, which reflects the user's revised intention and will be used to guide subsequent conversations and recommendations.

[0161] Preferably, the present invention further provides a language data processing system that executes the language data processing method as described above. The language data processing system includes:

[0162] A language feature extraction module, configured to extract the current round of conversation of the user to obtain the user's current round of conversation data; perform conversation language text processing based on the user's current round of conversation data to generate user conversation text feature data; perform audio / video language cross-modal text extraction on the user's current round of conversation data to obtain cross-modal conversation text feature data; perform weighted splicing on the cross-modal conversation text feature data and the user conversation text feature data to obtain the current round of conversation language feature data;

[0163] A referential chain dynamic parsing module, configured to detect pronouns in the current round of conversation language feature data and perform referential object component analysis to obtain the current candidate referential anchor data; obtain the historical round of conversation text data; perform dynamic historical anchor screening on the current candidate referential anchor data through the historical round of conversation text data to obtain cross-round screening anchor list data; perform pronoun-referential object link processing based on the cross-round screening anchor list data to generate cross-round context referential chain data;

[0164] A multi-round conversation intention reasoning module, configured to perform multi-round conversation directed graph fusion based on the cross-round context referential chain data to construct a user multi-round conversation directed graph; use the user multi-round conversation directed graph to perform multi-round conversation intention evolution reasoning to obtain preliminary conversation language intention vector data; track the user's subsequent interaction behaviors and perform negative feedback intention adjustment on the preliminary conversation language intention vector data to obtain an interaction adjustment intention representation strategy;

[0165] A cross-round key reference recognition module, configured to determine the key pronouns in the current round of conversation language according to the interaction adjustment intention representation strategy to achieve cross-round conversational language pronoun recognition.

[0166] Preferably, the present invention further provides a terminal device, and the terminal device includes:

[0167] A processor;

[0168] A memory for storing instructions executable by the processor;

[0169] Wherein, the processor is configured to implement the language data processing method as described in any one of the above.

[0170] Preferably, the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it implements the language data processing method as described in any one of the above.

[0171] This application aims to provide a method that can accurately understand and process complex language data in multi-turn conversations, showing significant advantages especially in colloquial expressions, indefinite references, and multi-modal information fusion. First, through the extraction and weighted splicing of cross-modal text features, this method can comprehensively capture the user's intentions and emotions. Even when the user's expression is vague or the semantics are unclear, it can rely on additional clues provided by non-verbal information such as audio and video to accurately understand the user's true needs. For example, when the user only says "This is good" to evaluate a product, the system can judge whether the user is slightly satisfied or very satisfied through cross-modal information such as the user's expression and intonation, so as to provide more accurate subsequent recommendations or services. By detecting pronouns and analyzing components in the current turn of the conversation, combined with historical turn conversation text data for dynamic historical anchor point screening, a cross-turn context reference chain is constructed, effectively solving the ambiguity problem of pronouns in multi-turn conversations. This enables the system to accurately link pronouns to their referents. For example, in a conversation, the user first mentions "I want to buy a mobile phone" and then says "Its color should be nice", the system can clearly identify that "it" refers to the "mobile phone" rather than other objects, thus improving the accuracy and coherence of conversation understanding. By constructing a directed graph of the user's multi-turn conversation, accurate capture and evolutionary reasoning of the dynamic changes in the user's intentions are achieved. The nodes and edges in the graph represent the key entities, concepts, and their relationships in the conversation. Through the analysis of the graph, the system can track the evolution of the user's intentions in real time and adjust the recommendation strategy in a timely manner. For example, if the user initially has requirements for the storage capacity of the mobile phone and then becomes interested in the camera function, the system can capture this change through graph analysis and quickly adjust the recommendation focus to meet the user's changing needs. By adaptively correcting the initial intention vector with negative feedback information, the intention representation strategy is further optimized. For example, if the user expresses dislike for the recommended red mobile phone, the system will adjust the intention vector in a timely manner according to this feedback, reducing the preference weight for red and avoiding recommending similar products again, thereby improving user satisfaction and the system's recommendation effect, and ensuring that cross-turn conversational language pronoun recognition always revolves around the user's actual needs.

[0172] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0173] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A language data processing method, characterized in that: The following steps are involved: Step S1: extract the current round of conversation of the user to obtain the current round of conversation data of the user; Process the conversation language text according to the user's current round of conversation data to generate user conversation text feature data; Perform audio / video language cross-modal text extraction on the user's current round of conversation data to obtain cross-modal conversation text feature data; Perform weighted concatenation on the cross-modal conversation text feature data and the user conversation text feature data to obtain the current round of conversation language feature data; Step S2: perform pronoun detection on the language feature data of the current round of dialogue, and perform referent object component analysis to obtain current candidate referent anchor point data; obtain historical round dialogue text data; perform dynamic historical anchor point screening on the current candidate referent anchor point data through the historical round dialogue text data to obtain cross-round screening anchor point list data; perform pronoun-referent object linking processing based on the cross-round screening anchor point list data to generate cross-round context referent chain data; Step S3: Based on the cross-round contextual reference chain data, multi-round dialogue directed graphs are integrated to construct a user multi-round dialogue directed graph; the multi-round dialogue intention evolution reasoning is performed using the user multi-round dialogue directed graph to obtain preliminary dialogue language intention vector data; the user's subsequent interactive behavior is tracked, and the preliminary dialogue language intention vector data is adjusted through negative feedback intention to obtain an interactive adjustment intention representation strategy; Step S4: Determine the key pronouns of the current turn of dialogue language according to the interaction adjustment intention representation strategy to achieve cross-turn dialogue language pronoun recognition.

2. The language data processing method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Obtain user ID data; Step S12: extracting the current round of conversation of the user based on the user ID data, and performing conversation timing processing to obtain the current round of conversation data of the user; Step S13: Processing the conversation language text according to the user's current round of conversation data to generate user conversation text feature data; Step S14: extracting audio / video language cross-modal text from the user's current round of conversation data through a preset audio / video language extraction network interface to obtain cross-modal conversation text feature data; Step S15: using the user conversation text feature data to perform relevance evaluation on the cross-modal conversation text feature data to generate cross-modal text relevance score data; Step S16: weighted concatenation of the cross-modal conversation text feature data and the user conversation text feature data is performed based on the cross-modal text relevance score data to obtain the current round conversation language feature data.

3. The language data processing method according to claim 2, characterized in that: Step S13 includes the following steps: Step S131: performing spoken text detection on the user's current round of dialogue data to generate user spoken dialogue text data; Step S132: segmenting the conversation text according to the user's spoken conversation text data, and removing stop words to obtain optimized spoken conversation text data; Step S133: performing context-aware part-of-speech tagging on the optimized spoken dialogue text data to obtain contextual part-of-speech tagging data; Step S134: performing multi-level nested named entity recognition on the optimized spoken dialogue text data through the contextual part-of-speech tagging data to obtain dialogue nested entity data; Step S135: semantic role recognition is performed based on the dialogue nested entity data, and dialogue text feature processing is performed on the optimized spoken dialogue text data to generate user dialogue text feature data.

4. The language data processing method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: performing pronoun detection on the language feature data of the current round of dialogue to generate pronoun data of the current round of dialogue; Step S22: performing referent component analysis based on the current round of dialogue pronoun data to obtain current candidate referent anchor point data; Step S23: dynamically retrieve the historical round conversation records according to the user ID data to obtain the historical round conversation text data; Step S24: dynamically filter the current candidate reference anchor point data through the historical round dialogue text data to obtain cross-round filtered anchor point list data; Step S25: using the cross-round screening anchor point list data to perform multi-level anchor point similarity matching on the current candidate reference anchor point data to generate cross-round anchor point matching data; Step S26: performing referential target resolution based on cross-round anchor point matching data and performing referential relationship reasoning to obtain a referential relationship graph of the current dialogue language; Step S27: Perform pronoun-pronoun object linking processing according to the current dialogue language reference relationship diagram, and perform anchor point context association to generate cross-round context reference chain data.

5. The language data processing method according to claim 4, characterized in that: Step S24 includes the following steps: Step S241: mining historical language dialogue objects based on historical turn dialogue text data to obtain historical language dialogue object data; Step S242: performing historical anchor point time sequence coding on the historical language dialogue object data to generate historical time sequence anchor point data; Step S243: performing anchor point time decay according to the historical time series anchor point data to obtain time decay anchor point data; performing anchor point type classification according to the historical time series anchor point data to obtain historical classification anchor point data; Step S244: extracting the referential component from the current candidate referential anchor point data to obtain the current referential component feature data; Step S245: performing cross-round anchor weighted correlation evaluation on the current reference component feature data through the time decay anchor data and the historical classification anchor data, and performing anchor activation calculation to obtain historical anchor activation data; Step S246: Based on a preset activation degree threshold, the historical anchor point activation degree data is used to dynamically filter the historical time series anchor point data to obtain cross-round filtered anchor point list data.

6. The language data processing method according to claim 5, characterized in that: Step S3 includes the following steps: Step S31: Perform multi-granularity semantic relationship analysis based on the cross-turn contextual reference chain data to generate multi-granularity semantic relationship edge data of the dialogue; Step S32: constructing a multimodal node according to the historical time series anchor point data and the current candidate reference anchor point data to generate multimodal node data of the dialogue language; Step S33: construct dynamic reference relationship edges for cross-round contextual reference chain data, and fuse multi-round dialogue directed graphs based on multi-granularity semantic relationship edge data and multi-modal node data of dialogue language to construct a multi-round dialogue directed graph of users; Step S34: performing graph embedding representation encoding on the user multi-round conversation directed graph to generate an encoded multi-round conversation directed graph; Step S35: using a preset graph neural network model to identify the current language intention node of the encoded multi-round dialogue directed graph, and perform multi-round dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data; Step S36: Track the user's subsequent interactive behavior, and make negative feedback intention adjustments to the preliminary dialogue language intention vector data to obtain an interactive adjustment intention representation strategy.

7. The language data processing method according to claim 6, characterized in that: Step S36 includes the following steps: Step S361: Tracking the user's interactive behavior to obtain user interactive behavior data; Step S362: performing negative feedback analysis according to the user interaction behavior data to obtain negative feedback information data; Step S363: locating negative objects according to the negative feedback information data, and performing negative information analysis to obtain negative keywords and negative object data respectively; Step S364: Based on the negative keywords and the negative object data, negative feedback is propagated in multiple levels through the directed graph of multiple rounds of user conversations to generate feedback propagation influence node data; Step S365: Adaptively correct the negative feedback weight of the preliminary dialogue language intention vector data by feedback propagation influencing node data to obtain an interactive adjustment intention representation strategy.

8. A language data processing system, characterized in that: For executing the language data processing method according to claim 1, the language data processing system comprises: The language feature extraction module is used to extract the current round of conversation of the user to obtain the current round of conversation data of the user; perform conversation language text processing according to the current round of conversation data of the user to generate user conversation text feature data; perform audio / video language cross-modal text extraction on the current round of conversation data of the user to obtain cross-modal conversation text feature data; perform weighted splicing on the cross-modal conversation text feature data and the user conversation text feature data to obtain the current round of conversation language feature data; The dynamic parsing module of the reference chain is used to detect the pronouns on the language feature data of the current round of dialogue, and to analyze the components of the reference object to obtain the current candidate reference anchor point data; obtain the text data of the historical round of dialogue; dynamically filter the current candidate reference anchor point data through the historical round of dialogue text data to obtain the cross-round filtered anchor point list data; perform pronoun-reference object link processing based on the cross-round filtered anchor point list data to generate the cross-round contextual reference chain data; The multi-round dialogue intention reasoning module is used to fuse the multi-round dialogue directed graph based on the cross-round contextual reference chain data to construct the user multi-round dialogue directed graph; use the user multi-round dialogue directed graph to perform multi-round dialogue intention evolution reasoning to obtain preliminary dialogue language intention vector data; track the user's subsequent interactive behavior, and perform negative feedback intention adjustment on the preliminary dialogue language intention vector data to obtain the interactive adjustment intention representation strategy; The cross-turn key pronoun recognition module is used to determine the key pronouns of the current round of dialogue language according to the interaction adjustment intention representation strategy, so as to realize the cross-turn dialogue language pronoun recognition.

9. A terminal device, characterized in that: The terminal device comprises: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the language data processing method described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the language data processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Electric power material purchasing scheme review optimization method and system based on artificial intelligence

    CN120764544A

  • Commodity information global association method and system in multi-round dialogue process

    CN120804275A

  • Method and system for globally associating commodity information in multi-round dialogue process

    CN120804275B

  • Intelligent customer service automatic reply generation method and system based on multi-modal learning

    CN120975248A

  • Big data processing analysis method, system and equipment based on business history and medium

    CN121117160A