A Chat Content Processing Method Based on Multi-level Classification and Associative Memory

Through multi-level classification and associative memory technology, the intelligent chat system can accurately extract and display relevant chat content, solving the problem of inaccurate extraction in existing technologies and improving service quality and user experience.

CN120470126BActive Publication Date: 2025-11-14BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510967150.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-14
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing intelligent chat systems are unable to accurately extract relevant chat content when processing long-term conversations, resulting in insufficient accuracy in understanding and responding to user needs.

Method used

We employ a multi-level classification and associative memory approach. Through a large model, we structure the dialogue content and classify it into multiple levels. We extract and sort associative memory data from the vector library using the association strength type, and dynamically update the vector library to ensure accurate information display.

Benefits of technology

It enables precise management and efficient utilization of chat content, improves the service quality and user experience of the intelligent chat system, and enhances the timeliness and personalized service capabilities of the conversation content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470126B_ABST
    Figure CN120470126B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, specifically to a chat content processing method based on multi-level classification and associative memory. Through steps such as structured processing of dialogue content, multi-level classification, associative memory extraction and sorting display, this application achieves accurate management and efficient utilization of chat content, effectively solving the problem of inaccurate extraction of relevant chat content in existing technologies, and significantly improving the service quality and user experience of intelligent chat systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a chat content processing method based on multi-level classification and associative memory. Background Technology

[0002] With the development of artificial intelligence technology, intelligent chatbots have been widely used in various fields, such as online customer service, health consultation, and personalized recommendations. However, existing intelligent chatbot systems face certain technical bottlenecks when processing long-term conversations. Especially during extended interactions between users and chatbots, traditional methods typically rely on only the most recent few chat messages to maintain the conversational context. This results in an inability to fully utilize important information contained in past conversations, thus affecting the accuracy of the intelligent chatbot's understanding and response to user needs.

[0003] To address the problems existing in traditional intelligent chat systems, current technologies process chat content by extracting dialogue semantics to update the user's long-term memory. However, this method still has certain limitations, particularly its inability to accurately extract relevant chat content.

[0004] Therefore, it is evident that when using long-term memory to process chat content in existing intelligent chat systems, there is a problem of not being able to accurately extract relevant chat content. Summary of the Invention

[0005] The purpose of this application is to provide a chat content processing method based on multi-level classification and associative memory, so as to solve the problem that when using long-term memory to process chat content in intelligent chat systems in the prior art, it is impossible to accurately extract relevant chat content.

[0006] To achieve the above objectives, the technical solution adopted in this application is as follows:

[0007] According to one aspect of the embodiments of this application, a chat content processing method based on multi-level classification and associative memory is provided, comprising: receiving dialogue content input by a user and converting the dialogue content into structured text data; using a large model to classify the text data into multiple levels according to a preset classification standard to obtain classification labels, wherein the classification labels are secondary classification labels containing main category labels and sub-category labels, and the classification standard is used to define multiple main category labels and sub-category labels under each main category label; using the large model to extract associative memory data from a vector library according to the classification labels and a preset association strength type, wherein the association strength type is used to indicate the semantic association strength between the dialogue content and the classification labels, and the vector library is dynamically updated based on the text vectors corresponding to the text data using the large model; sorting the associative memory data according to the semantic association strength indicated by the association strength type, and displaying the sorted associative memory data to the user through a visualization interface, wherein data with high association strength is displayed first.

[0008] Based on the aforementioned technical means, the processing format of unstructured dialogue content is unified through structured text conversion, providing a standardized data foundation for subsequent analysis. A two-level classification system of main and subcategories is adopted to achieve refined hierarchical organization of text data, ensuring the logical rigor of the classification framework while enhancing scenario coverage through subcategory expansion. Based on a dynamically updated vector library and association strength model, the intelligent chat system can capture semantic association features in real time and quantify the relevance level of the memorized data. Finally, through an association strength-driven sorting and display mechanism, high-value information is prioritized, thereby improving information retrieval efficiency while achieving accuracy, timeliness, and user orientation in dialogue content processing through hierarchical classification and dynamic memory association technology.

[0009] Furthermore, the large model is used to classify the text data in multiple levels according to the preset classification criteria to obtain classification labels. This includes: using the large model to classify the text data in the first classification according to the main classification criteria to obtain multiple main category labels, where the main classification criteria are used to define the main category labels; and using the large model to classify the text data in the second classification according to the sub-classification criteria to obtain multiple sub-category labels, where the sub-classification criteria are used to define the sub-category labels under each main category label.

[0010] Based on the aforementioned technical methods, the first classification using the primary classification standard can quickly categorize text data into predefined primary category labels, thereby establishing a clear information hierarchy framework and providing a structured foundation for subsequent analysis and processing. Building upon this, a second classification using sub-classification standards further refines the granularity of information within each primary category, resulting in more accurate classification results that better meet specific scenario requirements. This two-stage, multi-level classification method not only enhances the logical rigor and scenario coverage of the classification but also reduces the complexity of each individual classification step through layered processing, thereby improving overall processing efficiency and classification accuracy.

[0011] Furthermore, the main classification criteria include, but are not limited to, the following category tags: basic user information, health status, lifestyle habits, mental health, family member information, and others.

[0012] Based on the aforementioned technical means, by pre-setting a main classification standard covering core dimensions such as user basic information, health records, and lifestyle habits, the full-dimensional structured analysis of dialogue content is achieved. This classification framework not only improves the logical rigor of information organization, but also provides an scalable semantic navigation path for subsequent sub-classification and cross-domain association memory retrieval. Thus, while ensuring data processing efficiency, it significantly enhances the cognitive granularity and service professionalism of the intelligent chat system in complex scenarios.

[0013] Furthermore, the association strength types include three types: direct association, indirect association, and potential association. Direct association indicates that there is a direct relationship between the dialogue content and the category label, indirect association indicates that there is an indirect relationship between the dialogue content and the category label, and potential association indicates that there is a potential relationship between the dialogue content and the category label. The semantic association strength of the three types in the association strength type from high to low is: direct association, indirect association, and potential association.

[0014] Based on the aforementioned technical means, by introducing three types of semantic association strength—direct association, indirect association, and potential association—the intelligent chat system constructs a multi-layered association memory system. Direct association ensures the accurate matching of core relevant information, indirect association expands the boundaries of semantic understanding to capture implicit connections, and potential association uncovers unexpressed demand associations through deep semantic analysis. This hierarchical association mechanism not only makes information retrieval present a gradient feature from precise to exploratory, but also achieves dynamic adaptation between user focus and the knowledge reserves of the intelligent chat system through association strength ranking. While ensuring the priority display of core information, it provides a progressive information discovery path for complex dialogue scenarios, thereby significantly improving the semantic understanding depth and personalized service capabilities of the intelligent chat system.

[0015] Furthermore, after obtaining the classification labels, the method also includes: dynamically updating the vector library based on the text vectors corresponding to the text data using a large model. The step of dynamically updating the vector library includes: using the large model to vectorize the text data to obtain the corresponding text vectors; and storing the text vectors and classification labels in the vector library.

[0016] Based on the aforementioned technical means, by dynamically updating the vector library, the intelligent chat system can capture and store the text vectors and classification tags of each conversation in real time, thereby ensuring that the vector library always reflects the latest dialogue context and user characteristics. This mechanism not only enhances the real-time response capability of the intelligent chat system, but also provides a data foundation for deep semantic understanding and accurate memory retrieval by continuously accumulating personalized user information, thus improving the personalization and intelligence level of the dialogue service.

[0017] Furthermore, the text vectors and classification labels are stored in a vector library, including: performing similarity matching between the text vectors and all text vectors in the vector library to determine whether a target vector associated with the text vector exists in the vector library; if it is determined that no target vector associated with the text vector exists in the vector library, a summary of the text data is generated through a large model, and the summary, classification labels, and text vectors are stored as new records in the vector library; if it is determined that a target vector associated with the text vector exists in the vector library, the summary is updated through a large model based on the historical summaries and text data corresponding to the target vector, and the classification labels and text vectors are stored as new records in the vector library.

[0018] Based on the aforementioned technical means, a similarity matching mechanism is used to achieve dynamic deduplication and incremental updates of the vector library. This avoids the accumulation of redundant data and extracts the core semantics of the dialogue through summary generation technology. At the same time, when a related vector is detected, a historical summary fusion and update strategy is adopted to ensure that the vector library always remains concise, efficient, and contains the latest dialogue context. This dual-mode storage mechanism not only optimizes the storage efficiency of the knowledge base, but also provides structured and temporal data support for subsequent accurate memory retrieval based on semantic association strength through the association and binding of classification tags and summaries. This improves the response speed of the intelligent chat system while enhancing the coherence of dialogue history management and personalized service capabilities.

[0019] Furthermore, based on the similarity matching between the text vector and all text vectors in the vector library, it is determined whether there is a target vector associated with the text vector in the vector library. This includes: calculating the similarity between the text vector and each text vector in the vector library using a preset similarity calculation method; if the similarity is greater than or equal to a preset similarity threshold, it is determined that there is a target vector associated with the text vector in the vector library, and the text vector with the highest similarity is selected as the target vector; otherwise, it is determined that there is no target vector associated with the text vector in the vector library.

[0020] Based on the aforementioned technical means, through similarity calculation and threshold comparison mechanisms, the intelligent chat system can accurately identify the repetition or similarity of dialogue content, thereby avoiding the accumulation of redundant data in the vector library and effectively optimizing storage resources. At the same time, this mechanism ensures that new record creation is triggered only when the dialogue content is sufficiently novel, while when highly similar content is detected, the historical semantics are continuously updated through target vector association. This dynamic balancing strategy maintains the conciseness of the knowledge base and ensures the priority reuse of the most relevant historical context through similarity ranking, thereby improving the operating efficiency of the intelligent chat system while enhancing the coherence and semantic consistency of dialogue memory management.

[0021] According to another aspect of the embodiments of this application, a chat content processing apparatus based on multi-level classification and associative memory is also provided, comprising:

[0022] The acquisition module is used to receive the dialogue content input by the user and convert the dialogue content into structured text data;

[0023] The classification module is used to classify text data into multiple levels according to a preset classification standard using a large model, and obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification standard is used to define multiple main category labels and sub-category labels under each main category label.

[0024] The query module is used to extract associated memory data from the vector library based on the classification labels and preset association strength types using a large model. The association strength type is used to indicate the semantic association strength between the dialogue content and the classification labels. The vector library is dynamically updated based on the text vectors corresponding to the text data using a large model.

[0025] The display module sorts the associated memory data according to the semantic association strength indicated by the association strength type, and displays the sorted associated memory data to the user through a visual interface. Among them, the data with high association strength is displayed first.

[0026] According to another aspect of the embodiments of this application, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; wherein the memory is used to store computer programs; and the processor is used to execute the chat content processing method steps based on multi-level classification and associative memory in any of the above embodiments by running the computer programs stored in the memory.

[0027] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the chat content processing method steps based on multi-level classification and associative memory in any of the above embodiments when running.

[0028] The beneficial effects of this application are:

[0029] This application achieves precise management and efficient utilization of chat content through steps such as structured processing of dialogue content, multi-level classification, associative memory extraction, and sorted display. It effectively solves the problem of the inability to accurately extract relevant chat content in existing technologies, and significantly improves the service quality and user experience of intelligent chat systems. Attached Figure Description

[0030] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of the hardware environment for an optional chat content processing method based on multi-level classification and associative memory provided in an embodiment of this application;

[0033] Figure 2 This is a flowchart illustrating an optional chat content processing method based on multi-level classification and associative memory provided in an embodiment of this application.

[0034] Figure 3 This is a structural block diagram of an optional chat content processing device based on multi-level classification and associative memory provided in an embodiment of this application;

[0035] Figure 4 This is a structural block diagram of an optional electronic device provided in an embodiment of this application. Detailed Implementation

[0036] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0038] According to one aspect of the embodiments of this application, a chat content processing method based on multi-level classification and associative memory is provided. Optionally, in this embodiment, the above-mentioned chat content processing method based on multi-level classification and associative memory can be applied to a hardware environment consisting of a terminal and a server. The server is connected to the terminal via a network and can be used to provide services to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services to the server.

[0039] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal may not be limited to PC, mobile phone, tablet computer, etc.

[0040] The chat content processing method based on multi-level classification and associative memory in this application can be executed by a server, a terminal, or both. Specifically, the execution of the chat content processing method based on multi-level classification and associative memory in this application can also be performed by a client installed on the terminal.

[0041] Taking the chat content processing method based on multi-level classification and associative memory in this embodiment, executed by the server, as an example, please refer to... Figure 1 , Figure 1 This is a schematic diagram of the hardware environment for an optional chat content processing method based on multi-level classification and associative memory, as provided in an embodiment of this application. Figure 1 As shown, the hardware environment of the chat content processing method based on multi-level classification and associative memory includes: a terminal 102, and a server 104 connected to the terminal 102 via a network. The server 104 is used to deploy a large model, which executes the chat content processing method based on multi-level classification and associative memory of this application embodiment to perform multi-level classification and associative memory processing on the chat content; the terminal 102 is used to display the processed associative memory data, wherein the associative memory data can be obtained by classifying, associating, and sorting by the large model deployed on the server 104.

[0042] The chat content processing method based on multi-level classification and associative memory in this embodiment can be applied to scenarios such as intelligent customer service, intelligent assistants, and psychological counseling. For example, in intelligent customer service scenarios, multi-level classification and associative memory processing can more accurately understand user needs and provide personalized services; in psychological counseling scenarios, associative memory processing can uncover users' potential emotional needs and provide more considerate psychological support. This embodiment uses an intelligent customer service scenario as an example to illustrate the above-mentioned chat content processing method based on multi-level classification and associative memory.

[0043] With the development of artificial intelligence technology, intelligent chatbots have been widely used in the field of intelligent customer service. However, when processing long-term dialogue content, existing intelligent chat systems process chat content by extracting dialogue semantics to update the user's long-term memory. This method cannot accurately extract relevant chat content.

[0044] To address the aforementioned issues, this embodiment provides a chat content processing method based on multi-level classification and associative memory, running on the aforementioned server. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating an optional chat content processing method based on multi-level classification and associative memory provided in an embodiment of this application, as shown below. Figure 2 As shown, the chat content processing method based on multi-level classification and associative memory in this application embodiment specifically includes the following steps:

[0045] Step S201: Receive the dialogue content input by the user and convert the dialogue content into structured text data;

[0046] Step S202: Use the large model to classify the text data into multiple levels according to the preset classification criteria to obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification criteria are used to define multiple main category labels and sub-category labels under each main category label.

[0047] Step S203: Using a large model, association memory data is extracted from the vector library based on the classification labels and the preset association strength type. The association strength type is used to indicate the semantic association strength between the dialogue content and the classification labels. The vector library is dynamically updated based on the text vectors corresponding to the text data using the large model.

[0048] Step S204: Sort the association memory data according to the semantic association strength indicated by the association strength type, and display the sorted association memory data to the user through a visualization interface, wherein data with high association strength is displayed first.

[0049] Through steps S201 to S204, the system achieves precise management and efficient utilization of chat content by structuring dialogue content, multi-level classification, associative memory extraction, and sorting display. This effectively solves the problem of not being able to accurately extract relevant chat content in existing technologies, and significantly improves the service quality and user experience of the intelligent chat system.

[0050] The following is combined with Figure 2 The chat content processing method based on multi-level classification and associative memory in the embodiments of this application will be explained.

[0051] In the technical solution of step S201, the dialogue content input by the user is received and the dialogue content is converted into structured text data.

[0052] In the dialogue between an intelligent chatbot and a user, the user first asks a question or initiates communication. The dialogue content refers to the user's input, which can include any form of text, voice, or other multimodal information. The intelligent chatbot transforms the dialogue content into structured text data for further analysis. To enhance the intelligent chatbot's understanding capabilities, it not only needs to perform basic lexical analysis on the text content but also employs large models (such as chatGPT and deepseek) to analyze the user's underlying intentions and emotional tone, thus laying a solid foundation for subsequent classification and memory storage.

[0053] In intelligent chat scenarios, receiving user-input dialogue and transforming it into structured text data is a crucial first step in providing precise services. This step is not simply text entry, but involves a series of intelligent processing steps aimed at converting natural language into structured information that machines can understand and analyze.

[0054] Specifically, when a user inputs dialogue content through the intelligent chat system, the system first receives this information. Next, the intelligent chat system performs structured processing on the received dialogue content, transforming unstructured natural language text into a structured data format. It is understood that structured data has a clear data type and format, facilitating efficient processing and analysis by computers. In this embodiment, structured processing is achieved through large models such as chatGPT and deepseek, and this structured processing typically includes the following aspects:

[0055] Entity Recognition and Extraction: The intelligent chat system uses natural language processing technology to identify key entities in the dialogue, such as names, locations, times, and product names, and extracts them. For example, when a user says, "I want to check the weather in City A tomorrow," the intelligent chat system can recognize that "City A" is a location name and "tomorrow" is a time.

[0056] Intent recognition: Intelligent chat systems analyze the semantics of the conversation to understand the user's true intent. For example, if a user says, "I've been having trouble sleeping lately," the intelligent chat system can recognize that the user's intent is to seek advice on how to solve the insomnia problem.

[0057] Sentiment analysis: Intelligent chat systems can determine the emotional tone of conversations, such as positive, negative, or neutral. This helps intelligent customer service better understand users' emotional states and provide more personalized service.

[0058] In this embodiment, the structured data is saved as a JSON structure. It is understood that using JSON format to store structured data can significantly improve data processing efficiency, ensure data consistency, enable efficient querying and retrieval, thereby improving classification accuracy and supporting the application of complex classification models, bringing great convenience and advantages to data processing and analysis.

[0059] Understandably, through structured processing, the original dialogue content is transformed into structured text data rich in information. This data not only facilitates computer understanding and processing but also provides a solid foundation for subsequent multi-level classification, associative memory processing, and personalized services. In intelligent customer service scenarios, structured text data is the cornerstone of intelligent chat systems, enabling them to more accurately understand user needs and provide more efficient and personalized services.

[0060] In the technical solution of step S202, the text data is classified into multiple levels according to the preset classification criteria using a large model to obtain classification labels.

[0061] To achieve multi-level classification of text data and obtain secondary classification labels containing main category labels and subcategory labels, we fully leverage the powerful semantic understanding and intelligent classification capabilities of large-scale models, while strictly adhering to predefined classification criteria. The classification labels are secondary classification labels containing main category labels and subcategory labels. The classification criteria are used to define multiple main category labels and the subcategory labels under each main category label.

[0062] In this embodiment, the classification criteria are predefined, hierarchical criteria, consisting of primary classification criteria and sub-classification criteria. The primary classification criteria define the main category labels, which include multiple primary category labels. The sub-classification criteria define the sub-category labels under each primary category label, which also include multiple sub-category labels. The primary category label is the highest-level classification, summarizing the core category attributes of the text data. The sub-category labels are further subdivisions of the primary category labels, used to more precisely describe the specific content of the text data.

[0063] As an optional implementation, the main classification criteria include, but are not limited to, the following category tags: basic user information, health status, lifestyle habits, mental health, family member information, and others.

[0064] In practice, each primary classification criterion is further subdivided into multiple subcategories, and all the subcategories constitute the subclassification criteria. For example:

[0065] User basic information is used to describe the user's basic attributes, and its sub-category tags may include: name, age, gender, etc.

[0066] Health status is used to describe a user's health-related information, and its subcategories can include: physical examination results, medical history, medication history, etc.

[0067] Lifestyle habits are used to describe a user's lifestyle, and its subcategories can include: exercise habits, eating habits, sleep habits, entertainment habits, etc.

[0068] Mental health is used to describe a user's psychological state, and its subcategories can include: emotional state, stress level, etc.

[0069] Family member information is used to describe the user's family member information, and its sub-category tags may include: family member relationships, family member health status, etc.

[0070] In practical applications, each subcategory can be further subdivided according to needs. For example, medical history can be subdivided into chronic medical history, acute medical history, etc.

[0071] Understandably, by pre-setting a main classification standard that covers core dimensions such as user basic information, health records, and lifestyle habits, the system achieves a full-dimensional structured analysis of the dialogue content. This classification framework not only improves the logical rigor of information organization but also provides an scalable semantic navigation path for subsequent sub-classification and cross-domain association memory retrieval. Thus, while ensuring data processing efficiency, it significantly enhances the cognitive granularity and service professionalism of the intelligent chat system in complex scenarios.

[0072] As an optional implementation, a large model is used to classify text data in multiple levels according to a preset classification standard to obtain classification labels. This includes: using the large model to classify the text data in the first stage according to the main classification standard in the classification standard to obtain multiple main category labels, wherein the main classification standard is used to define the main category labels; and using the large model to classify the text data in the second stage according to the sub-classification standard in the classification standard to obtain multiple sub-category labels, wherein the sub-classification standard is used to define the sub-category labels under each main category label.

[0073] Upon receiving text data, the large model deeply understands the semantics and contextual information of the text, intelligently classifying the text data under the most appropriate main category label and subcategory label. The large model classification process is as follows:

[0074] The large model first performs a deep semantic understanding of the text data, grasping the core meaning and contextual information of the text. For example, for the text "I feel I prefer playing badminton," the large model can understand that "playing badminton" reflects the user's lifestyle habits.

[0075] The large model then uses its understanding of text semantics to categorize text data into corresponding main categories based on predefined classification criteria. For example, "playing badminton" is categorized under the main category of "lifestyle habits".

[0076] Furthermore, the large model will classify text data into more granular subcategories based on the complex information identified in the context, achieving multi-level, fine-grained classification. For example, "playing badminton" will be categorized under the "exercise habits" subcategory.

[0077] The above categorized content is displayed in JSON format as follows:

[0078] [{'id': '271fe04c-f262-4ac9-9e20-d488302e3a56', 'memory': 'I feel I prefer playing badminton', 'hash': '6facbeb0b6a40fc6e822aca174807f9c', 'metadata': {'business_line': 'lifestyle habits', 'main_category': 'exercise habits', 'sub_category': 'exercise methods'}, 'created_at': '2025-01-15T19:52:15.328509-08:00', 'updated_at':None, 'user_id': '53454', 'agent_id': 'MasterAgent'}]

[0079] Among them, memory is used to record text data, and metadata is used to record category labels.

[0080] In this embodiment, the large model has a certain degree of adaptability, and can flexibly adjust the classification method according to the dynamic changes in the dialogue content to ensure the real-time performance and accuracy of the classification.

[0081] Understandably, the first classification using the main classification criteria can quickly categorize text data into predefined main category labels, thus establishing a clear information hierarchy framework and providing a structured foundation for subsequent analysis and processing. Based on this, a second classification using sub-classification criteria further refines the information granularity under each main category, making the classification results more accurate and tailored to specific scenario needs. This two-stage, multi-level classification method not only improves the logical rigor and scenario coverage of the classification but also reduces the complexity of a single classification through layered processing, thereby improving overall processing efficiency and classification accuracy. In the technical solution of step S203, a large model is used to extract association memory data from the vector library based on the classification labels and preset association strength types. The association strength type indicates the semantic association strength between the dialogue content and the classification labels, and the vector library is dynamically updated using the large model based on the text vectors corresponding to the text data.

[0082] A vector library is a database that stores text vectors and category labels. It uses a vectorized approach to store data, enabling computers to perform similarity calculations and fast retrieval.

[0083] To accurately extract the information users need from massive long-term memories, we leverage a large model, combining classification labels and predefined association strength types, to efficiently retrieve relevant memory data from a dynamically updated vector library. To more accurately understand user needs and extract the most relevant memory data from the vector library, we predefined association strength types. These types indicate the semantic association strength between dialogue content and classification labels, measuring the relevance between memory data and user needs.

[0084] As an optional embodiment, the association strength types include three types: direct association, indirect association, and potential association. Direct association indicates that there is a direct association between the dialogue content and the category label, indirect association indicates that there is an indirect relationship between the dialogue content and the category label, and potential association indicates that there is a potential relationship between the dialogue content and the category label. The semantic association strength of the three types in the association strength type from high to low is: direct association, indirect association, and potential association.

[0085] Direct association indicates the most direct and obvious semantic connection between memory data and the user's query intent. This type of memory data usually directly answers the user's question or is highly relevant to the user's question. For example, when a user asks "treatment methods for hypertension," the directly associated memory data is "medical treatment options for hypertension" that the user previously consulted.

[0086] Indirect associations indicate that there is a semantic connection between the remembered data and the user's query intent, but this connection is not direct or obvious and may require logical reasoning or background knowledge to establish. This type of remembered data may provide users with background information, relevant cases, or extended knowledge to help them understand the question more comprehensively. For example, when a user asks "treatment methods for hypertension," the indirectly associated remembered data might be "dietary restrictions for hypertension" that the user previously knew.

[0087] Potential associations refer to the unspoken, semantic connections between remembered data and a user's query intent. This type of remembered data may seem unrelated to the user's question on the surface, but through in-depth analysis and mining, a potential link may be discovered, providing unexpected insights or assistance to the user. For example, when a user asks about "treatment methods for high blood pressure," the potentially associated remembered data might be information the user previously inquired about regarding "stress management and mental health," as stress management can also play a supporting role in controlling high blood pressure.

[0088] The association strength type of the current text data is determined by matching the corresponding association relationship according to the classification. This association relationship can be determined manually or automatically matched by a large model.

[0089] When a user submits a query, the large model combines category labels and pre-defined association strength types to accurately extract association memory data from a dynamically updated vector library. The process of the large model extracting association memory data is as follows:

[0090] The large model first performs a deep semantic understanding of the user's query to accurately grasp the user's true intent and needs. For example, when a user asks, "I've been having trouble sleeping lately, what should I do?", the large model can understand that the user's core need is to find a solution to the insomnia problem.

[0091] Based on its understanding of the user's query intent, the large model determines the long-term memory category to be queried, which is the category label. For example, for the question of "insomnia", the large model may match category labels such as "mental health" or "lifestyle habits > sleep habits".

[0092] The large model determines the strength of the association between the stored memory data in the vector library and the user's query intent based on the preset association strength type. For example, for the question of "insomnia", the large model will comprehensively consider factors such as semantic similarity, logical relationship, and background knowledge to extract content such as "mental health" or "lifestyle habits > sleep habits".

[0093] Understandably, by introducing three types of semantic association strength—direct association, indirect association, and potential association—the intelligent chat system constructs a multi-layered associative memory system. Direct association ensures accurate matching of core relevant information, indirect association expands the boundaries of semantic understanding to capture implicit connections, and potential association uncovers unexpressed needs through deep semantic analysis. This hierarchical association mechanism not only makes information retrieval present a gradient from precise to exploratory but also achieves dynamic adaptation between user focus and the knowledge reserves of the intelligent chat system through association strength ranking. While ensuring the priority display of core information, it provides a progressive information discovery path for complex dialogue scenarios, thereby significantly improving the semantic understanding depth and personalized service capabilities of the intelligent chat system.

[0094] In the technical solution of step S204, the associated memory data is sorted according to the semantic association strength indicated by the association strength type, and the sorted associated memory data is displayed to the user through a visual interface, wherein data with high association strength is displayed first.

[0095] After extracting the associative memory data related to the user's query intent, to further improve the user experience and ensure that users can efficiently obtain the information they need, this associative memory data will be sorted and presented to the user intuitively through a visual interface. The entire process of associative memory data sorting and visualization is as follows:

[0096] The associated memory data extracted from the vector library is prioritized according to the strength of the association, ensuring that data with high association strength is displayed to users first.

[0097] Using a large model, the association strength of each extracted associated memory data is evaluated. The evaluation is based on multiple dimensions, including semantic similarity, logical relevance, and contextual connection. For example, for a user query "treatment methods for hypertension," a memory data point that directly describes "hypertension drug treatment plan" has a significantly stronger association than a memory data point that describes "healthy eating advice."

[0098] Based on the results of the association strength assessment, the intelligent chat system pre-defines a set of ranking rules. Typically, directly related memory data receives the highest ranking priority, followed by indirectly related memory data, and finally, potentially related memory data. This ranking rule ensures that users see the most relevant and direct information to their query intent first. It's important to note that the ranking rules are not static; the intelligent chat system dynamically adjusts them based on the specific query scenario and user needs. For example, in some complex query scenarios, indirectly or potentially related memory data may be more valuable to the user, and the intelligent chat system will accordingly increase the ranking priority of this type of memory data.

[0099] The sorted associative memory data ultimately needs to be presented to the user through an intuitive and easy-to-use visual interface. The visual interface acts like an "information visualization window," transforming abstract textual information into an intuitive visual presentation, helping users understand and digest the information more easily. The specific implementation or design details of the visual interface are based on general technical implementation and user experience principles, and are not specifically limited in this embodiment.

[0100] Understandably, the structured text conversion unifies the processing format of unstructured dialogue content, providing a standardized data foundation for subsequent analysis; the adoption of a main-subcategory two-level classification system enables refined hierarchical organization of text data, ensuring the logical rigor of the classification framework while enhancing scenario coverage through subclass expansion; based on a dynamically updated vector library and association strength model, the intelligent chat system can capture semantic association features in real time and quantify the relevance level of the memorized data; finally, through an association strength-driven sorting and display mechanism, high-value information is prioritized, thereby improving information retrieval efficiency while achieving accuracy, timeliness, and user orientation in dialogue content processing through hierarchical classification and dynamic memory association technology.

[0101] The vector library is the "memory repository" of the entire long-term memory system. It is not static but in dynamic update. Whenever new conversation content is generated, the intelligent chat system uses a large model to vectorize this text data, converting it into a vector form that the computer can efficiently understand and process, and stores it in the vector library. This dynamic update mechanism ensures that the memory data stored in the vector library is always the latest and most relevant. In this embodiment, after obtaining the classification label, the vector library is dynamically updated based on the text vector corresponding to the text data using the large model. As an optional embodiment, after obtaining the classification label, the method further includes: dynamically updating the vector library based on the text vector corresponding to the text data using the large model, where the step of dynamically updating the vector library includes: vectorizing the text data using the large model to obtain the corresponding text vector; storing the text vector and the classification label in the vector library.

[0102] A text vector is a vector form obtained by converting text data through mathematical transformation. Vectorization enables the computer to understand and process text data. It captures the semantic information of the text, making texts with similar semantics close to each other in the vector space. A classification label is a label used to identify the category to which a text vector belongs. It helps the intelligent chat system manage the text data in a structured manner, enabling quick location of relevant memory data during retrieval.

[0103] Vectorization is the process of converting text data into numerical vectors that the computer can understand and process. This process is usually completed by a large model, which can capture the semantic information of the text and encode it into vector form. The general steps of vectorization are as follows:

[0104] Text preprocessing, including operations such as cleaning, tokenizing, and removing stop words from the original text data. Cleaning can remove noise in the text, such as special symbols and emojis; tokenizing is to split continuous text into meaningful word units; removing stop words is to filter out words that contribute less to semantic understanding, such as "de", "le", "zai", etc. The preprocessed text is more standardized and clean, which is beneficial to improving the quality of vectorization.

[0105] The preprocessed text data is input into the large model. The large model converts the words, phrases, and sentences in the text into high-dimensional numerical vectors. These vectors capture the semantic information of the text, making texts with similar semantics close to each other in the vector space. The vectorized text data can then be efficiently processed and analyzed by the computer.

[0106] Finally, the obtained text vector and the predefined classification label are stored in the vector library together. In this way, the vector library is dynamically updated with a new memory data, preparing for subsequent memory retrieval and application.

[0107] Understandably, by dynamically updating the vector library, intelligent chat systems can capture and store the text vectors and classification labels of each conversation in real time, ensuring that the vector library always reflects the latest conversation context and user characteristics. This mechanism not only enhances the real-time responsiveness of intelligent chat systems but also provides a data foundation for deep semantic understanding and accurate memory retrieval by continuously accumulating personalized user information, thereby improving the personalization and intelligence of dialogue services.

[0108] As an optional embodiment, storing text vectors and classification labels in a vector library includes: performing similarity matching between the text vector and all text vectors in the vector library to determine whether a target vector associated with the text vector exists in the vector library; if it is determined that no target vector associated with the text vector exists in the vector library, a summary of the text data is generated using a large model, and the summary, classification labels, and text vector are stored as new records in the vector library; if it is determined that a target vector associated with the text vector exists in the vector library, the summary is updated using a large model based on the historical summary and text data corresponding to the target vector, and the classification labels and text vector are stored as new records in the vector library.

[0109] Storing text vectors and category labels in a vector library is not simply a matter of data accumulation, but a sophisticated process that integrates intelligent matching, information refinement, and dynamic updates. This process ensures that the stored data in the vector library is both comprehensive and accurate, and can continuously evolve as information is updated.

[0110] Before storing a new text vector into the vector library, the intelligent chat system first performs similarity matching. The purpose of similarity matching is to determine whether a "target vector" with semantically similar meaning to the current text vector already exists in the vector library.

[0111] Based on the similarity matching results, the intelligent chat system will adopt different strategies to process text data, one of the core operations being summary generation.

[0112] If it is determined that no target vector associated with the current text vector exists in the vector library, it indicates that the current text data contains new information. At this point, the intelligent chat system uses a large model to summarize and refine the original text data, generating a concise summary. This summary encapsulates the core content of the text data while significantly reducing the data volume. Subsequently, the intelligent chat system stores the generated summary, category labels, and text vector as a new record in the vector library.

[0113] If a target vector associated with the current text vector is found in the vector library, it indicates that the current text data is highly similar to or repeats information already existing in the vector library, potentially containing updates or supplements to existing information. In this case, the intelligent chat system will not directly store a new summary. Instead, it will utilize a large model to perform a comprehensive analysis, combining the historical summaries corresponding to the target vector with the current text data. The large model will determine whether the current text data supplements, corrects, or updates existing information, and generate an updated summary accordingly. This updated summary integrates information from historical summaries and the current text data, ensuring the accuracy and timeliness of the summary. Subsequently, the intelligent chat system will overwrite the historical summaries corresponding to the target vector with the updated summary, while storing the category label and text vector as a new record in the vector library.

[0114] Understandably, the dynamic deduplication and incremental updates of the vector library are achieved through a similarity matching mechanism. This dynamic update mechanism avoids the accumulation of redundant data, extracts the core semantics of the dialogue through summary generation technology, and adopts a historical summary fusion update strategy when related vectors are detected. This ensures that the vector library always remains concise, efficient, and contains the latest dialogue context. This dual-mode storage mechanism not only optimizes the storage efficiency of the knowledge base, but also provides structured and temporal data support for subsequent accurate memory retrieval based on semantic association strength through the association and binding of classification tags and summaries. This improves the response speed of the intelligent chat system while enhancing the coherence of dialogue history management and personalized service capabilities.

[0115] As an optional embodiment, the method involves performing similarity matching between the text vector and all text vectors in the vector library to determine whether a target vector associated with the text vector exists in the vector library. This includes: calculating the similarity between the text vector and each text vector in the vector library using a preset similarity calculation method; if the similarity is greater than or equal to a preset similarity threshold, then it is determined that a target vector associated with the text vector exists in the vector library, and the text vector with the highest similarity is selected as the target vector; otherwise, it is determined that no target vector associated with the text vector exists in the vector library.

[0116] The intelligent chat system calculates the similarity between the current text vector and all text vectors in the vector library. Similarity calculation can use common vector similarity measures such as cosine similarity and Euclidean distance.

[0117] After calculating the similarity, the intelligent chat system compares it with a preset similarity threshold. If the similarity is greater than or equal to the threshold, it is considered that there is a target vector in the vector library associated with the current text vector, indicating that the information expressed by the current text vector is highly similar to or repeats the information already in the vector library; if the similarity is less than the threshold, it is considered that there is no target vector in the vector library associated with the current text vector, indicating that the current text vector contains new information.

[0118] It is understandable that, through similarity calculation and threshold comparison mechanisms, intelligent chat systems can accurately identify the repetition or similarity of dialogue content, thereby avoiding the accumulation of redundant data in the vector library and effectively optimizing storage resources. Simultaneously, this mechanism ensures that new record creation is triggered only when the dialogue content is sufficiently novel, while when highly similar content is detected, historical semantic continuity is achieved through target vector association. This dynamic balancing strategy maintains the conciseness of the knowledge base and ensures the priority reuse of the most relevant historical context through similarity ranking, thus improving the operational efficiency of the intelligent chat system while enhancing the coherence and semantic consistency of dialogue memory management. It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0119] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0120] According to another aspect of the embodiments of this application, a chat content processing apparatus based on multi-level classification and associative memory is also provided for implementing the above-described method. Please refer to... Figure 3 , Figure 3This is a structural block diagram of an optional chat content processing device based on multi-level classification and associative memory provided in an embodiment of this application, such as... Figure 3 As shown, the chat content processing device 300 based on multi-level classification and associative memory may include:

[0121] The acquisition module 301 is used to receive the dialogue content input by the user and convert the dialogue content into structured text data;

[0122] The classification module 302 is used to classify text data into multiple levels according to a preset classification standard using a large model to obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification standard is used to define multiple main category labels and sub-category labels under each main category label.

[0123] The query module 303 is used to extract associated memory data from the vector library using a large model based on the classification labels and preset association strength types. The association strength type is used to indicate the semantic association strength between the dialogue content and the classification labels. The vector library is dynamically updated based on the text vectors corresponding to the text data using a large model.

[0124] Display module 304 is used to sort the associated memory data according to the semantic association strength indicated by the association strength type, and to display the sorted associated memory data to the user through a visual interface, wherein data with high association strength is displayed first.

[0125] It should be noted that the acquisition module 301 in this embodiment can be used to perform the above step S201, the classification module 302 in this embodiment can be used to perform the above step S202, the query module 303 in this embodiment can be used to perform the above step S203, and the display module 304 in this embodiment can be used to perform the above step S204.

[0126] Regarding the chat content processing device based on multi-level classification and associative memory in this embodiment, the specific manner in which the acquisition module 301, classification module 302, query module 303, and display module 304 execute the chat content processing method based on multi-level classification and associative memory has been described in detail in the embodiments of the chat content processing method based on multi-level classification and associative memory, and will not be elaborated here.

[0127] It is understood that the technical solution provided in this embodiment, the various modules in the chat content processing device based on multi-level classification and associative memory, through steps such as structured processing of dialogue content, multi-level classification, associative memory extraction and sorting display, achieve precise management and efficient utilization of chat content, effectively solve the problem of not being able to accurately extract relevant chat content in the prior art, and significantly improve the service quality and user experience of the intelligent chat system.

[0128] In addition to the modules described above, the apparatus in this embodiment may also include modules that execute any method in any of the aforementioned embodiments of the chat content processing method based on multi-level classification and associative memory.

[0129] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can operate in ways such as... Figure 1 The method shown can be implemented in either software or hardware within a hardware environment, where the hardware environment includes a network environment.

[0130] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described chat content processing method based on multi-level classification and associative memory is also provided. The electronic device may be a server, a terminal, or a combination thereof.

[0131] According to another embodiment of this application, an electronic device is also provided; please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a structural block diagram of an optional electronic device provided in an embodiment of this application, such as... Figure 4 As shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.

[0132] Memory 1503 is used to store computer programs;

[0133] When processor 1501 executes the program stored in memory 1503, it performs the following steps:

[0134] Step S201: Receive the dialogue content input by the user and convert the dialogue content into structured text data;

[0135] Step S202: Use the large model to classify the text data into multiple levels according to the preset classification criteria to obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification criteria are used to define multiple main category labels and sub-category labels under each main category label.

[0136] Step S203: Using a large model, association memory data is extracted from the vector library based on the classification labels and the preset association strength type. The association strength type is used to indicate the semantic association strength between the dialogue content and the classification labels. The vector library is dynamically updated based on the text vectors corresponding to the text data using the large model.

[0137] Step S204: Sort the association memory data according to the semantic association strength indicated by the association strength type, and display the sorted association memory data to the user through a visualization interface, wherein data with high association strength is displayed first.

[0138] It is understood that the technical solution provided in this embodiment, through the processor of the electronic device, achieves precise management and efficient utilization of chat content by performing steps such as structured processing of dialogue content, multi-level classification, associative memory extraction and sorting display. This effectively solves the problem of the inability to accurately extract relevant chat content in the prior art, and significantly improves the service quality and user experience of the intelligent chat system.

[0139] Optionally, in this embodiment, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic device and other devices.

[0140] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0141] The processor mentioned above can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0142] This application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the method steps of the above method embodiments when it runs.

[0143] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.

[0144] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0145] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.

[0146] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0147] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or the indirect coupling or communication connection of units or modules may be electrical or other forms.

[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the solution provided in this embodiment, depending on actual needs.

[0149] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0150] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A chat content processing method based on multi-level classification and associative memory, applied to an intelligent chat system, characterized in that, include: The system receives dialogue content input by the user and transforms the dialogue content into structured text data by performing entity recognition, intent recognition, and sentiment analysis on the dialogue content. The text data is classified into multiple levels using a large model according to a preset classification standard to obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification standard is used to define multiple main category labels and sub-category labels under each main category label. The main category labels include at least one of the following: user basic information, health status, living habits, mental health, and family member information. The large model, based on the classification labels and preset association strength types, determines the association strength between the memory data stored in the vector library and the user's query intent, and extracts associated memory data from the vector library. The association strength types include three types: direct association, indirect association, and potential association, used to indicate the semantic association strength between the dialogue content and the classification labels. Whenever new dialogue content is generated, the vector library is dynamically updated using the large model based on the text vectors corresponding to the text data. This includes: using the large model to vectorize the text data to obtain corresponding text vectors; determining whether to generate a new summary or merge historical summaries based on the similarity matching results between the text vectors and text vectors in the vector library; and storing the summary, classification labels, and text vectors in the vector library. The associated memory data are sorted according to the semantic association strength indicated by the association strength type. Directly associated memory data receives the highest sorting priority, followed by indirectly associated memory data, and finally potentially associated memory data. The sorting rules are dynamically adjusted according to the specific query scenario and user needs. When indirectly associated or potentially associated memory data has higher value to the user, the sorting priority of the corresponding associated memory data is increased. The sorted associated memory data is then displayed to the user through a visual interface, with data with high association strength being displayed first.

2. The chat content processing method based on multi-level classification and associative memory according to claim 1, characterized in that, The text data is classified into multiple levels using a large model according to a preset classification standard to obtain classification labels, including: Using a large model, the text data is first classified according to the primary classification criteria in the classification standard to obtain multiple primary category labels, wherein the primary classification criteria are used to define the primary category labels; The text data is classified a second time using a large model according to the sub-classification criteria in the classification criteria, resulting in multiple sub-category labels. The sub-classification criteria are used to define the sub-category labels under each main category label.

3. The chat content processing method based on multi-level classification and associative memory according to claim 2, characterized in that, The main classification criteria include, but are not limited to, the following category tags: basic user information, health status, lifestyle habits, mental health, family member information, and others.

4. The chat content processing method based on multi-level classification and associative memory according to claim 1, characterized in that, The direct association indicates that there is a direct relationship between the dialogue content and the category label; the indirect association indicates that there is an indirect relationship between the dialogue content and the category label; and the potential association indicates that there is a potential relationship between the dialogue content and the category label. The semantic association strength of the three types in the association strength type is in the following order from high to low: direct association, indirect association, and potential association.

5. The chat content processing method based on multi-level classification and associative memory according to claim 1, characterized in that, Based on the similarity matching results between the text vector and text vectors in the vector library, a decision is made on whether to generate a new summary or merge historical summaries, including: By performing similarity matching between the text vector and all text vectors in the vector library, it is determined whether there is a target vector in the vector library that is associated with the text vector. If it is determined that there is no target vector associated with the text vector in the vector library, a new summary of the text data is generated through a large model; If it is determined that a target vector associated with the text vector exists in the vector library, the summary is updated based on the historical summary and text data corresponding to the target vector through the large model.

6. The chat content processing method based on multi-level classification and associative memory according to claim 5, characterized in that, Based on the similarity matching between the text vector and all text vectors in the vector library, determine whether there exists a target vector in the vector library associated with the text vector, including: The similarity between the text vector and each text vector in the vector library is calculated using a preset similarity calculation method. If the similarity is greater than or equal to a preset similarity threshold, then it is determined that there is a target vector in the vector library associated with the text vector, and the text vector with the highest similarity is selected as the target vector; otherwise, it is determined that there is no target vector in the vector library associated with the text vector.

7. A chat content processing device based on multi-level classification and associative memory, applied in an intelligent chat system, characterized in that, include: The acquisition module is used to receive dialogue content input by the user and convert the dialogue content into structured text data by performing entity recognition, intent recognition and sentiment analysis on the dialogue content. The classification module is used to classify the text data into multiple levels according to a preset classification standard using a large model to obtain classification labels. The classification labels are secondary classification labels that include main category labels and sub-category labels. The classification standard is used to define multiple main category labels and sub-category labels under each main category label. The main category labels include at least one of the following: user basic information, health status, living habits, mental health, and family member information. The query module is used to utilize the large model to determine the association strength between the memory data stored in the vector library and the user's query intent based on the classification labels and preset association strength types. It then extracts associated memory data from the vector library. The association strength types include three categories: direct association, indirect association, and potential association, used to indicate the semantic association strength between the dialogue content and the classification labels. Whenever new dialogue content is generated, the vector library is dynamically updated using the large model based on the text vectors corresponding to the text data. This includes: vectorizing the text data using the large model to obtain corresponding text vectors; determining whether to generate a new summary or merge historical summaries based on the similarity matching results between the text vectors and text vectors in the vector library; and storing the summary, classification labels, and text vectors in the vector library. The display module is used to sort the associated memory data according to the semantic association strength indicated by the association strength type. Directly associated memory data will receive the highest sorting priority, followed by indirectly associated memory data, and finally potentially associated memory data. At the same time, the sorting rules will be dynamically adjusted according to the specific query scenario and user needs. When indirectly associated or potentially associated memory data has higher value to the user, the sorting priority of the corresponding associated memory data will be increased, and the sorted associated memory data will be displayed to the user through a visual interface. Among the associated memory data, data with high association strength will be displayed first.

8. An electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein, The processor, the communication interface, and the memory communicate with each other via the communication bus, characterized in that... The memory is used to store computer programs; The processor is configured to execute the steps of the chat content processing method based on multi-level classification and associative memory as described in any one of claims 1 to 6 by running the computer program stored in the memory.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the steps of the chat content processing method based on multi-level classification and associative memory as described in any one of claims 1 to 6 when it runs.

Citation Information

Patent Citations

  • Dialogue processing method and device based on large model, dialogue method and device and electronic equipment

    CN118708688A

  • Systems and methods for hierarchical multi-label multi-class intent classification

    US20240054298A1

  • Systems and methods of determining genre information

    US8204883B1