Data processing method and device, storage medium and electronic equipment
By acquiring memory information from current and historical sessions, as well as target business knowledge, and using the target model to generate response information, the problem of models forgetting background details in intelligent customer service systems is solved, thereby improving the accuracy of problem solving and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing intelligent customer service systems often forget important background details, affecting the coherence of subsequent conversations and the accuracy of problem-solving.
By receiving user question information, the system obtains the first type of memory information of the current session, the second type of memory information of the historical sessions, and the target business knowledge information. It then uses the target model to perform comprehensive processing to generate response information.
It improves the accuracy of problem solving, enhances the personalization and satisfaction of the user experience, and avoids inconsistencies or inaccuracies caused by forgetting important details.
Smart Images

Figure CN121412359B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a data processing method and device, a storage medium and an electronic device. BACKGROUND
[0002] In the current intelligent customer service system technical field, large language models have become the core tool for improving interactive experience. In actual application scenarios, users expect accurate, timely and personalized solutions, while enterprises need efficient, stable and continuously learning and evolving intelligent customer service systems to handle large-scale concurrent user requests while ensuring service quality does not decline. However, the current industry technology shows some significant defects, and the model in the related technology is prone to "forgetting" important background details, affecting the coherence of subsequent conversations and the accuracy of problem solving. SUMMARY
[0003] The present application provides a data processing method, device, storage medium and electronic device to at least solve the technical problem of low model reply accuracy in related technologies.
[0004] The present application provides a data processing method, comprising: receiving problem information input by a target object; obtaining target data information according to the problem information, wherein the target data information includes first type memory information, second type memory information related to the problem information and target knowledge information related to the problem information, the first type memory information is obtained based on a current conversation corresponding to the problem information, the second type memory information is obtained based on historical conversations of the target object, and the target knowledge information is knowledge information corresponding to a target business related to the problem information; processing the target data information through a target model to obtain reply information corresponding to the problem information.
[0005] The present application provides a data processing device, comprising: a receiving unit configured to receive problem information input by a target object; a first obtaining unit configured to obtain target data information according to the problem information, wherein the target data information includes first type memory information, second type memory information related to the problem information and target knowledge information related to the problem information, the first type memory information is obtained based on a current conversation corresponding to the problem information, the second type memory information is obtained based on historical conversations of the target object, and the target knowledge information is knowledge information corresponding to a target business related to the problem information; and a first processing unit configured to process the target data information through a target model to obtain reply information corresponding to the problem information.
[0006] The application further provides an electronic device, comprising a memory for storing a computer program, and a processor for executing the computer program to implement the steps of any of the data processing methods.
[0007] The application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of any of the data processing methods.
[0008] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of any of the data processing methods.
[0009] According to the application, the first type of memory information is obtained according to the current conversation context with the user, the model can understand the background of the current question, the second type of memory information, i.e., the historical conversation record of the user, is used to supplement the question background, the business knowledge information related to the question, i.e., the target knowledge information, is retrieved, and the target model generates a comprehensive reply according to the above information, which not only considers the context of the instant conversation, but also combines the personalized preference of the historical conversation and the current business rule, so that the reply information is more comprehensive, the incoherence or inaccuracy caused by forgetting important details is avoided, and the technical effects of improving the accuracy of question solving and enhancing the personalization of user experience and the satisfaction are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0011] Figure 1 is a hardware structure block diagram of plug-in processing provided according to an embodiment of the application;
[0012] Figure 2 is a flowchart of a data processing method provided according to an embodiment of the application;
[0013] Figure 3 is a schematic diagram of a data processing method provided according to an embodiment of the application Figure 1 ;
[0014] Figure 4 is a schematic diagram of a data processing method provided according to an embodiment of the application Figure 2 ;
[0015] Figure 5 is a schematic diagram of a data processing method provided according to an embodiment of the application Figure 3 ;
[0016] Figure 6 This is a schematic diagram of the data processing method provided in the embodiments of this application. Figure 4 ;
[0017] Figure 7 This is a schematic diagram of a data processing apparatus provided according to an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] The specific application environment architecture or specific hardware architecture on which the execution of the device control method depends is described here.
[0022] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram illustrating the execution of system functions according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1more or less components than those shown, or configurations with different configurations and / or Figure 1
[0023] The memory 104 can be used to store computer programs, such as software programs of application software and modules, for example, a computer program corresponding to the execution method of the system function in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, and these remote memories can be connected to a server device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0024] The transmission device 106 is used to receive or send data via a network. The specific examples of the above network can include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, RF) module which is used to communicate with the Internet in a wireless manner.
[0025] The embodiments of the present application provide a data processing method, which is described in detail in combination with the execution flow of the data processing method.
[0026] In the embodiments, a data processing method is provided, Figure 2 is a flowchart of the data processing method provided by the embodiments of the present application, as Figure 2 shown, the data processing method comprises the following steps.
[0027] In step S201, question information input by a target object is received.
[0028] Optionally, the target object (e.g., a user or a customer) inputs question information through the intelligent customer service system. The intelligent customer service system can receive and process question information in various formats, including but not limited to text input, transcribed text after voice recognition, OCR recognized text in pictures or videos, etc. After receiving the question information, the real intention and emotional state of the user can be understood in depth through natural language processing techniques (such as syntax analysis, semantic understanding, sentiment analysis, etc.), and the question information can also be pre-processed, including removing redundant information, correcting spelling errors, converting colloquial expressions into standard formats, etc., to ensure that the question information is in the best state for the subsequent answer link and improve the efficiency of question identification and solution.
[0029] For ambiguous or incomplete questions, the intelligent customer service system can also help users express their needs or doubts more accurately and completely through a series of pre-set guiding questions or clarifying questions, and then send the optimized question information into the processing flow, avoiding ineffective answers or misunderstandings caused by question understanding bias.
[0030] In step S202, target data information is obtained according to the question information, wherein the target data information includes first type of memory information, second type of memory information related to the question information, and target knowledge information related to the question information, the first type of memory information is obtained based on a current session corresponding to the question information, the second type of memory information is obtained based on historical sessions of the target object, and the target knowledge information is knowledge information corresponding to a target business related to the question information.
[0031] Optionally, based on the received user question information, multi-level and multi-angle data information acquisition and integration are performed to build a more comprehensive and accurate question answering basis. The first type of memory information is obtained, which is the first-hand information based on the immediate context information of the current session, used to understand the immediate needs of the user and locate the problem background, which is crucial for generating coherent and targeted answers.
[0032] The second type of memory information is obtained, which is the historical session record of the user, reflecting the user's preferences, habits, historical problems and solutions, and any long-term interaction patterns. By searching the long-term memory of the user's historical sessions, the intelligent customer service system can extract the historical session records associated with the current question.
[0033] The target knowledge information refers to the business knowledge directly related to the user's question, including product manuals, service guidelines, frequently asked questions, and the latest rules. This information comes from an internal knowledge base and is authoritative and timely. The most relevant content can be retrieved from the pre-set knowledge base according to the keywords and topics of the question to provide accurate and professional answers. The acquisition of target knowledge information ensures the professionalism and accuracy of the model's answers and avoids customer service failures caused by knowledge lag or misunderstanding.
[0034] The target data information is constructed by the above three types of data, including the current context, historical background, and professional business knowledge.
[0035] In step S203, the target data information is processed by the target model to obtain the reply information corresponding to the question information.
[0036] Optionally, the target data information is processed by the target model, for example, the target model integrates the instant conversation context, user historical preferences, and business knowledge together for deep understanding and reasoning, and outputs specific reply information for the user's question.
[0037] Through the above steps, the limitations of traditional customer service can be overcome, such as information "forgetting", context fault, and knowledge base update delay, etc., to provide users with comprehensive, coherent, and highly professional service experience.
[0038] According to the current conversation context with the user, the first type of memory information is acquired, which can ensure that the model understands the background of the current question, and the second type of memory information, i.e., the user's historical conversation records, is used to supplement the question background and retrieve the business knowledge information related to the question, i.e., the target knowledge information. The target model generates a comprehensive reply based on the above information, considering not only the context of the instant conversation but also the personalized preferences of the historical conversation and the current business rules, so that the generated reply information is more comprehensive, avoiding the incoherence or inaccuracy caused by "forgetting" important details, thereby improving the accuracy of problem solving and enhancing the personalization of user experience and satisfaction.
[0039] Optionally, in the data processing method provided by the embodiments of the present application, the target data information is obtained according to the question information, including: determining the current conversation corresponding to the question information, and obtaining the first type of memory information based on the conversation information in the current conversation; matching the question information with the memory information in the first database to obtain the second type of memory information; and matching the question information with the knowledge information in the second database to obtain the target knowledge information.
[0040] In an optional embodiment, it is determined that the question corresponds to the current session, and first type memory information (also referred to as short-term memory information) is obtained based on the dialogue information in the current session, that is, the short-term memory information is obtained based on the instant dialogue content with which the user interacts. The short-term memory information is directly derived from the recent communication between the two parties and is the basis for understanding and responding to the user's demand.
[0041] The question information is compared and matched with historical memory information stored in the first database (long-term memory database), and second type memory information (also referred to as long-term memory information) related to the question information is extracted. The long-term memory information can include information such as the user's past service requests and their solutions. Through the combination of the short-term memory information, more personalized and coherent replies can be provided.
[0042] Based on the keywords and topics of the question, professional knowledge and rules covered in the second database (knowledge information database) are queried, and target knowledge information is obtained. This part of data is derived from the enterprise internal knowledge base and includes product details, operation guides, rule clauses, etc.
[0043] Through the above steps, the embodiment can deeply mine the background of the question, comprehensively utilize the details of the instant session, the user's historical behavior and the special knowledge information, construct a comprehensive and personalized information set, and then generate a reply with both depth and precision, thereby improving the quality of interaction and user satisfaction, effectively solving the problem that the model easily ignores important background information, and ensuring the coherence of the dialogue and the efficiency of the problem solving.
[0044] Optionally, in the data processing method provided in the embodiment of the application, based on the dialogue information in the current session, the first type memory information includes: obtaining a target dialogue round number that the first type memory of the target model can accommodate; and selecting the first type memory information from the current session according to the target dialogue round number.
[0045] In an optional embodiment, the target dialogue round number that the first type memory (i.e., short-term memory) of the target model can accommodate is obtained. After the target dialogue round number is determined, the most recent several rounds of dialogue from the dialogue information in the current session are selected as the first type memory information, that is, the short-term memory. This information set integrates the user's latest inquiry and the instant response of the system and is an important basis for the model to perform real-time reasoning and generate a reply.
[0046] Through the above steps, the target model can focus on the most recent context information, avoiding reasoning delay and resource waste caused by processing too many historical dialogues. At the same time, this ensures that the model can analyze and predict based on the most relevant and latest data when generating a reply, thereby improving the relevance and accuracy of the reply.
[0047] Optionally, in the data processing method provided in the embodiments of the present application, the method further comprises: determining a current session corresponding to the question information; obtaining dialogue information in the current session, and updating the memory information in the first database according to the dialogue information.
[0048] In an optional embodiment, the specific dialogue information in the current session is extracted, including the user's question, the system's previous response and any related tool invocation result. The dialogue information in the current session is compared with and updated to the long-term memory information in the first database. It should be noted that the memory information in the first database can be updated in real time according to the dialogue information in the current session.
[0049] Through the above steps, the long-term memory database of each session can be updated and optimized in real time or efficiently while processing each session, ensuring the timeliness, accuracy and integrity of the memory information, and further improving the response speed, personalized service ability and user satisfaction of the intelligent customer service.
[0050] Optionally, in the data processing method provided in the embodiments of the present application, the matching of the question information and the memory information in the first database to obtain the second type of memory information comprises: performing vectorization processing on the question information to obtain a first vector; calculating the similarity between the first vector and a second vector corresponding to the second type of memory information in the first database to obtain a target similarity; and obtaining the second type of memory information according to the target similarity.
[0051] In an optional embodiment, the question information can be converted into a first vector through a pre-trained Embedding model. This process converts text information into a numerical vector in a high-dimensional space, capturing the semantic features and contextual meanings of the text. All stored second vectors (i.e., vectors corresponding to long-term memory information) in the first database are retrieved.
[0052] For example, a semantic similarity calculation method (such as cosine similarity, Euclidean distance, etc.) is used to compare the first vector (vector representation of the question information) with each second vector (vector representation of historical memory information) to obtain multiple target similarity values. Based on the calculated target similarity, the most relevant long-term memory information to the current question is selected as the second type of memory information. For example, the memory fragment with the highest similarity score is selected, or a threshold is set, and the memory information exceeding the threshold is included in the second type of memory information. The acquisition of the second type of memory information helps to review the user's historical behavior patterns and preferences, so as to reflect individualization and contextualization in the reply.
[0053] Through the above steps, the information most relevant to the current problem can be effectively screened from a large amount of historical data, unnecessary computing overhead is avoided, and meanwhile, the accuracy and personalization of the reply are ensured, and the user experience is improved.
[0054] Optionally, in the data processing method provided in the embodiments of the present application, the target data information is processed through the target model to obtain the reply information corresponding to the question information, including: encoding the target data information through the embedding layer of the target model to obtain a third vector; calculating the third vector through the multi-head self-attention mechanism of the target model to obtain weight information; weighting and aggregating the third vector according to the weight information to obtain a fourth vector; processing the fourth vector through the decoding layer of the target model to obtain the reply information.
[0055] In an optional embodiment, the target data information (including user questions, short-term memory, long-term memory, and knowledge base information, etc.) is input into the embedding layer of the target model for encoding. This process converts text information into vectors, which facilitates model understanding and processing. The embedding layer generates a vector representation for each word according to its semantics and contextual relationships, i.e., the third vector.
[0056] Then, the third vector is passed to the multi-head self-attention (Multi-head Self-Attention) layer of the model. Through multi-head self-attention, the model can focus on different parts of the input sequence and calculate the corresponding weights based on the relationships between words, i.e., the importance of each word in the current context.
[0057] Using the calculated weights, the third vector is weighted and aggregated to generate a new vector representation, i.e., the fourth vector. Finally, the fourth vector is input into the decoding layer of the target model, and the decoding layer generates a specific text reply based on the fourth vector, i.e., the reply information corresponding to the question information.
[0058] By utilizing the generative capability of the model and combining the weight information, it is ensured that the generated reply content is both accurate and consistent with the context of the current conversation.
[0059] Optionally, in the data processing method provided in the embodiments of the present application, the target model corresponds to the first type of memory that can accommodate the target number of dialogue rounds, including: obtaining the maximum context length corresponding to the target model and obtaining the reference context length; obtaining a preset step-up amplitude and the maximum number of dialogue rounds of the target model in a conversation; calculating the target number of dialogue rounds according to the maximum context length, the reference context length, the step-up amplitude, and the maximum number of dialogue rounds.
[0060] In an optional embodiment, the maximum context length of the target model is obtained, and a standard length value (i.e., a reference context length) is preset as a reference for measuring the relative level of the context processing capability of the target model, facilitating subsequent calculation.
[0061] To adapt to the capabilities of different models, an ascending amplitude is preset for dynamically adjusting the number of dialogue rounds in short-term memory, ensuring a balance between model load and context demand. To avoid processing delay and resource waste caused by excessively long context, an upper limit of the number of dialogue rounds (i.e., the maximum number of dialogue rounds mentioned above) is preset to ensure the smoothness of real-time interaction.
[0062] Based on the above parameters, the number of dialogue rounds to be included in short-term memory is calculated, which takes into account the context length supported by the model, the reference length, the step amplitude, and the maximum allowed rounds, ensuring that short-term memory can carry enough information without exceeding the processing capability of the model.
[0063] In an optional embodiment, the context length parameter supported by the target model is automatically read during deployment, and based on this, the number of multi-round dialogue rounds that can be accommodated by short-term memory is dynamically calculated (e.g., the number of dialogue rounds in short-term memory is calculated based on the context length supported by the target model, the reference context length, the step amplitude, and the maximum allowed rounds) , to achieve the optimal balance between memory depth and reasoning performance.
[0064] The specific calculation formula is as follows:
[0065]
[0066] wherein, : the context length supported by the model, the larger the value, the longer the input supported by the model, and the number of dialogue rounds selected by short-term memory can also be increased accordingly; : the standard reference context length, used to measure the step scale of the context processing capability of LLM; : the upward rounding symbol; : the ascending amplitude of each step (e.g., 2); : the maximum number of multi-round dialogues (e.g., set to 5 to prevent excessive context from causing time consumption and memory divergence).
[0067] Calculation example: the context length supported by model 1 is 10, calculated by the formula as follows: : the context length supported by model 2 is 20, calculated by the formula as follows: . .
[0068] The "step-by-step growth" design can: make full use of the contextual ability of the model, automatically extend the dialogue window in models with large context; avoid excessive stacking of context, automatically shorten the dialogue depth on models with limited computing power; dynamically balance response speed and semantic coherence, improve model inference efficiency and reduce token consumption.
[0069] Through the above steps, in different model deployment scenarios, the management of short-term memory can both utilize the most relevant context information for accurate inference and avoid processing delays and excessive resource consumption, thereby improving the efficiency, accuracy, and user experience of real-time question answering.
[0070] Optionally, in the data processing method provided by the embodiments of the present application, after obtaining the target data information according to the question information, the method further comprises: obtaining a target compression rate; compressing the target data information according to the target compression rate to obtain compressed target data information, wherein the compressed target data information is processed by the target model to obtain the reply information corresponding to the question information.
[0071] In an optional embodiment, the target compression rate can be a preset value or a dynamically calculated value according to the current model load. The compression rate determines the degree of compression of the target data information to reduce unnecessary information redundancy while maintaining key context and semantic clues.
[0072] The obtained target data information is compressed according to the calculated target compression rate. For example, intelligent summarization of text content, keyword extraction, or semantic vector simplification, etc., aims to reduce the amount of information input into the model, thereby speeding up the inference process, reducing computational resource consumption, while ensuring the accuracy and coherence of the output results are not affected. After compression processing, the obtained original data information is simplified into compressed target data information.
[0073] In an optional embodiment, based on the target compression rate, the target data information related to the question information is intelligently compressed. During the compression process, key information is retained, while redundant or secondary details are simplified, thereby reducing the input amount of the model processing, reducing computational resource consumption, and speeding up the inference speed. The compressed target data information is input into the target model for processing, and the target model generates an effective response to the question information, i.e., the reply information, based on the compressed data information. Through data compression, the volume is reduced while maintaining the key information, which is conducive to improving the inference speed and resource usage efficiency of the model, especially in scenarios of processing large-scale data or high-concurrency requests, which can significantly improve the service quality and response speed.
[0074] Optionally, in the data processing method provided in the embodiments of the present application, the matching of the question information and the memory information in the first database obtains the second type of memory information, which includes: matching the question information and the memory information in the vector database of the first database to obtain first initial memory information; matching the question information and the memory information in the graph database of the first database to obtain second initial memory information; and obtaining the second type of memory information according to the first initial memory information and the second initial memory information.
[0075] In an optional embodiment, the following steps can be used to obtain the second type of memory information described above: encoding the user's question information through an embedding model to convert it into a semantic vector, and then matching in the vector database layer of the first database. By calculating the cosine similarity or Euclidean distance between the user question vector and the memory vector stored in the database, the most similar memory information to the question information is determined, and the first initial memory information is obtained.
[0076] At the same time, the user's question is matched in the graph database layer. For example, the entities and relationships involved in the question are identified, and then the related entities and edges in the graph database are searched, such as querying the adjacent nodes of the entity, path retrieval based on entity relationship, etc., to obtain the second initial memory information related to the question.
[0077] After obtaining the first initial memory information and the second initial memory information, comprehensive analysis and memory fusion can be performed to obtain the second type of memory information. For example, the information obtained from the vector database and the graph database is semantically integrated, redundant information is removed, key information is extracted, and a second type of memory information that is related to the user's question, rich in content and clear in structure is generated.
[0078] Through this double-layer database matching mechanism, the data processing method of the embodiments of the present application can effectively retrieve the most relevant historical interaction information to the user's question, not only improving the response speed and accuracy of the intelligent customer service system, but also enhancing its understanding ability for complex and diversified questions, providing a more comprehensive and personalized customer service experience.
[0079] Optionally, in the data processing method provided in the embodiments of the present application, obtaining the target compression rate includes: obtaining a current load value of the target model and a total number of characters corresponding to the target data information; calculating a weighted value according to the current load value and the total number of characters; and calculating the target compression rate according to the weighted value and a preset compression rate range corresponding to the target data information.
[0080] In an optional embodiment, the target compression rate can be dynamically calculated by the following steps: monitoring the computational resource usage of the target model (such as CPU, GPU occupancy or memory usage) to obtain its load level when processing the current task. Determine the text length of the target data information, the total number of characters reflects the size of the information to be processed, which is another important basis for formulating the compression strategy. Longer information may require a higher degree of compression to reduce the occupation of model resources.
[0081] Combined with the load value and the information length, a weighted value is calculated through a weighted function, which reflects the urgency and necessity of current data processing. The weighted value combines the real-time pressure of the model and the size of the data information, providing a quantitative reference for subsequent compression rate calculation.
[0082] According to the calculated weighted value and the preset compression rate range (for example, the compression rate range of short-term memory is 0.4-0.6, and the compression rate range of long-term memory is 0.3-0.5), the final target compression rate is determined. The setting of the target compression rate not only considers the real-time load of the model, but also ensures that the compressed information can still meet the needs of inference, maintain the coherence and integrity of the context.
[0083] Through the above steps, the real-time status of the model and the characteristics of the data information can be intelligently responded to, and the compression strategy can be dynamically adjusted to ensure efficient and accurate information processing in different scenarios, improving the response speed and service quality of the enterprise-level intelligent customer service system.
[0084] In an optional embodiment, the dynamic adjustment formula of the compression rate is as follows:
[0085]
[0086] Wherein: represents the compression rate of the selected data type (short-term, long-term, knowledgebase); : the lower and upper limits of the compression rate range of this type of data; : the load of the current model (for example, CPU / GPU usage); : real-time token consumption, representing the number of tokens input before compression in the current conversation. : a dynamic weighted function according to the model load and token consumption, which can be specifically represented as:
[0087]
[0088] Wherein, is the weight coefficient, , The maximum value of the load and the token consumption, respectively.
[0089] When the model has a high computational load and a large token consumption, a compression ratio close to the maximum compression ratio is selected (i.e., a high compression ratio is selected). At this time, redundant content is reduced to improve response speed and system stability. When the model has a low load and a small token consumption, the compression ratio is appropriately reduced to ensure the reasoning accuracy of the model and the full use of the context. In the case of moderate load and token consumption, the system sets the compression ratio to the median value in the compression range to balance accuracy and efficiency.
[0090] By dynamically adjusting the compression strategy of the memory, the system can continuously provide a smooth and accurate dialogue experience under different loads. It also provides flexible adaptation capabilities for deployments of different sizes and different hardware configurations.
[0091] Optionally, in the data processing method provided by the embodiments of the present application, for the first type of memory information in the target data information, obtaining the target compression ratio includes: obtaining the current load value of the target model and the total number of characters corresponding to the first type of memory information; calculating the weighted value corresponding to the first type of memory information according to the current load value and the total number of characters corresponding to the first type of memory information; and calculating the target compression ratio corresponding to the first type of memory information according to the weighted value corresponding to the first type of memory information and the preset compression ratio range corresponding to the first type of memory information.
[0092] In an optional embodiment, for the first type of memory information (i.e., short-term memory information) in the target data information, obtaining the target compression ratio includes: monitoring the real-time usage of the computing resources of the target model in real time, including but not limited to CPU usage, GPU occupancy, memory consumption, etc., and counting the text length (i.e., token number) of the short-term memory information. The total number of characters reflects the size of this part of information and is an important basis for formulating the compression strategy. Combined with the real-time pressure of the model and the size of the short-term memory information, a quantitative index, i.e., a weighted value, is calculated through a weighting function. This index is used to balance the model performance and information processing efficiency, and to ensure that the compression strategy can both relieve the computational pressure and retain key context information. Based on the calculated weighted value and the preset compression ratio range (such as between 0.4 and 0.6) for the short-term memory information, the final target compression ratio is determined.
[0093] It should be noted that different compression rate ranges can be set according to different data types in the target data information. Different compression rate intervals are set for short-term memory, long-term memory and knowledge base, and are adjusted in real time according to the running state. Short-term memory: compression rate range is 0.4-0.6; long-term memory: compression rate range is 0.3-0.5; knowledge base: compression rate range is 0.3-0.6. When setting the compression rate range, the importance of information, the coherence of context and the accuracy of reasoning should be considered comprehensively to achieve effective management of resources and improvement of user experience. In addition, these compression rate ranges can be dynamically adjusted according to specific business scenarios and model processing capacity to ensure the best compression effect under different conditions.
[0094] The data processing method provided by the embodiments of the present application can intelligently adapt to different data types and business needs, effectively reduce the input data amount of the model, improve processing efficiency, and at the same time ensure that the information after data compression is still sufficient to support accurate decision-making and high-quality user interaction.
[0095] Optionally, in the data processing method provided by the embodiments of the present application, the target data information is compressed according to the target compression rate to obtain compressed target data information, comprising: calculating the matching degree of the target data information and the problem information to obtain the target value of each data in the target data information; and compressing the target data information according to the target value and the target compression rate to obtain the compressed target data information.
[0096] In an optional embodiment, the relevance of each part of the target data information to the current problem information is evaluated by natural language processing technology (such as semantic similarity algorithm or machine learning model). Information with high matching degree is more critical to solving the problem, while information with low matching degree may be considered as redundant or irrelevant information.
[0097] Based on the above matching degree calculation, a target value index is assigned to each piece of data information, which reflects the contribution of the data information to solving the current problem. Data with high target value should be retained first, while data with low target value can be considered for compression or elimination. Combined with the target value of each data and the target compression rate calculated in advance, a compression strategy is developed. For example, all data with a target value higher than a certain threshold are retained, and the remaining data is compressed until the target compression rate is met. This strategy can not only ensure the integrity of key information, but also effectively reduce the overall data amount and improve processing speed.
[0098] Through compression processing, the target data information that retains key details and reduces information amount is finally obtained, which provides optimized input for the subsequent reasoning process, helps to improve response speed and resource utilization efficiency, and at the same time ensures the accuracy of the answer and the coherence of the context.
[0099] Optionally, in the data processing method provided by the embodiments of the present application, the target data information is compressed according to the target value and the target compression rate to obtain compressed target data information, which comprises: sorting each data in the target data information according to the target value to obtain sorted data information; determining a retention ratio according to the target compression rate; and screening the sorted data information according to the retention ratio to obtain the compressed target data information.
[0100] In an optional embodiment, the importance or relevance of each data in the target data information to the solution of the problem, i.e. the target value, is calculated. Then, all data are sorted in descending order according to these value indicators to form a sorted data information list. This sorting process ensures that high-value data are processed and retained first.
[0101] The target compression rate determines the proportion of the amount of data information after final compression to the original data. For example, if the target compression rate is 0.5, the top 50% of the sorted data information is retained, and so on. According to the determined retention ratio, the corresponding proportion of data is screened from the sorted data information list, and these data constitute the compressed target data information. In the screening process, high-value data are retained, while low-value data are simplified, so as to achieve the purpose of reducing data volume without losing key information.
[0102] Through the above steps, the data processing method provided by the embodiments of the present application not only can effectively compress data and reduce the burden of model processing, but also can ensure that the compressed data information can still fully support the solution of the problem and maintain the accuracy of the answer and the coherence of the context.
[0103] In an optional embodiment, a self-attention mechanism can also be used to compress the target data information according to the target compression rate. First, the target data information is preprocessed, including removing stop words, punctuation marks, formatting, etc., to ensure the purity and readability of the information. A pre-trained self-attention model, such as part of the Transformer model, is used to encode the preprocessed data information. The self-attention mechanism can help the model identify the most important parts of the text, laying the foundation for subsequent compression processing. The weights of each word or phrase in the target data information are calculated by the self-attention model, which reflect the contribution of each part to the overall information. According to the set target compression rate, the percentage of the final retained information is calculated. For example, if the target compression rate is 0.3, 30% of the highest weight information is retained. Based on the calculated self-attention weights, the parts with weights higher than the threshold are selected, and the selected high-weight information is recombined to generate the compressed data information.
[0104] Optionally, in the data processing method provided in the embodiments of the present application, the method further comprises: determining a current session corresponding to the question information; judging whether a dialogue round of the current session reaches a target threshold; if the dialogue round of the current session reaches the target threshold, triggering a second type of memory updating task, wherein the target threshold is less than a target dialogue round number that the first type of memory corresponding to the target model can accommodate, and the second type of memory updating task is used for performing an updating operation on the memory information in the first database according to the dialogue information in the current session.
[0105] In an optional embodiment, after the system receives the question information of the user, the session to which the question information belongs, i.e., the current interactive process, is first determined. A target threshold is set as an updating trigger point of the dialogue round, for example, when the dialogue round reaches 10 rounds, the long-term memory updating is prepared to be performed. The setting of the threshold is based on the short-term memory processing capacity of the model and the business requirement, which can timely update the long-term memory information while ensuring the coherence of the dialogue, and avoid outdated or missing information.
[0106] In an optional embodiment, the target threshold is cross-redundant design: the updating interval is slightly shorter than the short-term memory capacity (i.e., the target threshold is less than the target dialogue round number that the first type of memory corresponding to the target model can accommodate).
[0107] In an optional embodiment, in order to reduce the computing pressure in the high concurrency scenario, the long-term memory no longer uses the real-time updating mode, but triggers the updating operation periodically through the fixed round interval strategy. In this mechanism, the updating interval can be set according to the short-term memory window of the specific model, so that the long-term memory updating and the short-term memory form a reasonable cross. For example, for the target model, the short-term memory capacity is about 15 rounds of dialogue, and the updating interval of the long-term memory can be set to 10 rounds. The updating interval is slightly shorter than the short-term memory capacity (10<15), which can ensure that the relevant information is written into the long-term memory before the short-term memory is covered by the new content, so as to avoid memory loss or information gap.
[0108] When it is detected that the dialogue round of the current session reaches or exceeds the target threshold, the second type of memory updating task is immediately started. The target of the task is to update the dialogue information accumulated in the current session, especially the key facts related to the user behavior, preference or business, to the first database, as part of the long-term memory.
[0109] It should be noted that when the short-term memory and the long-term memory are contradictory, the target model is automatically reminded to follow the short-term memory. This mechanism ensures that the system maintains long-term consistency while giving priority to reflecting the latest context changes, avoiding reply errors caused by outdated memory.
[0110] By setting a target threshold for the dialogue round to trigger long-term memory updates, the overconsumption of computing resources caused by real-time updating of long-term memory in high-concurrency scenarios is avoided. This strategy reduces the amount of data processing by the model during real-time inference, significantly improving the response speed and stability of the system to user requests. The fixed round interval update mechanism ensures the continuous updating of long-term memory and the accurate storage of information. When the conversation reaches the preset round threshold, key facts and user preferences are automatically written into the long-term memory database, avoiding outdated or missing information, thereby improving the timeliness and accuracy of the intelligent customer service system's memory content.
[0111] Optionally, in the data processing method provided by the embodiments of the present application, after triggering the second type of memory update task, the method further includes: detecting that the target model is in a first state, and adding the second type of memory update task to a task queue; after adding the second type of memory update task to the task queue, recording dialogue information in the current conversation to a first database, and setting a target identifier for the dialogue information, wherein the target identifier is used to represent that the dialogue information is candidate memory information to be processed; and executing tasks in the task queue according to the task order in the task queue.
[0112] In an optional embodiment, after triggering the second type of memory update task, the current state of the target model is first detected to determine whether it is in a first state that is not suitable for long-term memory updating. The first state can refer to situations such as high load of the model, resource shortage, and ongoing critical conversations. When it is detected that the target model is in the first state, the second type of memory update task is added to a special task queue. The task queue organizes and schedules tasks according to certain rules (such as first-come-first-served, etc.), ensuring that each update task can be processed at the appropriate time.
[0113] In an optional embodiment, when it is detected that the current conversation reaches the preset target threshold, the triggered long-term memory update task is not immediately executed, but is first added to a task queue. The task queue acts as a buffer, allowing these update tasks to be processed collectively at the appropriate time, rather than being executed immediately during high concurrency, thereby avoiding the additional computational pressure caused by real-time updating. The scheduling of the task queue follows certain rules to ensure the priority and efficiency of the tasks. According to the task order in the task queue, the first-in-first-out principle can be used, or dynamic scheduling can be performed based on comprehensive indicators such as the urgency of the task, the frequency of invocation, and the time interval. In this way, the system can prioritize tasks that are more urgent, frequently invoked, or have been waiting for a longer time, ensuring that the updating of long-term memory is both efficient and reasonable.
[0114] In an optional embodiment, after detecting that the dialogue turn reaches the preset target threshold, all dialogue information in the current session is recorded into the first database. This step ensures that all dialogue content is retained for subsequent processing and analysis.
[0115] For each item of dialogue information recorded in the first database, a target identifier (for example, flag = 1) is set. The target identifier marks these dialogue information as candidate memory information to be processed, indicating that they have not been analyzed and confirmed and need to be checked and updated in the next stage.
[0116] For example, in the subsequent idle stage or according to the preset update interval, the dialogue information with the target identifier in the first database is scanned to determine whether to add to the long-term memory, update the existing information, or delete the conflicting content, ensuring the accuracy and timeliness of the long-term memory.
[0117] In a high-concurrency environment, the task execution order in the task queue can be dynamically adjusted according to the real-time load. For example, if the computing resources are tight, the execution of some non-urgent tasks can be delayed, and the update tasks that directly affect the user service are prioritized. This flexible task scheduling strategy can maintain good response speed and user satisfaction in the case of limited resources.
[0118] Through the above strategy, the data processing method of the embodiment of the present application can not only effectively manage the update of the long-term memory and avoid the computing burden that the real-time update may bring, but also ensure that the execution of the update task is orderly and efficient through the scheduling mechanism of the task queue.
[0119] Through the target identifier mechanism, all dialogue information to be processed is properly handled, avoiding information omission or improper processing. At the same time, this mechanism also allows the system to temporarily suspend the long-term memory update work in a high-concurrency environment and continue processing when the system load is reduced, thereby balancing the real-time service quality and resource consumption of background data analysis, improving the overall system performance and user experience.
[0120] Optionally, in the data processing method provided by the embodiment of the present application, after triggering the second type of memory update task, the method further includes: detecting that the target model is in the second state, and executing the second type of memory update task.
[0121] In an optional embodiment, after triggering the second type of memory update task, the current state of the target model is first detected to determine whether it is in a second state suitable for long-term memory update. It should be noted that the second state can be a low-load state: the CPU and GPU usage of the target model is at a low level, ensuring that the update process will not significantly affect the real-time processing of user requests. Idle time: when the user request volume is relatively small, such as at night or off-peak hours, selecting to perform long-term memory update can effectively avoid resource competition with real-time services. Memory resources are abundant: sufficient memory space is ensured for processing memory update tasks to avoid update failure or delay due to insufficient resources. High stability: the intelligent customer service platform runs stably and there is no ongoing important maintenance or major update to prevent any potential interference or data loss risk.
[0122] When the target model meets the conditions of the above-mentioned second state, the execution of the second type of memory update task is started, the dialogue information is deeply analyzed to identify entities, relationships and key facts, and then compared and integrated with existing memory information in the first database to ensure that the update of long-term memory is both accurate and efficient.
[0123] By executing the second type of memory update task, not only the accuracy and timeliness of long-term memory can be ensured, but also the demand for real-time service and background maintenance can be effectively balanced. Through intelligent scheduling, complex and time-consuming memory processing can be performed at times when resources are sufficient, thereby optimizing the overall performance and user experience of the entire intelligent customer service system.
[0124] Optionally, in the data processing method provided in the embodiments of the present application, after adding the second type of memory update task to the task queue, the method further comprises: calculating the priority of the second type of memory update task to obtain a target priority; and determining whether to adjust the execution order of the second type of memory update task according to the target priority.
[0125] In an optional embodiment, in order to improve the flexibility of the second type of memory (i.e. long-term memory), the second type of memory update task is calculated according to preset priority evaluation indicators, including but not limited to urgency (U), call frequency (C) and time interval (T). These indicators reflect the urgency, practicality and timeliness of the task, and the target priority of the second type of memory update task is calculated by urgency (U), call frequency (C) and time interval (T). According to the calculated target priority, the tasks in the task queue are reordered. High-priority tasks will be placed at the front of the queue to ensure that they can be executed first. This priority-based sorting strategy can prioritize long-term memory update tasks that have a greater impact on user service and business operation in the case of limited resources, improving resource utilization efficiency and system response speed.
[0126] During task execution, the system continuously monitors changes in resource status and task queue, dynamically adjusting the weight coefficients in the calculation formula of target priority to adapt to different business scenarios and system states. For example, during high concurrency, the system may increase the weight of urgency (U) to ensure that tasks that directly affect service quality and user experience are prioritized for execution; while the system is idle, it may focus more on call frequency (C) and time interval (T) to optimize the maintenance and updating of long-term memory.
[0127] Through the above steps, the data processing method of the embodiments of the present application can intelligently manage the updating tasks of long-term memory, ensuring timely processing of key information and improving the overall operation efficiency of the system and the rationality of resource utilization.
[0128] Optionally, in the data processing method provided by the embodiments of the present application, the target priority of the second type of memory updating task is calculated, including: determining the urgency of the current session according to the type of the current session; calculating the practical value of the second type of memory information according to the number of references to the second type of memory information in the reply information output by the target model; calculating the time difference between the time when the second type of memory updating task is added to the task queue and the current time; and calculating the target priority according to the urgency, the practical value and the time difference.
[0129] In an optional embodiment, the urgency (U) of the current session is determined according to the type of the current session. This index reflects the importance and urgency of the session, for example, business interruption, complaint type problem is classified as high urgency (U=2), account inquiry, after-sales consultation is classified as medium urgency (U=1), and regular question and answer, general consultation is classified as low urgency (U=0). The determination of the urgency can be realized by business rules, user labels or artificial intelligence model analysis to adapt to different business scenarios.
[0130] Secondly, the practical value (C) of the second type of memory information is calculated according to the number of references to the second type of memory information in the reply information output by the target model. This index reflects the frequency of use of long-term memory in multiple rounds of dialogue, and embodies the actual value of user habits and long-term memory. If the long-term memory is not hit, the number of uses remains unchanged, and the number of uses is increased by 1 for each hit memory.
[0131] Further, the time difference (T) between the time when the second type of memory updating task is added to the task queue and the current time is calculated, which is an index used to measure the length of time the task waits for execution, reflecting the timeliness and urgency of the task.
[0132] Finally, the target priority (P) is calculated by comprehensively considering the urgency (U), the practical value (C) and the time interval (T). In this way, the memory updating tasks can be intelligently sorted, ensuring that those tasks with high urgency, practical value and long waiting time are given priority, thereby improving resource utilization efficiency and system response speed while ensuring service quality.
[0133] In an optional embodiment, the priority evaluation index includes:
[0134] Urgency Level (U), representing the importance and urgency of the current user session, such as High (U=2): business interruption, complaint type problem, Medium (U=1): account inquiry, after-sales consultation, Low (U=0): regular question and answer, general consultation; In the initial stage, the urgency level can be scored by the user or the large model; In the later stage, a lightweight embedding classification model is fine-tuned based on the collected historical data to realize automatic and rapid discrimination.
[0135] Call Frequency (C), representing the usage frequency of long-term memory in multiple rounds of dialogue, reflecting the user's habits and the practical value of long-term memory:
[0136]
[0137] If the search for long-term memory does not hit (no result), the usage frequency remains unchanged. Each time the memory is hit, the usage frequency is increased by 1. The more frequently the user accesses certain information, the more important the long-term memory is, and the higher the updating priority is; on the contrary, the user who prefers to ask new questions can delay the updating of long-term memory.
[0138] Time Interval (T), representing the time difference between the task arrival time and the current time, measured in hours. The longer the time, the longer the memory has been queued, and the priority should be correspondingly improved.
[0139] To achieve efficient scheduling in a multi-task concurrent scenario, a task sorting mechanism based on a priority queue is adopted. The comprehensive priority of each long-term memory updating task is defined as:
[0140]
[0141] wherein: U: Urgency Level; C: Call Frequency; T: Time Interval; : Tunable hyperparameters; for example, the parameters are set as follows: .
[0142] Tasks are sorted by priority in descending order, and high-priority tasks are updated first. Through this method, long-term memory update tasks can be adaptively scheduled and dynamically allocated resources in different business scenarios. It should be noted that if a new dialogue is generated during queuing, the information in the corresponding queue is updated in a timely manner.
[0143] Optionally, in the data processing method provided in the embodiments of the present application, the method further includes: determining candidate memory information with the target identifier, and converting the candidate memory information into a target vector; retrieving a memory vector most similar to the target vector from the first database; and processing the memory vector according to the target vector.
[0144] In an optional embodiment, the dialogue information with the target identifier, i.e., the candidate memory information to be processed, is filtered out from the first database. In order to perform efficient retrieval in the vector database, a pre-trained embedding model can be used to convert the candidate memory information into a high-dimensional vector, i.e., a target vector. This process captures the semantic features of the information, enabling the database to quickly match based on semantic similarity. Based on the cosine similarity or Euclidean distance between vectors and other metrics, a memory vector most similar to the target vector is retrieved from the first database.
[0145] After retrieving the similar memory vector, the memory vector in the first database is processed according to the similarity and content difference between the target vector and the memory vector. For example, the memory vector is added, deleted, modified, etc. according to the target vector.
[0146] Optionally, in the data processing method provided in the embodiments of the present application, processing the memory vector according to the target vector includes: if the similarity between the target vector and the memory vector is lower than a preset threshold, storing the target vector into the first database; if the target vector is a supplement or update of the memory vector, updating the memory vector according to the target vector; and if the target vector and the memory vector are contradictory, deleting the memory vector and storing the target vector into the first database.
[0147] In an optional embodiment, the similarity between the target vector (the vector of the user's current session or new information) and the memory vector stored in the first database (the long-term memory database) is first calculated. For example, the similarity is obtained by calculating the cosine similarity or Euclidean distance between the two vectors.
[0148] A preset similarity threshold is set to determine whether the target vector and the memory vector are similar enough to be considered as the same concept or information. If the similarity between the target vector and any memory vector is lower than the threshold, indicating that they represent independent information or different perspectives, the target vector is stored as a new memory vector in the first database, expanding the information base of long-term memory.
[0149] If the similarity between the target vector and the memory vector is higher than the preset threshold, but the target vector contains additional information or supplements the original memory, the memory vector is updated to reflect a more complete or up-to-date situation. For example, the reconstruction of the vector or the modification of its metadata (such as updating the timestamp, information source, etc.) ensures the timeliness and accuracy of the memory vector.
[0150] On the basis of similarity judgment, if the target vector and the memory vector have conflicts in content, i.e. they represent mutually exclusive or non-existent facts at the same time, the memory vector is deleted and the target vector is stored in the first database.
[0151] In an optional embodiment, the processing of the target vector on the memory vector can be as follows:
[0152] ADD (add): If the target vector is completely new compared to the most similar memory vector, the information is added as a new memory to the database.
[0153] UPDATE (update): If the target vector contains new details or modifications to the original information compared to the most similar memory vector, the memory vector is updated to reflect the latest information.
[0154] DELETE (delete): If the target vector contradicts a certain memory vector, indicating that the original information is outdated, the original memory vector is deleted and the updated information is retained.
[0155] NOOP (null operation): If the target vector is completely identical to a certain memory vector, it is determined that there is no need to update, avoiding unnecessary operations and data redundancy.
[0156] Through the above steps, the data processing method of the embodiments of the present application can intelligently manage and optimize long-term memory dynamically, ensuring the accuracy and timeliness of the memory content, while avoiding data duplication and conflicts, improving the performance and user experience of the enterprise-level intelligent customer service system.
[0157] Optionally, in the data processing method provided in the embodiments of the present application, the method further comprises: determining candidate memory information with the target identifier, and performing entity and relationship identification on the candidate memory information to obtain a target entity and a target relationship; obtaining a memory graph corresponding to the target object from the first database; and processing the memory graph according to the target entity and the target relationship.
[0158] In an optional embodiment, the first database includes a graph database layer, and the user, product, order and the like are organized in the form of nodes and edges based on the entity and relationship knowledge graph. The entity and relationship identification is performed on the candidate memory information with the target identifier, and the user, product, order status and the like in the text are abstracted into nodes in the graph database, and the relationships between the entities such as purchase, inquiry and feedback are identified and abstracted into edges in the graph database. The memory graph related to the target entity is retrieved from the graph database layer of the first database. If the target entity already exists in the graph, the existing graph corresponding to the entity is obtained; if the target entity appears for the first time, a new memory graph is created based on the initialization rule of the graph database, which is used to record all future interactions and information related to the entity.
[0159] Optionally, in the data processing method provided in the embodiments of the present application, the processing of the memory graph according to the target entity and the target relationship comprises: if the memory graph does not exist in association with the target entity and the target relationship, the target entity and the target relationship are added to the memory graph; and if the memory graph exists in association with the target entity or the target relationship, the entities and relationships in the memory graph are updated according to the target entity and the target relationship.
[0160] In an optional embodiment, it is queried in the memory graph whether there is a record related to the target entity and the target relationship. If there is no any record related to the target entity and the target relationship in the memory graph, it means that the current entity or relationship is brand new, and in this case, the target entity and the target relationship are added to the memory graph as new nodes and edges to enrich the content of the knowledge graph and enhance its understanding and response ability to user demand.
[0161] If the associated record of the target entity or the target relationship is found in the memory graph, it is further judged whether the target entity and the target relationship are a supplement or an update to the existing record. For example, if the target relationship is the latest evaluation of a user on a product, and there is an early evaluation of the user on the same product in the memory graph, the existing relationship edge is updated to reflect the latest user feedback.
[0162] In an optional embodiment, although an associated record of the target entity or target relationship is found in the memory graph, the target entity and target relationship contradict the existing information in the memory graph, for example, if the user's evaluation of a product recorded in the memory graph is positive, but the target relationship indicates that the user's recent evaluation of the product is negative, it is contradictory. If the new target entity and target relationship are considered more accurate or more up-to-date, update the relevant entity or relationship in the memory graph to reflect the latest situation.
[0163] In an optional embodiment, according to the target entity and the target relationship, the processing of the memory graph includes: adding new nodes or edges: if the target entity or target relationship is completely new (i.e. the memory graph has no association with the target entity and the target relationship), add corresponding nodes or edges in the memory graph to expand the knowledge coverage of the graph. Update the existing edge: if the target relationship is associated with the existing edge information in the graph, update the state, timeliness or other metadata information of these edges to ensure the timeliness and accuracy of the memory graph. Delete outdated or conflicting edges: if the target memory contradicts the existing information in the graph, according to the preset conflict processing rules, the edges may be deleted or marked as outdated to maintain the logical consistency and information timeliness of the memory graph.
[0164] By converting entities such as users, products, order status, and their relationships such as purchase, inquiry, feedback, etc. into nodes and edges in a graph database, the user's historical behavior and preferences can be more accurately understood, thereby providing more personalized, contextual and coherent customer service. This not only improves the user experience, but also improves the satisfaction and efficiency of the service.
[0165] In an optional embodiment, an enterprise-level intelligent customer service system based on long short-term memory optimization, such as Figure 3 , is divided into the following four modules: management module, storage module, compression module and output optimization module. The management module includes short-term memory management and long-term memory management. The storage module is a layered heterogeneous storage: combined with a relational database (storing multi-round dialogue original data, user index and operation log), a vector database and a graph database. The vector database is divided into vector database 1 (storing long-term memory), vector database 2 (preset knowledge base) and vector database 3 (incremental database). After receiving the user's question, the short-term memory, long-term memory and related knowledge are obtained from the storage module, and the short-term memory, long-term memory and related knowledge are input to the compression module. The compressed fusion data is output to the target model in the output optimization module, and the reply result is output to the front-end interaction interface. The interactive session is updated to the storage module through the management module.
[0166] In an optional embodiment, the memory update process is as follows Figure 4As shown, the user generates context in multiple rounds of conversation, combined with the prompt to form the instruction of the input natural language model; the natural language model analyzes the content of the instruction and identifies two types of core information: entity / relationship information, key fact information. Graph database update: determine whether there is the same entity: if there is, further determine whether there is the same name relationship. If there is the relationship, no operation is performed; if there is no such relationship, the relationship is added. If the entity does not exist, the entity and the relationship are added, that is, the new entity and its relationship are added into the graph database.
[0167] Vector database update: determine whether the same vector (i.e. semantically repeated information) already exists. If it exists, no operation is performed; if it does not exist, the vector is added.
[0168] In an optional embodiment, the memory update process is as shown in Figure 5 All memories to be processed (i.e. flag = 1) are collected to form a candidate memory set, which is arranged in ascending order of recording time. When the candidate memory set is not empty, each candidate memory is traversed, and the following operations are performed: the most similar existing memories (existing memories, i.e. flag = 0) are retrieved from the vector database; the candidate memory and these similar memories are submitted to the natural language model together, and the natural language model decides to perform one of the following four operations: add (ADD): if the candidate memory is completely new information, the candidate memory is retained. Update (UPDATE): if the candidate memory is a supplement or update to an existing memory, the existing information is updated, and the candidate memory is deleted. Delete (DELETE): if the candidate memory contradicts an existing memory, the old information is deleted, and the candidate memory is retained. No operation (NOOP): if the candidate memory is repeated with an existing memory, the candidate memory is deleted. The above steps are repeated until the candidate memory set is empty.
[0169] In an optional embodiment, the memory update process is as shown in Figure 6 All relationships to be verified are collected to form a candidate relationship set, which is arranged in ascending order of recording time. Each candidate relationship is traversed, and the following operations are performed: the corresponding existing relationships (corresponding to the entity pair connected by the candidate relationship) are retrieved from the graph database; the candidate relationship and these corresponding existing entities / relationships are submitted to the natural language model together, and the natural language model determines whether the candidate relationship conflicts with the existing relationship: if not, the candidate relationship is retained; if so, the conflicting existing relationship is marked as obsolete, the invalid time is recorded, and the graph database is updated.
[0170] In an optional embodiment, to further enhance the explainability and interactivity of the intelligent customer service system, during the answer generation process, not only is the knowledge base reference source added, but also the memory reference source is introduced. This innovative design enables each answer to not only explicitly identify its knowledge or information source, but also trace back to relevant memories in past interactions with the user, thereby enhancing the transparency and trustworthiness of the system.
[0171] Introduction of knowledge base reference source: During the generation of answers, the information source in the knowledge base is automatically identified and referenced, and the specific reference source is explicitly indicated in the answer. In this way, the user can directly view the knowledge source behind the reply when receiving the answer, enhancing the trust in the accuracy and completeness of the answer.
[0172] Reference method: For example: "According to our help center [document number XYZ-2025], your problem can be solved by the following way……".
[0173] Introduction of memory reference source: In addition to the knowledge base reference, the memory reference source is introduced. When answering the user, the system not only relies on real-time query of knowledge base information, but also references and integrates past dialogue history, user preferences, and previous interaction records. This approach makes the response of the intelligent customer service system more personalized and contextual, and the answer accurately reflects the understanding of the system for the user's specific needs or preferences. Each referenced memory source is closely related to the user's historical dialogue record or user-specific preference settings.
[0174] Reference method: For example: "According to your mention in the last conversation [session record on October 25, 2025], your previous choice is……".
[0175] According to the content of each response, the appropriate reference source is automatically selected. For example, if the answer involves standardized product information, the knowledge base is preferred; if the answer is based on the user's personalized needs or historical behavior, the memory library is referenced, and the memory source is indicated in the answer.
[0176] Embodiments of the present application also provide a data processing apparatus, as shown in Figure 7 The data processing apparatus includes a receiving unit 701, a first obtaining unit 702, and a first processing unit 703.
[0177] The receiving unit 701 is configured to receive problem information input by a target object;
[0178] The first obtaining unit 702 is configured to obtain target data information according to the question information, wherein the target data information comprises first memory information, second memory information related to the question information, and target knowledge information related to the question information, the first memory information is obtained based on a current session corresponding to the question information, the second memory information is obtained based on historical sessions of a target object, and the target knowledge information is knowledge information corresponding to a target service related to the question information.
[0179] The first processing unit 703 is configured to process the target data information by using a target model to obtain reply information corresponding to the question information.
[0180] Optionally, in the data processing apparatus provided in the embodiments of the present application, the first obtaining unit comprises: a first determining subunit, configured to determine a current session corresponding to the question information, and obtain the first memory information based on dialogue information in the current session; a first matching subunit, configured to match the question information with memory information in the first database to obtain the second memory information; and a second matching subunit, configured to match the question information with knowledge information in the second database to obtain the target knowledge information.
[0181] Optionally, in the data processing apparatus provided in the embodiments of the present application, the first determining subunit comprises: an obtaining module, configured to obtain a target number of dialogue rounds that the first memory can accommodate corresponding to the target model; and a selecting module, configured to select the first memory information from the current session according to the target number of dialogue rounds.
[0182] Optionally, in the data processing apparatus provided in the embodiments of the present application, the apparatus further comprises: a first determining unit, configured to determine a current session corresponding to the question information; and a second obtaining unit, configured to obtain dialogue information in the current session, and update memory information in the first database according to the dialogue information.
[0183] Optionally, in the data processing apparatus provided in the embodiments of the present application, the first matching subunit comprises: a processing module, configured to perform vectorization processing on the question information to obtain a first vector; a calculation module, configured to calculate a similarity between the first vector and a second vector corresponding to the second memory information in the first database to obtain a target similarity; and a first determining module, configured to obtain the second memory information according to the target similarity.
[0184] Optionally, in the data processing apparatus provided in the embodiment of the present application, the first processing unit comprises: an encoding subunit, configured to encode the target data information through an embedding layer of the target model to obtain a third vector; a first calculation subunit, configured to calculate the third vector through a multi-head self-attention mechanism of the target model to obtain weight information; an aggregation subunit, configured to aggregate the third vector according to the weight information to obtain a fourth vector; and a processing subunit, configured to process the fourth vector through a decoding layer of the target model to obtain the reply information.
[0185] Optionally, in the data processing apparatus provided in the embodiment of the present application, the obtaining module comprises: a first obtaining sub-module, configured to obtain a maximum context length corresponding to the target model and a reference context length; a second obtaining sub-module, configured to obtain a preset step-up amplitude and a maximum dialogue round number of the target model in one session; and a calculation sub-module, configured to calculate according to the maximum context length, the reference context length, the step-up amplitude and the maximum dialogue round number to obtain the target dialogue round number.
[0186] Optionally, in the data processing apparatus provided in the embodiment of the present application, the apparatus further comprises: a third obtaining unit, configured to obtain a target compression rate after obtaining the target data information according to the question information; and a second processing unit, configured to compress the target data information according to the target compression rate to obtain compressed target data information, wherein the compressed target data information is processed by the target model to obtain the reply information corresponding to the question information.
[0187] Optionally, in the data processing apparatus provided in the embodiment of the present application, the first matching subunit comprises: a first matching module, configured to match the question information and the memory information in the vector database of the first database to obtain first initial memory information; a second matching module, configured to match the question information and the memory information in the graph database of the first database to obtain second initial memory information; and a second determination module, configured to obtain the second type of memory information according to the first initial memory information and the second initial memory information.
[0188] Optionally, in the data processing apparatus provided in the embodiment of the present application, the third obtaining unit comprises: a first obtaining subunit, configured to obtain a current load value of the target model and a total number of characters corresponding to the target data information; a second calculation subunit, configured to calculate according to the current load value and the total number of characters to obtain a weighting value; and a third calculation subunit, configured to calculate according to the weighting value and a preset compression rate range corresponding to the target data information to obtain the target compression rate.
[0189] Optionally, in the data processing apparatus provided by the embodiment of the present application, for the first type of memory information in the target data information, the third obtaining unit comprises: a second obtaining subunit, configured to obtain a current load value of the target model and a total number of characters corresponding to the first type of memory information; a fourth calculating subunit, configured to calculate the total number of characters corresponding to the first type of memory information according to the current load value, to obtain a weighting value corresponding to the first type of memory information; and a fifth calculating subunit, configured to calculate the weighting value corresponding to the first type of memory information according to a preset compression rate range corresponding to the first type of memory information, to obtain a target compression rate corresponding to the first type of memory information.
[0190] Optionally, in the data processing apparatus provided by the embodiment of the present application, the second processing unit comprises: a sixth calculating subunit, configured to calculate a matching degree between the target data information and the question information, to obtain a target value of each data in the target data information; and a compression subunit, configured to perform compression processing on the target data information according to the target value and the target compression rate, to obtain compressed target data information.
[0191] Optionally, in the data processing apparatus provided by the embodiment of the present application, the compression subunit comprises: a sorting module, configured to sort each data in the target data information according to the target value, to obtain sorted data information; a third determining module, configured to determine a retention ratio according to the target compression rate; and a screening module, configured to screen the sorted data information according to the retention ratio, to obtain the compressed target data information.
[0192] Optionally, in the data processing apparatus provided by the embodiment of the present application, the apparatus further comprises: a second determining unit, configured to determine a current session corresponding to the question information; a judging unit, configured to judge whether a dialogue round of the current session reaches a target threshold value, wherein the target threshold value is less than a target dialogue round number of the first type of memory that can be accommodated by the target model; and a triggering unit, configured to trigger a second type of memory updating task if the dialogue round of the current session reaches the target threshold value, wherein the second type of memory updating task is used to perform an updating operation on the memory information in the first database according to dialogue information in the current session.
[0193] Optionally, in the data processing apparatus provided by the embodiment of the present application, the apparatus further comprises: an adding unit, configured to add the second type of memory updating task to a task queue after triggering the second type of memory updating task, when the target model is in the first state; a recording unit, configured to record the dialogue information in the current session to the first database and set a target identifier to the dialogue information after adding the second type of memory updating task to the task queue, wherein the target identifier is used to represent that the dialogue information is candidate memory information to be processed; and a first executing unit, configured to execute tasks in the task queue according to a task order in the task queue.
[0194] Optionally, in the data processing apparatus provided in the embodiments of the present application, the apparatus further comprises a second execution unit configured to, after triggering the second type of memory updating task, detect that the target model is in the second state, and execute the second type of memory updating task.
[0195] Optionally, in the data processing apparatus provided in the embodiments of the present application, the apparatus further comprises a calculation unit configured to, after adding the second type of memory updating task to the task queue, calculate a priority of the second type of memory updating task to obtain a target priority; and a third determination unit configured to determine whether to adjust an execution order of the second type of memory updating task according to the target priority.
[0196] Optionally, in the data processing apparatus provided in the embodiments of the present application, the calculation unit comprises a second determination sub-unit configured to determine an urgency degree corresponding to the current session according to a type of the current session; a seventh calculation sub-unit configured to calculate a practical value corresponding to the second type of memory information according to a number of times that the second type of memory information is referenced in the reply information output by the target model; an eighth calculation sub-unit configured to calculate a time difference between a time when the second type of memory updating task is added to the task queue and a current time; and a ninth calculation sub-unit configured to calculate the target priority according to the urgency degree, the practical value and the time difference.
[0197] Optionally, in the data processing apparatus provided in the embodiments of the present application, the apparatus further comprises a fourth determination unit configured to determine candidate memory information with the target identifier, and convert the candidate memory information into a target vector; a retrieval unit configured to retrieve a memory vector most similar to the target vector from the first database; and a third processing unit configured to process the memory vector according to the target vector.
[0198] Optionally, in the data processing apparatus provided in the embodiments of the present application, the third processing unit comprises a first storage sub-unit configured to, if a similarity between the target vector and the memory vector is lower than a preset threshold, store the target vector into the first database; a first updating sub-unit configured to, if the target vector is a supplement or an update to the memory vector, update the memory vector according to the target vector; and a deletion sub-unit configured to, if the target vector is in conflict with the memory vector, delete the memory vector and store the target vector into the first database.
[0199] Optionally, in the data processing apparatus provided in the embodiments of the present application, the apparatus further comprises: a fifth determining unit, configured to determine candidate memory information with the target identifier, and perform entity and relationship identification on the candidate memory information to obtain a target entity and a target relationship; a fourth obtaining unit, configured to obtain a memory graph corresponding to the target object from the first database; and a fourth processing unit, configured to process the memory graph according to the target entity and the target relationship.
[0200] Optionally, in the data processing apparatus provided in the embodiments of the present application, the fourth processing unit comprises: an adding sub-unit, configured to add the target entity and the target relationship into the memory graph if the memory graph does not have an association with the target entity and the target relationship; and a second updating sub-unit, configured to update entities and relationships in the memory graph according to the target entity and the target relationship if the memory graph has an association with the target entity or the target relationship.
[0201] The features of the embodiments corresponding to the data processing apparatus can be referred to the related descriptions of the embodiments of the data processing method, which will not be repeated here.
[0202] The embodiments of the present application further provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-mentioned data processing method embodiments.
[0203] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned data processing method embodiments when running.
[0204] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0205] The embodiments of the present application further provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned data processing method embodiments.
[0206] The embodiments of the present application further provide another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned data processing method embodiments.
[0207] Those skilled in the art will further appreciate that the units and algorithms described in connection with the examples disclosed herein can be embodied directly in hardware, in software, or in a combination of the two. For the sake of brevity, descriptions of a method of operation of the examples will be presented in general terms, with reference to the attached figures. Those skilled in the art will understand that the descriptions are not limited to the examples described and that the examples can be practiced in a variety of ways. Those skilled in the art will further appreciate that the descriptions and examples are not intended to limit the scope of the claims to the specific forms disclosed. Rather, the descriptions and examples are intended to cover all functional equivalents within the scope of the claims.
[0208] The above has carried out detailed introduction to the data processing method provided by the application. The principle and implementation of the application are described by applying specific examples in the present text. The above example description is only for helping to understand the method and its core idea of the application. It should be pointed out that, for the ordinary skilled in the art, some improvements and modifications can be made to the application without departing from the principle of the application, and these improvements and modifications also fall within the protection scope of the claims of the application.
Claims
1. A data processing method, characterized by, The method comprises: receiving question information input by a target object; obtaining target data information according to the question information, wherein the target data information comprises first type memory information, second type memory information related to the question information, and target knowledge information related to the question information, the first type memory information is obtained based on a current session corresponding to the question information, the second type memory information is obtained based on historical sessions of the target object, and the target knowledge information is knowledge information corresponding to a target business related to the question information; processing the target data information through a target model to obtain reply information corresponding to the question information; selecting the first type memory information from the current session according to a target dialogue round number; matching the question information with memory information in a first database to obtain the second type memory information; matching the question information with knowledge information in a second database to obtain the target knowledge information; processing the target data information through the target model to obtain the reply information corresponding to the question information comprises: encoding the target data information through an embedding layer of the target model to obtain a third vector; calculating the third vector through a multi-head self-attention mechanism of the target model to obtain weight information; weighting and aggregating the third vector according to the weight information to obtain a fourth vector; and processing the fourth vector through a decoding layer of the target model to obtain the reply information; wherein a maximum context length corresponding to the target model and a reference context length are obtained; a preset step-up amplitude and a maximum dialogue round number of the target model in one session are obtained; the maximum context length, the reference context length, the step-up amplitude, and the maximum dialogue round number are calculated to obtain a target dialogue round number of first type memory information that can be accommodated by the target model, and the target dialogue round number is calculated according to the following formula: ; wherein, is a target number of dialogue turns, is a maximum context length, is a reference context length, is a ladder up magnitude, is a maximum number of dialogue turns.
2. The data processing method according to claim 1, characterized in that, The method further comprises: determining a current session corresponding to the question information; obtaining dialogue information in the current session and updating memory information in the first database according to the dialogue information.
3. The data processing method of claim 1, wherein, The matching of the question information with memory information in the first database to obtain the second type memory information comprises: vectorizing the question information to obtain a first vector; calculating a similarity between the first vector and a second vector corresponding to the second type memory information in the first database to obtain a target similarity; obtaining the second type memory information according to the target similarity.
4. The data processing method of claim 1, wherein, After obtaining the target data information according to the question information, the method further comprises: obtaining a target compression rate; compressing the target data information according to the target compression rate to obtain compressed target data information, wherein the compressed target data information is processed through the target model to obtain the reply information corresponding to the question information.
5. The data processing method of claim 1, wherein, The matching of the question information with memory information in the first database to obtain the second type memory information comprises: Matching the problem information and memory information in a vector database of the first database obtains first initial memory information; Matching the problem information and memory information in a graph database of the first database obtains second initial memory information; According to the first initial memory information and the second initial memory information, the second type of memory information is obtained.
6. The data processing method according to claim 4, characterized in that, The target compression rate includes: Obtaining the current load value of the target model and the total number of characters corresponding to the target data information; According to the current load value and the total number of characters, a weighted value is calculated; According to the weighted value and the preset compression rate range corresponding to the target data information, the target compression rate is calculated.
7. The data processing method according to claim 5, characterized in that, For the first type of memory information in the target data information, the target compression rate includes: Obtaining the current load value of the target model and the total number of characters corresponding to the first type of memory information; According to the current load value and the total number of characters corresponding to the first type of memory information, a weighted value corresponding to the first type of memory information is calculated; According to the weighted value corresponding to the first type of memory information and the preset compression rate range corresponding to the first type of memory information, the target compression rate corresponding to the first type of memory information is calculated.
8. The data processing method according to claim 4, characterized in that, According to the target compression rate, the target data information is compressed to obtain compressed target data information includes: Calculating the matching degree of the target data information and the problem information to obtain the target value of each data in the target data information; According to the target value and the target compression rate, the target data information is compressed to obtain compressed target data information.
9. The data processing method according to claim 8, characterized in that, According to the target value and the target compression rate, the target data information is compressed to obtain compressed target data information includes: According to the target value, the data in the target data information is sorted to obtain sorted data information; According to the target compression rate, a retention ratio is determined; According to the retention ratio, the sorted data information is screened to obtain the compressed target data information.
10. The data processing method of claim 1, wherein, The method further includes: Determining the current session corresponding to the problem information; Judging whether the dialogue round of the current session reaches a target threshold, wherein the target threshold is less than the target dialogue round number that the first type of memory can accommodate corresponding to the target model; If the dialogue round of the current session reaches the target threshold, triggering a second type of memory update task, wherein the second type of memory update task is used to update the memory information in the first database according to the dialogue information in the current session.
11. The data processing method according to claim 10, characterized in that, After triggering the second type of memory update task, the method further includes: When the target model is in a first state, adding the second type of memory update task to a task queue; After adding the second type of memory update task to the task queue, recording the dialogue information in the current session to the first database and setting a target identifier for the dialogue information, wherein the target identifier is used to represent that the dialogue information is candidate memory information to be processed; According to an order of tasks in the task queue, the tasks in the task queue are executed.
12. The data processing method of claim 10, wherein, After triggering the second type of memory updating task, the method further comprises: detecting that the target model is in a second state, and executing the second type of memory updating task.
13. The data processing method of claim 11, wherein, After adding the second type of memory updating task to the task queue, the method further comprises: calculating a priority of the second type of memory updating task to obtain a target priority; determining whether to adjust an execution order of the second type of memory updating task according to the target priority.
14. The data processing method according to claim 13, characterized in that, The calculating of the priority of the second type of memory updating task to obtain the target priority comprises: determining an urgency degree corresponding to the current session according to a type of the current session; calculating a practical value corresponding to the second type of memory information according to a number of times of referring to the second type of memory information in the reply information output by the target model; calculating a time difference between a time when the second type of memory updating task is added to the task queue and a current time; calculating the target priority according to the urgency degree, the practical value and the time difference.
15. The data processing method according to claim 12, characterized in that, The method further comprises: determining candidate memory information with a target identifier, and converting the candidate memory information into a target vector; retrieving a memory vector most similar to the target vector in the first database; processing the memory vector according to the target vector.
16. The data processing method according to claim 15, characterized in that, The processing of the memory vector according to the target vector comprises: if a similarity between the target vector and the memory vector is lower than a preset threshold, storing the target vector into the first database; if the target vector is a supplement or an update of the memory vector, updating the memory vector according to the target vector; if the target vector is in conflict with the memory vector, deleting the memory vector and storing the target vector into the first database.
17. The data processing method of claim 12, wherein, The method further comprises: determining candidate memory information with a target identifier, and performing entity and relationship recognition on the candidate memory information to obtain a target entity and a target relationship; obtaining a memory graph corresponding to the target object from the first database; processing the memory graph according to the target entity and the target relationship.
18. The data processing method according to claim 17, characterized in that, The processing of the memory graph according to the target entity and the target relationship comprises: if the memory graph is not associated with the target entity and the target relationship, adding the target entity and the target relationship to the memory graph; if the memory graph is associated with the target entity or the target relationship, updating entities and relationships in the memory graph according to the target entity and the target relationship.
19. A data processing apparatus, characterized by comprises: a receiving unit configured to receive question information input by a target object; The first obtaining unit is configured to obtain target data information according to the question information, wherein the target data information comprises first memory information, second memory information related to the question information, and target knowledge information related to the question information, the first memory information is obtained based on a current session corresponding to the question information, the second memory information is obtained based on historical sessions of the target object, and the target knowledge information is knowledge information corresponding to a target service related to the question information; The first processing unit is configured to process the target data information by using a target model to obtain reply information corresponding to the question information; The device is further configured to select the first memory information from the current session according to a target dialogue round number; The question information and memory information in the first database are matched to obtain the second memory information; The question information and knowledge information in the second database are matched to obtain the target knowledge information; The first processing unit comprises an encoding subunit configured to encode the target data information by using an embedding layer of the target model to obtain a third vector, a first calculation subunit configured to calculate the third vector by using a multi-head self-attention mechanism of the target model to obtain weight information, an aggregation subunit configured to aggregate the third vector by using the weight information to obtain a fourth vector, and a processing subunit configured to process the fourth vector by using a decoding layer of the target model to obtain the reply information; The maximum context length corresponding to the target model and the reference context length are obtained; A preset step-up amplitude and a maximum dialogue round number of the target model in one session are obtained; The maximum context length, the reference context length, the step-up amplitude, and the maximum dialogue round number are used for calculation to obtain a target dialogue round number of the first memory information that can be accommodated by the target model, and the target dialogue round number is calculated according to the following formula: ; wherein, is a target number of dialogue turns, is a maximum context length, is a reference context length, is a ladder up magnitude, is a maximum number of dialogue turns.
20. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 18. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 18. The computer program is executed by the processor to implement the steps of the data processing method according to any one of claims 1 to 18.
21. A computer-readable storage medium, characterized in that, 22. A computer program product comprising a computer program, characterised in that,
Citation Information
Patent Citations
Machine question and answer dialogue method and device
CN120705284A
Multi-round dialogue semantic analysis method and system based on long short-term memory network
WO2021042543A1