Data processing method and device based on human-computer interaction, equipment and medium

By performing standard natural language processing, basic intent classification, contextual memory, and fine-grained analysis on multimodal data, and combining it with an external knowledge base to generate personalized responses, the problem of insufficient accuracy of human-computer interaction platforms in processing complex business logic and image recognition in existing technologies is solved, achieving more efficient and accurate user responses and interactive experiences.

CN120596716APending Publication Date: 2025-09-05CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510675438.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing human-computer interaction platforms have limited capabilities in processing complex business logic, unstructured data, and deeply understanding user intent. Image recognition systems have insufficient accuracy in recognizing complex or fuzzy images, resulting in inaccurate and inefficient responses.

Method used

By performing standard natural language processing, basic intent classification, contextual memory processing, and fine-grained parsing on multimodal data, a multi-dimensional query vector is constructed and combined with an external knowledge base to generate personalized intermediate responses, which are then optimized and adjusted through a pre-acquired feedback mechanism.

Benefits of technology

It significantly improves user satisfaction and convenience, provides more accurate and timely responses, enhances the interactive experience, reduces the need for manual intervention, improves business processing speed and accuracy, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596716A_ABST
    Figure CN120596716A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a data processing method, device, equipment and medium based on man-machine interaction. Performing basic intention classification on the standardized data to obtain classified data; performing context memory processing on the classified data to obtain memory data; performing fine-grained analysis on the memory data to obtain analysis data; constructing a multi-dimensional query vector for the analysis data, and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate reply; and performing optimization adjustment on the intermediate reply according to a pre-obtained feedback mechanism to obtain an optimization result. By deeply understanding the intention of the user and providing personalized services, the satisfaction and convenience of the user are remarkably improved, the user can obtain more accurate, more timely and more demand-meeting response in the use process, and the interaction experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a data processing method, device, equipment and medium based on human-computer interaction. Background Art

[0002] Data processing based on human-computer interaction is a working mode that combines robotics technology with process orchestration concepts to achieve automated and efficient task processing.

[0003] In the field of healthcare, robots can collect multi-dimensional data such as patients' symptoms, test results, medical history, etc., analyze them using machine learning and artificial intelligence algorithms, provide doctors with diagnostic advice and references, assist doctors in making more accurate diagnostic decisions, and reduce the possibility of missed diagnosis and misdiagnosis.

[0004] In the field of financial technology, robots can automatically collect customer information, complete a series of operations such as identity verification and document review, and quickly open bank accounts or securities accounts for customers, greatly shortening account opening time and improving business processing efficiency. In terms of account management, robots can regularly monitor accounts for risks, update information, and ensure safe and compliant account operations.

[0005] However, while existing human-computer interaction platforms provide graphical interfaces that allow users to build processes, they are limited in their ability to process complex business logic, unstructured data (such as images and voice), and deeply understand user intent. Furthermore, while natural language processing technology can generate basic text responses, it lacks a deep understanding of context, resulting in potentially inaccurate responses.

[0006] At the same time, the existing image recognition system has good processing quality in specific scenarios, but the recognition accuracy of complex or blurred images still needs to be improved, especially in the claims scenario, where there is a lack of accurate recognition of vehicle characteristics of different models and brands. Summary of the Invention

[0007] The present invention provides a data processing method, device, equipment and medium based on human-computer interaction, which significantly improves user satisfaction and convenience by deeply understanding user intentions and providing personalized services. Users can obtain more accurate, timely and demand-oriented responses during use, thereby enhancing the interactive experience.

[0008] In a first aspect, a data processing method based on human-computer interaction is provided, comprising:

[0009] Perform natural language standard processing on multimodal data to obtain standardized data;

[0010] Performing basic intent classification on the standardized data to obtain classified data;

[0011] Performing context memory processing on the classified data to obtain memory data;

[0012] Performing fine-grained analysis on the memory data to obtain analyzed data;

[0013] Constructing a multi-dimensional query vector for the parsed data, and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response;

[0014] The intermediate response is optimized and adjusted according to the feedback mechanism obtained in advance to obtain an optimization result, which is used as the target result.

[0015] In a second aspect, a data processing device based on human-computer interaction is provided, comprising:

[0016] A processing module is used to perform natural language standard processing on multimodal data to obtain standardized data;

[0017] A classification module, configured to perform basic intent classification on the standardized data to obtain classified data;

[0018] A memory module, configured to perform context memory processing on the classified data to obtain memory data;

[0019] A parsing module, configured to perform fine-grained parsing on the memory data to obtain parsed data;

[0020] A construction module, configured to construct a multi-dimensional query vector for the parsed data;

[0021] A generation module, configured to combine the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response;

[0022] An adjustment module, configured to optimize and adjust the intermediate response according to a pre-acquired feedback mechanism to obtain an optimization result;

[0023] As a module, it is used to take the optimization result as the target result.

[0024] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned data processing method based on human-computer interaction are implemented.

[0025] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned data processing method based on human-computer interaction are implemented.

[0026] In the above-mentioned scheme implemented by the data processing method, device, computer equipment and storage medium based on human-computer interaction, the user intention is deeply understood and personalized services are provided through basic intent classification, context memory, and fine-grained analysis of multimodal data, which significantly improves user satisfaction and convenience. Users can obtain more accurate, timely and demand-oriented responses during use, enhance the interactive experience, automatically process images and voice information, and optimize processes through intelligent decision-making, which greatly reduces the need for manual intervention and improves the speed and accuracy of business processing; the analyzed data is combined with the preset external knowledge base to generate intermediate responses to automatically complete the task, reducing the time and error rate of manual operations, and significantly improving the efficiency of the overall business process. At the same time, automated and intelligent process processing reduces dependence on manpower, reduces operating costs, and improves work efficiency, bringing significant economic benefits to the enterprise. By reducing labor costs and increasing processing speed, the system saves a lot of resources for the enterprise, while increasing the scale and speed of business processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0028] Figure 1 This is a schematic diagram of an application environment of a data processing method based on human-computer interaction in one embodiment of the present invention;

[0029] Figure 2 This is a flow chart of a data processing method based on human-computer interaction in one embodiment of the present invention;

[0030] Figure 3 This is a structural diagram of a data processing device based on human-computer interaction in one embodiment of the present invention;

[0031] Figure 4 is a structural diagram of a computer device according to an embodiment of the present invention;

[0032] Figure 5 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0034] The data processing method based on human-computer interaction provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the client communicates with the server through a network. The server can perform natural language standard processing on multimodal data to obtain standardized data; perform basic intent classification on the standardized data to obtain classified data; perform context memory processing on the classified data to obtain memory data; perform fine-grained parsing on the memory data to obtain parsed data; construct a multi-dimensional query vector for the parsed data, and combine the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response; optimize and adjust the intermediate response according to a pre-acquired feedback mechanism to obtain an optimized result, use the optimized result as the target result, and feed the target result back to the client. The present invention provides a data processing device based on human-computer interaction. For target result business, it significantly improves user satisfaction and convenience by deeply understanding user intentions and providing personalized services. Users can obtain more accurate, timely, and more demand-oriented responses during use, thereby enhancing the interactive experience. Among them, the client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0035] See also Figure 2 As shown, Figure 2 A flowchart of a data processing method based on human-computer interaction provided by an embodiment of the present invention includes the following steps:

[0036] S1. Perform natural language standard processing on multimodal data to obtain standardized data.

[0037] In the embodiment of the present invention, the natural language standard processing refers to the process of cleaning, word segmentation, entity recognition, and initial intent screening of multimodal data.

[0038] Specifically, natural language processing technology is used to process multimodal data in various forms according to unified standards and specifications, and convert them into a structured data form that is easy for computers to understand, store, analyze and apply.

[0039] In specific healthcare scenarios, in medical research, standardized multimodal data facilitates large-scale data mining and analysis. Researchers can more efficiently extract valuable information, discover the occurrence and development patterns of diseases, the relationship between treatment effects and related factors, etc., provide strong data support for medical research, and promote the advancement of medical technology.

[0040] In the FinTech scenario, converting multimodal data into standardized data helps intelligent customer service systems better understand customer needs. Intelligent customer service can perform entity recognition and initial screening of intent based on the customer's text consultation content, answer questions quickly and accurately, and at the same time combine customer transaction records, identity authentication and other non-text data to provide more personalized services and enhance customer experience.

[0041] In the embodiment of the present invention, performing natural language standard processing on the multimodal data to obtain standardized data includes:

[0042] Divide multimodal data into textual data and non-textual data;

[0043] Performing entity recognition on the text data to obtain recognition data;

[0044] Performing a preliminary intention screening on the text data to obtain preliminary screening data, and combining the recognition data and the preliminary screening data into comprehensive text data;

[0045] extracting structured tags and content categories of the non-text data, and merging the structured tags and content categories with the non-text data into comprehensive non-text data;

[0046] Normalized data is generated based on the combination of the integrated text data and the integrated non-text data.

[0047] In an embodiment of the present invention, the division refers to dividing a complex data set containing multiple modalities into two major categories: text data with text symbols as the main form of expression, and other non-text data in the form of non-text symbols based on the expression form and nature characteristics of the data. The entity recognition refers to automatically identifying entities with specific meanings from text data and classifying them into predefined categories. The initial screening of intentions refers to the process of making preliminary judgments and screening on the intentions expressed by text data. The extraction refers to the use of specific technologies and methods to extract valuable information from non-text data, namely structured labels and content categories, so that computers can understand and process these data. The combined generation refers to the fusion of processed comprehensive text data with comprehensive non-text data, and generating standardized data with unified format, specifications and semantics according to certain rules and methods.

[0048] Specifically, multimodal data refers to data that contains multiple forms, such as text, images, audio, video, etc. When dividing, it is distinguished according to the form of data presentation. Data that exists in text form and can be directly read and understood by humans, such as documents, web page text, chat records, etc. are classified as text data; and other non-text forms of data such as images, audio, video, sensor data, etc. are classified as non-text data.

[0049] Furthermore, the entity recognition algorithm or model in natural language processing technology is used to process text data; first, the text data is segmented to split the continuous text into individual words or phrases, and then, based on a pre-trained model, such as a named entity recognition model (NER) based on deep learning, combined with part-of-speech tagging, syntactic analysis and other information, each word or phrase is judged to identify entities with specific meanings. Finally, the identified entities and their related information (such as entity category, position in the text, etc.) are organized into recognition data.

[0050] Furthermore, a rule set for intent classification is constructed or a simple intent classification model is trained. For the rule set method, a series of rules are formulated to correspond between keywords and intents. For example, the appearance of keywords such as "buy" and "place an order" corresponds to "shopping intent"; the appearance of keywords such as "query" and "understand" corresponds to "information query intent", etc. For the intent classification model, annotated text data is used as the training set, and the annotated information is the intent category corresponding to the text. A model that can classify the input text by intent is trained, such as a naive Bayes model or a decision tree model. Then, the text data to be processed is input into the intent classification model based on the rule set or the trained intent classification model. The model will output the intent category to which the text may belong based on the content and features of the text. This intent category information is the preliminary screening data. The recognition data and the preliminary screening data obtained previously are then integrated. The entity information in the recognition data can be associated with the intent category information in the preliminary screening data to form a more comprehensive data set.

[0051] Specifically, different types of non-text data are processed using corresponding technical methods. For image data, image recognition algorithms, such as object detection models based on convolutional neural networks (CNNs), are used to identify objects, scenes, and other information in the image, generating structured labels such as "cat" and "car." The content category is determined based on the overall content of the image, such as "animal picture" and "traffic scene picture." For audio data, audio processing techniques, such as Mel-Frequency Cepstral Coefficients (MFCCs), are used to extract audio features. A trained audio classification model is then used to identify the speech content and sound categories in the audio, generating structured labels such as "speaker's name" and "music genre," and determining the content category. For video data, the video can be first decomposed into frames, and then image recognition techniques are used to process each frame, extracting key frames and information such as objects and scenes within the video, generating structured labels, and determining the content category based on the video's theme and content, such as "teaching video" and "advertising video." The extracted structured labels and content categories are then associated with the original non-text data.

[0052] This approach first determines a standardized data format and structure. For example, a table format is used, with columns representing different information fields, such as "text content," "entity information," "intent category," "non-text data type," "structured label," and "content category." The combined text and non-text data are then organized and populated according to this unified format, ultimately generating standardized data that meets specific standards and requirements.

[0053] In an embodiment of the present invention, data of different modalities have different characteristics and representation methods. Through natural language standard processing, data of various modalities are converted into a unified natural language representation, which facilitates the fusion of data of different modalities, breaks down barriers between data, realizes data sharing and interaction, and helps to more comprehensively mine the information in the data.

[0054] S2. Perform basic intent classification on the standardized data to obtain classified data.

[0055] In an embodiment of the present invention, the basic intent classification refers to the preliminary and relatively broad classification of the intent contained in the data using a specific classification method or model based on the already generated standardized data, so as to clarify the main intent direction of the data.

[0056] Specifically, a certain classification method or model is used to process the generated standardized data according to a pre-set basic intent category system, and each piece of data is classified into the corresponding basic intent category, thereby obtaining classified data with clear intent classification information.

[0057] In specific medical and health scenarios, relevant standardized data is classified into intent categories such as "emergency," "common disease," and "health consultation" based on the content of patient consultations, helping the hospital's triage system quickly guide patients to appropriate departments and doctors, improving medical service efficiency and ensuring that critically ill patients receive timely treatment.

[0058] In the specific scenarios of financial technology, relevant standardized data are classified into intent categories such as "account inquiry", "business processing", "complaints and suggestions" based on the content of customer consultations and feedback, enabling financial institutions to quickly respond to customer needs, provide personalized services, and improve customer satisfaction and loyalty.

[0059] In the embodiment of the present invention, performing basic intent classification on the standardized data to obtain classified data includes:

[0060] Performing feature extraction on the standardized data to obtain feature data;

[0061] Mapping the feature data to a preset intent category space to obtain an intent classification corresponding to the feature data;

[0062] Calculate the probability of each intent classification and select the intent classification with the highest probability;

[0063] The selected intent classification is combined with the normalized data to form classification data.

[0064] In an embodiment of the present invention, the feature extraction refers to the process of extracting the essential features and key information that can represent the data from standardized data; the mapping refers to establishing a correspondence between the feature data extracted from the standardized data and each intent category in a pre-set intent category space according to certain rules or methods, so as to determine the intent classification corresponding to the feature data; the calculation refers to using a specific algorithm, model or mathematical formula to quantitatively evaluate the possibility of each intent classification, so as to obtain the probability value corresponding to each intent classification; the combination formation refers to integrating the intent classification with the highest probability selected after calculating the probability with the original standardized data to construct a new data set, namely, classified data.

[0065] Specifically, specific algorithms and techniques are used to extract key features of the data. In text processing, the Bag of Words model and TF-IDF (Term Frequency-Inverse Document Frequency) algorithm can be used to extract word frequency-related features. For image data, common feature extraction methods include SIFT (Scale-Invariant Feature Transform), HOG (Histogram of Oriented Gradients), and feature maps extracted by convolutional neural networks (CNN). When processing audio data, features such as Mel-Frequency Cepstral Coefficients (MFCC) may be extracted. The final set of features that can represent the essential information of the data is feature data.

[0066] In detail, a series of intent categories are pre-defined based on specific application scenarios and needs. For example, in the field of healthcare, intent categories may include disease diagnosis, treatment plan consultation, drug side effect inquiry, etc.; in the field of financial technology, intent categories can be set as account information inquiry, loan application, investment product recommendation, etc., and the extracted feature data is matched with the preset intent categories using existing classification models or mapping rules. For example, classification models such as support vector machines (SVM) and decision trees have learned the feature patterns corresponding to different intent categories during the training phase. When feature data is input, the model will determine which intent category the feature data is most likely to belong to based on these learned patterns, or by formulating a rule-based mapping method, determine the corresponding intent classification based on the presence or absence of specific features in the feature data or the range of feature values.

[0067] Furthermore, for each intent classification, a probability model or algorithm is used to calculate the probability that the feature data belongs to that intent classification. In a neural network model, an activation function is used in the model's output layer to convert the output value into a probability distribution, thereby obtaining a probability value for each intent classification. The calculated probabilities for each intent classification are then compared to find the one with the highest probability value, which is then determined as the most likely intent classification for the feature data.

[0068] Furthermore, the selected intent classification information is added to the standardized data. For example, if the standardized data is stored in a table format, with each row representing a data record, a new column can be added to store the intent classification. If the data is stored in a dictionary format, a new key-value pair can be added, with the key being "intent classification" and the value being the selected intent classification name. In this way, intent classification is integrated with the standardized data.

[0069] In this embodiment of the present invention, a predefined intent category space is defined based on business needs and scenarios. Mapping feature data into this space imbues the data with clear semantic information, linking the previously abstract feature data with specific business intent, facilitating data understanding and processing. Calculating the probability of intent classification also provides a quantitative reliability indicator for each classification result. The probability value reflects the likelihood that feature data belongs to a particular intent category, making the classification results more scientific and reliable.

[0070] In this embodiment of the present invention, classified data can be stored, managed, and retrieved according to different intent categories. Compared to unclassified standardized data, this organization is more organized and facilitates long-term data maintenance and utilization. When searching for data with a specific intent, the corresponding category can be quickly located, improving the efficiency of data query and access.

[0071] S3. Perform context memory processing on the classified data to obtain memory data.

[0072] In an embodiment of the present invention, the contextual memory processing means that when processing data, not only the information of the current data itself is paid attention to, but also the contextual environment information of the data is considered, and this information is integrated and stored to form more comprehensive and meaningful memory data.

[0073] Specifically, when processing categorical data, the contextual information of the data is considered and this information is combined with the categorical data to form memory data with richer connotations and relevance, so as to better understand and utilize the data and provide more comprehensive support for subsequent tasks or analysis.

[0074] In the specific scenario of medical health, doctors can obtain more comprehensive information after the classified data such as patients' symptom descriptions, examination reports, medical history, etc. in the medical data are processed through contextual memory.

[0075] In financial scenarios, financial institutions can gain a deeper understanding of customers' financial needs and goals by performing contextual memory on classified data such as their investment preferences, asset allocation, and transaction history.

[0076] In the embodiment of the present invention, performing context memory processing on the classified data to obtain memory data includes:

[0077] Performing context encoding on the classification data to obtain an encoding vector;

[0078] Memory-outputting the encoding vector to obtain a memory vector;

[0079] The encoding vector and the memory vector are weightedly combined to obtain memory data.

[0080] In an embodiment of the present invention, the context encoding refers to a technical method for converting classified data and its related context information into vector representation, presenting the semantics, structure and other information contained in the classified data in the form of a numerical vector that can be processed and understood by a computer. The memory output refers to further processing the encoding vector obtained by context encoding, extracting and integrating key information from it, and outputting it in the form of a new vector. The weighted combination refers to a method of linearly combining two or more vectors according to different weights, aiming to comprehensively utilize the information of the encoding vector and the memory vector to generate more expressive and targeted memory data.

[0081] Specifically, an appropriate contextual encoding model is selected based on the data type and task requirements. In natural language processing, common models include Transformers and recurrent neural networks. Preprocessed categorical data is fed into the selected encoding model. Based on the input data and its context, the model generates a vector representation for each element (e.g., a word in text, a pixel block in an image). These vectors not only incorporate the element's own characteristics but also incorporate its contextual information. Based on the task objectives and requirements, the model determines which information to extract and retain from the encoding vector as memory content. Specific methods are used to extract key information from the encoding vector and integrate it into a new vector. This can be achieved through fully connected layers, pooling operations, or attention mechanisms. For example, an attention mechanism can be used to assign a weight to each dimension of the encoding vector, reflecting the importance of the information in that dimension. The weighted sum of the dimensions of the encoding vector is then calculated based on these weights to produce a compressed memory vector. Alternatively, a max pooling operation can be used to select the maximum value of each dimension of the encoding vector to form a memory vector, ultimately resulting in a memory vector containing the key information.

[0082] Furthermore, weights are assigned to the encoding vector and the memory vector. The weight value range is usually between 0 and 1, and the sum of the two weights is 1. The weights can be determined based on experience, experiments or model training results. According to the determined weights, the corresponding elements of the encoding vector and the memory vector are weightedly calculated. After weighted calculation, the final memory data vector is obtained.

[0083] In an embodiment of the present invention, through context encoding, the model can consider the relationships between data elements, not just the information of a single element, thereby more comprehensively understanding the semantics of the data, extracting key information from the encoding vector and forming a memory vector, and highlighting the most valuable part of the data for the task. Through weighted combination, the detailed information of the current data contained in the encoding vector is organically combined with the key information or prior knowledge in the memory vector.

[0084] In the embodiment of the present invention, since the memory data integrates contextual information, when the task requirements or goals change, the model can adjust and adapt based on this rich information.

[0085] S4. Perform fine-grained analysis on the memory data to obtain analyzed data.

[0086] In the embodiment of the present invention, the fine-grained analysis refers to an in-depth, detailed and comprehensive analysis of the memory data, breaking down and understanding the data from multiple dimensions and multiple levels to obtain more accurate, specific and targeted information.

[0087] Specifically, fine-grained analysis no longer focuses solely on the overall characteristics of the data or macro-level information, but examines the data from multiple different angles, decomposing the data according to a certain hierarchical structure, and gradually deepening from coarser granularity to finer granularity, providing a richer and more accurate information basis for subsequent decision-making, prediction, knowledge discovery and other tasks.

[0088] In specific healthcare scenarios, doctors can obtain more comprehensive and accurate information about a patient's condition by performing fine-grained analysis of patient records, examination reports, imaging data, and other stored data. For example, in medical imaging analysis, fine-grained analysis can identify subtle pathological features in the image, such as the tumor's margins, density, and internal structure. This helps doctors more accurately determine the nature and stage of the tumor, providing strong support for accurate diagnosis.

[0089] In FinTech scenarios, fine-grained analysis of customer data, including financial transaction records, credit reports, and consumer behavior, enables a more comprehensive and accurate assessment of a customer's credit risk and probability of default. By analyzing detailed information such as a customer's transaction flow, repayment history, and consumer preferences, financial institutions can develop personalized credit ratings for their customers, rationally determine credit limits and interest rates, and mitigate credit risk.

[0090] In the embodiment of the present invention, the performing fine-grained parsing on the memory data to obtain parsed data includes:

[0091] Dividing the memory data into multiple independent segments, and setting timestamps and confidence labels for the independent segments one by one;

[0092] performing logical conflict correction on the independent segment according to the timestamp and the confidence tag to obtain a corrected segment;

[0093] Construct a causal relationship map based on the correction fragments;

[0094] The causal relationship map and the correction fragment are fused and analyzed to obtain analysis data.

[0095] In an embodiment of the present invention, the division refers to cutting or decomposing the overall memory data into multiple relatively independent and non-overlapping parts based on certain rules or standards, the setting refers to assigning a specific timestamp and confidence label to each independent fragment, and the logical conflict correction refers to checking, analyzing and adjusting the information in the independent fragment based on the timestamp and confidence label to eliminate the logical contradictions or irrationalities therein, so that the information of each independent fragment remains logically consistent and reasonable. The construction refers to analyzing the causal relationship between each piece of information therein according to the corrected independent fragment, and displaying these causal relationships in a graphical manner to form a causal relationship map. The fusion analysis refers to the process of organically combining and deeply analyzing the causal relationship map and the corrected fragment to fully explore the information value of both, thereby obtaining more comprehensive, accurate and in-depth analytical data.

[0096] Specifically, based on the type (such as text, images, audio, numbers, etc.) and structural characteristics of the memory data, a reasonable division method is determined, and the corresponding tools or algorithms are used to implement the division. For example, for text, relevant functions in the natural language processing library can be used, and for time series data, segmentation by time window can be achieved through programming; if the memory data itself carries time information, directly extract and assign an accurate time identifier to each independent fragment; if the data does not have clear time information, estimate and add a timestamp based on the data generation order, the time correlation of related events, etc., and comprehensively consider factors such as the source of the data (such as the confidence of data released by authoritative agencies is usually higher), the integrity of the data (whether there are missing values, etc.), and the consistency of the data (whether it is consistent with other related data), and determine a confidence value for each independent fragment. The value generally ranges from 0 to 1, and the confidence label can be determined through manual evaluation, machine learning model prediction, etc.

[0097] Furthermore, the system checks the chronological logical consistency of different independent fragments based on timestamps. Information in different fragments is compared based on confidence tags. Any inconsistencies between high-confidence and low-confidence fragments are considered potential conflict points. For example, in financial data, if high-confidence transaction records are inconsistent with low-confidence financial statement data, further analysis and in-depth investigation of the detected logical conflicts are required to determine whether they are caused by data entry errors, data transmission issues, data misunderstandings, and other reasons. Appropriate corrective measures are then taken based on the different causes of the conflict.

[0098] Furthermore, domain knowledge and experience are used, combined with the information in the correction fragments, to determine the causal relationship between events or factors. For example, in the medical field, the causal relationship between disease symptoms and causes is determined based on medical theory and clinical experience. In the financial field, the causal relationship between market fluctuations and economic policies is analyzed based on economic principles. Data mining algorithms (such as association rule mining, Bayesian networks, etc.) are used to mine potential causal relationship patterns from the correction fragments. The causes and results of the identified causal relationships are used as nodes in the causal relationship graph, and directed edges are used to represent the direction of the causal relationship (from cause to result). Attribute information is added to the nodes and edges. For example, nodes can contain detailed descriptions of events, time ranges, confidence levels, etc., and edges can be labeled with the strength of the causal relationship, correlation coefficients, etc. Graph databases or visualization tools (such as Graphviz, Neo4j, etc.) are used to combine nodes and edges into a causal relationship graph, and the layout is adjusted and beautified to make it easier to understand and analyze.

[0099] In detail, the nodes and edges in the causal relationship graph are associated and matched with the specific data in the correction fragment. The correspondence between the graph and the correction fragment is established through timestamps, key events and other information to ensure the consistency and coherence of the information. The structure and logic of the causal relationship graph are used to conduct in-depth analysis and interpretation of the data in the correction fragment. More detailed information is obtained from the correction fragment to verify and improve the causal relationship graph, discover new causal relationships or correct existing relationships, organize and summarize the results of the fusion analysis, extract key information and conclusions, and form structured analysis data.

[0100] In an embodiment of the present invention, a timestamp is set for each independent fragment, which can clarify the time sequence of the data, help analyze the changing trend of the data over time, and the time correlation between different fragments, and provide important time clues for subsequent causal relationship analysis. A causal relationship map is constructed to clearly display the causal relationship in the corrected fragment in a graphical manner, making complex causal relationships intuitive and easy to understand, making it easier for people to quickly understand and grasp the internal connection between various factors in the data and discover potential laws and patterns.

[0101] In the embodiments of the present invention, the in-depth fine-grained analysis of memory data can better meet the personalized needs of individuals and provide personalized products, services or suggestions based on the unique memory data characteristics of each user or object.

[0102] S5. Construct a multi-dimensional query vector for the parsed data, and combine the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response.

[0103] In an embodiment of the present invention, the construction refers to the process of creating a query vector that can describe and represent data from multiple dimensions based on the analyzed data through a specific method or algorithm, and the combined generation refers to the process of fusing and processing the analyzed data with the information of a preset external knowledge base, and constructing an intermediate response that can specifically meet user needs.

[0104] Specifically, core elements are extracted from the analyzed data to reflect the key features of the current situation. Based on the extracted and associated information, customized processing is performed according to the user's specific background, preferences and other factors. The real-time information from the analyzed data and the general knowledge in the knowledge base are seamlessly integrated. Natural language generation technology is used or certain speech templates are followed to convert the integrated information into coherent and fluent text.

[0105] In an embodiment of the present invention, constructing a multi-dimensional query vector for the parsed data and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response includes:

[0106] Extracting user intent chains, risk markers, and decision dependencies from the parsed data, and constructing a multi-dimensional query vector based on the user intent chains, risk markers, and decision dependencies;

[0107] Performing a multi-level search on the multi-dimensional query vector according to a preset external knowledge base to obtain a search result;

[0108] Calculating the scene relevance and dynamic allocation weight of the search results respectively;

[0109] The scene relevance is combined with the dynamically assigned weight to generate an intermediate response.

[0110] In an embodiment of the present invention, the extraction refers to finding and separating specific types of useful information from the parsed data, namely, user intent chains, risk markers, and decision dependencies; the construction refers to integrating these different types of information on the basis of extracting user intent chains, risk markers, and decision dependencies to form a multi-dimensional query vector; the multi-level retrieval refers to performing hierarchical and phased retrieval operations on the multi-dimensional query vector based on a preset external knowledge base, so as to gradually narrow the scope and obtain more accurate and relevant information; the calculation refers to analyzing the relevance of each retrieval result to the current user's scenario, and dynamically assigning weights to each retrieval result based on the different characteristics of the retrieval results and the user's real-time needs and circumstances; the combination generation refers to the process of integrating and processing the retrieval results based on the calculated scenario relevance and the dynamically assigned weights, thereby generating an intermediate reply.

[0111] Specifically, we conduct in-depth analysis of the parsed data to tease out the coherent intentions behind user behaviors, operations, or expressions. Natural language processing techniques (such as semantic analysis and intent recognition algorithms) are used to assist in extraction. Especially when the parsed data contains a large amount of textual information, we analyze the keywords and sentence structure in the text to infer user intent. We carefully review the parsed data, identify and label various possible risk factors, and determine the factors in the parsed data that play a key role in user decision-making. In the parsed data of investment decisions, decision dependencies may include market trend analysis, return on investment forecasts, and industry competition. The extracted user intent chains, risk markers, and decision dependencies are quantified and encoded, and this encoded information is combined into a multi-dimensional vector, with each dimension corresponding to a type of information (user intent chain, risk markers, decision dependencies) and its specific characteristics.

[0112] In detail, primary retrieval is based on the main features of the multi-dimensional query vector (such as the key intentions of the user intent chain) and conducts an extensive search in the preset external knowledge base. Intermediate retrieval combines the risk marker information in the multi-dimensional query vector to further screen the primary retrieval results. Advanced retrieval uses the decision dependencies in the multi-dimensional query vector to deeply filter and accurately match the intermediate retrieval results; for each retrieval result, the degree of matching between its content and the user's current scenario (reflected by the multi-dimensional query vector) is analyzed. Similarity calculation algorithms (such as cosine similarity, edit distance, etc.) can be used to quantify this matching degree, compare the information in the retrieval results with the user intent chain, risk markers and decision dependencies, and dynamically assign weights to them based on multiple factors of the retrieval results. Specifically, a weighted algorithm or machine learning model is used to determine the weight value of each retrieval result, making the weight distribution more scientific and reasonable.

[0113] Furthermore, the search results are screened based on the calculated scenario relevance and weight, giving priority to information with high scenario relevance and large weight, and excluding information with low relevance and small weight to ensure the quality and relevance of the reply content. The screened search results are integrated to remove duplicate or redundant parts. In addition, clauses and case fragments related to the current request are extracted from the external knowledge base, and the extracted content is embedded in the model input prompt word to generate a reply containing citation annotations. The source credibility and version timeliness of the cited content are verified, and finally an intermediate reply is generated.

[0114] In the embodiment of the present invention, a multi-dimensional query vector is constructed to convert complex user information into a computer-processable form, facilitate subsequent matching and retrieval with an external knowledge base, and improve the efficiency and accuracy of information processing.

[0115] In an embodiment of the present invention, the parsed data is combined with an external knowledge base to filter out the most relevant information from the knowledge base based on these personalized factors, so that the generated response is more tailored to the user's specific situation and avoids giving overly general and generalized answers.

[0116] S6. Optimize and adjust the intermediate response according to the pre-acquired feedback mechanism to obtain an optimization result.

[0117] In the embodiment of the present invention, the optimization adjustment refers to improving and perfecting the generated intermediate responses in various aspects according to a pre-set feedback mechanism, so as to make them of higher quality and better effect, and accurately meet user needs.

[0118] Specifically, a set of feedback rules and processes (feedback mechanism) set and prepared in advance are used to comprehensively review, improve and enhance the intermediate responses that have been initially generated, and ultimately obtain optimization results that are of higher quality, more practical, and more in line with user expectations.

[0119] In the embodiment of the present invention, optimizing and adjusting the intermediate response according to the pre-acquired feedback mechanism to obtain the optimization result includes:

[0120] Performing sensitivity screening on the intermediate responses to obtain initial screening results;

[0121] Cross-validating the initial screening results to obtain validation results;

[0122] Providing optimization suggestions on the verification results to obtain suggested results;

[0123] The intermediate response is optimized and adjusted according to the suggestion result to obtain an optimized result.

[0124] In an embodiment of the present invention, the sensitivity screening refers to a comprehensive inspection of the content of the intermediate replies, scanning the replies to see whether they contain words, expressions or topics involving sensitive information based on pre-set sensitive word libraries, sensitive topic classifications, and relevant policies, regulations, ethical standards and other standards. The cross-validation refers to the use of a variety of different methods or tools based on the initial screening results obtained by the sensitivity screening, or re-checking and confirming the initial screening results from multiple different angles.

[0125] Specifically, sensitive fields in the output results (such as ID card number and bank card number) are checked through regular expression matching, and the risk control model is called to evaluate the potential legal risks of decision recommendations. When high-risk operations are detected, the manual review process is triggered and automatic execution is frozen. The presentation format of the output content is adjusted according to the user terminal type (mobile / PC), and multi-option guided questions and answers are generated based on user interaction history. An interactive explanation module is attached to complex terms (such as clicking to expand the original text of the terms).

[0126] Furthermore, real-time collection is conducted of user ratings, error correction marks or manual review results of reply content, and user interaction behavior data, including the duration of the reply content, click hot zone distribution and subsequent operation paths, to reversely evaluate the effectiveness of the reply based on the actual business results triggered by the reply (such as claim approval rate, dispute resolution time limit).

[0127] Furthermore, high-value feedback samples (such as low-scoring replies) are input into the generative model for incremental training, focusing on optimizing the knowledge retrieval weight distribution strategy. When similar risks are detected repeatedly in business feedback (such as complaints caused by misunderstanding of terms), constraints are automatically added to the rule engine (such as forcibly inserting citations of the original terms), and the feedback scores of knowledge fragments after being cited are counted, and their sorting priority in the retrieval results is dynamically adjusted.

[0128] Furthermore, the differences between the original response and the optimized version are recorded, along with the basis for key optimization decisions and the quantitative evaluation of the optimized results, to generate optimization suggestions, thus providing a basis for the optimization results.

[0129] In this embodiment of the present invention, the feedback mechanism can help capture users' specific needs, preferences, and opinions on previous replies. Based on this information, intermediate replies can be optimized and adjusted to more accurately address user issues and better suit the language style of user preferences, thereby improving user satisfaction and service recognition, and strengthening the bond between users and service providers.

[0130] It can be seen that in the above scheme, for the target result business, the multimodal data is processed according to natural language standards to obtain standardized data; the standardized data is classified according to basic intent to obtain classified data; the classified data is contextually memorized to obtain memorized data; the memorized data is fine-grainedly parsed to obtain parsed data; a multi-dimensional query vector is constructed for the parsed data, and the multi-dimensional query vector is combined with a preset external knowledge base to generate an intermediate response; the intermediate response is optimized and adjusted according to a pre-acquired feedback mechanism to obtain an optimized result, and the optimized result is used as the target result. By deeply understanding user intentions and providing personalized services, user satisfaction and convenience are significantly improved. Users can obtain more accurate, timely, and demand-oriented responses during use, thereby enhancing the interactive experience.

[0131] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0132] In one embodiment, a data processing device based on human-computer interaction is provided, and the data processing device based on human-computer interaction corresponds one-to-one to the data processing method based on human-computer interaction in the above embodiment. Figure 3 As shown, the data processing device based on human-computer interaction includes a processing module 101, a classification module 102, a memory module 103, a parsing module 104, a construction module 105, a generation module 106, an adjustment module 107, and a serving module 108. The functional modules are described in detail as follows:

[0133] The processing module 101 is used to perform natural language standard processing on the multimodal data to obtain standardized data;

[0134] A classification module 102 is configured to perform basic intent classification on the standardized data to obtain classified data;

[0135] The memory module 103 is used to perform context memory processing on the classified data to obtain memory data;

[0136] The parsing module 104 is used to perform fine-grained parsing on the stored data to obtain parsed data;

[0137] A construction module 105 is configured to construct a multi-dimensional query vector for the parsed data;

[0138] A generating module 106, configured to combine the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response;

[0139] An adjustment module 107 is configured to optimize and adjust the intermediate response according to a pre-acquired feedback mechanism to obtain an optimization result;

[0140] As module 108, it is used to take the optimization result as the target result.

[0141] In one embodiment, when performing natural language standard processing on multimodal data to obtain standardized data, the processing module 101 is configured to:

[0142] Divide multimodal data into textual data and non-textual data;

[0143] Performing entity recognition on the text data to obtain recognition data;

[0144] Performing a preliminary intention screening on the text data to obtain preliminary screening data, and combining the recognition data and the preliminary screening data into comprehensive text data;

[0145] extracting structured tags and content categories of the non-text data, and merging the structured tags and content categories with the non-text data into comprehensive non-text data;

[0146] Normalized data is generated based on the combination of the integrated text data and the integrated non-text data.

[0147] In one embodiment, when the classification module 102 performs basic intent classification on the standardized data to obtain classified data, it is configured to:

[0148] Performing feature extraction on the standardized data to obtain feature data;

[0149] Mapping the feature data to a preset intent category space to obtain an intent classification corresponding to the feature data;

[0150] Calculate the probability of each intent classification and select the intent classification with the highest probability;

[0151] The selected intent classification is combined with the normalized data to form classification data.

[0152] In one embodiment, when the memory module 103 performs context memory processing on the classified data to obtain memory data, it is configured to:

[0153] Performing context encoding on the classification data to obtain an encoding vector;

[0154] Memory-outputting the encoding vector to obtain a memory vector;

[0155] The encoding vector and the memory vector are weightedly combined to obtain memory data.

[0156] In one embodiment, when the parsing module 104 performs fine-grained parsing on the memory data to obtain parsed data, it is configured to:

[0157] Dividing the memory data into multiple independent segments, and setting timestamps and confidence labels for the independent segments one by one;

[0158] performing logical conflict correction on the independent segment according to the timestamp and the confidence tag to obtain a corrected segment;

[0159] Construct a causal relationship map based on the correction fragments;

[0160] The causal relationship map and the correction fragment are fused and analyzed to obtain analysis data.

[0161] In one embodiment, when combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response, the generation module 106 is configured to:

[0162] Performing a multi-level search on the multi-dimensional query vector according to a preset external knowledge base to obtain a search result;

[0163] Calculating the scene relevance and dynamic allocation weight of the search results respectively;

[0164] The scene relevance is combined with the dynamically assigned weight to generate an intermediate response.

[0165] In one embodiment, when the adjustment module 107 optimizes and adjusts the intermediate response according to the pre-acquired feedback mechanism to obtain an optimization result, it is configured to:

[0166] Performing sensitivity screening on the intermediate responses to obtain initial screening results;

[0167] Cross-validating the initial screening results to obtain validation results;

[0168] Providing optimization suggestions on the verification results to obtain suggested results;

[0169] The intermediate response is optimized and adjusted according to the suggestion result to obtain an optimized result.

[0170] The present invention provides a data processing device based on human-computer interaction. For target result business, the device performs natural language standard processing on multimodal data to obtain standardized data; performs basic intent classification on the standardized data to obtain classified data; performs context memory processing on the classified data to obtain memory data; performs fine-grained analysis on the memory data to obtain parsed data; constructs a multidimensional query vector for the parsed data, and combines the multidimensional query vector with a preset external knowledge base to generate an intermediate response; optimizes and adjusts the intermediate response according to a pre-acquired feedback mechanism to obtain an optimized result, and uses the optimized result as the target result. By deeply understanding user intentions and providing personalized services, user satisfaction and convenience are significantly improved. Users can obtain more accurate, timely, and demand-oriented responses during use, thereby enhancing the interactive experience.

[0171] For the specific definition of a data processing device based on human-computer interaction, please refer to the definition of a data processing method based on human-computer interaction above, which will not be repeated here. The various modules in the above-mentioned data processing device based on human-computer interaction can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0172] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a data processing method based on human-computer interaction.

[0173] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a data processing method based on human-computer interaction.

[0174] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0175] Perform natural language standard processing on multimodal data to obtain standardized data;

[0176] Performing basic intent classification on the standardized data to obtain classified data;

[0177] Performing context memory processing on the classified data to obtain memory data;

[0178] Performing fine-grained analysis on the memory data to obtain analyzed data;

[0179] Constructing a multi-dimensional query vector for the parsed data, and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response;

[0180] The intermediate response is optimized and adjusted according to the feedback mechanism obtained in advance to obtain an optimization result, which is used as the target result.

[0181] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0182] Perform natural language standard processing on multimodal data to obtain standardized data;

[0183] Performing basic intent classification on the standardized data to obtain classified data;

[0184] Performing context memory processing on the classified data to obtain memory data;

[0185] Performing fine-grained analysis on the memory data to obtain analyzed data;

[0186] Constructing a multi-dimensional query vector for the parsed data, and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response;

[0187] The intermediate response is optimized and adjusted according to the feedback mechanism obtained in advance to obtain an optimization result, which is used as the target result.

[0188] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0189] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0190] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0191] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. If software tools or components other than those of the company appear in the application embodiments, they are merely used for illustration and do not represent actual use. Although the present invention has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above-mentioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A data processing method based on human-computer interaction, characterized in that: include: Perform natural language standard processing on multimodal data to obtain standardized data; Performing basic intent classification on the standardized data to obtain classified data; Performing context memory processing on the classified data to obtain memory data; Performing fine-grained analysis on the memory data to obtain analyzed data; Constructing a multi-dimensional query vector for the parsed data, and combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response; The intermediate response is optimized and adjusted according to the feedback mechanism obtained in advance to obtain an optimization result, which is used as the target result.

2. The data processing method based on human-computer interaction according to claim 1, characterized in that: The performing of natural language standard processing on the multimodal data to obtain standardized data includes: Divide multimodal data into textual data and non-textual data; Performing entity recognition on the text data to obtain recognition data; Performing a preliminary intention screening on the text data to obtain preliminary screening data, and combining the recognition data and the preliminary screening data into comprehensive text data; extracting structured tags and content categories of the non-text data, and merging the structured tags and content categories with the non-text data into comprehensive non-text data; Normalized data is generated based on the combination of the integrated text data and the integrated non-text data.

3. The data processing method based on human-computer interaction according to claim 1, characterized in that: The performing basic intent classification on the standardized data to obtain classified data includes: Performing feature extraction on the standardized data to obtain feature data; Mapping the feature data to a preset intent category space to obtain an intent classification corresponding to the feature data; Calculate the probability of each intent classification and select the intent classification with the highest probability; The selected intent classification is combined with the normalized data to form classification data.

4. The data processing method based on human-computer interaction according to claim 1, characterized in that: The performing context memory processing on the classified data to obtain memory data includes: Performing context encoding on the classification data to obtain an encoding vector; Memory-outputting the encoding vector to obtain a memory vector; The encoding vector and the memory vector are weightedly combined to obtain memory data.

5. The data processing method based on human-computer interaction according to claim 1, characterized in that: The performing fine-grained parsing on the memory data to obtain parsed data includes: Dividing the memory data into multiple independent segments, and setting timestamps and confidence labels for the independent segments one by one; performing logical conflict correction on the independent segment according to the timestamp and the confidence tag to obtain a corrected segment; Construct a causal relationship map based on the correction fragments; The causal relationship map and the correction fragment are fused and analyzed to obtain analysis data.

6. The data processing method based on human-computer interaction according to claim 1, characterized in that: Combining the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response includes: Performing a multi-level search on the multi-dimensional query vector according to a preset external knowledge base to obtain a search result; Calculating the scene relevance and dynamic allocation weight of the search results respectively; The scene relevance is combined with the dynamically assigned weight to generate an intermediate response.

7. The data processing method based on human-computer interaction according to claim 1, characterized in that: The optimizing and adjusting the intermediate response according to the pre-acquired feedback mechanism to obtain an optimization result includes: Performing sensitivity screening on the intermediate responses to obtain initial screening results; Cross-validating the initial screening results to obtain validation results; Providing optimization suggestions on the verification results to obtain suggested results; The intermediate response is optimized and adjusted according to the suggestion result to obtain an optimized result.

8. A data processing device based on human-computer interaction, characterized in that: include: A processing module is used to perform natural language standard processing on multimodal data to obtain standardized data; A classification module, configured to perform basic intent classification on the standardized data to obtain classified data; A memory module, configured to perform context memory processing on the classified data to obtain memory data; A parsing module, configured to perform fine-grained parsing on the memory data to obtain parsed data; A construction module, configured to construct a multi-dimensional query vector for the parsed data; A generation module, configured to combine the multi-dimensional query vector with a preset external knowledge base to generate an intermediate response; An adjustment module, configured to optimize and adjust the intermediate response according to a pre-acquired feedback mechanism to obtain an optimization result; As a module, it is used to take the optimization result as the target result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the data processing method based on human-computer interaction as claimed in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the data processing method based on human-computer interaction as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data processing method and system for metal material simulation database

    CN121560877A

  • A data processing method and system for a metal material simulation database

    CN121560877B