A construction and application method of a multi-layer semantic understanding large model intelligent agent

Through multi-layer semantic understanding of large model agents, the problems of insufficient information fusion depth, poor model flexibility and low storage retrieval efficiency in multimodal data processing are solved, and high-precision, coherent and personalized semantic understanding and answer generation are achieved, which improves the real-time performance of various application scenarios.

CN119089931BActive Publication Date: 2025-07-11SHANGHAI ENTROPY INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411112965.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-07-11
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient depth of information fusion, poor model flexibility and adaptability, insufficient utilization of context information, lack of coherence and depth of generated answers, limited semantic differences processing capabilities of heterogeneous data, and low storage and retrieval efficiency in storage and retrieval.

Method used

Multi-layer semantic understanding of large model agents is adopted, and through multi-level semantic feature extraction, deep semantic understanding, dynamic hyperparameter optimization and hybrid neural symbol inference networks, combined with large language models and adaptive context generation algorithms, multi-modal data preprocessing, feature extraction, structured representation and vectorized storage are carried out to achieve high-precision semantic understanding and fusion.

Benefits of technology

It significantly improves the depth and accuracy of semantic understanding of multimodal data, enhances the flexibility and adaptability of the model, generates coherent and personalized answers, improves storage and retrieval efficiency, and meets the real-time needs of various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089931B_ABST
    Figure CN119089931B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing and applying a multi-layer semantic understanding large model intelligent agent, which includes the following steps: S1. Preprocess the input data; S2. Use a large language model and conversation history to analyze the preprocessed data, identify the user's intention and extract deep semantic features; S3. Extract key information from the data identified by intention recognition, perform structured representation, and adjust the model configuration through dynamic hyperparameter optimization; S4. Query the knowledge graph, associate and organize the extracted key information with existing data to generate a structured knowledge representation; S5. Vectorize the structured knowledge representation and store it in a vector database to support fast similarity search; S6. Combine the generated unified semantic representation and the large model to deeply analyze the user input, generate a coherent and logical personalized answer, and return it to the user. The present invention can support various application scenarios such as intelligent question answering, content generation, sentiment analysis, and multi-modal search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for constructing and applying a large model intelligent agent for multi-layer semantic understanding. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the continuous progress in fields such as natural language processing, computer vision, and audio processing, how to achieve high-precision semantic understanding and fusion in multi-modal data has become a current research hotspot and difficulty. Traditional single-modal processing methods can no longer meet the increasingly complex application requirements, and multi-modal fusion technology has emerged. By integrating information from multiple modalities such as text, images, and audio, it realizes a comprehensive understanding of complex scenarios. However, there are still many deficiencies in the existing technology in the fusion and semantic understanding of multi-modal data.

[0003] Although certain progress has been made in the field of multi-modal data processing and semantic understanding in the existing technology, there are still many defects. First, when processing multi-modal data, the existing technology usually adopts simple feature splicing or early fusion strategies, directly splicing features from different modalities together. This method ignores the deep interaction information between different modalities, resulting in insufficient depth of information fusion and inability to fully explore the potential associations of multi-modal data. Many existing technologies rely on fixed model architectures and hyperparameter settings, lacking flexibility and adaptability, and it is difficult to adapt to the changes in different tasks and data distributions. This makes the model show poor generalization ability when facing diverse actual application scenarios. The fixed model architecture cannot dynamically adjust hyperparameters to adapt to new data distributions, resulting in inconsistent performance of the model on different tasks and inability to effectively process complex data in actual applications.

[0004] In intelligent question-answering and dialogue systems, the existing technology often ignores the full utilization of context information and historical conversations, and the generated answers lack coherence and logic. This makes it difficult for the system to maintain consistent semantic fluency and logical continuity in multi-round conversations. The answers generated by the dialogue system are often isolated and do not fully utilize context information, resulting in unnatural and unintelligent answers that cannot meet the personalized needs of users. In the process of information extraction and semantic representation, there are often problems of information loss and insufficient expression in the existing technology. This is because traditional methods mostly rely on rule matching and shallow feature extraction, and cannot capture deep semantic relationships and complex entity associations. Rule matching methods are overwhelmed when facing changing and complex actual applications, and it is difficult to process a large amount of unstructured data, resulting in the accuracy and comprehensiveness of information extraction being affected.

[0005] In the storage and retrieval of high-dimensional vectors, existing technologies often exhibit low efficiency. Traditional databases and storage technologies are difficult to efficiently process large-scale high-dimensional vector data, resulting in slow similarity search and retrieval speeds and unable to meet the requirements of real-time applications. The storage and retrieval of high-dimensional data require more efficient indexing and optimization technologies. The research and application of existing methods in this regard are relatively lagging, affecting the overall performance of the system. In content generation and question-answering systems, the answers generated by existing technologies often lack depth and relevance and cannot provide valuable and meaningful content. This is mainly due to the lack of understanding of deep semantics and the effective integration of multimodal information, making it difficult for the generated content to meet the actual needs of users. The answers generated by existing methods are mostly templated and superficial, lacking a profound understanding of user questions and personalized responses, resulting in a poor user experience.

[0006] Multimodal data often comes from different sources and has heterogeneity. When existing technologies process heterogeneous data, they often lack effective fusion and parsing technologies, resulting in difficulty in eliminating semantic differences between data from different sources. The semantic differences in heterogeneous data affect the unified representation and comprehensive utilization of multimodal data. The processing capabilities of existing technologies in this regard are limited, and it is difficult to achieve comprehensive and consistent semantic understanding.

[0007] Therefore, how to provide a construction and application method for a large model intelligent agent with multi-layer semantic understanding is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose a construction and application method for a large model intelligent agent with multi-layer semantic understanding. The present invention makes full use of multimodal data fusion, deep learning, and semantic understanding technologies, and details the algorithms for achieving high-precision semantic understanding and fusion in multimodal data, with the following advantages: First, through multi-level semantic feature extraction and deep semantic understanding, the comprehensive processing capabilities of text, image, and audio data are effectively improved; Second, through dynamic hyperparameter optimization and hybrid neuro-symbolic reasoning networks, the flexibility and adaptability of the model are enhanced; In addition, the present invention also demonstrates wide applicability and high efficiency in various application scenarios such as intelligent question answering, content generation, sentiment analysis, and multimodal search.

[0009] A construction and application method for a large model intelligent agent with multi-layer semantic understanding according to an embodiment of the present invention includes the following steps:

[0010] S1. Preprocess the input data to adapt to the subsequent semantic understanding process. The preprocessing includes data cleaning and annotation, multimodal data alignment and fusion, data enhancement and extension;

[0011] S2. Use a semantic understanding agent to analyze the preprocessed data, identify the user's intentions and needs through a large language model and conversation history, and extract deep semantic features using a hybrid neuro-symbolic reasoning network;

[0012] S3. Extract key information from the identified intention data and the extracted deep semantic features, use information extraction technology to identify entities, relationships, and events in the data, form structured data, and adjust the model hyperparameter configuration to adapt to different data distributions and task requirements through dynamic hyperparameter optimization based on evolutionary strategies, generating optimized model parameters;

[0013] S4. Query the knowledge graph based on the extracted key information and optimized model parameters, associate and organize the information, and generate a structured knowledge representation;

[0014] S5. Vectorize the generated structured knowledge representation, store it in a vector database, and generate a unified semantic representation by weighted fusion of different modality features through a unified attention mechanism;

[0015] S6. Combine the generated unified semantic representation and the large model to deeply analyze the user input, generate relevant answers based on context information and conversation history, apply an adaptive context generation algorithm to provide coherent and logical outputs, and return the results to the user.

[0016] Optionally, the S1 specifically includes:

[0017] S11. Perform preliminary processing on the input data to remove noise and redundant information;

[0018] S12. Perform label annotation on the preliminarily processed data to provide a dataset for model training;

[0019] S13. Align data from different modalities in terms of time and space, enabling text, image, and audio data to be processed within a unified time axis and spatial range;

[0020] S14. Use a fusion algorithm to fuse the aligned multi-modal data and generate a comprehensive data representation of multi-modal features;

[0021] S15. Through data augmentation techniques, generate more training samples based on the generated comprehensive data representation to enrich the dataset;

[0022] S16. On the basis of data augmentation, further expand the scale of the dataset, optimize the model's adaptability under different data distributions by increasing diverse data samples.

[0023] Optionally, the S2 specifically includes:

[0024] S21. Receive and process the generated comprehensive data representation;

[0025] S22. Use a large language model to deeply analyze the comprehensive data representation and extract deep semantic features;

[0026] S23. Combine the conversation history and context information to perform context - correlation analysis on the comprehensive data representation, and identify the user's intentions and needs;

[0027] S24. Use a hybrid neuro - symbolic reasoning network to reason and analyze the deep semantic features, and extract high - level semantic information. The hybrid neuro - symbolic reasoning network combines the representation ability of neural networks and the reasoning ability of symbolic logic:

[0028]

[0029] where P(Q|F) represents the probability of query Q given feature F, λ i is a weight parameter, f i is a feature function, and Z is a normalization factor;

[0030] S25. Match the extracted high - level semantic information with the identified user intentions and needs to generate a structured semantic representation;

[0031] S26. Optimize the structured semantic representation through an adaptive knowledge embedding and dynamic logic reasoning mechanism to generate the final data representation of deep semantic features and user intentions.

[0032] Optionally, the specific steps of S3 include:

[0033] S31. Extract key information from the identified intention data and the extracted deep semantic features;

[0034] S32. Use information extraction techniques to identify entities, relationships, and events in the data:

[0035]

[0036] where Score(e i ,r,e j ) represents the score of entities e i and e j under relationship r, h r represents the embedding vector of relationship r, W r represents the weight matrix of relationship r, b r represents the bias vector of relationship r, g(e i ) and g(e j ) represent the embedding vectors of entities e i and e j respectively, [g(ei ) ; g(e j )] represents the concatenation operation of entity vectors;

[0037] S33. Structurally represent the identified entities, relationships, and events to form structured data;

[0038] S34. Use dynamic hyperparameter optimization based on evolutionary strategies to adjust the hyperparameter configuration of the model to adapt to different data distributions and task requirements:

[0039]

[0040] where θ t represents the model parameters of the t-th generation, α is the learning rate, β is the regularization parameter, and J(θ; z i ) is the loss function of the model on the sample z i and N is the total number of samples, and M is the number of samples for the regularization term;

[0041] S35. Generate the final optimized model parameters according to the structured data and the optimized model parameters.

[0042] Optionally, the specific steps of S4 include:

[0043] S41. Query the knowledge graph according to the extracted key information and the optimized model parameters, and obtain the data related to the key information from it;

[0044] S42. Associate and organize the newly extracted key information with the information obtained from the knowledge graph to generate a preliminary structured knowledge representation;

[0045] S43. Use semantic similarity calculation and graph matching algorithms to find the association points between the preliminary structured knowledge representation and the nodes and relationships in the existing knowledge graph:

[0046]

[0047] where Sim(x, y) represents the semantic similarity between two entities x and y, and α i represents the weight of the i-th feature, represents the cosine similarity of the i-th feature vector;

[0048] S44. Perform semantic fusion of information based on the found association points to unify the information representations from different sources. During the semantic fusion process, use a multi-layer perceptron MLP for non-linear transformation:

[0049]

[0050] where v fused represents the fused semantic vector, and βi represents the weight of the i-th information source, v i represents the semantic vector of the i-th information source;

[0051] S45. Employ ontology matching and semantic parsing techniques to address the semantic differences in heterogeneous data:

[0052]

[0053] Among them, Match(O1, O2) represents the matching degree between ontology O1 and ontology O2, C represents the set of concepts in the ontology, and sim ontology (c1, c1) represents the similarity between concepts c1 and c2, and depth(c) represents the depth of concept c in the ontology tree, which is used for weight adjustment;

[0054] S46. Associate the newly extracted information with the nodes and relationships in the existing knowledge graph after parsing and matching to form a unified structured knowledge representation, and use the generated unified structured knowledge representation as the final structured knowledge representation.

[0055] Optionally, the S5 specifically includes:

[0056] S51. Vectorize the generated structured knowledge representation. The vectorization process includes embedding generation, feature extraction, and dimensionality reduction. Embedding generation uses a pre-trained language model to represent the knowledge as a high-dimensional vector:

[0057] v embed = BCE_Embedding(K);

[0058] Among them, v embed represents the high-dimensional vector representation of knowledge K;

[0059] S52. Perform feature extraction. Through a unified attention mechanism, different modality features are weighted and fused to generate a unified semantic representation:

[0060]

[0061] Among them, h unified represents the feature vector after weighted fusion, α l is the weight coefficient in the attention mechanism, W l is the weight matrix of the l-th layer, h l is the feature vector of the l-th layer, and σ is the activation function;

[0062] S53. Perform dimensionality reduction. Through multi-level principal component analysis technology, the high-dimensional vector is reduced to a low-dimensional vector suitable for storage and retrieval:

[0063]

[0064] Among them, v reduced represents the vector after dimensionality reduction, and W i represents the principal component analysis matrix of the i-th level;

[0065] S54. Use a vector database to store vectorized knowledge data, optimize the storage and retrieval of high-dimensional vectors for similarity search and processing;

[0066] S55. Generate the final unified semantic representation by weighted fusion of different modality features through a unified attention mechanism:

[0067]

[0068] Among them, v final represents the final unified semantic representation, β i represents the weight coefficient in the attention mechanism, v i represents the vectors of different modality features, and σ is the activation function.

[0069] Optionally, the specific steps of S6 include:

[0070] S61. Combine the generated unified semantic representation v final and a large model to deeply analyze the user input. The analysis process includes comprehensive processing of the user input, context information, and conversation history to generate a context semantic vector:

[0071] h context = f Qwen (input, context, history);

[0072] Among them, h context represents the semantic representation combined with the user input, context information, and conversation history, and f Qwen represents the semantic vector generation function of the large model;

[0073] S62. Generate a relevant answer based on h context using an adaptive context generation algorithm for answer generation. The adaptive context generation algorithm uses the combination of a dynamic attention mechanism and a recurrent neural network to generate high-quality answers:

[0074]

[0075] Among them, r gen represents the generated answer, g context represents the adaptive context generation algorithm, W is the weight matrix, b is the bias vector, α t is the dynamic attention weight, is the hidden state of the RNN;

[0076] S63. Generate coherent and logical outputs through multi-level processing of the adaptive context generation algorithm:

[0077]

[0078] Among them, r final represents the finally generated coherent and logical output, α i represents the weight coefficient in the adaptive context generation algorithm, W i and b i represent the weight matrix and bias vector of the i-th layer respectively, σ is the activation function, and ELU is the exponential linear unit activation function;

[0079] S64. Return the finally generated coherent and logical output r final to the user to complete the entire semantic understanding and response generation process.

[0080] The beneficial effects of the present invention are:

[0081] Through multi-level semantic feature extraction and deep semantic understanding, the present invention can deeply understand the semantics of text, image, and audio data layer by layer, greatly improving the depth and accuracy of semantic understanding. By using a hybrid neural-symbolic reasoning network, which combines the representation ability of neural networks and the reasoning ability of symbolic logic, high-level semantic information is extracted, effectively solving the problems of information loss and insufficient feature expression. In addition, by adopting a dynamic hyperparameter optimization technique based on evolutionary strategies, the hyperparameter configuration of the model can be dynamically adjusted according to different data distributions and task requirements, significantly enhancing the flexibility and adaptability of the model, improving the generalization ability on different tasks and datasets, and solving the problem of poor generalization ability of traditional methods in dealing with variable application scenarios.

[0082] The present invention combines a large model and an adaptive context generation algorithm to deeply analyze user inputs, fully utilize context information and conversation history to generate personalized and context-related answers. Through the adaptive context generation algorithm, coherent and logical outputs are generated, improving the relevance and naturalness of the generated content and solving the problem of lack of coherence and logic in existing answers. In addition, through an innovative attention mechanism and a multi-level interaction network, features of different modalities are weighted and fused to generate a unified semantic representation, ensuring the deep fusion and complementarity of different modality information and solving the problem of insufficient inter-modal association in existing methods.

[0083] The present invention uses a vector database to store vectorized knowledge data, optimizing the storage and retrieval efficiency of high-dimensional vectors for fast similarity search and efficient processing. This not only improves the storage and retrieval efficiency of the system but also enhances the performance of real-time applications. Through these innovative technologies, it demonstrates wide applicability and high efficiency in various application scenarios such as intelligent question answering, content generation, sentiment analysis, and multimodal search, significantly improving the semantic understanding effect and practical application effect of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0085] Figure 1 is a flowchart of a method for constructing and applying a multi-layer semantic understanding large model agent proposed by the present invention;

[0086] Figure 2 is a flowchart of data preprocessing and semantic understanding in the present invention;

[0087] Figure 3 is a flowchart of vectorization of structured knowledge representation and response generation in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0088] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0089] Refer to Figures 1 - 3 , a method for constructing and applying a multi-layer semantic understanding large model agent, includes the following steps:

[0090] S1. Preprocess the input data to adapt to the subsequent semantic understanding process. The preprocessing includes data cleaning and annotation, multimodal data alignment and fusion, data enhancement and extension;

[0091] S2. Use a semantic understanding agent to analyze the preprocessed data. Through a large language model and conversation history, identify the user's intentions and needs, and use a hybrid neural-symbolic reasoning network to extract deep semantic features;

[0092] S3. Extract key information from the identified intention data and the extracted deep semantic features. Use information extraction technology to identify entities, relationships, and events in the data to form structured data. Through dynamic hyperparameter optimization based on evolutionary strategies, adjust the model hyperparameter configuration to adapt to different data distributions and task requirements, and generate optimized model parameters;

[0093] S4. Query the knowledge graph based on the extracted key information and optimized model parameters, associate and organize the information, and generate a structured knowledge representation;

[0094] S5. Vectorize the generated structured knowledge representation, store it in a vector database, and generate a unified semantic representation by weighted fusion of different modality features through a unified attention mechanism;

[0095] S6. Combine the generated unified semantic representation and the large model to deeply analyze the user input, generate relevant answers based on the context information and conversation history, apply an adaptive context generation algorithm to provide a coherent and logical output, and return the result to the user.

[0096] In this embodiment, the specific steps of S1 include:

[0097] S11. Perform preliminary processing on the input data to remove noise and redundant information;

[0098] S12. Perform label annotation on the preliminarily processed data to provide a dataset for model training;

[0099] S13. Align the data from different modalities in terms of time and space, so that text, image, and audio data can be processed within a unified time axis and spatial range;

[0100] S14. Use a fusion algorithm to fuse the aligned multi-modal data and generate a comprehensive data representation of multi-modal features;

[0101] S15. Through data augmentation techniques, generate more training samples based on the generated comprehensive data representation to enrich the dataset;

[0102] S16. On the basis of data augmentation, further expand the scale of the dataset, optimize the adaptability of the model under different data distributions by adding diverse data samples.

[0103] In this embodiment, the specific steps of S2 include:

[0104] S21. Receive and process the generated comprehensive data representation;

[0105] S22. Use a large language model to deeply analyze the comprehensive data representation and extract deep semantic features;

[0106] S23. Combine the conversation history and context information to perform context correlation analysis on the comprehensive data representation, and identify the user's intentions and needs;

[0107] S24. Use a hybrid neuro-symbolic reasoning network to reason about and analyze the deep semantic features, and extract high-level semantic information. The hybrid neuro-symbolic reasoning network combines the representation ability of neural networks and the reasoning ability of symbolic logic:

[0108]

[0109] Among them, P(Q|F) represents the probability of query Q given feature F, and λ i is a weight parameter, f i is a feature function, and Z is a normalization factor;

[0110] S25. Match the extracted high-level semantic information with the identified user intentions and requirements to generate a structured semantic representation;

[0111] S26. Optimize the structured semantic representation through an adaptive knowledge embedding and dynamic logical reasoning mechanism to generate the final data representation of the deep semantic features and user intentions.

[0112] In this embodiment, the specific steps of S3 are as follows:

[0113] S31. Extract key information from the identified intention data and the extracted deep semantic features;

[0114] S32. Use information extraction technology to identify entities, relationships, and events in the data:

[0115]

[0116] Among them, Score(e i , r, e j ) represents the score of entities e i and e j under relationship r. h r represents the embedding vector of relationship r, W r represents the weight matrix of relationship r, b r represents the bias vector of relationship r, g(e i ) and g(e j ) represent the embedding vectors of entities e i and e j . [g(e i ) ; g(e j )] represents the concatenation operation of entity vectors;

[0117] S33. Structurally represent the identified entities, relationships, and events to form structured data;

[0118] S34. Use dynamic hyperparameter optimization based on evolutionary strategies to adjust the hyperparameter configuration of the model to adapt to different data distributions and task requirements:

[0119]

[0120] Among them, θ t represents the model parameters of the t-th generation, α is the learning rate, β is the regularization parameter, and J(θ; z i ) is the loss function of the model on the sample z i , N is the total number of samples, and M is the number of samples of the regularization term;

[0121] S35. Generate the final optimized model parameters based on the structured data and the optimized model parameters.

[0122] In this embodiment, the S4 specifically includes:

[0123] S41. Query the knowledge graph according to the extracted key information and the optimized model parameters, and obtain the data related to the key information therefrom;

[0124] S42. Associate and organize the newly extracted key information with the information obtained from the knowledge graph to generate a preliminary structured knowledge representation;

[0125] S43. Use semantic similarity calculation and graph matching algorithms to find the association points between the preliminary structured knowledge representation and the nodes and relationships in the existing knowledge graph:

[0126]

[0127] Among them, Sim(x, y) represents the semantic similarity between two entities x and y, and α i represents the weight of the i-th feature, represents the cosine similarity of the i-th feature vector;

[0128] S44. Perform semantic fusion of information based on the found association points, unify the information representations from different sources, and use a multi-layer perceptron MLP for non-linear transformation during the semantic fusion process:

[0129]

[0130] Among them, v fused represents the fused semantic vector, β i represents the weight of the i-th information source, and v i represents the semantic vector of the i-th information source;

[0131] S45. Adopt ontology matching and semantic parsing technologies to solve the semantic differences of heterogeneous data:

[0132]

[0133] Among them, Match(O1, O2) represents the matching degree between ontology O1 and ontology O2, C represents the set of concepts in the ontology, and sim ontology (c1, c2) represents the similarity between concepts c1 and c2, and depth(c) represents the depth of concept c in the ontology tree, which is used for weight adjustment;

[0134] S46. Associate the newly extracted information with the nodes and relationships in the existing knowledge graph after parsing and matching to form a unified structured knowledge representation, and use the generated unified structured knowledge representation as the final structured knowledge representation.

[0135] In this embodiment, S5 specifically includes:

[0136] S51. Vectorize the generated structured knowledge representation. The vectorization process includes embedding generation, feature extraction, and dimensionality reduction. Embedding generation uses a pre-trained language model to represent knowledge as a high-dimensional vector:

[0137] v embed = BCE_Embedding(K);

[0138] Among them, v embed represents the high-dimensional vector representation of knowledge K;

[0139] S52. Perform feature extraction. By using a unified attention mechanism to weightedly fuse different modality features, generate a unified semantic representation:

[0140]

[0141] Among them, h unified represents the weighted fused feature vector, α l is the weight coefficient in the attention mechanism, W l is the weight matrix of the l-th layer, h l is the feature vector of the l-th layer, and σ is the activation function;

[0142] S53. Perform dimensionality reduction. By using a multi-level principal component analysis technique, reduce the high-dimensional vector to a low-dimensional vector suitable for storage and retrieval:

[0143]

[0144] Among them, v reduced i represents the matrix of the i-th level of principal component analysis;

[0145] S54. Use a vector database to store the vectorized knowledge data, optimize the storage and retrieval of high-dimensional vectors for similarity search and processing;

[0146] S55. Generate the final unified semantic representation by weighted fusion of different modality features through a unified attention mechanism:

[0147]

[0148] Among them, v final represents the final unified semantic representation, β i represents the weight coefficient in the attention mechanism, v i represents the vector of different modality features, and σ is the activation function.

[0149] In this embodiment, the S6 specifically includes:

[0150] S61. Combine the generated unified semantic representation v final and the large model to deeply analyze the user input. The analysis process includes comprehensive processing of the user input, context information, and dialogue history to generate a context semantic vector:

[0151] h context = f Qwen (input, context, history);

[0152] Among them, h context represents the semantic representation that combines the user input, context information, and dialogue history, and f Qwen represents the semantic vector generation function of the large model;

[0153] S62. Generate relevant answers based on h context using an adaptive context generation algorithm for answer generation. The adaptive context generation algorithm uses the combination of a dynamic attention mechanism and a recurrent neural network to generate high-quality answers:

[0154]

[0155] Among them, r gen represents the generated answer, g context represents the adaptive context generation algorithm, W is the weight matrix, b is the bias vector, α t is the dynamic attention weight, is the hidden state of the RNN;

[0156] S63. Generate a coherent and logical output through multi-level processing of the adaptive context generation algorithm:

[0157]

[0158] Among them, r final represents the finally generated coherent and logical output, α i represents the weight coefficient in the adaptive context generation algorithm, Wi and b i respectively represent the weight matrix and bias vector of the i-th layer, σ is the activation function, and ELU is the exponential linear unit activation function;

[0159] S64. Return the finally generated coherent and logical output r final to the user to complete the entire semantic understanding and response generation process.

[0160] Example 1:

[0161] To verify the feasibility of the present invention in implementation, the present invention is applied to a specific application process in a medical consultation scenario, aiming to solve the problems of multi-modal data fusion and high-precision semantic understanding. The data input by the user is natural language text, such as "How to reimburse for seeing a doctor in People's Hospital?". First, the system performs standardization processing on the input text, including synonym normalization and lemmatization. For example, "seeing a doctor" can be converted to "seeking medical treatment". This preprocessing helps to eliminate the ambiguity caused by language diversity and improve the accuracy of subsequent semantic understanding.

[0162] The preprocessed text is input into the semantic understanding agent. The agent uses a large language model for analysis, combines the conversation history and context information, and identifies that the user's intention is to consult the medical insurance policy. From the result of semantic understanding, the system extracts key information. In this embodiment, the named entity recognition technology is used to identify the entity of "People's Hospital". According to the extracted entity, the system queries the knowledge graph. Since there are many records associated with "People's Hospital" in the existing medical institution knowledge graph, the agent will consider the information incomplete at this time and trigger multi-round questions and answers. The agent will combine the relevant records in the knowledge graph and ask the user: "Which specific hospital do you mean by 'People's Hospital', such as xx People's Hospital, xx People's Hospital..." to further clarify the user's specific intention.

[0163] When the user determines the specific name of a medical institution, the system obtains the detailed information of that institution from the knowledge graph. For example, after the user enters "xx People's Hospital", the system confirms that the institution is a "tertiary medical institution". Subsequently, the agent updates the previous query processing result to "How to reimburse for medical treatment at xx People's Hospital (tertiary medical institution)?". After each step, the agent will evaluate whether the current information is complete enough. If the agent detects missing information, it will continue to guide the user to supplement the necessary information. For example, when the "medical insurance type" is missing, the agent will ask: "Do you want to consult urban employee medical insurance or urban and rural resident medical insurance?" After the user answers "urban employee medical insurance", the agent finds that the information is still not complete enough and continues to ask: "The reimbursement treatment for urban employee medical insurance varies according to different personnel categories. The personnel categories of urban employee medical insurance are divided into on-the-job personnel (including flexible employees), retired personnel, and veteran workers before the founding of the People's Republic of China. Which category of personnel do you want to consult?"

[0164] If the user answers "on-the-job employee", the agent will consider the information to be basically complete at this time. The system processes the current query statement as "How to reimburse for medical treatment of on-the-job employees' urban employee medical insurance at xx People's Hospital (tertiary medical institution)?". Next, the system quantizes the updated query statement to generate a corresponding vector representation and stores it in the vector database for quick retrieval. Finally, combining the medical insurance policy knowledge stored in the vector database, the agent generates the final summary text information and returns it to the user.

[0165] In a pilot application of a large hospital information system, the system can complete the whole process from user input to final response in an average time of less than 2 seconds, showing a significant improvement compared to the average response time of traditional methods (about 5 seconds). At the same time, the user satisfaction survey shows that more than 90% of the users are satisfied with the answers generated by the agent, believing that the answers are coherent and logically clear. Specific data show that during the pilot period of a certain top-three hospital in Beijing, the system processes about 500 user consultations per day, and the system response accuracy rate reaches more than 95%, greatly improving the information service efficiency of the hospital. Through this method of multi-modal fusion and high-precision semantic understanding, not only the workload of artificial customer service is reduced, but also the user experience is improved. In the future, this system can be promoted to more medical institutions and application scenarios to provide efficient and accurate intelligent services for more users. In summary, this embodiment demonstrates the superior performance and wide applicability of the present invention in practical applications, proving the significant advantages of the present invention in multi-modal data processing and semantic understanding.

[0166] Table 1 Data table of the overall performance of the system and user satisfaction

[0167]

[0168]

[0169] Table 2 Multi-round Q&A Processing Statistics

[0170] Information missing type Number of inquiries Accuracy of user answers Name of medical institution 150 times 98% Medical insurance type 200 times 95% Personnel category 150 times 97%

[0171] Table 3 Performance Data of Each Module of the System

[0172] Module Average processing time Processing accuracy Data preprocessing 0.5 seconds 99% Semantic understanding 1 second 95% Key information extraction 0.3 seconds 97% Knowledge graph query 0.5 seconds 96% Structured knowledge representation vectorization 0.2 seconds 98% Final response generation 0.5 seconds 94%

[0173] Based on the summary of the above table data, the present invention performs excellently in terms of overall performance and user satisfaction in actual applications. The average response time of the system is less than 2 seconds, the number of daily processed consultations reaches more than 500 times, the accuracy rate of the system response exceeds 95%, and the user satisfaction exceeds 90%. During the multi-round Q&A process, the number of inquiries about medical institution names, medical insurance types, and personnel categories are 150 times, 200 times, and 150 times respectively, and the accuracy rates of user answers are 98%, 95%, and 97% respectively. In terms of the performance of each module of the system, the average processing times of data preprocessing, semantic understanding, key information extraction, knowledge graph query, vectorization of structured knowledge representation, and final response generation are 0.5 seconds, 1 second, 0.3 seconds, 0.5 seconds, 0.2 seconds, and 0.5 seconds respectively, and the processing accuracy rates are 99%, 95%, 97%, 96%, 98%, and 94% respectively. These data indicate that the method of the present invention has significant advantages and innovation points in multi-modal data processing and semantic understanding, can efficiently process information and generate responses, and improves the user experience and system performance.

[0174] Through the above embodiments, the present invention realizes the whole process from user input to final response generation, and verifies the effectiveness and practicality of the method of the present invention. In actual applications, the processing speed and accuracy of the present invention have been significantly improved. In the embodiment, the natural language text input by the user such as "How to reimburse for seeing a doctor in People's Hospital?" is first subjected to standardization processing, including synonym normalization and lemmatization, to improve the accuracy of subsequent semantic understanding. The preprocessed text is input into the semantic understanding agent, analyzed using a large language model, combined with the conversation history and context information, and the user's intention to consult medical insurance policies is identified. The system uses named entity recognition technology to extract key information such as "People's Hospital", and queries the knowledge graph for multi-round Q&A, and finally confirms the specific medical institution and medical insurance type to generate an accurate consultation answer.

[0175] In summary, through innovative technologies, the present invention achieves deep integration of multimodal data and high-precision semantic understanding, effectively solving the problems existing in the prior art, such as insufficient depth of modal fusion, lack of flexibility and adaptability, insufficient utilization of context and historical information, information loss and inadequate expression, low storage and retrieval efficiency, and lack of depth and relevance in generating answers. The practical applications and data in the embodiments verify the effectiveness and practicality of the method of the present invention, demonstrating its wide applicability and high efficiency in various application scenarios such as intelligent question answering, content generation, sentiment analysis, and multimodal search, providing a more efficient and accurate solution for future intelligent systems.

[0176] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A method for constructing and applying a multi-layer semantic understanding large model intelligent agent, characterized in that, It includes the following steps: S1. Preprocess the input data to adapt to the subsequent semantic understanding process. The preprocessing includes data cleaning and annotation, multi-modal data alignment and fusion, data augmentation and extension; S2. Use a semantic understanding agent to analyze the preprocessed data. Through a large language model and conversation history, identify the user's intentions and needs, and use a hybrid neural-symbolic reasoning network to extract deep semantic features; S3. Extract key information from the identified intention data and the extracted deep semantic features. Adopt information extraction technology to identify entities, relationships and events in the data, form structured data, and adjust the model hyperparameter configuration to adapt to different data distributions and task requirements through dynamic hyperparameter optimization based on evolutionary strategies, and generate optimized model parameters; S4. According to the extracted key information and optimized model parameters, query the knowledge graph, associate and organize the information, and generate a structured knowledge representation; S5. Vectorize the generated structured knowledge representation and store it in a vector database. Generate a unified semantic representation by weighted fusion of different modal features through a unified attention mechanism; S6. Combine the generated unified semantic representation and a large model to deeply analyze the user input, generate relevant answers based on context information and conversation history, apply an adaptive context generation algorithm to provide coherent and logical outputs, and return the results to the user; The specific content of S4 includes: S41. According to the extracted key information and optimized model parameters, query the knowledge graph and obtain data related to the key information from it; S42. Associate and organize the newly extracted key information with the information obtained from the knowledge graph to generate a preliminary structured knowledge representation; S43. Use semantic similarity calculation and graph matching algorithms to find the association points between the preliminary structured knowledge representation and the nodes and relationships in the existing knowledge graph; Among them, Sim(x, y) represents the semantic similarity between two entities x and y, and α i represents the weight of the i-th feature, represents the cosine similarity of the i-th feature vector; S44. Based on the found association points, perform semantic fusion of information to unify the information representations from different sources. During the semantic fusion process, use a multi-layer perceptron (MLP) for non-linear transformation; Among them, v fused represents the fused semantic vector, and β i represents the weight of the i-th information source, and v i represents the semantic vector of the i-th information source; S45. Adopt ontology matching and semantic parsing techniques to solve the semantic differences of heterogeneous data; Among them, Match(O1, O2) represents the matching degree between ontology O1 and ontology O2, C represents the set of concepts in the ontology, and sim ontology (c1, c2) represents the similarity between concepts c1 and c2, and depth(c) represents the depth of concept c in the ontology tree, which is used for weight adjustment; S46. Associate the newly extracted information with the nodes and relationships in the existing knowledge graph after parsing and matching to form a unified structured knowledge representation, and use the generated unified structured knowledge representation as the final structured knowledge representation.

2. The intelligent agent construction and application method of a multi-layer semantic understanding large model according to claim 1, characterized in that The specific content of S1 includes: S11. Perform preliminary processing on the input data to remove noise and redundant information; S12. Perform label annotation on the preliminarily processed data to provide a dataset for model training; S13. Align data from different modalities in terms of time and space, so that text, image and audio data are processed within a unified time axis and spatial range; S14. Adopt a fusion algorithm to fuse the aligned multi-modal data and generate a comprehensive data representation of multi-modal features; S15. Through data extension technology, generate more training samples based on the generated comprehensive data representation to enrich the dataset; S16. On the basis of data expansion, further expand the scale of the dataset. By adding diverse data samples, optimize the adaptability of the model under different data distributions.

3. The construction and application method of a multi-layer semantic understanding large model intelligent agent according to claim 1, characterized in that The specific steps of S2 include: S21. Receive and process the generated comprehensive data representation; S22. Use a large language model to deeply analyze the comprehensive data representation and extract deep semantic features; S23. Combine the conversation history and context information to conduct context correlation analysis on the comprehensive data representation and identify the user's intentions and needs; S24. Use a hybrid neural-symbolic reasoning network to reason and analyze the deep semantic features and extract high-level semantic information. The hybrid neural-symbolic reasoning network combines the representation ability of neural networks and the reasoning ability of symbolic logic: Among them, P(Q|F) represents the probability of query Q given feature F, and λ i is the weight parameter, f i is the feature function, and Z is the normalization factor; S25. Match the extracted high-level semantic information with the identified user intentions and needs to generate a structured semantic representation; S26. Optimize the structured semantic representation through an adaptive knowledge embedding and dynamic logic reasoning mechanism to generate the final data representation of deep semantic features and user intentions.

4. A method for constructing and applying a multi-layer semantic understanding large model intelligent agent according to claim 1, characterized in that, The specific steps of S3 include: S31. Extract key information from the identified intention data and the extracted deep semantic features; S32. Use information extraction techniques to identify entities, relationships, and events in the data: Among them, Score(e i ,r,e j ) represents the score of entity e i and e j under the relation r, h r represents the embedding vector of relation r, W r represents the weight matrix of relation r, b r represents the bias vector of relation r, g(e i ) and g(e j ) represent the embedding vectors of entity e i and e j , [g(e i ) ; g(e j )] represents the concatenation operation of entity vectors; S33. Structurally represent the identified entities, relationships, and events to form structured data; S34. Use dynamic hyperparameter optimization based on evolutionary strategies to adjust the hyperparameter configuration of the model to adapt to different data distributions and task requirements: Among them, θ t represents the model parameters of the t-th generation, α is the learning rate, β is the regularization parameter, and J(θ; z i ) is the loss function of the model on the sample z i , N is the total number of samples, and M is the number of samples of the regularization term; S35. Generate the final optimized model parameters according to the structured data and the optimized model parameters.

5. The intelligent agent construction and application method of a multi-layer semantic understanding large model according to claim 1, wherein The specific steps of S5 include: S51. Vectorize the generated structured knowledge representation. The vectorization process includes embedding generation, feature extraction, and dimensionality reduction. Embedding generation uses a pre-trained language model to represent knowledge as high-dimensional vectors: vembed = BCE_Embedding(K); Among them, v embed represents the high-dimensional vector representation of knowledge K; S52. Conduct feature extraction. Through a unified attention mechanism, weight and fuse different modality features to generate a unified semantic representation: Among them, h unified represents the feature vector after weighted fusion, α l is the weight coefficient in the attention mechanism, W l is the weight matrix of the l-th layer, h l is the feature vector of the l-th layer, and σ is the activation function; S53. Conduct dimensionality reduction. Through multi-level principal component analysis technology, reduce the high-dimensional vectors to low-dimensional vectors suitable for storage and retrieval: Among them, v reduced represents the vector after dimensionality reduction, and W i represents the principal component analysis matrix of the i-th level; S54. Use a vector database to store the vectorized knowledge data and optimize the storage and retrieval of high-dimensional vectors for similarity search and processing; S55. Through a unified attention mechanism, weight and fuse different modality features to generate the final unified semantic representation: Among them, v final represents the final unified semantic representation, and β i represents the weight coefficient in the attention mechanism, and v i represents the vector of different modality features, and σ is the activation function.

6. The construction and application method of a multi-layer semantic understanding large model intelligent agent according to claim 1, characterized in that, The specific steps of S6 include: S61. Combine the generated unified semantic representation v final and the large model to deeply analyze the user input. The analysis process includes comprehensive processing of the user input, context information, and conversation history to generate a context semantic vector: hcontext = f Q wen(input, context, history); Among them, h context represents the semantic representation that combines user input, context information, and conversation history, and f Qwen represents the semantic vector generation function of the large model; S62. Based on h context Generate relevant answers using an adaptive context generation algorithm for answer generation. The adaptive context generation algorithm combines a dynamic attention mechanism and a recurrent neural network to generate high-quality answers: Among them, r gen represents the generated answer, g context represents the adaptive context generation algorithm, W is the weight matrix, b is the bias vector, α t is the dynamic attention weight, is the hidden state of the RNN; S63. Generate coherent and logical outputs through multi-level processing of an adaptive context generation algorithm: Among them, r final represents the finally generated coherent and logical output, and α i represents the weight coefficient in the adaptive context generation algorithm. W i and b i respectively represent the weight matrix and the bias vector of the i-th layer, σ is the activation function, and ELU is the exponential linear unit activation function; S64. Return the finally generated coherent and logical output r final to the user to complete the entire semantic understanding and response generation process.

Citation Information

Patent Citations

  • Information retrieval method and device

    CN117033657A

  • Customer service searching method and system based on intelligent large model

    CN118296132A