Intelligent reply method and device based on multiple modes, electronic equipment and medium

By integrating multimodal data features and knowledge graphs, the problem of insufficient utilization of image and audio information in traditional intelligent response methods is solved, resulting in more accurate intelligent responses and enhanced understanding and answering capabilities in professional domains.

CN120832397APending Publication Date: 2025-10-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510774349.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Traditional intelligent response methods in financial scenarios lack the utilization of multimodal information such as images and audio, resulting in low accuracy of intelligent responses and difficulty in accurately understanding the specific meaning of metaphorical concepts in professional fields.

Method used

By acquiring multimodal data on the target question, including text, image, and audio data, feature extraction and fusion are performed, and knowledge fusion is combined with a pre-built knowledge graph to generate accurate response content.

Benefits of technology

It improves the accuracy of intelligent responses, enabling precise answers to professional questions in different fields, providing rich background knowledge and visual and auditory aids to enhance users' understanding of professional knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832397A_ABST
    Figure CN120832397A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent reply method and device based on multiple modes, electronic equipment and a medium, belongs to the technical field of artificial intelligence, and is applied to financial scenes and medical scenes. The method comprises the steps of obtaining at least two kinds of modal data in a target problem text, a target problem image and target problem audio data of target problem information, performing feature extraction and feature fusion on the multi-modal data, performing knowledge fusion on the fused multi-modal features, problem information features and a target knowledge graph, and obtaining a target knowledge graph; and performing question reply generation on the target question based on the target fusion knowledge features. According to the embodiment of the invention, the target question information is replied based on the fused target question information, the fused multi-modal features and the target fused knowledge features of the target knowledge graph, so that rich professional background knowledge and visual and auditory auxiliary information can be provided, the understanding of a user on professional knowledge is enhanced, and the accuracy of intelligent reply is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and is applied to financial scenarios and medical scenarios, and in particular relates to a multi-modal based intelligent reply method and device, an electronic device and a medium. BACKGROUND

[0002] Traditional intelligent reply methods usually generate corresponding reply texts for professional field questions raised by users through natural language generation (NLG) technology. In a financial scenario, if the question text of the user is to explain the impact of the compound interest effect on long-term savings, the NLG technology retrieves the mathematical model, application cases and savings suggestions related to the "compound interest effect" and "long-term savings" from the financial knowledge base to generate a reply text that the compound interest effect accumulates through continuous initial investment and income, and the income becomes more significant over time. However, this method only refers to the text information in the knowledge base, lacks the use of information such as images and audios, and is also difficult to accurately understand the specific meanings of different field professional concept metaphors in the question, resulting in low accuracy of intelligent reply. Therefore, how to improve the accuracy of intelligent reply has become a problem to be solved. SUMMARY

[0003] The main purpose of the embodiments of the present application is to provide a multi-modal based intelligent reply method and device, an electronic device and a medium, which aims to improve the accuracy of intelligent reply.

[0004] To achieve the above purpose, a first aspect of the embodiments of the present application provides a multi-modal based intelligent reply method, which comprises:

[0005] Obtaining target question information and performing feature extraction on the target question information to obtain question information features;

[0006] Obtaining multi-modal question data of the target question information; wherein the multi-modal question data includes at least two modal data in target question text, target question image and target question audio data;

[0007] Performing text feature extraction on the target question text to obtain question text features;

[0008] Performing image feature extraction on the target question image to obtain question image features;

[0009] Performing speech recognition on the target question audio data to obtain target question audio text, and performing audio text feature extraction on the target question audio text to obtain audio text features;

[0010] fusing the multi-modal features to obtain a fusion multi-modal feature;

[0011] fusing the fusion multi-modal feature and the question information feature to obtain a fusion question feature;

[0012] fusing the fusion question feature and a pre-constructed target knowledge graph to obtain a target fusion knowledge feature;

[0013] generating a target reply content based on the target fusion knowledge feature and the target question information.

[0014] In some embodiments, the generating the target reply content based on the target fusion knowledge feature and the target question information comprises:

[0015] analyzing a question demand of the target question information to obtain a question demand category; the question demand category comprises a question text demand category, a question image demand category, a question animation demand category, and a question knowledge system demand category;

[0016] if the question demand category is the question text demand category, generating a target reply text based on the target fusion knowledge feature and the target question information to obtain a target reply text, and determining the target reply text as the target reply content;

[0017] if the question demand category is the question image demand category, screening a target reply image from the target question image according to the target reply text, and determining the target reply image and the target reply text as the target reply content;

[0018] if the question demand category is the question animation demand category, generating a target reply animation based on the target reply text to obtain a target reply animation, and determining the target reply animation and the target reply text as the target reply content;

[0019] if the question demand category is the question knowledge system demand category, visualizing a knowledge graph based on the target reply text to obtain a target reply knowledge graph, and determining the target reply knowledge graph and the target reply text as the target reply content.

[0020] In some embodiments, the generating the target reply text based on the target fusion knowledge feature and the target question information comprises:

[0021] obtaining user information of a target user, and obtaining historical dialogue semantic data of the target user;

[0022] perform feature extraction on the user information to obtain user information features, and perform feature extraction on the historical dialogue semantic data to obtain historical dialogue semantic features;

[0023] optimize the target fusion knowledge features based on the user information features and the historical dialogue semantic features to obtain optimized fusion knowledge features;

[0024] generate a question reply text based on the optimized fusion knowledge features to obtain the target reply text.

[0025] In some embodiments, the generating a question reply text based on the optimized fusion knowledge features to obtain the target reply text comprises:

[0026] perform reasoning on the target question information based on the optimized fusion knowledge features to obtain an initial reply text;

[0027] perform update detection on the target question information to obtain an updated question text;

[0028] update the historical dialogue semantic features based on the updated question text and the initial reply text to obtain updated historical dialogue semantic features;

[0029] update the optimized fusion knowledge features based on the updated historical dialogue semantic features to obtain updated fusion knowledge features;

[0030] perform sentiment recognition on the updated question text to obtain a question text sentiment;

[0031] update the initial reply text based on the updated fusion knowledge features and the question text sentiment to obtain the target reply text.

[0032] In some embodiments, the knowledge fusion of the fusion question features and a pre-constructed target knowledge graph to obtain target fusion knowledge features comprises:

[0033] obtaining target entities and target entity relationships of the target knowledge graph;

[0034] performing embedding processing on the target entities to obtain target entity vectors, and performing embedding processing on the target entity relationships to obtain target entity relationship vectors;

[0035] aligning the target entity vectors with the fusion question features to obtain aligned entity features, and aligning the target entity relationship vectors with the fusion question features to obtain aligned relationship features;

[0036] The aligned entity features, the aligned relation features, and the fused question features are fused to obtain the target fused knowledge features.

[0037] In some embodiments, the text feature extraction on the target question text obtains question text features, including:

[0038] The target question text is processed by word segmentation to obtain a segmented question text.

[0039] The segmented question text is subjected to entity recognition to obtain question text entities.

[0040] The question text entities are subjected to semantic dependency analysis to obtain question entity semantic relations.

[0041] A question text semantic network is constructed based on the question text entities and the question entity semantic relations.

[0042] The segmented question text is subjected to embedding processing to obtain segmented question features, the question text entities are subjected to embedding processing to obtain question text entity features, and the question text semantic network is subjected to embedding processing to obtain question semantic network features.

[0043] The segmented question features, the question text entity features, and the question semantic network features are fused to obtain the question text features.

[0044] In some embodiments, the image feature extraction on the target question image obtains question image features, including:

[0045] The target question image is subjected to visual basic feature extraction to obtain question visual basic features.

[0046] The target question image is subjected to visual semantic feature extraction to obtain question visual semantic features.

[0047] The question visual basic features and the question visual semantic features are fused to obtain the question image features.

[0048] To achieve the above object, a second aspect of the embodiments of the present application proposes an intelligent reply device based on multi-modal, which comprises:

[0049] A question information feature extraction module is configured to acquire target question information and extract features from the target question information to obtain question information features.

[0050] A multi-modal question data acquisition module is configured to acquire multi-modal question data of the target question information; wherein the multi-modal question data comprises at least two modal data in target question text, target question image and target question audio data;

[0051] A question text feature extraction module is configured to perform text feature extraction on the target question text to obtain question text features;

[0052] A question image feature extraction module is configured to perform image feature extraction on the target question image to obtain question image features;

[0053] A question audio feature extraction module is configured to perform speech recognition on the target question audio data to obtain target question audio text, and perform audio text feature extraction on the target question audio text to obtain audio text features;

[0054] A multi-modal feature fusion module is configured to perform multi-modal fusion on at least two features in the question text features, question image features and audio text features to obtain fused multi-modal features;

[0055] A multi-modal and question feature fusion module is configured to perform feature fusion on the fused multi-modal features and the question information features to obtain fused question features;

[0056] A knowledge fusion module is configured to perform knowledge fusion on the fused question features and a pre-constructed target knowledge graph to obtain target fused knowledge features;

[0057] A question reply generation module is configured to perform question reply generation on the target question information based on the target fused knowledge features to obtain target reply content.

[0058] To achieve the above object, a third aspect of embodiments of the present application provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.

[0059] To achieve the above object, a fourth aspect of embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.

[0060] The multi-modal based intelligent reply method and device, the electronic device and the medium provided in the application first acquire multi-modal question data of target question information, and perform feature extraction and multi-modal feature fusion on at least two modal data in the target question text, the target question image and the target question audio data, so that the fusion of different modal features is realized, rich and comprehensive multi-modal question data support is provided for subsequent intelligent reply, and the accuracy of the intelligent reply is facilitated to be improved subsequently; secondly, through the fusion of question information features and the fusion of multi-modal features, it can be ensured that the content generated subsequently can accurately solve the problem of the user, and through the knowledge fusion of the fusion question features and the target knowledge graph, the specific meanings of different field professional concept metaphors in the question are accurately understood by combining professional knowledge and structured information, and the depth of the subsequent intelligent reply is improved; finally, the target question information is replied and generated based on the target fusion knowledge features, which can accurately answer professional problems in different fields, also provides rich professional field background knowledge, visual and auditory auxiliary information, enhances the understanding of the user on the professional field knowledge, and significantly improves the accuracy of the intelligent reply. BRIEF DESCRIPTION OF DRAWINGS

[0061] Figure 1 is a flowchart of the multi-modal based intelligent reply method provided by the embodiment of the application;

[0062] Figure 2 is a flowchart of step S103 in Figure 1

[0063] Figure 3 is a flowchart of step S104 in Figure 1

[0064] Figure 4 is a flowchart of step S108 in Figure 1

[0065] Figure 5 is a flowchart of step S109 in Figure 1

[0066] Figure 6 is a flowchart of step S502 in Figure 5

[0067] Figure 7 is a flowchart of step S604 in Figure 6

[0068] Figure 8 is a structural schematic diagram of the multi-modal based intelligent reply device provided by the embodiment of the application;

[0069] Figure 9 is a hardware structure schematic diagram of the electronic device provided by the embodiment of the application.​​​​​​ DETAILED DESCRIPTION

[0070] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application.

[0071] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in a manner different from the module division in the device or the sequence in the flowchart. The terms "first", "second", and the like in the specification and claims and the above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0073] First, the meanings of several terms involved in the present application are analyzed:

[0074] Artificial intelligence (AI): is a new technical science to study, develop, simulate, extend and expand human intelligence, and is a branch of computer science. Artificial intelligence aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0075] The embodiments of the present application provide a multi-modal based intelligent reply method and device, an electronic device and a medium, aiming to improve the accuracy of intelligent reply.

[0076] The multi-modal based intelligent reply method and device, the electronic device and the medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the multi-modal based intelligent reply method in the embodiments of the present application is described.

[0077] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system of using digital computers or machine controlled by digital computers to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.

[0078] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0079] The multi-modal based intelligent reply method provided by the embodiments of the present application relates to the field of artificial intelligence. The multi-modal based intelligent reply method provided by the embodiments of the present application can be applied in a terminal, can also be applied in a server end, and can also be software running in a terminal or a server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server end can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, can also be configured as a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform; the software can be an application that implements the multi-modal based intelligent reply method, etc., but is not limited to the above forms.

[0080] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0081] It should be noted that in various specific embodiments of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the user, such as user information, the user's permission or consent will be obtained first, and the collection, use and processing of such data will comply with relevant laws, regulations and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the user, the separate permission or separate consent of the user will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the user, the necessary user-related data for enabling the embodiments of the present application to normally operate will be obtained.

[0082] Figure 1 is an optional flowchart of the intelligent reply method based on multi-modal provided by the embodiments of the present application, Figure 1 The method in can include but is not limited to steps S101-S109.

[0083] Step S101, obtaining target question information and performing feature extraction on the target question information to obtain question information features.

[0084] Step S102, obtaining multi-modal question data of the target question information; wherein the multi-modal question data includes at least two modal data in the target question text, target question image and target question audio data.

[0085] Step S103, performing text feature extraction on the target question text to obtain question text features.

[0086] Step S104, performing image feature extraction on the target question image to obtain question image features.

[0087] Step S105, performing speech recognition on the target question audio data to obtain target question audio text, and performing audio text feature extraction on the target question audio text to obtain audio text features.

[0088] Step S106, performing multi-modal fusion on at least two features in the question text features, question image features and audio text features to obtain fused multi-modal features.

[0089] Step S107, performing feature fusion on the fused multi-modal features and the question information features to obtain fused question features.

[0090] Step S108, performing knowledge fusion on the fused question features and a pre-constructed target knowledge graph to obtain target fused knowledge features.

[0091] Step S109, generating a question reply based on the target fused knowledge features to obtain target reply content.

[0092] The steps S101 to S109 shown in the embodiments of the present application first acquire multi-modal problem data of target problem information, and perform feature extraction and multi-modal feature fusion on at least two modal data in the target problem text, target problem image and target problem audio data, so as to realize fusion of different modal features, provide rich and comprehensive multi-modal problem data support for subsequent intelligent reply, facilitate subsequent improvement of the accuracy of intelligent reply; secondly, through fusion of problem information features and fusion of multi-modal features, it can be ensured that the generated content can accurately solve the user's problem, and through knowledge fusion of the fusion problem features and the target knowledge graph, the specific meanings of different field professional concept metaphors in the problem are accurately understood, and the depth of subsequent intelligent reply is improved; finally, based on the target fusion knowledge features, the target problem information is replied to generate, which can accurately answer professional problems in different fields, also provides rich professional field background knowledge, visual and auditory auxiliary information, enhances the user's understanding of professional field knowledge, and significantly improves the accuracy of intelligent reply.

[0093] In step S101 of some embodiments, specifically, the target problem information refers to the question text related to the professional field currently input by the target user.

[0094] For example, in the financial scenario, the target problem information can be: what is the specific effect of "compound interest effect"; in the medical scenario, the target problem information can be: why are the specific symptoms of "high blood pressure" and "diabetes" different.

[0095] Specifically, the problem information features refer to the problem semantic feature vector extracted from the target problem information.

[0096] Specifically, the target problem information can be segmented to obtain problem segmentation tokens, and the problem segmentation tokens can be subjected to word embedding operation through the Bert model to convert the target problem information into a structured vector representation, so as to capture the semantic information features of the target problem information.

[0097] In step S102 of some embodiments, specifically, the multi-modal problem data refers to a data set containing different types of data, and the multi-modal problem data can contain at least two modal data in the target problem text, target problem image and target problem audio data, which can be determined based on actual application scenarios.

[0098] Specifically, the target problem text refers to various types of text materials in the professional field, which can be collected from various professional field related websites, digital libraries and other platforms through network crawler technology.

[0099] For example, in the financial field, the target problem text can include but is not limited to financial regulations, financial research reports, financial news reports, financial analysis articles, financial papers, financial lecture texts, and financial cases; in the medical field, the target problem text can include but is not limited to medical research papers, clinical diagnosis and treatment guidelines, drug instructions, health popular science articles, and patient medical records.

[0100] Specifically, the target problem image refers to visual materials in the professional field, which can be collected by a camera, a drone, or directly obtained from a professional field website, a professional field database, or social media.

[0101] For example, in the financial field, the target problem image can include but is not limited to financial charts, financial trading places, financial conferences, and financial activities; in the medical field, the target problem image can include but is not limited to medical images (such as X-rays, CT, MRI, etc.), medical equipment images, medical surgery scenes, medical conferences, and medical activities.

[0102] Specifically, the target problem audio data refers to relevant audio information in the professional field, which can be collected by a recording device or directly obtained from a professional field website or social media.

[0103] For example, in the financial field, the target problem audio data can include but is not limited to financial news broadcasts, financial expert interviews, investment analysis audio, market trend interpretation, and financial lecture recordings; in the medical field, the target problem audio data can include but is not limited to doctor's diagnosis recordings, medical lecture audio, health popular science podcasts, heart and lung sound auscultation recordings, and patient symptom description recordings.

[0104] Specifically, for the target problem information, it is necessary to determine the multi-modal data that may be related to the target problem information, so as to facilitate subsequent generation of problem reply content combined with different modal data.

[0105] For example, in the financial field, for the financial problem of "how to understand the volatility of the financial market", the related problem text (such as research report text describing the volatility of the financial market), problem image (such as chart picture of the volatility of the financial market), and financial lecture audio (such as expert explanation of the volatility of the financial market) can be obtained; in the medical field, for the medical problem of "how to diagnose hypertension and its complications", the related problem text (such as clinical diagnosis and treatment guidelines), problem image (such as cardiovascular image), and medical expert interpretation audio (such as heart doctor's explanation of hypertension treatment) can be obtained.

[0106] In this embodiment, by obtaining the multi-modal problem data of the target problem information, the target problem can be analyzed from multiple modalities such as text, image and audio, solving the problem that the traditional method only analyzes from the text level and lacks the use of image and audio information related to professional field concepts, and providing rich and multi-dimensional data support for the accurate reply generated subsequently for the target problem information.

[0107] Please refer to Figure 2 In some embodiments, step S103 includes but is not limited to steps S201 to S206:

[0108] Step S201, performing word segmentation processing on the target problem text to obtain a segmented problem text.

[0109] Step S202, performing entity recognition on the segmented problem text to obtain a problem text entity.

[0110] Step S203, performing semantic dependency analysis on the problem text entity to obtain a problem entity semantic relationship.

[0111] Step S204, constructing a problem text semantic network based on the problem text entity and the problem entity semantic relationship.

[0112] Step S205, performing embedding processing on the segmented problem text to obtain a segmented problem feature, and performing embedding processing on the problem text entity to obtain a problem text entity feature, and performing embedding processing on the problem text semantic network to obtain a problem semantic network feature.

[0113] Step S206, fusing the segmented problem feature, the problem text entity feature and the problem semantic network feature to obtain a problem text feature.

[0114] In step S201 of some embodiments, specifically, the continuous target problem text string is segmented into individual independent lexical units.

[0115] For example, in a financial scenario, the target problem text "credit card overdue record on personal credit" is segmented into the segmented problem text "credit card / overdue / record / on / personal / credit"; in a medical scenario, the target problem text "relationship between hypertension and diabetes" is segmented into the segmented problem text "hypertension / and / diabetes / relationship".

[0116] Further, since the problem text often contains professional terms, the accuracy of word segmentation can be further optimized by combining a domain dictionary.

[0117] For example, in a financial scenario, for the text of "credit report", the domain dictionary can be combined to avoid incorrect segmentation of "credit report" into "credit / report", but to retain its complete semantics; in a medical scenario, for "how to prevent cardiovascular disease", the domain dictionary can be combined to avoid incorrect segmentation of "cardiovascular disease" into "heart / blood vessel / disease", but to retain its complete semantics.

[0118] In step S202 of some embodiments, specifically, the question text entity refers to an object or concept that has meaning and exists independently in the target question text, such as professional terms, characters, places, and events in a professional field.

[0119] Specifically, the part-of-speech tagging is performed on the segmented question text to obtain the part-of-speech of the segmented question text, and the named entity recognition model based on deep learning (such as the BERT model) is used to perform entity recognition on the segmented question text according to the part-of-speech of the segmented question text, to identify professional field concepts, characters, events, and other entities in the segmented question text.

[0120] For example, in a financial scenario, the part-of-speech tagging of the segmented text "credit card / overdue / record / to / personal / credit / the / impact" can obtain: credit card (N), overdue (V), record (N), to (P), personal (N), credit (N), the (U), and impact (N), and the entity recognition of the part-of-speech tagged text can obtain two credit evaluation entities: credit card overdue and personal credit; in a medical scenario, the part-of-speech tagging of the segmented text "high blood pressure / and / diabetes / relationship" can obtain: high blood pressure (N), and (C), diabetes (N), and relationship (N), and the entity recognition of the part-of-speech tagged text can obtain two medical concept entities: high blood pressure and diabetes. Wherein, N represents noun, C represents conjunction, V represents verb, P represents preposition, and U represents auxiliary word.

[0121] In step S203 of some embodiments, specifically, the question entity semantic relationship refers to the semantic relationship between the target entities, which can include but is not limited to causal relationship, equivalence relationship, hierarchical relationship, parallel relationship, metaphorical relationship, and structural relationship (such as equivalence relationship).

[0122] Specifically, the semantic dependency analysis is realized by the graph-based dependency syntax analysis technology combined with the semantic relationship rules of the professional field.

[0123] For example, in the financial field, the semantic relationship between credit card overdue and personal credit represents a causal relationship, and credit card overdue can be regarded as the cause and personal credit as the effect, because the semantic representation is that credit behavior (credit card overdue) will have a negative impact on credit evaluation (personal credit); in the medical field, the semantic relationship between hypertension and diabetes represents a correlation relationship, because they often coexist and influence each other, and they can be regarded as a parallel relationship, and the semantic representation of the relationship between hypertension and diabetes is that both are common chronic diseases, and there is a mutual influence between hypertension and diabetes.

[0124] In this embodiment, the semantic relationship between the question entity can reveal the deep semantic relationship implied in the question text by performing semantic dependency analysis on the question text entity, and is not limited to the literal meaning of the surface text, which helps to further understand the true meaning conveyed by the question text and facilitate subsequent improvement of the accuracy of intelligent reply.

[0125] In step S204 of some embodiments, specifically, the question text semantic network takes the question text entity as the node and the question entity semantic relationship as the edge to form a complete knowledge representation graph.

[0126] Specifically, the question text semantic network can be constructed based on a graph database (such as Neo4j).

[0127] For example, in the financial scenario, taking credit influencing factors as an example, a complete credit evaluation influencing factor network can be formed by connecting four core nodes of credit card overdue, loan default, query times, and debt ratio through relationship edges such as influence degree, duration, repair difficulty, and weight proportion; in the medical scenario, taking diseases such as hypertension, diabetes, and cardiovascular disease as an example, a complete medical health management network can be formed by connecting three core nodes of hypertension, diabetes, and cardiovascular disease through relationship edges such as correlation and complications.

[0128] In this embodiment, the question text semantic network is constructed based on the question text entity and the question entity semantic relationship, which can automatically establish the association between different doctrines through the semantic network, realize the systematic organization of professional field knowledge, and form a complete question knowledge system.

[0129] In step S205 of some embodiments, specifically, the word segmentation question feature refers to the question text semantic feature vector extracted from the word segmentation question text.

[0130] Specifically, the word segmentation token can be subjected to word embedding operation by the Bert model to convert the word segmentation question text into a structured vector representation, so as to capture the semantic information features of the word segmentation question text.

[0131] Specifically, the question text entity feature refers to the question text entity represented by a low-dimensional vector.

[0132] Specifically, the question text entity can be embedded into a word using the Bert model to map the question text entity into a low-dimensional vector space, thereby capturing the semantic information characteristics of the question text entity.

[0133] Specifically, the question semantic network feature refers to the semantic network of the question text represented by a vector, which is used to reflect the semantic connection and overall structural information between entities in the question text.

[0134] Specifically, the nodes (question text entities) and edges (question entity semantic relationships) in the question text semantic network can be encoded through a graph neural network (such as GNN) to obtain a feature vector representation of the question semantic network.

[0135] In this embodiment, by embedding the word segmentation question text, question text entities and question text semantic network respectively, the different levels of information in the question text can be converted into a unified vector form. The vector not only contains the semantic information of the question vocabulary, but also contains the semantic associations of the entities and the structural information of the semantic network. It can realize the subtle differences between professional field concepts in different contexts, which helps to improve the accuracy of the question text features.

[0136] In step S206 of some embodiments, specifically, question text features can be obtained by vectoring word segmentation question features, question text entity features, and question semantic network features.

[0137] In this embodiment, by integrating word segmentation features, entity features and semantic network features, the lexical information, semantic information and structural information of the text can be captured, which helps to more accurately understand the actual meaning of the question text.

[0138] Through steps S201 to S206, through a series of text segmentation, entity recognition, semantic dependency segmentation, semantic network construction and feature extraction processes, rich semantic information and semantic structure information can be extracted from the question text, providing professional domain semantic support for subsequent multimodal fusion and knowledge fusion.

[0139] See also Figure 3 In some embodiments, step S104 includes but is not limited to steps S301 to S303:

[0140] Step S301 : extracting basic visual features of the target problem image to obtain basic visual features of the problem.

[0141] Step S302: extract visual semantic features from the target question image to obtain question visual semantic features.

[0142] Step S303, the question visual basis feature and the question visual semantic feature are fused to obtain a question image feature.

[0143] In step S301 of some embodiments, specifically, the question visual basis feature is low-level visual information extracted from the question image, including but not limited to color, texture, shape, and composition, etc., which is used to reflect the basic visual attributes of the question image.

[0144] Specifically, the target question image can be convoluted by different size convolution kernels of a convolutional neural network (such as a CNN model) to extract multi-scale features of color, texture, shape, and composition of the target question image, and the features are respectively subjected to a maximum pooling operation to obtain a pooling feature, which is further activated by an activation function (such as a Softmax function) to output the question visual basis feature.

[0145] For example, in a financial scenario, a 7x7 convolution kernel is used to obtain the color scale feature of a credit score distribution graph (such as green for good and red for high risk), and a 3x3 convolution kernel is used to extract the fluctuation feature of a default rate curve (such as a steep rising pattern), then a pooling operation is performed to reduce the size of the feature map while still retaining the image color distribution feature and image texture feature information, and finally the pooled feature is activated by a softmax function to output the question visual basis feature containing the credit score color scale distribution and the default curve pattern; in a medical scenario, for analysis of a lung CT, a 7x7x7 convolution kernel is used to obtain the lung tissue density distribution feature, and a 3x3x3 convolution kernel is used to extract the fine texture feature of ground glass nodules, then a pooling operation is performed to retain the key image features, and finally the pooled feature is activated by a softmax function to output the question visual basis feature containing the lung tissue density and the nodular texture.

[0146] In this embodiment, by extracting the visual basis feature of the target question image, multi-scale and multi-level features of the target question image can be extracted, reflecting the basic visual attributes of the target question image.

[0147] In step S302 of some embodiments, specifically, the question visual semantic feature refers to a high-level visual feature vector extracted from the target question image, which can represent the connotation, meaning, and context in the professional field.

[0148] Specifically, in the pre-training stage, the target problem image is block encoded by the multi-head attention of the visual semantic understanding model (such as the ViT model) fine-tuned by the 500,000 labeled problem image dataset through transfer learning, to obtain a plurality of patch sequences, and the plurality of patch sequences are linearly projected to obtain an image feature vector. Further, different weights are assigned to different patches through the cross-patch attention mechanism, and the feature vectors with different weights are classified through a fully connected layer to obtain the visual semantic features of the problem.

[0149] For example, in a financial scenario, the credit score heat map is block encoded by the ViT model to obtain a 16x16 patch sequence, and the 16x16 patch sequence is linearly projected to obtain a 768-dimensional feature vector. Further, the attention weight of the high-risk credit aggregation area is increased by 0.3 through the cross-patch attention mechanism, while the interference of the background grid line is suppressed, and the image classification semantic label (such as high-risk credit user) of the feature vector with different weights is output through the fully connected layer. In a medical scenario, the CT image is block encoded by the ViT model to obtain a 16x16 patch sequence, and the 16x16 patch sequence is linearly projected to obtain a 768-dimensional feature vector. Further, the attention weight of the lesion area of the CT image is increased by 0.3 through the cross-patch attention mechanism, while the interference of the image background area is suppressed, and the image classification semantic label (such as tumor, inflammation) of the feature vector with different weights is output through the fully connected layer.

[0150] In this embodiment, by extracting the visual semantic features of the target problem image, the semantic information contained in the target problem image can be accurately recognized, which helps to further understand the deep semantics of the problem image in subsequent intelligent reply.

[0151] In step S303 of some embodiments, the visual basic features and the visual semantic features are spliced into a long vector, and then nonlinearly transformed by a multilayer perceptron (MLP, Multilayer Perceptron) to obtain the fused problem image features. The problem image features contain both low-level visual information and high-level semantic information of the problem image, and can more comprehensively represent the meaning of the target problem image.

[0152] Through steps S301 to S303, by extracting and fusing the visual basic features and the visual semantic features of the target problem image, rich visual information and visual semantic information contained in the target problem image can be extracted, which helps to further understand the deep connotation of the problem image in subsequent problem reply generation.

[0153] In step S105 of some embodiments, specifically, the target problem audio text refers to the text content extracted from the target problem audio data.

[0154] For example, in the financial scenario, the target question audio text can be the recognized text in the audio in which a financial consultant explains personal investment strategy, recording the detailed explanation and suggestion of the financial consultant on different investment products; in the medical scenario, the target question audio text can be the recognized text in the audio in which a medical expert explains Huangdi Neijing, recording the detailed explanation of the expert on Huangdi Neijing.

[0155] Specifically, the audio text feature refers to an audio text semantic feature vector extracted from the target question audio text.

[0156] Specifically, first, the pitch frequency, timbre, and prosody information of the target question audio signal are analyzed, and according to the pitch frequency, timbre, and prosody information, the target question audio signal is subjected to speech recognition through an end-to-end speech recognition model based on a Transformer architecture, to obtain the target question audio text.

[0157] Further, the target question audio signal can be segmented into multiple short time windows through short-time Fourier transform, and Fourier transform is performed on each short time window to analyze the pitch frequency of the audio signal; the audio signal can be converted to a mel frequency through a mel frequency cepstral coefficient, and the mel frequency cepstral coefficient is calculated to extract the timbre information of the target question audio signal; the speech rate, pause, and stress of the target question audio signal can be analyzed through wavelet transform to determine the prosodic features of the target question audio signal.

[0158] For example, in the financial scenario, by analyzing the intonation changes, speech rate, and pause positions of the financial consultant when explaining, combined with the text content of speech recognition, the key content emphasized by the financial consultant can be recognized; in the medical scenario, by analyzing the intonation changes, speech rate, and pause positions of the medical expert when explaining, combined with the text content of speech recognition, the key content emphasized by the expert can be recognized.

[0159] In this embodiment, the target question audio text is obtained by performing speech recognition on the target question audio data, and the audio text feature is extracted from the target question audio text, which not only converts the question audio data into a processable text form, but also completely retains the semantic information of the target question audio data, improving the accuracy of the target question audio feature extraction.

[0160] In step S106 of some embodiments, specifically, the fusion question feature refers to a composite feature vector that comprehensively represents the concept of a professional field by integrating the semantics of the question text, the visual symbol of the question image, and the question audio feature.

[0161] Specifically, at least two of the question text features, question image features and audio text features are weighted and summed through a cross-modal attention mechanism to obtain fused multi-modal features.

[0162] For example, in a financial scenario, the question text feature is a 256-dimensional vector [0.1, 0.2, …, 1.0] of the impact of overdue records in a personal credit report on credit scores, the weight is 0.6, the question image feature is a chart showing a personal credit report, and the chart shows detailed data of credit scores, overdue records and repayment history in a 1024-dimensional vector [0.2, 0.3, …, 0.9], and the weight is 0.3; the question audio text feature is a 512-dimensional vector [0.3, 0.4, …, 0.8] of a financial expert explaining “the impact of overdue records in a personal credit report on credit scores”, and the weight is 0.1. The question text feature, question image feature and audio text feature are normalized to obtain normalized question text features, question image features and audio text features, all of which are 1024-dimensional vectors. The normalized features are weighted and summed: 0.6x[0.1, 0.2, …, 1.0]+0.3x[0.2, 0.3, …, 0.9]+0.1x[0.3, 0.4, …, 0.8]=[0.15, 0.23, …, 0.83](1024 dimensions).

[0163] For example, in a medical scenario, the question text feature is a 256-dimensional vector [0.1, 0.2, …, 1.0] of “hypertension is a common chronic disease that needs long-term management”, the weight is 0.3, the question image feature is a medical image depicting a blood pressure measurement scene of a hypertension patient, and the image shows a sphygmomanometer and the patient's state in a 1024-dimensional vector [0.2, 0.3, …, 0.9], and the weight is 0.6; the question audio text feature is a 512-dimensional vector [0.3, 0.4, …, 0.8] of a medical expert explaining “hypertension is a common chronic disease that needs long-term management”, and the weight is 0.1. The question text feature, question image feature and audio text feature are normalized to obtain normalized question text features, question image features and audio text features, all of which are 1024-dimensional vectors. The normalized features are weighted and summed: 0.3x[0.1, 0.2, …, 1.0]+0.6x[0.2, 0.3, …, 0.9]+0.1x[0.3, 0.4, …, 0.8]=[0.18, 0.28, …, 0.92](1024 dimensions).

[0164] Specifically, if the core content is expressed in text, the semantic information weight of the text modality will be relatively large, the image modality also plays a role in auxiliary understanding, and the audio modality can help better understand the key points.

[0165] In this embodiment, by fusing at least two modal data in the target question text, the target question image and the target question audio data, the fusion of different modal features can be realized, the problem of lack of utilization of question image and question audio information in the traditional method is solved, and rich, comprehensive and multi-modal question data support is provided for subsequent intelligent reply.

[0166] In step S107 of some embodiments, specifically, the fusion question feature refers to the integrated feature vector representation of the fusion multi-modal feature and the question information feature.

[0167] Specifically, the fusion question feature can be represented by the following formula:

[0168] G = a x F + (1-a) x Q

[0169] Wherein, G represents the fusion question feature, a represents the weight coefficient, F represents the fusion multi-modal feature, and Q represents the question information feature.

[0170] Wherein, a can be dynamically adjusted according to the specific application scene, and a can be 0.7, which is not limited here.

[0171] In this embodiment, the fusion multi-modal feature and the question information feature are fused, which can make the fused feature more suitable for the question itself and provide more accurate feature basis for subsequent knowledge fusion.

[0172] Please refer to Figure 4 In some embodiments, step S108 includes but is not limited to steps S401 to S404:

[0173] Step S401, obtaining a target entity and a target entity relationship of a target knowledge graph.

[0174] Step S402, embedding processing is performed on the target entity to obtain a target entity vector, and embedding processing is performed on the target entity relationship to obtain a target entity relationship vector.

[0175] Step S403, aligning the target entity vector with the fusion question feature to obtain an aligned entity feature, and aligning the target entity relationship vector with the fusion question feature to obtain an aligned relationship feature.

[0176] Step S404, knowledge feature fusion is performed on the aligned entity feature, the aligned relationship feature and the fusion question feature to obtain a target fusion knowledge feature.

[0177] In step S401 of some embodiments, specifically, the target knowledge graph refers to a semantic network containing a target entity and a target entity relationship.

[0178] For example, in the financial scenario, taking the concept of "credit card overdue" as an example, the target knowledge graph not only stores the definition of "the behavior of not repaying the credit card bill on time", but also contains the direct influence relationship between "credit card overdue" and "credit score reduction", and the subordinate relationship with "bad credit record"; in the medical scenario, taking "diabetes" as an example, the target knowledge graph not only defines it as "a metabolic disease caused by insulin secretion deficiency", but also stores the complication relationship between "diabetes" and "retinopathy", and the subordinate relationship with "endocrine disease", etc.

[0179] Specifically, the target entity refers to an independent unit with clear semantics in the problem knowledge system, which is represented as a node with a type label in the knowledge graph.

[0180] For example, in the financial scenario, the target entity can be a credit report, a credit score, a debt-income ratio, etc.; in the medical scenario, the target entity can be hypertension, diabetes, retinopathy, endocrine disease, aspirin, etc.

[0181] Specifically, the target entity relationship refers to the logical association and semantic connection between entities, which is embodied as a directed edge connecting nodes in the knowledge graph.

[0182] For example, in the financial scenario, the target entity relationship can be the causal relationship between credit overdue and credit rating reduction; in the medical scenario, the target entity relationship can be the treatment relationship between aspirin for preventing cardiovascular disease or the causal relationship between high-salt diet and hypertension, etc.

[0183] In step S402 of some embodiments, specifically, the embedding processing refers to mapping the target entity and the target entity relationship into a low-dimensional vector space, so that similar target entities or target entity relationships are closer in the vector space.

[0184] For example, in the financial scenario, through a knowledge graph embedding model (such as TransE), "credit report" and "credit score" can be mapped into target entity vectors, and the "contains" entity relationship can also be mapped into a target entity relationship vector; in the medical scenario, through TransE, "hypertension" and "aspirin" can be mapped into target entity vectors, and the "treatment" entity relationship can also be mapped into a target entity relationship vector. In step S403 of some embodiments, specifically, by linearly projecting the target entity vector into a space consistent with the dimension of the fused problem feature vector, the aligned entity feature is obtained.

[0185] For example, if the fused problem feature is a 1024-dimensional vector and the target entity vector is a 512-dimensional vector, a 1024x512 matrix can be used to convert the 512-dimensional target entity vector into a 1024-dimensional aligned entity feature.

[0186] Specifically, the target entity relation vector is linearly projected into a space consistent with the dimension of the fusion question feature vector to obtain an aligned entity relation feature.

[0187] For example, if the target entity relation vector is a 256-dimensional vector, a 1024*256 matrix can be used to convert the 256-dimensional target entity relation vector into a 1024-dimensional aligned entity relation vector.

[0188] In this embodiment, by aligning the target entity vector and the target entity relation vector with the fusion question feature respectively, the feature vectors of different dimensions can be adjusted to the same space, avoiding information loss or fusion difficulty due to inconsistent dimensions in subsequent fusion.

[0189] In step S404 of some embodiments, specifically, the target fusion knowledge feature is a comprehensive vector representation integrating the aligned entity feature, the aligned relation feature, and the fusion question feature.

[0190] Specifically, the aligned entity feature, the aligned relation feature, and the fusion question feature can be vector spliced by a multi-layer perception (MLP) to obtain the target fusion knowledge feature.

[0191] In this embodiment, by fusing the aligned entity feature, the aligned relation feature, and the fusion question feature, the professional knowledge and structured information of different fields are combined, which can accurately understand the specific meaning of the metaphor of professional concepts in different fields.

[0192] Through steps S401 to S404, by fusing the question feature and the target knowledge graph, the professional knowledge and structured information of different fields are combined, which can accurately understand the specific meaning of the metaphor of professional concepts in different fields, solve the problem of accurately understanding the specific meaning of the metaphor of professional concepts in different fields in traditional methods, and improve the depth of subsequent intelligent reply.

[0193] Please refer to Figure 5 In some embodiments, step S109 includes but is not limited to steps S501 to S505:

[0194] Step S501, the target question information is analyzed for question demand to obtain a question demand category; the question demand category includes a question text demand category, a question image demand category, a question animation demand category, and a question knowledge system demand category.

[0195] Step S502, if the question demand category is the question text demand category, the target question information is replied based on the target fusion knowledge feature to generate a target reply text, and the target reply text is determined as the target reply content.

[0196] Step S503, if the question demand category is the question image demand category, a target reply image is filtered from the target question image according to the target reply text, and the target reply image and the target reply text are determined as the target reply content.

[0197] Step S504, if the question demand category is the question animation demand category, an animation is generated based on the target reply text to obtain a target reply animation, and the target reply animation and the target reply text are determined as the target reply content.

[0198] Step S505, if the question demand category is the question knowledge system demand category, a knowledge graph visualization is performed based on the target reply text to obtain a target reply knowledge graph, and the target reply knowledge graph and the target reply text are determined as the target reply content.

[0199] In step S501 of some embodiments, specifically, the question demand category can include a question text demand category, a question image demand category, a question animation demand category, and a question knowledge system demand category.

[0200] For example, in a financial scenario, if the target question information involves credit score standard interpretation, it is determined that the question demand category is a text reply demand; if the target question information involves credit score distribution trend, it is determined that the question demand category is a question image demand category, and a text-image reply is needed; if the target question information involves a complex credit evaluation process, it is determined that the question demand category is a question image demand category, and a credit evaluation process demonstration is needed; if the target question information involves a question knowledge system or concept structure association, it is determined that the question demand category is a target knowledge graph demonstration.

[0201] For example, in a medical scenario, if the target question information involves disease symptom description or drug usage instructions, it is determined that the question demand category is a text reply demand; if the target question information involves medical image interpretation or test index visualization analysis, it is determined that the question demand category is a medical image demand, and a text-image combined reply is needed; if the target question information involves a complex surgical procedure or treatment scheme deduction, it is determined that the question demand category is a diagnosis and treatment process demonstration demand, and a three-dimensional animation demonstration is needed; if the target question information involves a disease knowledge system or pathology mechanism association, it is determined that the question demand category is a medical knowledge graph demand, and a target knowledge graph demonstration of the relationship network from disease cause to symptom to treatment is needed.

[0202] In step S502 of some embodiments, if the question demand category is the question text demand category, it means that only a question text reply is needed, and the specific reply text generation process is as follows.

[0203] Please refer to Figure 6In some embodiments, step S502 includes but is not limited to steps S601-S604.

[0204] In step S601, user information of a target user is obtained, and historical dialogue semantic data of the target user is obtained.

[0205] In step S602, feature extraction is performed on the user information to obtain user information features, and feature extraction is performed on the historical dialogue semantic data to obtain historical dialogue semantic features.

[0206] In step S603, the target fusion knowledge features are optimized based on the user information features and the historical dialogue semantic features to obtain optimized fusion knowledge features.

[0207] In step S604, question reply text generation is performed on the target question information based on the optimized fusion knowledge features to obtain target reply text.

[0208] In step S601 of some embodiments, specifically, the user information can include but is not limited to basic information such as age, gender, and professional field interest preference of the user.

[0209] Specifically, the historical dialogue semantic data is a dialogue record of the user in a historical interaction period, and the historical dialogue semantic data is used to reflect the professional field interest and professional field theory demand of the user.

[0210] In step S602 of some embodiments, specifically, the user information features refer to vector representations extracted from the user information.

[0211] Specifically, the historical dialogue semantic features refer to semantic vectors of all professional field questions and corresponding reply contents of the user in the historical period extracted from the historical dialogue semantic data.

[0212] Specifically, the user information can be subjected to word embedding operation by a Bert model to map the user information into a low-dimensional vector space, and the historical dialogue semantic data can also be subjected to word embedding operation by the Bert model to map the historical dialogue semantic data into a low-dimensional vector space, so as to capture the historical dialogue semantic features.

[0213] In step S603 of some embodiments, specifically, the optimized fusion knowledge features refer to features of the target fusion knowledge features adjusted in combination with the user information features and the historical dialogue semantic features, which are used as a basis for subsequent reply data.

[0214] Specifically, different weights can be assigned to the user information features and the historical dialogue semantic features, and then weighted summation is performed on the target fusion knowledge features to obtain the optimized fusion knowledge features.

[0215] For example, if the weight of the user information feature is 0.4, the weight of the historical dialogue semantic feature is 0.3, and the weight of the target fusion knowledge feature is 0.3, then 0.4 x user information feature + 0.3 x historical dialogue semantic feature + 0.3 x target fusion knowledge feature = optimized fusion knowledge feature.

[0216] For example, in the financial field, if the user is a beginner and the historical dialogue semantic feature indicates that the user often asks basic financial noun concepts, then in subsequent replies, some complex financial concepts need to be simplified to make it easier for the user to understand; if the user is a researcher in the financial field, then more in-depth and professional answers need to be provided.

[0217] In this embodiment, the target fusion knowledge feature is optimized based on the user information feature and the historical dialogue semantic feature, which can adjust the current professional knowledge according to the user background and the historical dialogue semantic feature, helping to further understand the real needs of the user and facilitating the subsequent improvement of the accuracy of the target reply text generation.

[0218] Please refer to Figure 7 In some embodiments, step S604 includes but is not limited to steps S701 to S706:

[0219] Step S701, reasoning based on the optimized fusion knowledge feature to obtain an initial reply text.

[0220] Step S702, updating detection is performed on the target question information to obtain an updated question text.

[0221] Step S703, updating the historical dialogue semantic feature based on the updated question text and the initial reply text to obtain an updated historical dialogue semantic feature.

[0222] Step S704, updating the optimized fusion knowledge feature according to the updated historical dialogue semantic feature to obtain an updated fusion knowledge feature.

[0223] Step S705, performing sentiment recognition on the updated question text to obtain a question text sentiment.

[0224] Step S706, updating the initial reply text based on the updated fusion knowledge feature and the question text sentiment to obtain a target reply text.

[0225] In step S701 of some embodiments, specifically, the initial reply text refers to a preliminary textual answer to the user's question generated based on the target fusion knowledge feature.

[0226] Specifically, the optimized fusion knowledge feature is input into an intelligent reply model (such as a Transformer model) to generate possible reply texts in an autoregressive manner.

[0227] For example, in the financial scenario, the target question information is how to understand the credit score in the credit report. When the question is input into the Transformer model, the financial knowledge and user background information in the optimized fused knowledge features are used to generate the initial reply text: The credit score is a comprehensive evaluation value calculated based on the user's or enterprise's repayment records, debt level, credit history length, new credit account opening, and credit portfolio dimensions. The score range is usually between 300 and 850 points. The higher the score, the lower the credit risk. Good repayment records and moderate debt levels can help improve the credit score, while overdue records or excessive debt can lower the score.

[0228] For example, in the medical scenario, the target question information is how to manage high blood pressure. When the question is input into the Transformer model, the medical knowledge and user background information in the optimized fused knowledge features are used to generate the initial reply text: "High blood pressure" is a chronic disease characterized by persistently elevated blood pressure beyond the normal range. High blood pressure management includes lifestyle adjustments and drug treatment. Lifestyle adjustments include reducing salt intake, maintaining a healthy weight, and regular physical activity. Drug treatment involves selecting appropriate antihypertensive drugs based on the patient's specific condition. The importance of managing high blood pressure lies in reducing the risk of complications such as heart disease, stroke, and kidney disease.

[0229] In step S702 of some embodiments, specifically, cosine similarity calculation is performed on the new question features and question information features proposed by the user to detect whether the target question information has changed. If the similarity is less than the preset threshold, it indicates that the target question information has been updated.

[0230] For example, in the financial scenario, the user's initial question is "how to understand the credit score in the credit report." When the user further proposes a new financial question "how does the credit score affect loan approval," the similarity between the two questions is calculated to be 0.5, which is less than the preset threshold of 0.65, indicating that the question has been updated. The updated question text is "how does the credit score affect loan approval." In the medical scenario, the user's initial question is "how to manage high blood pressure." When the user further proposes a new medical question "how does high blood pressure management affect patients' daily life," the similarity between the two questions is calculated to be 0.55, which is less than the preset threshold of 0.65, indicating that the question has been updated. The updated question text is "how does high blood pressure management affect patients' daily life."

[0231] In step S703 of some embodiments, specifically, the decay rate of historical dialogue semantic features can be controlled by a gating recurrent unit (such as a Gated Recurrent Unit, GRU), such as a decay coefficient of 0.7 for financial or medical concepts not mentioned in more than 3 rounds of dialogue, and a memory reinforcement weight (such as 1.2-1.5 times) is set for core professional terms (such as “credit evaluation”, “cardiovascular disease” or “coronary heart disease”, etc.), further semantic association is suggested (such as automatically establishing concept association if the updated question text involves “risk management” and the historical dialogue semantic representation is “credit risk”, or automatically establishing concept association if the updated question text involves “cardiovascular disease” and the historical dialogue semantic representation is “coronary heart disease”), to finally determine the updated historical dialogue semantic features.

[0232] In step S704 of some embodiments, specifically, the updated fusion knowledge features refer to the integrated vector representation of the updated historical dialogue semantic features and the optimized fusion knowledge features.

[0233] Specifically, the updated historical dialogue semantic features and the optimized fusion knowledge features are vector-spliced to obtain the updated fusion knowledge features.

[0234] In step S705 of some embodiments, specifically, the question text emotion refers to the emotional tendency of the user in the inquiry process, including but not limited to the emotional states of seeking knowledge, confusion, anxiety, satisfaction, and excitement, etc.

[0235] For example, in a financial scenario, if the user asks “Will credit card overdue affect loan approval?”, the attention mechanism can distinguish the multi-dimensional emotions of the user's question, including worry emotion weight 0.88, anxiety emotion weight 0.82, and seeking knowledge emotion weight 0.63, and the highest weight worry emotion is taken as the question text emotion; in a medical scenario, if the user asks “How to manage high blood pressure?”, the attention mechanism can distinguish the multi-dimensional emotions of the user's question, including confusion emotion weight 0.82, seeking knowledge emotion weight 0.72, and anxiety emotion weight 0.6, and the highest weight confusion emotion is taken as the question text emotion.

[0236] Specifically, after determining the emotional state of the user, the tone and manner of the reply are adjusted, if the user shows confusion, a more patient and detailed manner is needed for answering; if the user shows satisfaction, encouragement and affirmation are needed; if the user has confusion in interpretation, more examples and explanations are needed to be provided actively.

[0237] In step S706 of some embodiments, the target reply text refers to the final text reply to the professional field question after professional knowledge and emotional adjustment, and the target reply text can be used as the target reply content.

[0238] Specifically, the updated fusion knowledge features, the sentiment of the question text, and the initial reply text can be vector-concatenated to obtain the target reply text.

[0239] For example, in the financial field, if the updated question is: As a credit novice, how do you understand the relationship between credit score and loan approval (the emotional characteristics detected are curiosity 0.72 and worry 0.65), first, a basic explanation is generated based on the target knowledge graph: credit score is the core indicator for banks to assess loan risks (the higher the score, the higher the pass rate), and it is positively correlated with loan approval (correlation 0.82); secondly, if it is detected that the user is a beginner in credit knowledge, the density of professional terms is reduced (such as from 40% to 25% of professional terms), and life metaphors are added: credit score is like a personal credit report card, and the bank decides whether to admit the student based on this score; further practical suggestions are inserted to target the worry emotion: Even if the current credit score is not high, it can be significantly improved by repaying on time for 6 months; finally, the target reply text is generated: It is recommended that you take the following measures: check your personal credit report to understand the current situation, and set repayment reminders to avoid overdue payments.

[0240] For example, in a medical scenario, if the updated question is: As a newly diagnosed patient, how do you understand the relationship between blood glucose monitoring and insulin dosage (the emotional characteristics detected are confusion 0.80 and anxiety 0.70), first, a basic explanation is generated based on the medical knowledge graph: blood glucose value is the key basis for adjusting insulin dosage (such as the ideal value of fasting blood glucose is 4.4-7.0mmol / L), and there is a dose-response relationship between the two (correlation 0.85); secondly, considering that the patient is newly diagnosed, the proportion of professional terms is reduced from 45% to 25%, and explained in a daily way: blood glucose monitoring is like a car dashboard, and insulin dosage is the throttle control; soothing content is added to address anxiety: initial fluctuations in values ​​are normal; and finally, the target response text is generated: it is recommended that you take the following measures: measure fasting and postprandial blood glucose at fixed times every day, record the correspondence between the values ​​and insulin dosage, and communicate with the attending physician every week to adjust the plan.

[0241] Through steps S701 to S706, first, the knowledge features are integrated to generate basic answers, and then the contextual semantics of the conversation are further captured through dynamic question update detection and conversation semantic tracking, which can respond to the user's new needs in a timely manner and ensure the contextual consistency of multiple rounds of conversation interactions; secondly, combined with emotion recognition technology, the reply can not only accurately convey professional knowledge, but also fit the user's emotional state and understanding level, which not only maintains the systematicness and integrity of the intelligent reply, but also achieves a personalized teaching effect similar to the master-apprentice dialogue, significantly improving the accuracy of the intelligent reply.

[0242] Through steps S601 to S604, the user information and user historical conversation semantics are comprehensively considered, the contextual semantic features of the conversation are further captured, a more accurate and personalized reply can be generated, and the accuracy of the intelligent reply is improved.

[0243] In step S503 of some embodiments, specifically, if the question demand category is a question image demand category, it indicates that the reply needs to be combined with image content.

[0244] Specifically, the target reply image includes but is not limited to professional field images, charts, and schematic diagrams, etc.

[0245] Specifically, the keywords of the target reply text are analyzed, images matching the keywords are searched from the target question image as the target reply image, the image position of the target reply image is identified, the target reply image is inserted into the target reply text, and the final target reply content is determined to realize visual display of the text and image.

[0246] For example, in a financial scenario, if the target question information is: how to improve my credit score, first, the reply text about credit score improvement is generated, and the credit score composition pie chart, credit score trend chart, and optimization repayment plan examples related to credit score improvement are determined based on the reply text; in a medical scenario, if the target question information is: how does a hypertension patient manage his own condition, first, the reply text about hypertension management is generated, and medical images related to hypertension management, such as a sphygmomanometer use schematic diagram, are determined based on the reply text to visually show how to measure blood pressure, and a lifestyle adjustment schematic diagram for hypertension management is made to help users more intuitively understand the specific measures for hypertension management.

[0247] In this embodiment, if the question demand category is a question image demand category, the target reply image is selected from the target question image according to the target reply text, and the target reply image and the target reply text are determined as the target reply content, which can provide diversified answer forms according to the user's demand, and the reply mode combining text and image not only improves the user experience, but also makes the dissemination of professional knowledge more accurate and vivid.

[0248] In step S504 of some embodiments, specifically, if the question demand category is a question animation demand category, it indicates that the reply needs to be combined with animation demonstration as the reply content.

[0249] Specifically, for the professional field content that needs to be demonstrated, the professional field content animation demonstration and VR / AR immersive experience content can be made through animation making software, virtual reality (VR) and augmented reality (AR) technology, and the animation demonstration and experience content and the target reply text are taken as the target reply content to realize the visualization of the animation demonstration and the reply text.

[0250] For example, in a financial scenario, for the question of how to make a monthly budget, a 3D animation demonstration about making a monthly budget is made to show how to classify income and expenditure, set savings goals, and adjust the steps and points of the budget, and an immersive experience of a virtual budget management scene can also be generated based on VR to visually display abstract financial concepts and budget management methods; in a medical scenario, for the medical question of how to perform cardiopulmonary resuscitation, a 3D animation demonstration about the operation steps of cardiopulmonary resuscitation is made to show the specific steps and points of cardiopulmonary resuscitation, and an immersive experience of a hospital emergency scene can also be generated based on VR to enable users to experience the emergency process and operation points in a virtual environment, and to visually display abstract medical operations and first aid knowledge.

[0251] In this embodiment, if the question demand category is the question animation demand category, animation generation is performed based on the target reply text to obtain a target reply animation, and the target reply animation and the target reply text are determined as the target reply content, which can further provide diversified answering forms according to the user's demand, and the reply mode combining text and animation not only improves the user experience, but also generates more accurate and vivid professional knowledge reply content.

[0252] In step S505 of some embodiments, specifically, if the question demand category is the question knowledge system demand category, it indicates that the question knowledge system demonstration needs to be combined as the target reply content.

[0253] Specifically, the target entity knowledge node is extracted from the target reply text, the associated subgraph is retrieved in the target knowledge graph, the associated graph is set in the professional field concept hierarchy vertically, the core professional knowledge node is automatically centered and enlarged, the direct derivation is realized by using the solid line table, and the indirect association relationship is realized by using the dashed line table, which can display the entity and the entity relationship in the form of nodes and edges, the user can view detailed information by clicking the node, and different parts of the knowledge graph can be browsed by zooming and dragging operation.

[0254] For example, in the financial scenario, taking personal credit assessment as an example, the generated target reply knowledge graph presents the main path of credit report, credit score and risk management, and the user can view the information of credit score definition, calculation method, influencing factors and the relationship with other financial concepts such as credit report, overdue record and repayment history by clicking the "credit score" node; in the medical scenario, taking hypertension management as an example, the generated target reply knowledge graph presents the main path of hypertension, lifestyle adjustment and drug treatment, and the user can view the information of hypertension definition, symptoms, treatment methods and the relationship with other medical concepts such as lifestyle, drug treatment and complications by clicking the "hypertension" node.

[0255] In this embodiment, if the question demand category is a question knowledge system demand category, knowledge graph visualization is performed based on the target reply text to obtain a target reply knowledge graph, and the target reply knowledge graph and the target reply text are determined as the target reply content, which can combine the reply text and the visualized knowledge graph system to display the positions of different professional field concepts in the entire professional system, so that the user can more easily learn the overall structural theory of the professional field concepts, and it is helpful to further improve the accuracy of the target reply content.

[0256] Through steps S501 to S505, first, the target question information is analyzed for question demand, and it is determined whether the question demand requires question text, image, animation or explanation of knowledge system; second, according to the specific demand category, the target fusion knowledge features are used to generate reply content in corresponding forms such as text, image, animation or knowledge graph visualization, which not only can obtain personalized professional field knowledge according to the user demand, but also can intuitively display the professional field knowledge by combining multiple media forms, enhance the user's understanding of the professional field knowledge, realize the personalized reply of the professional field knowledge, and further improve the accuracy of the intelligent reply.

[0257] The embodiment of the application first extracts and fuses features of at least two modalities in target question text, target question image and target question audio data by acquiring multi-modal question data of target question information, so as to realize fusion of different modalities, provide rich and comprehensive multi-modal question data support for subsequent intelligent reply, and facilitate subsequent improvement of accuracy of intelligent reply; secondly, by fusing question information features and multi-modal features, it can be ensured that the generated content can accurately solve the user's problem, and by fusing the fused question features and the target knowledge graph, the specific meaning of the metaphor of different field professional concepts in the question can be accurately understood, and the depth of subsequent intelligent reply is improved; finally, the target question information is replied to by generating question reply based on the target fused knowledge features, which can accurately answer professional questions in different fields, also provides rich professional field background knowledge, visual and auditory auxiliary information, enhances the user's understanding of professional field knowledge, and significantly improves the accuracy of intelligent reply.

[0258] Please refer to Figure 8 The embodiment of the application also provides an intelligent reply device based on multi-modalities, which can realize the intelligent reply method based on multi-modalities, and the device comprises:

[0259] A question information feature extraction module is configured to acquire target question information and extract features of the target question information to obtain question information features.

[0260] A multi-modal question data acquisition module is configured to acquire multi-modal question data of target question information, wherein the multi-modal question data comprises at least two modalities in target question text, target question image and target question audio data.

[0261] A question text feature extraction module is configured to extract text features of target question text to obtain question text features.

[0262] A question image feature extraction module is configured to extract image features of target question image to obtain question image features.

[0263] A question audio feature extraction module is configured to perform speech recognition on target question audio data to obtain target question audio text, and perform audio text feature extraction on the target question audio text to obtain audio text features.

[0264] A multi-modal feature fusion module is configured to fuse at least two features in question text features, question image features and audio text features to obtain fused multi-modal features.

[0265] The multi-modal and question feature fusion module is configured to fuse the fused multi-modal features and question information features to obtain fused question features.

[0266] The knowledge fusion module is configured to fuse the fused question features and a pre-constructed target knowledge graph to obtain target fused knowledge features.

[0267] The question reply generation module is configured to generate a question reply for the target question information based on the target fused knowledge features to obtain target reply content.

[0268] The specific implementation of the multi-modal based intelligent reply device is basically the same as that of the multi-modal based intelligent reply method described above, and thus will not be repeated here.

[0269] Embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the multi-modal based intelligent reply method described above. The electronic device can be any intelligent terminal, such as a tablet computer or a vehicle-mounted computer.

[0270] Please refer to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which includes:

[0271] The processor 901 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0272] The memory 902 can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 902 can store a processing system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 902 and called and executed by the processor 901 to implement the multi-modal based intelligent reply method of the embodiments of the present application.

[0273] The input / output interface 903 is configured to realize information input and output.

[0274] The communication interface 904 is configured to realize the communication interaction between the device and other devices. The communication can be realized in a wired manner (for example, a USB, a network cable, or the like) or in a wireless manner (for example, a mobile network, WIFI, Bluetooth, or the like).

[0275] The bus 905 is configured to transmit information between various components (for example, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904) of the device.

[0276] The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other through the bus 905.

[0277] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to realize the multi-modal based intelligent reply method.

[0278] The memory is a non-transitory computer readable storage medium, and can be used to store a non-transitory software program and a non-transitory computer executable program. In addition, the memory can include a high-speed random access memory, and can further include a non-transitory memory, for example, at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor. The remote memory can be connected to the processor through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0279] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0280] Those skilled in the art can understand that the technical solutions shown in the drawings do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than those shown in the drawings, or combine certain steps or different steps.

Claims

1. A multimodal intelligent reply method, characterized in that: The method comprises: acquiring target question information and performing feature extraction on the target question information to obtain question information features; acquiring multi-modal question data of the target question information; wherein the multi-modal question data comprises at least two modal data in target question text, target question image and target question audio data; performing text feature extraction on the target question text to obtain question text features; performing image feature extraction on the target question image to obtain question image features; performing speech recognition on the target question audio data to obtain target question audio text, and performing audio text feature extraction on the target question audio text to obtain audio text features; performing multi-modal fusion on at least two features in the question text features, question image features and audio text features to obtain fused multi-modal features; performing feature fusion on the fused multi-modal features and the question information features to obtain fused question features; performing knowledge fusion on the fused question features and a pre-constructed target knowledge graph to obtain target fused knowledge features; generating a question reply based on the target fused knowledge features to obtain target reply content.

2. The method of claim 1, wherein, The method comprises: performing question demand analysis on the target question information to obtain question demand categories; the question demand categories comprise question text demand categories, question image demand categories, question animation demand categories and question knowledge system demand categories; if the question demand category is the question text demand category, generating a reply text based on the target fused knowledge features to obtain a target reply text, and determining the target reply text as the target reply content; if the question demand category is the question image demand category, filtering a target reply image from the target question image according to the target reply text, and determining the target reply image and the target reply text as the target reply content; if the question demand category is the question animation demand category, generating an animation based on the target reply text to obtain a target reply animation, and determining the target reply animation and the target reply text as the target reply content; if the question demand category is the question knowledge system demand category, performing knowledge graph visualization based on the target reply text to obtain a target reply knowledge graph, and determining the target reply knowledge graph and the target reply text as the target reply content.

3. The method of claim 2, wherein, The method comprises: acquiring user information of a target user and historical dialogue semantic data of the target user; performing feature extraction on the user information to obtain user information features, and performing feature extraction on the historical dialogue semantic data to obtain historical dialogue semantic features; optimizing the target fusion knowledge feature based on the user information feature and the historical dialogue semantic feature, to obtain an optimized fusion knowledge feature; generating a question reply text based on the optimized fusion knowledge feature and the target question information, to obtain the target reply text.

4. The method of claim 3, wherein, The generating a question reply text based on the optimized fusion knowledge feature and the target question information, to obtain the target reply text, comprises: reasoning based on the optimized fusion knowledge feature and the target question information, to obtain an initial reply text; updating detection on the target question information, to obtain an updated question text; updating the historical dialogue semantic feature based on the updated question text and the initial reply text, to obtain an updated historical dialogue semantic feature; updating the optimized fusion knowledge feature based on the updated historical dialogue semantic feature, to obtain an updated fusion knowledge feature; emotional recognition on the updated question text, to obtain a question text emotion; updating the initial reply text based on the updated fusion knowledge feature and the question text emotion, to obtain the target reply text.

5. The method of claim 1, wherein, The knowledge fusion of the fusion question feature and a pre-constructed target knowledge graph, to obtain a target fusion knowledge feature, comprises: obtaining target entities and target entity relationships of the target knowledge graph; embedding processing on the target entities, to obtain target entity vectors, and embedding processing on the target entity relationships, to obtain target entity relationship vectors; aligning the target entity vectors with the fusion question feature, to obtain aligned entity features, and aligning the target entity relationship vectors with the fusion question feature, to obtain aligned real relationship features; knowledge feature fusion of the aligned entity features, the aligned real relationship features and the fusion question feature, to obtain the target fusion knowledge feature.

6. The method of claim 1, wherein, The text feature extraction on the target question text, to obtain a question text feature, comprises: word segmentation processing on the target question text, to obtain a segmented question text; entity recognition on the segmented question text, to obtain a question text entity; semantic dependency analysis on the question text entity, to obtain a question entity semantic relationship; construction of a question text semantic network based on the question text entity and the question entity semantic relationship; embedding processing on the segmented question text, to obtain a segmented question feature, embedding processing on the question text entity, to obtain a question text entity feature, and embedding processing on the question text semantic network, to obtain a question semantic network feature; fusion of the segmented question feature, the question text entity feature and the question semantic network feature, to obtain the question text feature.

7. The method according to any one of claims 1 to 5, characterized in that, The image feature extraction on the target question image, to obtain a question image feature, comprises: visual basic feature extraction on the target question image, to obtain a question visual basic feature; visual semantic feature extraction on the target question image, to obtain a question visual semantic feature; fusion of the question visual basic feature and the question visual semantic feature, to obtain the question image feature.

8. A multi-modal based intelligent reply device, characterized in that, The device comprises: a question information feature extraction module configured to obtain target question information and perform feature extraction on the target question information to obtain question information features; a multi-modal question data acquisition module configured to acquire multi-modal question data of the target question information, wherein the multi-modal question data comprises at least two modal data among target question text, target question image and target question audio data; a question text feature extraction module configured to perform text feature extraction on the target question text to obtain question text features; a question image feature extraction module configured to perform image feature extraction on the target question image to obtain question image features; a question audio feature extraction module configured to perform speech recognition on the target question audio data to obtain target question audio text, and perform audio text feature extraction on the target question audio text to obtain audio text features; a multi-modal feature fusion module configured to perform multi-modal fusion on at least two features among the question text features, the question image features and the audio text features to obtain fused multi-modal features; a multi-modal and question feature fusion module configured to perform feature fusion on the fused multi-modal features and the question information features to obtain fused question features; a knowledge fusion module configured to perform knowledge fusion on the fused question features and a pre-constructed target knowledge graph to obtain target fused knowledge features; a question reply generation module configured to generate a question reply for the target question information based on the target fused knowledge features to obtain target reply content.

9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal based intelligent reply method of any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-9. The computer program is executed by the processor to implement the multi-modal based intelligent reply method of any one of claims 1 to 7.

Citation Information

Cited By

  • Interaction method and device based on large model, intelligent agent and storage medium

    CN121501956A