Intelligent interaction system for multi-modal data fusion

Through the intelligent interactive system of multimodal data fusion, the problem of insufficient multimodal data fusion in the existing system is solved, rapid information positioning and AI-generated text detection are realized, the practicality and flexibility of the system are improved, and the interaction needs of different users are met.

CN120764701AInactive Publication Date: 2025-10-10STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511293564.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing intelligent interactive systems lack the ability to fuse multimodal data, cannot effectively integrate multi-source data, have poor information acquisition capabilities, lack the ability to reverse monitor AI-generated text, and cannot meet the interaction needs of different users in different scenarios.

Method used

An intelligent interactive system that uses multimodal data fusion, including a user question-and-answer module, an AI text detection module, and a meeting minutes generation module, quickly locates information through a multimodal knowledge base, uses an AI text detection model to determine whether the text is AI-generated text, and generates meeting minutes.

Benefits of technology

It improves the practicality and flexibility of intelligent interaction, quickly locates answers and provides the source of answers, reduces manpower input, avoids omissions and errors, and meets the interaction needs of different users in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764701A_ABST
    Figure CN120764701A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion intelligent interaction system, and the system comprises a user question and answer module which is used for obtaining a target question input by a user, accurately retrieving a question answer and an answer source of the target question from a pre-constructed knowledge base, and feeding back the question answer and the answer source to the user; the AI text detection module is used for acquiring a to-be-detected text input by a user, judging whether the text input by the user is an AI generation text or not by utilizing a pre-constructed AI text detection model, and outputting a judgment result; the conference summary generation module is used for obtaining the conference record input by the user and generating the conference summary corresponding to the conference record, the practicability and flexibility of intelligent interaction can be improved, and the interaction requirements of different users in different scenes can be met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent interaction, and in particular to an intelligent interaction system based on multi-modal data fusion. BACKGROUND

[0002] With the continuous development of artificial intelligence technology, intelligent interaction systems are becoming more and more popular, which can help users or enterprises solve problems and obtain information.

[0003] The current intelligent interaction system still has some technical immaturity problems, for example, the knowledge base is stored in a single mode (only text or only file), which cannot effectively fuse multi-source data, has poor information acquisition ability, and usually does not have AI-generated text reverse monitoring capability, lacks universality and flexibility, and cannot meet the interaction needs of different users in different scenarios. SUMMARY

[0004] The embodiment of the present application provides an intelligent interaction system based on multi-modal data fusion, which can improve the practicability and flexibility of intelligent interaction and meet the interaction needs of different users in different scenarios.

[0005] The intelligent interaction system based on multi-modal data fusion of the embodiment of the present application comprises: A user question and answer module is configured to obtain a target question input by a user, accurately retrieve a question answer and an answer source of the target question from a pre-constructed knowledge base, and feed back the question answer and the answer source to the user. An AI text detection module is configured to obtain a to-be-detected text input by a user, determine whether the text input by the user is an AI-generated text by using a pre-constructed AI text detection model, and output a determination result. A conference minutes generation module is configured to obtain a conference recording input by a user, and generate conference minutes corresponding to the conference recording.

[0006] Further, the construction process of the knowledge base comprises: Collecting historical consultation records and knowledge original files within a preset time limit; the historical consultation records comprise historical consultation questions and historical question answers; Uploading the historical consultation records and the knowledge original files to the knowledge base; Extracting file information of the knowledge original files, generating a metadata table corresponding to each type according to the file type, and storing the file information in the metadata table; By recording the unique identifier of the knowledge original file in the metadata table, an association between the metadata table and the knowledge original file is established, so that the knowledge original file can be quickly located by the unique identifier during knowledge retrieval.

[0007] Further, the target question of the user input is acquired, and the question answer and answer source of the target question are accurately retrieved from a pre-constructed knowledge base, including: Acquiring the question input of the user, preprocessing the question input to generate a target question, and the form of the question input includes but is not limited to text input, picture input, file input and video input; Determine the question and answer prompt word for the target question, and accurately retrieve the question answer and answer source of the target question from the pre-constructed knowledge base in combination with the question and answer prompt word.

[0008] Further, the user question and answer module is also used for: Receiving the feedback mark of the user to the question answer, and triggering the corresponding processing logic for different feedback marks; the feedback mark is used to indicate the accuracy of the question answer; When the accuracy does not reach the preset accuracy requirement, the target question is pushed to the artificial answering module, and the retrieval algorithm is adjusted accordingly; When the accuracy reaches the preset accuracy requirement, no processing is performed.

[0009] Further, the AI text detection module includes: The model detection unit is used for inputting the to-be-detected text into a pre-trained AI generation detection model, so that the AI generation detection model detects the to-be-detected text, and outputs a probability value of the to-be-detected text being an AI generated text; The comprehensive judgment unit is used for calculating the text entropy value anomaly degree and the professional term density difference value of the to-be-detected text, and combining the probability value to obtain a judgment result of whether the to-be-detected text is an AI generated text.

[0010] Further, the training process of the AI generation detection model includes: An initial model is constructed by using a convolutional neural network and a bidirectional long short-term memory network; A certain number of positive samples and negative samples are divided into a training set, a validation set and a test set according to a preset ratio; wherein the positive sample is a text written by an artificial, and the negative sample is a text generated by using an AI tool; The training set is used to train the initial model, and the parameters of the initial model are adjusted by using a back propagation algorithm, so that the model can accurately distinguish AI generated text and artificial written text; The validation set is used to evaluate and optimize the initial model, and the model parameters with the best performance are selected; The test set is used to test the initial model, and finally a trained AI generation detection model is obtained.

[0011] Further, the computing the text entropy value anomaly degree and the professional term density difference value of the to-be-detected text comprises: segmenting the to-be-detected text into a plurality of paragraphs according to a preset character interval, and computing an information entropy of each paragraph; based on the information entropy and a pre-set reference entropy value, computing a paragraph entropy value anomaly degree of each paragraph and a text entropy value anomaly degree of the to-be-detected text; calculating a professional term density measured value of the to-be-detected text according to a professional term library, and calculating a professional term density difference value between the professional term density measured value and a pre-set minimum threshold value of the professional term density.

[0012] Further, the combining the probability value to obtain a judgment result of whether the to-be-detected text is an AI-generated text comprises: based on the text entropy value anomaly degree, the professional term density difference value, and the probability value, and pre-allocated anomaly degree weight, density difference value weight, and probability value weight, computing a comprehensive judgment value; according to the comprehensive judgment value, determining the judgment result of whether the to-be-detected text is an AI-generated text.

[0013] Further, the meeting minutes generation module comprises: a character conversion unit, configured to, when the question input is a meeting recording, invoke a speech-to-text model to convert the meeting recording into a character record, and proofread and polish the character record; a template matching unit, configured to match a target template from a meeting minutes template library in the knowledge base according to a pre-set meeting minutes template matching rule; an information extraction unit, configured to extract meeting key information from the processed character record by using a natural language processing technology; a minutes generation unit, configured to fill the meeting key information into the target template to generate a meeting minutes corresponding to the meeting recording.

[0014] Further, the meeting minutes template matching rule comprises: a local matching rule, configured to, according to local information of a meeting, preferentially match a template with completely consistent local information in the meeting minutes template library, and if there is no such template, sequentially match templates of a superior local; a version matching rule, configured to, when a plurality of templates meeting the local requirement are matched, preferentially select a template with the latest version; a content matching rule, configured to, when a plurality of templates with the same version are matched, compute a similarity between the character record and the template content, and select a template with the highest similarity.

[0015] Compared with the prior art, the intelligent interaction system of multi-modal data fusion provided by the embodiment of the application has the beneficial effects that: the user question and answer module is used to acquire a target question input by a user, accurately retrieve a question answer and an answer source of the target question from a pre-constructed knowledge base, and feed back the question answer and the answer source to the user; the module uses a multi-modal knowledge base, can quickly locate relevant information, greatly shortens the time for the user to acquire the answer, provides the answer source while giving the question answer, and meets the requirement of the user for information accuracy; the AI text detection module is used to acquire a text to be detected input by a user, judge whether the text input by the user is an AI generated text by using a pre-constructed AI text detection model, and output a judgment result; the module can quickly detect a large amount of text, mark out the content possibly generated by the AI, and improve the detection efficiency; the conference minutes generation module is used to acquire a conference recording input by a user, and generate conference minutes corresponding to the conference recording; the module can quickly generate the conference minutes, reduces the human input, and can avoid the omission and errors possibly occurring when the conference minutes are recorded by the human; the intelligent interaction system of multi-modal data fusion can improve the practicability and flexibility of the intelligent interaction, and meets the interaction requirements of different users in different scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical features of the embodiments of the application, the drawings needed to be used in the embodiments of the application will be briefly introduced as follows. Obviously, the drawings described below are only some of the embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0017] Figure 1 is a structural schematic diagram of one embodiment of the intelligent interaction system of multi-modal data fusion provided by the application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the application.

[0019] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a manner different from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to be limiting.

[0021] Referring to Figure 1 FIG. 1 is a structural schematic diagram of an embodiment of a multi-modal data fusion intelligent interaction system provided by the present application. As shown in Figure 1 The multi-modal data fusion intelligent interaction system comprises: a user question and answer module 11 configured to obtain a target question input by a user, accurately retrieve a question answer and an answer source of the target question from a pre-constructed knowledge base, and feed back the question answer and the answer source to the user; an AI text detection module 12 configured to obtain a text to be detected input by the user, determine whether the text to be detected is an AI generated text by using a pre-constructed AI text detection model, and output a determination result; a conference minutes generation module 13 configured to obtain a conference recording input by the user, and generate a conference minutes corresponding to the conference recording.

[0022] Specifically, the user question and answer module obtains a target question input by a user, accurately retrieves a question answer and an answer source of the target question from a pre-constructed knowledge base, and feeds back the question answer and the answer source to the user. The module uses a multi-modal knowledge base, can quickly locate relevant information, greatly shortens the time for the user to obtain the answer, provides the answer source while giving the question answer, and meets the requirement of the user for the accuracy of information.

[0023] The AI text detection module is configured to obtain a text to be detected input by the user, determine whether the text to be detected is an AI generated text by using a pre-constructed AI text detection model, and output a determination result. The module can quickly detect a large amount of text, mark out the content that may be AI generated, and improve the detection efficiency.

[0024] The conference minutes generation module is configured to obtain a conference recording input by the user, and generate a conference minutes corresponding to the conference recording. The module can quickly generate the conference minutes, reduces the human input, and can avoid the omission and errors that may occur when manually recording. The multi-modal data fusion intelligent interaction system of the present application can improve the practicability and flexibility of the intelligent interaction, and meet the interaction requirements of different users in different scenarios.

[0025] In an optional embodiment, the construction process of the knowledge base comprises: collecting historical consultation records and knowledge original files within a preset time limit, wherein the historical consultation records include historical consultation questions and historical question answers; uploading the historical consultation records and the knowledge original files to a knowledge base; extracting file information of the knowledge original files, generating a metadata table corresponding to each type according to the file type, and storing the file information into the metadata table; establishing an association between the metadata table and the knowledge original files by recording a unique identifier of the knowledge original files in the metadata table, so that the knowledge original files can be quickly located through the unique identifier during knowledge retrieval.

[0026] Specifically, historical consultation records and knowledge original files within a preset time limit are collected. The historical consultation records cover historical consultation questions and corresponding historical question answers, reflect the past users' demands for knowledge acquisition and the provided solutions, and are valuable information sources in the knowledge base. The knowledge original files can be various forms of documents, such as product manuals, technical research and development documents, and training materials. The preset time limit can be determined according to actual needs and knowledge update frequency, ensuring that the collected knowledge has a certain timeliness and covers enough historical information.

[0027] The historical consultation records and the knowledge original files are uploaded to the knowledge base by selecting a suitable knowledge base storage system. Key information of the knowledge original files of different types (such as documents, pictures, and videos) is extracted as metadata, including file name, file size, creation time, modification time, author, and keywords. For example, for an enterprise system document, the system name, issuing department, issuing date, and applicable scope can be extracted as metadata. Metadata tables corresponding to different file types are generated according to the file types. Each metadata table has a specific field structure for storing metadata of the corresponding type of file. For example, the metadata table of the document type may include file name, file path, document type, and page number. The extracted file information is accurately stored in the corresponding metadata table.

[0028] A unique identifier is generated for each knowledge original file, which can be a number, a letter, or a combination thereof, ensuring its uniqueness in the entire knowledge base. The unique identifier of each file is recorded in the metadata table, and the unique identifier is also recorded in the storage path or attribute of the knowledge original file. During knowledge retrieval, when relevant file information is found through metadata, the specific knowledge original file can be quickly located according to the unique identifier.

[0029] In an optional embodiment, the target question input by the user is obtained, and the question answer and answer source of the target question are accurately retrieved from the pre-constructed knowledge base, including: Obtaining the user's question input, preprocessing the question input to generate a target question; the form of the question input includes but is not limited to text input, picture input, file input and video input; Determine the question and answer prompt word for the target question, and accurately retrieve the question answer and answer source of the target question from the pre-constructed knowledge base in combination with the question and answer prompt word.

[0030] Specifically, first, the diversified question input of the user is obtained, including but not limited to text input, picture input, file input and video input, the question input is preprocessed, for text input, irrelevant characters, punctuation marks, special symbols and the like are removed, the text format is unified, for picture input, the OCR technology is used to accurately recognize the text in the picture, which is converted into an editable text format, for file input, according to different file formats, the corresponding analysis tool is used to extract the text content in the file, for video input, the speech in the video is first converted into text through speech recognition technology, and then the converted text is preprocessed, which is similar to the processing mode of the text input, after the preprocessing, the target question in a unified format is generated, which is convenient for subsequent retrieval operation.

[0031] The core words capable of accurately summarizing the theme and key points of the question are extracted from the target question, the natural language processing technology is used for semantic analysis of the target question, the intention and context relationship of the question are understood, according to the semantic analysis result, the question and answer prompt word more targeted is generated, the determined question and answer prompt word is used as a retrieval condition, and the pre-constructed knowledge base is searched, the knowledge base stores a large amount of knowledge information, including historical consultation records, knowledge original files and the like, through matching with the prompt word, the related knowledge content is found, in the retrieved related knowledge, the question answer most matched with the target question is extracted, and the source information of the answer is recorded, such as which knowledge original file the answer comes from, which part of the file, so as to enable the user to evaluate the reliability and authority of the answer.

[0032] The multi-modal data fusion intelligent interaction system of the embodiment of the application can accept question input in various forms such as text, picture, file and video, meets the use habits and scene requirements of different users, whether the question picture is uploaded by taking a picture through a mobile phone or the document file containing the question is uploaded, the system can process, improves the use convenience of the user, various forms of input are converted into a unified target question through preprocessing operation, the influence of the difference in input form on retrieval is reduced, at the same time, the question and answer prompt word is determined, the related information in the knowledge base can be more accurately positioned, so as to improve the accuracy of the retrieved answer, the retrieval result not only contains the question answer, but also provides the source information of the answer, so that the user can know the source of the answer, and meets the requirement of the user on the accuracy of knowledge.

[0033] In an alternative embodiment, the user question and answer module is further configured to: receive a feedback mark of the user on the question and answer, and trigger corresponding processing logic for different feedback marks; the feedback mark is used to indicate the accuracy of the question and answer; when the accuracy does not meet the preset accuracy requirement, push the target question to the artificial answering module, and adjust the retrieval algorithm accordingly; when the accuracy meets the preset accuracy requirement, no processing is performed.

[0034] Specifically, after the user question and answer module completes the question and answer retrieval and presents it to the user, it further receives the feedback mark of the user on the accuracy of the answer, and according to the accuracy indicated by the feedback mark, the module triggers different processing logic. The feedback mark can have multiple forms, such as a simple binary mark, for example, "accurate" and "inaccurate", or a multi-element hierarchical mark, for example, "very accurate", "relatively accurate", "generally accurate", "inaccurate", "very inaccurate", etc. The multi-element hierarchical mark can more accurately reflect the user's evaluation of the accuracy of the answer, and a clear feedback button or interactive area can be provided on the answer display interface to allow the user to conveniently and quickly perform feedback operations.

[0035] When the user feedback on the accuracy of the answer meets the preset standard, the system does not perform additional processing. When the user feedback on the accuracy of the answer does not meet the standard, the system forwards the target question to the artificial answering module, which can arrange professional customer service personnel, domain experts, etc. to answer the question and provide more accurate and detailed answers. The system will optimize and adjust the retrieval algorithm based on the user feedback and the results of the artificial answering, analyze whether the key word extraction is inaccurate, the semantic understanding is biased, or the knowledge in the knowledge base is incomplete, etc., and improve the retrieval algorithm, for example, optimize the key word weight distribution, enhance the ability of the semantic analysis model, and expand the knowledge base, etc., to improve the accuracy of subsequent retrieval answers.

[0036] Through user feedback, the embodiments of the present application can timely discover problems in the answers and push inaccurate questions to artificial processing, while adjusting the retrieval algorithm. Through such continuous optimization, the quality of the answers provided by the system will gradually improve, better meeting the needs of users, and the system can improve the retrieval algorithm, the knowledge base, etc., promoting the continuous development and improvement of the system.

[0037] In an alternative embodiment, the AI text detection module comprises: a model detection unit configured to input the text to be detected into a pre-trained AI-generated detection model, so that the AI-generated detection model detects the text to be detected and outputs a probability value of the text to be detected being an AI-generated text; The comprehensive judgment unit is used for calculating the text entropy abnormality degree and the professional term density difference value of the text to be detected, and combining the probability value to obtain a judgment result of whether the text to be detected is an AI generated text.

[0038] Specifically, the AI text detection module mainly consists of a model detection unit and a comprehensive judgment unit, which cooperate with each other to jointly complete the judgment task of whether the text to be detected is an AI generated text. The text to be detected is input into a pre-trained AI generation detection model. The AI generation detection model uses its internal complex neural network structure and the features learned from a large amount of training data to analyze and detect the text to be detected, and outputs a probability value of the text to be detected being an AI generated text. The probability value is a value between 0 and 1. The larger the value, the higher the possibility that the text is AI generated.

[0039] Text entropy is an index for measuring the uncertainty of text information. AI generated text may differ from human written text in terms of vocabulary distribution and information density, resulting in abnormal text entropy value. The professional term density refers to the ratio of the number of professional terms in the text to the total number of words in the text. The comprehensive judgment unit is used for calculating the text entropy abnormality degree and the professional term density difference value of the text to be detected. The probability value output by the model detection unit is comprehensively analyzed with the calculated text entropy abnormality degree and professional term density difference value to obtain a judgment result of whether the text to be detected is an AI generated text.

[0040] The embodiment of the application can analyze the text from different angles by comprehensively judging the probability value, the text entropy abnormality degree and the professional term density difference value of the model detection, thereby improving the accuracy of the detection result.

[0041] In an alternative embodiment, the training process of the AI generation detection model includes: An initial model is constructed using a convolutional neural network and a bidirectional long short-term memory network. A certain number of positive samples and negative samples are divided into a training set, a validation set and a test set according to a predetermined ratio. The positive samples are human written texts, and the negative samples are texts generated using AI tools. The training set is used to train the initial model, and the parameters of the initial model are adjusted through a back propagation algorithm to enable the model to accurately distinguish between AI generated texts and human written texts. The validation set is used to evaluate and optimize the initial model, and the model parameters with the best performance are selected. The test set is used to test the initial model, and a trained AI generation detection model is finally obtained.

[0042] Specifically, an initial model is constructed using a convolutional neural network (CNN) and a bidirectional long short-term memory network (BiLSTM). The CNN is good at capturing local features of text, such as specific word combinations and phrase patterns, and can quickly identify some prominent features in the text. The BiLSTM can consider the context information of the text and better understand the semantics and grammatical structure of the text by processing the text sequence in both directions. By combining the two, the advantages of each can be fully utilized to improve the model's ability to extract text features.

[0043] A certain number of positive and negative samples are collected, the positive samples are artificially written texts, and the negative samples are texts generated using AI tools. The positive and negative samples are divided into a training set, a validation set, and a test set according to a predetermined ratio, which can be 7:1:2 or 8:1:1. The training set is used for model training. The text data in the training set is input into the initial model, the model calculates the probability value of the input text being an AI-generated text, and then the loss function of the model is calculated through the backpropagation algorithm, and the parameters of the model are adjusted according to the gradient of the loss function. The backpropagation algorithm can effectively pass the error from the output layer to the input layer, allowing the model's parameters to be continuously optimized, thereby improving the model's classification ability. The training process usually requires multiple rounds, and each round uses all samples in the training set for a complete training. During the training process, learning rate, batch size, and other hyperparameters can be set to control the learning speed and stability of the model.

[0044] The model during the training process is evaluated using the validation set. Based on the evaluation results of the validation set, the parameters of the model are adjusted, for example, if the model's accuracy on the validation set is low, the learning rate can be adjusted, the complexity of the model can be increased, and so on. By continuously adjusting the parameters, the optimal model parameters are selected, and the final optimized model is tested using the test set. The data in the test set is independent of the training set and the validation set, and can more truly reflect the generalization ability of the model. Similarly, the evaluation indicators of the model on the test set are calculated to determine the final performance of the model. If the performance of the model on the test set meets the requirements, the model is the trained AI-generated detection model. If the performance does not meet the requirements, the model structure or training parameters need to be adjusted again, and the training and testing are performed again until a model that meets the requirements is obtained.

[0045] The model construction, data set division, training, evaluation and optimization, and testing of the embodiments of the present application form a systematic training process, ensuring the scientificity and effectiveness of the model training. The trained AI-generated detection model fully combines the advantages of CNN and BiLSTM, improves the model's ability to extract text features, and thus more accurately distinguishes AI-generated text from artificially written text.

[0046] In an optional embodiment, the text entropy value anomaly degree and the professional term density difference value of the to-be-detected text are calculated, including: The to-be-detected text is divided into a plurality of paragraphs according to a preset character interval, and the information entropy of each paragraph is calculated. Based on the information entropy and a pre-set reference entropy value, the paragraph entropy value anomaly degree of each paragraph and the text entropy value anomaly degree of the to-be-detected text are calculated. According to the professional term library, the professional term density measured value in the to-be-detected text is calculated, and the professional term density difference value between the professional term density measured value and a pre-set minimum threshold value of the professional term density is calculated.

[0047] Specifically, the to-be-detected text is divided into a plurality of paragraphs according to a preset character interval. The character interval can be set according to actual needs, for example, the text is divided according to natural paragraphs, fixed word interval, etc. Reasonable division mode is helpful to more accurately calculate the information entropy of each paragraph, so as to reflect the information characteristics of different parts of the text. For each divided paragraph, the information entropy thereof is calculated, and the paragraph entropy value anomaly degree of each paragraph is calculated based on the comparison between the calculated information entropy of each paragraph and a pre-set reference entropy value. The text entropy value anomaly degree of the to-be-detected text is obtained by comprehensively calculating the entropy value anomaly degrees of all paragraphs. The formula is as follows: ; ; Wherein, is the text entropy value anomaly degree, is the total number of paragraphs after text division, is the paragraph entropy value anomaly degree of the i-th paragraph, is the i-th paragraph, is the paragraph information entropy, is the artificial text reference entropy value.

[0048] According to the professional term library, the number of professional terms in the to-be-detected text is counted, and then the professional term density measured value is calculated. The professional term density measured value = the number of professional terms / the total number of text words. The calculated professional term density measured value is compared with a pre-set minimum threshold value of the professional term density, and the professional term density difference value is calculated.

[0049] In an optional embodiment, the probability value is combined to obtain a judgment result of whether the to-be-detected text is an AI-generated text, including: Based on the text entropy value anomaly degree, the professional term density difference value, the probability value, and pre-allocated anomaly degree weight, density difference value weight, and probability value weight, a comprehensive judgment value is calculated. According to the comprehensive judgment value, a judgment result of whether the to-be-detected text is an AI-generated text is determined.

[0050] Specifically, according to the importance of the text entropy abnormality, the professional term density difference value and the probability value in judging whether the text is an AI-generated text, weights, i.e., an abnormality weight, a density difference value weight and a probability value weight, are assigned to the three indexes. The text entropy abnormality, the professional term density difference value and the probability value are weighted and summed according to the assigned weights to obtain a comprehensive judgment value. The formula is as follows: ; wherein, is the comprehensive judgment value, is the probability value output by the AI generation detection model, is the text entropy abnormality, is an adjustment parameter, is the professional term density actual measurement value, is a minimum threshold value of the professional term density, is the probability value weight, is the abnormality weight, is the density difference value weight.

[0051] In an optional embodiment, the meeting minutes generation module comprises: a character conversion unit configured to, when the question input is a meeting recording, invoke a speech-to-text model to convert the meeting recording into a character record, and proofread and polish the character record; a template matching unit configured to match a target template from a meeting minutes template library in the knowledge base according to a pre-set meeting minutes template matching rule; an information extraction unit configured to extract meeting key information from the processed character record by using a natural language processing technology; a minutes generation unit configured to fill the meeting key information into the target template to generate a meeting minutes corresponding to the meeting recording.

[0052] Specifically, when detecting that the question input is a conference recording, a voice-to-text model is automatically called to convert the recording content into a text record, the converted text record is checked and corrected for possible voice recognition errors such as misspelling, semantic inconsistency, and the like, and is polished to make the text expression more fluent, accurate and professional, a pre-set conference minutes template matching rule is used to search and match a target template that fits the current conference from a conference minutes template library in a knowledge base, natural language processing techniques such as morphological analysis, syntactic analysis, semantic understanding, and the like are used to deeply analyze the text record, extract key information of the conference including a conference theme, conference participants, conference time, discussion points, decision results, and the like, the key information obtained by the information extraction unit is filled in according to the format and requirements of the target template, and the key information is placed in the corresponding position, so that a conference minutes corresponding to the conference recording is generated, which is complete in structure and accurate in content.

[0053] The embodiment of the present application can ensure that the generated conference minutes meet the pre-set specifications and requirements, so that the structure and content of the minutes are consistent and standardized, facilitating reading, archiving, and subsequent review and use. The entire conference minutes generation process is automated, from conversion of the conference recording to extraction of the key information, and then to generation of the minutes, without the need for manual operation, greatly improving work efficiency and saving labor costs. Different template matching rules and template libraries can be pre-set according to different conference types and requirements, which has strong versatility and flexibility.

[0054] In an optional embodiment, the conference minutes template matching rule comprises: a territorial matching rule, used to match a template with completely consistent territory in the conference minutes template library according to the territory information of the conference, and if there is no such template, the superior territory is matched in turn; a version matching rule, used to select the latest version template when multiple templates meeting the territory requirement are matched; a content matching rule, used to calculate the similarity of the text record and the template content when multiple templates of the same version are matched, and select the template with the highest similarity.

[0055] Specifically, the meeting minutes template matching rule aims to accurately and efficiently select the most suitable template for the current meeting from the meeting minutes template library, taking into account multiple factors such as the meeting's territorial information, template version, and the similarity between the text record and the template content. Through step-by-step filtering and comparison, it ensures that the finally matched template can meet the needs of meeting minutes generation to the greatest extent. The territorial matching rule takes the meeting's territorial information as the primary matching basis. It first tries to find a template that is completely consistent with the meeting's territory in the meeting minutes template library. If not found, it matches to the superior territory in turn according to the hierarchical relationship of the territory, until a suitable template is found or all possible territorial levels are traversed. Through the territorial matching rule, the generated meeting minutes can better meet the actual situation and requirements of the local area, enhancing the applicability and standardization of the minutes.

[0056] When multiple templates that meet the territorial requirements are found according to the territorial matching rule, the version matching rule will preferentially select the template with the latest version, which can ensure that the meeting minutes comply with the latest standards and specifications, utilize the latest functions and formats, and improve the quality and readability of the minutes.

[0057] If multiple templates with the same version are found during the matching process, the content matching rule will calculate the similarity between the text record and the content of each template. Text similarity algorithms such as cosine similarity and Jaccard similarity can be used. The template with the highest similarity is finally selected as the target template, which can ensure that the selected template is most consistent with the actual content of the meeting, making the generated meeting minutes accurately reflect the key information such as the discussion points and decision results of the meeting, and improving the accuracy and relevance of the minutes.

[0058] The above is only the preferred embodiment of the present application, but the protection scope of the present application is not limited thereto. It should be noted that for those skilled in the art, without departing from the technical principles of the present application, a number of equivalent obvious variations and / or equivalent replacement methods can be made, which should also be considered as the protection scope of the present application.

Claims

1. An intelligent interactive system for multimodal data fusion, characterized in that: include: The user question-answering module is used to obtain the target question input by the user, accurately retrieve the answer to the target question and the source of the answer from the pre-built knowledge base, and feed back the answer to the user; The AI ​​text detection module is used to obtain the text to be detected input by the user, use the pre-built AI text detection model to determine whether the text input by the user is AI-generated text, and output the judgment result; The meeting minutes generation module is used to obtain the meeting recording input by the user and generate the meeting minutes corresponding to the meeting recording.

2. The multimodal data fusion intelligent interactive system according to claim 1, characterized in that: The process of building the knowledge base includes: Collect historical consultation records and original knowledge documents within a preset period of time; the historical consultation records include historical consultation questions and answers to historical questions; Uploading the historical consultation records and the original knowledge files to the knowledge base; Extracting file information of the original knowledge file, generating a metadata table corresponding to each file type according to the file type, and storing the file information in the metadata table; By recording the unique identifier of the knowledge original file in the metadata table, an association relationship between the metadata table and the knowledge original file is established, so that the knowledge original file can be quickly located through the unique identifier during knowledge retrieval.

3. The multimodal data fusion intelligent interactive system according to claim 1, characterized in that: The step of obtaining a target question input by a user and accurately retrieving the answer to the target question and the source of the answer from a pre-built knowledge base includes: Obtaining a user's question input, preprocessing the question input, and generating a target question; the question input may be in the form of, but not limited to, text input, image input, file input, and video input; Determine the question and answer prompt words for the target question, and accurately retrieve the answer to the target question and the source of the answer from a pre-built knowledge base based on the question and answer prompt words.

4. The multimodal data fusion intelligent interactive system according to claim 1, characterized in that: The user question and answer module is also used to: receiving feedback marks from the user on the answer to the question, and triggering corresponding processing logic for different feedback marks; the feedback marks are used to indicate the accuracy of the answer to the question; When the accuracy does not meet the preset accuracy requirement, the target question is pushed to the manual answer module and the retrieval algorithm is adjusted accordingly; When the accuracy reaches the preset accuracy requirement, no processing is performed.

5. The multimodal data fusion intelligent interactive system according to claim 1, characterized in that: The AI ​​text detection module includes: A model detection unit, configured to input the text to be detected into a pre-trained AI-generated detection model, so that the AI-generated detection model detects the text to be detected and outputs a probability value that the text to be detected is AI-generated text; The comprehensive judgment unit is used to calculate the text entropy value abnormality and the professional term density difference of the text to be detected, and combine the probability value to obtain a judgment result on whether the text to be detected is an AI-generated text.

6. The multimodal data fusion intelligent interactive system according to claim 5, characterized in that: The training process of the AI-generated detection model includes: Convolutional neural network and bidirectional long short-term memory network were used to build the initial model; Divide a certain number of positive samples and negative samples into training sets, validation sets, and test sets according to preset ratios; wherein the positive samples are manually written texts, and the negative samples are texts generated using AI tools; Training the initial model using the training set, and adjusting the parameters of the initial model through a backpropagation algorithm so that the model can accurately distinguish between AI-generated text and human-written text; Using the validation set to evaluate and tune the initial model, and select model parameters with optimal performance; The initial model is tested using the test set to ultimately obtain a trained AI-generated detection model.

7. The multimodal data fusion intelligent interactive system according to claim 5, characterized in that: The calculating of the text entropy abnormality and the professional term density difference of the text to be detected includes: Divide the text to be detected into several paragraphs according to the preset character intervals, and calculate the information entropy of each paragraph; Based on the information entropy and a preset benchmark entropy value, calculating the paragraph entropy abnormality of each paragraph and the text entropy abnormality of the text to be detected; The measured value of the professional terminology density in the text to be detected is calculated according to the professional terminology database, and the professional terminology density difference between the measured value of the professional terminology density and a preset professional terminology density minimum threshold is calculated.

8. The multimodal data fusion intelligent interactive system according to claim 5, characterized in that: Combining the probability value to obtain a judgment result on whether the text to be detected is AI-generated text includes: Calculate a comprehensive judgment value based on the text entropy abnormality, the professional term density difference, the probability value, and the pre-assigned abnormality weight, density difference weight, and probability value weight; According to the comprehensive judgment value, a judgment result of whether the text to be detected is an AI-generated text is determined.

9. The multimodal data fusion intelligent interactive system according to claim 1, characterized in that: The meeting minutes generation module includes: a text conversion unit, configured to, when the question input is a conference recording, call a speech-to-text model to convert the conference recording into a text record, and to proofread and polish the text record; A template matching unit, configured to match a target template from a meeting minutes template library in the knowledge base according to a preset meeting minutes template matching rule; An information extraction unit, configured to extract key meeting information from the processed text records using natural language processing technology; The minutes generating unit is used to fill the key information of the meeting into the target template to generate the meeting minutes corresponding to the meeting recording.

10. The multimodal data fusion intelligent interactive system according to claim 9, characterized in that: The meeting minutes template matching rules include: The location matching rule is used to prioritize matching templates with exactly the same location in the meeting minutes template library based on the location information of the meeting. If no templates exist, the matching is performed in the upper-level locations in sequence. Version matching rules are used to give priority to the template with the latest version when multiple templates that meet the local requirements are matched; The content matching rule is used to calculate the similarity between the text record and the template content when multiple versions of the same template are matched, and select the template with the highest similarity.

Citation Information

Patent Citations

  • Intelligent question and answer interaction system based on semantic analysis engine

    CN117056479A

  • Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

    CN118627628A

  • Conference summary generation method and device, terminal and computer readable storage medium

    CN119150814A

  • Artificial intelligence automatic response conference assistant

    CN119847389A

  • System and method for providing answers to questions

    US20090287678A1