Multimedia information processing method, system, and electronic device

By receiving inquiry information in a generative question-and-answer system, rewritten as candidate questions, and using a large language model to process multimedia data fragments, the problems that cannot be provided with personalized answers in the prior art are solved, and the generation of personalized answers is realized.

WO2025130162A1PCT designated stage expired Publication Date: 2025-06-26ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2024/116687
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-09-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The prior art is difficult to provide personalized answers to multimedia data in a generative question and answer system.

Method used

By receiving inquiry information, it is rewritten as a candidate question for custom questions, and querying relevant data fragments from pre-uploaded multimedia information, and processing these data fragments using a large language model to generate personalized reply content.

Benefits of technology

It realizes personalized answers to multimedia information in a generative question-and-answer system, and solves the problem that personalized answers cannot be provided for multimedia data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116687_26062025_PF_FP_ABST
    Figure CN2024116687_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the fields of large language models and machine learning. Disclosed are a multimedia information processing method, a system, and an electronic device. The method is applied to a generative question-answering system, and comprises: receiving inquiry information; rewriting the inquiry information into a candidate query used for executing customized querying for multimedia information, wherein the multimedia information is content uploaded to the generative question-answering system in advance; retrieving from a question-answering database a data chunk associated with the candidate query, wherein at least one set of data chunks is stored in the question-answering database, and a data source of the data chunks is the multimedia information; and using a large language model to process the data chunk to generate at least one piece of reply content matching the inquiry information.
Need to check novelty before this filing date? Find Prior Art

Description

Multimedia information processing method, system and electronic device

[0001] Cross-reference

[0002] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 2023117453192 and invention name “Method, system and electronic device for processing multimedia information”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0003] The present disclosure relates to large model technology and the field of machine learning, and more specifically, to a method, system, and electronic device for processing multimedia information. Background Art

[0004] Currently, generative question-answering systems typically use artificial intelligence (AI) assistants that learn through work to complete responses to inquiries. With the continuous development of generative question-answering systems, new requirements have been placed on them in different scenarios. Generative question-answering systems are required to be able to take notes on multimedia information, organize interviews, and so on. However, most current AI assistants are only called through built-in triggering opportunities, and users cannot actively interact with large language models, resulting in a technical problem of being unable to provide personalized answers based on the content of multimedia data.

[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.

[0006] Summary of the Invention

[0007] The embodiments of the present disclosure provide a method, system, and electronic device for processing multimedia information to at least solve the technical problem of being unable to provide personalized answers to the content of multimedia data.

[0008] According to one aspect of an embodiment of the present disclosure, a method for processing multimedia information is provided. This method, which can be applied to a generative question-answering system, may include: receiving a query; rewriting the query into candidate questions for performing customized questions on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; querying a question-answering database to obtain data segments associated with the candidate questions, wherein the question-answering database stores at least one set of data segments, the data segments being derived from the multimedia information; and processing the data segments using a large language model to generate at least one response matching the query.

[0009] According to another aspect of an embodiment of the present disclosure, another method for processing multimedia information is also provided. This method can be applied to a generative question-answering system and may include: in response to an input instruction on an operation interface, obtaining input query information; in response to a query instruction on the operation interface, rewriting the query information into a candidate question for performing a customized question on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; querying a question-answering database to obtain data segments associated with the candidate question, wherein the question-answering database stores at least one set of data segments, the data segments being derived from the multimedia information; processing the data segments using a large language model to generate at least one response matching the query information; and displaying the response content on the operation interface.

[0010] According to another aspect of an embodiment of the present disclosure, an electronic device is also provided, which may include a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions. When the above-mentioned computer-executable instructions are executed by the processor, the above-mentioned method of any one of the above-mentioned items is implemented.

[0011] According to another aspect of an embodiment of the present disclosure, a processor is further provided, which is used to run a program, wherein any one of the above methods is executed when the program is running.

[0012] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is further provided, the computer-readable storage medium including a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute any one of the above methods.

[0013] According to another aspect of an embodiment of the present disclosure, a computer program product is further provided, comprising a non-volatile computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any one of the above methods when executed by a processor.

[0014] In an embodiment of the present disclosure, a query message is received; the query message is rewritten into a candidate question for performing a customized question on multimedia information, wherein the multimedia information is content pre-uploaded to a generative question-answering system; a data segment associated with the candidate question is queried from a question-answering database, wherein the question-answering database stores at least one set of data segments, the data source of which is the multimedia information; the data segment is processed using a large language model to generate at least one reply content that matches the query message. That is, in this embodiment, when a query message is received, the query message is rewritten into a candidate question for performing a customized question, a data segment associated with the candidate question is determined from the multimedia information, and the data segment is processed using a large language model to obtain at least one reply content that matches the query message, thereby achieving the technical effect of providing personalized answers to the content of the multimedia information and solving the technical problem of not being able to provide personalized answers to the content of the multimedia data.

[0015] It is easy to note that the above general description and the following detailed description are only for the purpose of exemplifying and explaining the present disclosure, and do not constitute a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0017] FIG1 is a schematic diagram of an application scenario of a method for processing multimedia information according to an embodiment of the present disclosure;

[0018] FIG2 is a block diagram of a computing environment according to an embodiment of the present disclosure;

[0019] FIG3 is a flowchart of a method for processing multimedia information according to an embodiment of the present disclosure;

[0020] FIG4 is a flowchart of another method for processing multimedia information according to an embodiment of the present disclosure;

[0021] FIG5 is a schematic diagram of a multimedia information processing system according to an embodiment of the present disclosure;

[0022] FIG6 is a schematic diagram of constructing training data according to an embodiment of the present disclosure;

[0023] FIG7 is a schematic diagram of manual annotation according to an embodiment of the present disclosure;

[0024] FIG8 is a schematic diagram of constructing training data according to an embodiment of the present disclosure;

[0025] FIG9 is a schematic diagram of another embodiment of the present disclosure for constructing training data;

[0026] FIG10 is a schematic diagram of a meeting minutes according to an embodiment of the present disclosure;

[0027] FIG11 is a schematic diagram of an interaction scenario according to an embodiment of the present disclosure;

[0028] FIG12 is a schematic diagram of a reply content according to an embodiment of the present disclosure;

[0029] FIG13 is a schematic diagram of a recommendation question according to an embodiment of the present disclosure;

[0030] 14 is a hardware structure block diagram of a computer terminal (or mobile device) according to a method for processing multimedia information according to an embodiment of the present disclosure;

[0031] FIG15 is a structural block diagram of a service grid according to an embodiment of the present disclosure;

[0032] FIG16 is a schematic diagram of a multimedia information processing device according to an embodiment of the present disclosure;

[0033] FIG17 is a schematic diagram of a multimedia information processing device according to an embodiment of the present disclosure;

[0034] FIG18 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure;

[0035] FIG19 is a block diagram of an electronic device according to a method for processing multimedia information according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] The technical solution provided by the present disclosure can be implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. The large model can also be called a cornerstone model / foundation model (Foundation Model). The large model is pre-trained by large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large-scale language models (LLMs) and multi-modal pre-training models.

[0039] It should be noted that when the large model is actually applied, the pre-trained model can be fine-tuned through a small number of samples, so that the large model can be applied to different tasks. For example, the large model can be widely used in natural language processing (NLP), computer vision, speech processing and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, etc. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiment of the present disclosure, data processing through a machine learning model in a dialogue scenario is used as an example for explanation.

[0040] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:

[0041] Supervised fine-tuning (SFT) can refer to tuning a pre-trained language model on a small amount of labeled data to learn a supervised strategy for generating output from a given list of prompts;

[0042] Zero-shot refers to the situation in machine learning where the model has not been directly exposed to certain data during training, but needs to make predictions or classifications on this data during testing.

[0043] Elastic Search is a distributed, multi-user, full-text search engine that provides real-time search, data analysis, and visualization through a RESTful web interface.

[0044] Prompt: In the field of natural language generation, the model needs to generate corresponding text content based on the given prompt text, which can be a description of a question or task;

[0045] Chatbot GPT (ChatGPT for short) can be a chatbot based on a Generative Pre-trained Transformer (GPT) model that can output corresponding natural language responses based on the user's conversation history and context;

[0046] The Large Language Model (LLM) model is a very large deep learning model pre-trained based on a large amount of data. It can be used to answer tasks, translate languages, or complete sentences.

[0047] According to a method of an embodiment of the present disclosure, a method for processing multimedia information is provided. As an optional implementation, the multimedia information processing method may include, but is not limited to, application in the application scenario shown in Figure 1. Figure 1 is a schematic diagram of an application scenario of a multimedia information processing method according to an embodiment of the present disclosure. As shown in Figure 1, in the application scenario, a terminal device 102 may, but is not limited to, communicate with a server 106 via a network 104. For example, it may be used to transmit query information, response content, etc. The server 106 may, but is not limited to, perform operations on a database 108, such as write or read data operations. The terminal device 102 may, but is not limited to, include a human-computer interaction screen, a processor, and a memory. The human-computer interaction screen may, but is not limited to, display query information and response content on the terminal device 102. The processor may, but is not limited to, responding to the human-computer interaction operation, executing corresponding operations, or generating corresponding instructions and sending the generated instructions to the server 106. The memory is used to store relevant processed data, such as a large language model, multimedia data, etc.

[0048] As an optional method, the following steps in the method for processing multimedia information can be executed on the server 106: Step S102, receiving the query information; Step S104, rewriting the query information into a candidate question for performing a customized question on the multimedia information; Step S106, querying and obtaining a data segment associated with the candidate question from the question and answer database; Step S108, processing the data segment using a large language model to generate at least one reply content that matches the query information.

[0049] Using the above method, a query message is received; the query message is rewritten into a candidate question for performing a customized question on a multimedia message, wherein the multimedia message is content pre-uploaded to a generative question-answering system; data segments associated with the candidate question are retrieved from a question-answering database, wherein the question-answering database stores at least one set of data segments, the data segments being derived from the multimedia message; the data segments are processed using a large language model to generate at least one response matching the query message. That is, in this embodiment, when a query message is received, the query message is rewritten into a candidate question for performing a customized question, a data segment associated with the candidate question is determined from the multimedia message, and the data segment is processed using a large language model to obtain at least one response matching the query message, thereby achieving the technical effect of providing personalized responses based on the content of the multimedia message and resolving the technical problem of being unable to provide personalized responses based on the content of the multimedia data.

[0050] In another optional embodiment, FIG2 shows in a block diagram an embodiment of using a computer terminal (or mobile device) as a computing node in a computing environment 201. FIG2 is a structural block diagram of a computing environment according to an embodiment of the present disclosure. As shown in FIG2 , the computing environment 201 includes multiple (shown in the figure as 210-1, 210-2, ...) computing nodes (such as servers) running on a distributed network. The computing nodes all contain local processing and memory resources, and the end user 202 can remotely run applications or store data in the computing environment 201. The application can be provided as multiple services 220-1, 220-2, 220-3 and 220-4 in the computing environment 201, representing services "A", "D", "E" and "H" respectively.

[0051] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to the ingress gateway 230. The ingress gateway 230 may include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in the computing environment 201).

[0052] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly contacting any actual hardware resources. While the virtual machine virtualizes the machine, according to container-based virtualization, a container can be started to virtualize the entire operating system (OS) so that multiple workloads can run on a single operating system instance.

[0053] In one embodiment based on container virtualization, several containers of a service can be assembled into a computing unit (e.g., a Kubernetes Pod). For example, as shown in Figure 2, service 220-2 can be equipped with one or more computing units (Pods) Pod240-1, 240-2, ..., 240-N (collectively referred to as Pods). The Pod may include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in the Pod process requests related to one or more corresponding functions of the service, and the proxy 245 generally controls network functions related to the service, such as routing, load balancing, etc. Other services can also be equipped with Pods similar to Pods.

[0054] During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201. Executing one or more functions of one service may require invoking one or more functions of another service. As shown in FIG2 , service "A" 220-1 receives a user request from end user 202 from ingress gateway 230. Service "A" 220-1 may invoke service "D" 220-2, and service "D" 220-2 may request service "E" 220-3 to execute one or more functions.

[0055] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.

[0056] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure, such as weather forecast results and other data, are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0057] In the above operating environment, the present disclosure provides a method for processing multimedia information. FIG3 is a flow chart of a method for processing multimedia information according to an embodiment of the present disclosure. As shown in FIG3, the method can be applied to a generative question-answering system. The generative question-answering system can be deployed in the server 106 in FIG1 to perform the following steps:

[0058] Step S302: receiving inquiry information.

[0059] In the technical solution provided in step S302 of the present disclosure, query information obtained by the generative question-answering system can be received. The query information can be a user query input, a simple instruction input by the user, a customized question posed by the user of the generative system, or a question posed by the user in natural language, such as "What building is mentioned in the article?" or "When was the building built?" It should be noted that the query information is not specifically limited in its form.

[0060] Optionally, query information input by the user into the generative question answering system is obtained.

[0061] Furthermore, the use scenarios of the generative question-answering system are further expanded. The generative question-answering system can be used to process inquiry information in meeting scenarios, and can be used to extract meeting minutes, records of to-do items in meetings, and determination of meeting conclusions, etc. It can also process inquiry information in speech scenarios, and can be used to determine golden sentences in speeches, speech cases, and core ideas in speeches, etc. It can also process inquiry information in learning scenarios, and can be used to determine core knowledge points in learning, as well as reading suggestions, etc. It can also process inquiry information in general scenarios, and can be used to determine milestone events, key data, concept explanations, etc. It should be noted that this is only an example, and does not impose specific restrictions on the use scenarios of the generative question-answering system.

[0062] Step S304 : rewriting the query information into candidate questions for performing customized questioning on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system.

[0063] In the technical solution provided in the above step S304 of the present disclosure, the query information is processed and can be rewritten into candidate questions for performing customized questions on the multimedia information. The multimedia information can be a file pre-uploaded to the generative question-answering system, and can be audio information, text information, and the like. For example, it can be a video or audio recording of the entire meeting process, or a speech video, audio, or speech manuscript. This is only an example and does not impose specific restrictions on the type and content of the multimedia information. The candidate question can be a well-structured question that can be used to ask questions about the content in the multimedia information. For example, it can be a standard question adjusted according to a pre-set format.

[0064] Because the user omits the previously mentioned contextual information during multiple rounds of interaction with the generative question-answering system, and only gives a simple query information. For example, the first round of questions is: "Which buildings in Hangzhou are mentioned in the article?" The generative question-answering system answers: "It mentions the International Finance Center (IFC) and xxx in Hangzhou." Then the user asks: "When was this built?" It can be seen that the second round of query information itself does not bring in a subject. If the query information is used directly to answer, the content of the answer to the query information will be inaccurate. Therefore, in order to avoid the above problems, in this embodiment, after obtaining the query information, the query information can be rewritten to obtain a candidate question with more complete and rich semantic logic. By processing the candidate information, the accurate answer content of the query information is obtained, thereby achieving the technical effect of improving the accuracy of the answer content and solving the technical problem of poor accuracy of the answer content.

[0065] Optionally, the multimedia information can be information from a variety of scenarios, such as meetings, classes, interviews, training sessions, live broadcasts, podcasts, and so on. Multimedia information from these scenarios can be uploaded to the generative question-answering system in advance. After obtaining the query information, the query information can be rewritten to obtain at least one candidate question that can be used to ask questions about the multimedia information. It should be noted that the above scenarios are merely examples and do not impose any specific restrictions on the scenarios in which multimedia information is generated.

[0066] Step S306: query and obtain data segments associated with the candidate question from the question-answer database, wherein the question-answer database stores at least one set of data segments, and the data source of the data segments is multimedia information.

[0067] In the technical solution provided in step S306 above of the present disclosure, after obtaining a candidate question, a data segment associated with the candidate question can be queried from a question-and-answer database. The question-and-answer database can be a database that stores at least one set of data segments, or a retrieval module for retrieving associated data segments, such as a distributed search engine (Elastic Search, abbreviated as ES). This is merely an example, and the type of question-and-answer database is not specifically limited. The data segment can be a chunk obtained by processing multimedia information, can be used to answer the candidate question, and can be a text segment or an audio segment, etc. The type of data segment is not specifically limited.

[0068] Optionally, uploaded multimedia information can be pre-acquired and processed to obtain multiple data segments, which can then be stored in a question-and-answer database. After the query information is rewritten to obtain candidate questions, data segments associated with the candidate questions can be retrieved from the question-and-answer database. Association can mean similarity in content or a logical relationship, and the definition of association is not specifically limited herein.

[0069] For example, after a generative question-answering system receives multiple or a single piece of multimedia information, it can segment the information to produce multiple data segments. These segments can then be stored in a question-answering database. Once a candidate question is obtained, the question-answering database can be used to determine the data segments associated with the candidate question based on the candidate information.

[0070] Step S308: Process the data segment using the large language model to generate at least one response content that matches the query information.

[0071] In the technical solution provided in step S308 of the present disclosure, a data segment can be used as input and transmitted to a large language model, which can then process the data segment to generate at least one response that matches the query information. The large language model can be a Large Language Models (LLM) model, which can be used to determine the response content and quickly extract and precipitate the content in the data segment to generate the response content. The response content can be an answer to the query information.

[0072] For audio and video content scenarios with certain added value, since most artificial intelligence functions are only called through built-in trigger logic, the artificial intelligence functions are then used to directly display the answer content on the interface of the generative question-answering system. However, in this method, users cannot actively interact, so there is a technical problem that personalized audio and video questions and answers cannot be provided to users. In order to solve the above technical problems, this embodiment designs a generative question-answering system. The generative question-answering system can be used as an artificial intelligence assistant for work and study. When the user's inquiry information is obtained, the inquiry information can be rewritten to obtain candidate questions, and the data segments associated with the candidate questions can be determined. By processing the data segments, at least one answer content of the inquiry information is determined and provided to the user, thereby achieving the technical effect of providing users with personalized audio and video questions and answers based on the user's inquiry information.

[0073] Optionally, after determining multiple or one data segments associated with the candidate question, the acquired data segments and query information can be used as input and transmitted to the large language model, so as to generate at least one response content matching the query information through the large language model.

[0074] For example, a video introducing Shanghai Airport is uploaded to a generative question-answering system. The system processes the video and generates multiple data segments. When it receives the query "When was it built?" and determines that it lacks a subject, it can rewrite the query based on the multimedia data surrounding the query or previous conversations to generate a candidate question: "When was Shanghai Airport built?" The system then identifies the data segments associated with the query in the question-answering database. The system then uses a large prediction model to process the data segments to generate a response that matches the query: "Shanghai Airport was built in 1998."

[0075] In an embodiment of the present disclosure, a query message is received; the query message is rewritten into a candidate question for performing a customized question on multimedia information, wherein the multimedia information is content pre-uploaded to a generative question-answering system; a data segment associated with the candidate question is retrieved from a question-answering database, wherein the question-answering database stores at least one set of data segments, the data source of which is the multimedia information; the data segment is processed using a large language model to generate at least one reply content that matches the query message. That is, in this embodiment, when a query message is received, the query message is rewritten into a candidate question for performing a customized question, a data segment associated with the candidate question is determined from the multimedia information, and the data segment is processed using a large language model to obtain at least one reply content that matches the query message, thereby achieving the technical effect of providing personalized answers to the content of the multimedia information and solving the technical problem of not being able to provide personalized answers to the content of the multimedia data.

[0076] The above method of this embodiment is further introduced below.

[0077] As an optional implementation, step S304 rewrites the query information into candidate questions for performing customized questioning of multimedia information, including: obtaining multi-round dialogue samples stored by the generative question-answering system within a historical time period, wherein the multi-round dialogue samples record at least a plurality of semantically associated historical query information occurring within the historical time period; identifying at least one historical query information matching the query information from the multi-round dialogue samples; and rewriting the query information based on the at least one historical query information to generate at least one candidate question.

[0078] In this embodiment, after the query information is obtained, in order to improve the accuracy of the large language model in processing the query information, the query information can be rewritten first to obtain candidate questions that can be used to perform customized questioning on the multimedia information.

[0079] In an optional technical solution disclosed herein, candidate questions can be rewritten based on multimedia information according to a pre-designed question structure, or the query information can be rewritten using multi-turn conversation samples stored by the generative question-answering system within a historical time period. Furthermore, the multi-turn conversation samples record at least a plurality of semantically related historical query information that occurred within the historical time period. Therefore, after obtaining the query information, the multi-turn conversation samples stored by the generative question-answering system within the historical time period can be obtained. From the multi-turn conversation samples, at least one historical query information matching the query information can be identified. The query information can then be rewritten based on the historical query information to generate at least one candidate question.

[0080] For example, suppose the first round of dialogue samples is "What buildings are mentioned in the video?" - "Big Wild Goose Pagoda and Small Wild Goose Pagoda"; and the second round of dialogue samples is "How tall are the two buildings?" - "350 meters and 250 meters respectively." When the query "Which building is taller?" is obtained, the first and second rounds of dialogue samples can be obtained and, based on them, the query can be rewritten to produce candidate questions: "Which building is taller, the 350-meter Big Wild Goose Pagoda or the 250-meter Small Wild Goose Pagoda?" or "Which building is taller, the 350-meter building or the 250-meter building?"

[0081] As an optional implementation, in step S306, before querying and obtaining multiple groups of data segments associated with candidate questions from the question and answer database, the method may also include: obtaining uploaded multimedia information; converting the multimedia information into text to generate text data in text format; segmenting the text data to obtain at least one group of data segments; and constructing a question and answer database based on at least one group of data segments, wherein each group of data segments is stored in the question and answer database in the form of a feature vector.

[0082] In this embodiment, before querying the question-and-answer database to obtain multiple sets of question-and-answer data segments associated with candidate questions, it is necessary to first store the multiple sets of data segments in the question-and-answer database. The question-and-answer database storing the multiple sets of data segments can be constructed by the following steps: first, obtaining one or more uploaded multimedia information, converting the multimedia information into text to obtain text data in a text format. The text data can be segmented to obtain at least one set of data segments. Based on the at least one set of data segments, a question-and-answer database can be constructed. Each set of data segments can be stored in the question-and-answer database in the form of a feature vector, so that the data segment associated with the query information can be determined based on the feature vector.

[0083] Optionally, when the generative question-answering system obtains multimedia information, if the multimedia information is text data, the segmentation step can be performed directly without conversion. If the multimedia information is non-text data, the multimedia information can be converted into text data in text format first. It should be noted that there is no specific restriction on the conversion method here, and it can be selected according to actual conditions.

[0084] Optionally, after acquiring the multimedia information, the multimedia information may be transcribed to obtain text data in a text format, and the text data may be sliced ​​to obtain at least one set of data segments. The sliced ​​data segments may be vectorized and stored to construct a question-and-answer database. Based on the constructed question-and-answer database, data segments associated with candidate questions may be retrieved.

[0085] For example, discrete natural language data segments can be converted into continuous feature vectors through the model to facilitate subsequent semantic-based retrieval. For example, the Transformer-based encoder BERT can be used to vectorize the data segments. After obtaining the feature vectors of the data segment vectors, the data segments can be stored in the form of feature vectors in the question-and-answer database. The multimedia information can be uploaded by the user or recorded in real time. This is only an example and does not impose any specific restrictions on the method of obtaining multimedia information.

[0086] In this embodiment, multiple rounds of interaction can be performed through natural language and a large language model. Considering that the context mentioned already is easily omitted during communication, a rewriting step is introduced to take into account the current query information and the input and output in the historical time period (that is, multiple rounds of dialogue samples), and rewrite the query information into a candidate question with rich context and complete semantic information. By processing the candidate questions, the recall rate and accuracy of retrieval in the question and answer database can be improved, thereby achieving the technical effect of accurately determining the content of the reply and solving the technical problem of poor accuracy of the reply to the query information.

[0087] As an optional implementation, text data is segmented to obtain at least one set of data segments, including: identifying the length of the text data and / or the semantic information of the text data; and segmenting the text data based on the length and / or semantic information to obtain at least one set of data segments.

[0088] In this embodiment, since the large language model has a limit on the length of the input data, the text data obtained after conversion needs to be segmented to obtain data segments that do not exceed a specific size. However, when segmenting the data segments, it is also necessary to avoid segmenting data segments that cannot express complete information. Therefore, in this embodiment, the overall length of the text data and / or the semantic information of the text data are identified, and the text data can be segmented based on the length and / or semantic information to obtain at least one set of data segments. Among them, the length can be used to represent the amount of text in the text data, for example, it can be 300 words, 500 words, etc. This is only an example and does not impose a specific limit on the length. Semantic information can be used to represent the content that the text data wants to express. For example, the semantic information of "The pony ate grass for 100 consecutive days and was not tired of it" can be "The pony likes to eat grass."

[0089] Optionally, when dividing the text data, in addition to considering the input length during large language model inference, it is also necessary to consider the allowable input length in the question-and-answer data path during subsequent vectorization and sorting, as well as the integrity of the text data. Therefore, the text data is divided based on length and / or semantic information to obtain at least one group of data segments, for example, data segments of 300 to 500 words.

[0090] As an optional implementation, querying and obtaining data segments associated with candidate questions from a question and answer database includes: reading at least one group of data segments from the question and answer database; obtaining the semantic similarity between each group of data segments and the candidate question; and determining the screened data segments whose semantic similarity exceeds a similarity threshold as data segments associated with the candidate question.

[0091] In this embodiment, when it is necessary to query the question-answer database to obtain data segments associated with candidate questions, the semantic similarity between each group of data segments and the candidate questions can be determined separately, and at least one group of data segments whose semantic similarity exceeds the similarity threshold can be screened out to obtain data segments associated with the candidate questions. Semantic similarity can be used to determine the similarity between the semantics in the data segments and the semantics in the candidate questions, and can be used to characterize the correlation between the data segments and the candidate questions. The expression can be in the form of a numerical value, text, percentage, etc. For example, it can be a correlation score. This is only for illustration and does not impose specific restrictions on the expression of semantic similarity. The similarity threshold can be a value set in advance according to actual conditions.

[0092] Optionally, the candidate questions and data fragments are encoded and calculated through an encoder (BERT) based on a deep learning architecture (e.g., Transformer) of the question and answer database to determine the semantic similarity between each group of data fragments and the candidate question. Based on the semantic similarity, at least one group of data fragments associated with the candidate question can be screened out from the question and answer database to obtain the data fragments associated with the candidate question.

[0093] If the data segments most similar to the query information are determined only through the semantic indexing of paragraphs, there will be a technical problem of low accuracy of the determined data segments. Taking the above problem into consideration, in this embodiment, taking into account the high degree of colloquialism of the text data obtained after the multimedia information is transcribed, in order to better understand the content of the multimedia information, a multi-way recall based on a distributed search engine is introduced to retrieve data segments at the word level and the semantic level, covering as much content of the data segments as possible, thereby improving the technical effect of the accuracy of determining the data segments and solving the technical problem of low accuracy of the determined data segments.

[0094] As an optional implementation, obtaining the semantic similarity between each group of data segments and the candidate question includes: obtaining the feature vectors of each group of data segments and the feature vectors of the candidate question; performing similarity calculations on the feature vectors of each group of data segments and the feature vectors of the candidate question to obtain the semantic similarity between each group of data segments and the candidate question.

[0095] In this embodiment, when it is necessary to determine the semantic similarity between each group of data segments and the candidate question, the feature vector of each group of data segments and the feature vector of the candidate question can be obtained, and the similarity between the feature vector of each group of data segments and the feature vector of the candidate question can be calculated to obtain the semantic similarity between each group of data segments and the candidate question. Based on this similarity, the data segments associated with the candidate question in the question and answer database can be determined.

[0096] Optionally, the data segments can be vectorized and stored in a distributed search engine. For candidate information that has also been vectorized online, multiple related data segments can be found through multi-way recall, for example, using word matching based on a matching algorithm (Best Matching, abbreviated as BM25) and semantic similarity based on vector calculation (Proxima). The determined multiple data segments are used as reference inputs for the large language model to answer query information.

[0097] As an optional implementation, after determining the data segments associated with the candidate questions, the method may further include: sorting all screened data segments associated with the candidate questions according to the semantic similarity between the data segments and the candidate questions; extracting at least one group of data segments within a preset ranking according to the sorting result; reading the generation time of each group of data segments within the preset ranking; assembling each group of data segments in sequence according to the generation time of each group of data segments within the preset ranking; processing the data segments using a large language model to generate at least one reply content that matches the query information, including: processing the assembled data segments using a large language model to generate at least one reply content that matches the query information.

[0098] In this embodiment, after determining the data segments associated with the candidate questions, all the screened data segments can be sorted according to semantic similarity, and at least one group of data segments within the preset ranking can be extracted according to the sorting results. The generation time of each group of data segments within the preset ranking is determined, and each group of data segments is assembled in sequence according to the generation time of each group of data segments within the preset ranking. The assembled data segments can be processed using a large language model to obtain at least one reply content that matches the query information. The generation time can be used to characterize the time when the data segment was generated. For example, at 15:00, speaker 1 said data segment 1, and the generation time of data segment 1 is 3 pm. The preset ranking can be a pre-set ranking, for example, it can be from first to fourth. This is only for example, and there is no specific limitation on the size of the preset ranking.

[0099] Optionally, after determining the data segments associated with the candidate questions according to the semantic similarity between the data segments and the candidate questions, the determined data segments can be sorted based on the semantic similarity, and at least one group of data segments within a preset ranking can be extracted according to the sorting results. The data segments can be spliced ​​in sequence according to the generation time of at least one group of data segments within the preset ranking to obtain the input to be input into the large language model.

[0100] Since the input of a large language model usually has a certain length limit, the recalled multiple related fragments cannot all be included in the input. Therefore, a ranking model can be used to score the relevance of the candidate questions and data fragments. The data fragments can be arranged from high to low according to the degree of semantic similarity, and the data fragments within the preset ranking can be screened out. The data fragments within the preset ranking are placed one by one in the prompt information assembly according to the time of generation, and the assembly is stopped until the assembly result exceeds the input length limit. That is, in this embodiment, data fragments with high semantic similarity are obtained for splicing, and the reply content of the inquiry information is determined based on the spliced ​​data fragments, thereby achieving the technical effect of improving the accuracy of the reply content and solving the technical problem of poor accuracy of the reply content.

[0101] As an optional embodiment, in step S304, after the query information is rewritten into a candidate question for performing a customized question on the multimedia information, the method further includes: determining whether to perform a query operation on the data segments recalled from the question and answer database based on the type of the candidate question; if the query operation is performed, determining the number of retrievals when performing the query operation; based on the number of retrievals, determining the number of data segments expected to be obtained by the query, wherein the expected number of segments is used to determine the number of data segments that match the candidate question.

[0102] In this embodiment, considering that there are various types of candidate questions and different types of candidate information require different numbers of data segments, in order to improve data processing efficiency, after rewriting the query information into at least one candidate question for performing customized questioning of multimedia information, the type of the candidate question is determined, and based on the type of the candidate question, the number of data segments that need to be queried is determined, thereby reducing data processing time.

[0103] Optionally, after obtaining a candidate question, the type of the candidate question is determined. Based on the type of the candidate question, a determination is made as to whether to perform a query operation on the data segments in the question-and-answer database. If a query operation is performed, a corresponding retrieval quantity is determined based on the type of the candidate question. Based on the retrieval quantity, an estimated number of data segments to be obtained from the query is determined, and the estimated number of data segments is retrieved from the question-and-answer database according to the retrieval quantity. The estimated number of segments can be used to determine the number of data segments that match the candidate question.

[0104] Optionally, the type of the candidate question is determined, the candidate questions are divided based on the type of the candidate question, and the number of data segments that need to be retrieved for the candidate question is determined based on the result of the division.

[0105] For example, a question triage module can be provided to determine the number of data segments required when the retrieval module transmits the data segments to the large language model. Different numbers of data segments can be provided for different types of questions. In other words, in this embodiment, the number of retrieved data segments can be flexibly adjusted based on the type of candidate question. This ensures that the majority of the responses to the candidate questions are included in the multimedia information while avoiding the introduction of excessive noise, thereby improving the accuracy of the responses.

[0106] In an optional embodiment of the present disclosure, the types of candidate questions can be divided into partial knowledge question-answering type (also known as question-answering), partial summary and abstract type (also known as summary) and unanswerable type (also known as rejection type). If the candidate question is: What is the height of Mount Everest? It can be determined that the type of candidate question is partial knowledge question-answering type, and the number of data segments required for this type of candidate question is relatively small. For example, the data segments that need to be retrieved can be 3, which can be the top 3 (Top3) data segments with semantic similarity to the candidate question in the segmented data segments. If the candidate question is: What are the main points introduced in this article? It can be determined that the type of candidate question is partial summary and abstract type, and the number of data segments required for this type of candidate question is relatively large. For example, the data segments that need to be retrieved can be 10, which can be the top 10 (Top10) data segments with semantic similarity to the candidate question in the segmented data segments. If the candidate question is: What is the temperature in Hangzhou today? It can be determined that the type of candidate question is unanswerable type, and this type of candidate question does not require data segments and can be answered according to pre-set standard answers.

[0107] As an optional implementation, the method may further include: displaying the reply content, the data segments used when generating the reply content, and the timestamps corresponding to the used data segments in the operation interface.

[0108] In this embodiment, after the reply content is determined, the reply content, the data segments used to generate the reply content, and the timestamps corresponding to the data segments used can be displayed in a streaming manner on the operation interface. The operation interface can be an interface of a mobile terminal, such as a display interface of a mobile terminal, such as an interface of a mobile phone, an interface of a computer, etc. This is for illustrative purposes only and does not impose any specific limitation on the type of operation interface.

[0109] Optionally, the candidate question and the assembled data fragments are used as input and transmitted to the large language model for inference to obtain the reply content. The reply content can be returned in a streaming manner by calling the Dash scope's Application Programming Interface (API). The reply content, the data fragments used to generate the reply content, and the timestamps corresponding to the used data fragments can be displayed to the user through the front-end dialogue interface.

[0110] As an optional implementation, the method may further include: in response to a touch operation on a timestamp in the operation interface, locating a data segment corresponding to the timestamp in the operation interface.

[0111] In this embodiment, the large language model supports answer tracing while giving the reply content. It can respond to the touch operation of the timestamp in the operation interface and locate multiple timestamps of the multimedia information to help users better understand the detailed information.

[0112] As an optional implementation, the method may also include: obtaining at least one historical text information based on the type of candidate question; segmenting the historical text information to obtain multiple groups of segmented texts; constructing training data based on the multiple groups of segmented texts; and using the training data to train a large language model.

[0113] In this embodiment, considering that candidate questions can be divided into multiple types, in order to improve the accuracy of the large language model in processing different types of candidate questions, training data matching the types of candidate questions can be constructed in different ways to train the large language model. The training data can be constructed in the following ways. The types of candidate questions can include at least single-document knowledge question-answering type, multi-document knowledge question-answering type, and summary type. It should be noted that the types of candidate questions are not specifically limited here.

[0114] Optionally, the type of candidate question is determined, and based on the type of candidate question, at least one historical text information is obtained; the historical text information can be segmented to obtain multiple groups of segmented texts; training data is constructed based on the multiple groups of segmented texts; and a large language model is trained using the training data.

[0115] As an optional implementation, training data is constructed based on multiple groups of segmented texts, including: in response to the type of the candidate question being a single-document knowledge question-and-answer type, retrieving historical query information of the segmented text and historical response content of the historical query information; segmenting the segmented text to obtain multiple sub-segmented texts; and determining the historical query information, historical response results, and multiple sub-segmented texts as training data.

[0116] In this embodiment, considering that the generative question-answering system is primarily oriented toward consumers (C), candidate questions for knowledge-based question-answering generally have the following characteristics: high accuracy requirements for the answers, a wide range of sources, possible cross-slice or cross-text situations, and a complex overall algorithm chain. Therefore, for candidate questions for knowledge-based question-answering, data production schemes can be designed, using both single-text and multi-text sources. The selection of inquiry information and the writing of answers are all performed manually, and the entire offline chain simulates the construction of a real online scenario.

[0117] Optionally, in response to the candidate question being a single-document knowledge question-answering type, for the construction of single-text training data, any historical text information can be obtained and segmented to obtain multiple sets of segmented texts. Any segmented text can be selected, and manually written historical query information related to the selected segmented text can be retrieved, as well as historical responses to the manually written historical query information. The segmented text can be segmented to obtain multiple sub-segmented texts, and the historical query information, historical response results, and multiple sub-segmented texts can be used to construct training data.

[0118] For example, given a historical text information, the historical text information can be segmented according to the length of the historical text information to obtain multiple segmented texts with a length of 2000. The multiple segmented texts are randomly sampled, and one of the segmented texts is randomly selected. A generative pre-trained transformer model (GPT) can be used to generate a list of three candidate questions related to the selected segmented text. A question can be randomly selected and answered manually to obtain historical inquiry information and historical reply content. Furthermore, the selected segmented text can be segmented into sub-segmented texts with a smaller length to obtain sub-segmented text one, sub-segmented text two, and sub-segmented text three. Sub-segmented text one, sub-segmented text two, sub-segmented text three, the selected historical inquiry information, and the historical reply content can be used as training data.

[0119] Optionally, data can be annotated in an intelligent data annotation platform (Information TAG, abbreviated as ITAG). For example, the platform can be given the name of a selected historical text message and automatically generated multiple historical query messages. The user can manually select the historical query message to answer and select data segments within the historical text message to obtain the historical response content.

[0120] In this embodiment, training data can also be constructed that includes historical query information that is more summary-oriented. This type of historical query information refers to a general type of question for a specific audio or video scenario, such as a meeting or speech. It can tend to summarize and generalize the key points of an article. Compared to knowledge-based questions, this presents an additional challenge: how to define general questions for audio or video scenarios. Therefore, to construct training data for this scenario, a seed scenario can be set. For example, the seed scenario can include five scenarios, such as meetings and speeches. GPT-4 can then be used to expand to other scenarios, such as talk shows and news broadcasts. For each scenario, candidate questions are generated. These can be general questions. For example, for a meeting scenario, multiple standard questions can be derived, such as "What is the goal of this meeting?", "Who are the participants in the meeting?", and "What core agenda items were discussed in the meeting?" After generating candidate questions for each scenario, the generated historical query information can be manually filtered and modified to control its quality, ultimately forming a fixed set of "scenario and question lists."

[0121] Furthermore, considering that writing a summary is more difficult for humans than for machines, in this embodiment, the GPT model can be used to generate the final training data. The overall process is similar to the process of constructing training data containing partial knowledge-based historical inquiry information. Historical text information one can be selected and sliced ​​to obtain segmented text one, segmented text two, ... Segmented text one, segmented text two, ... can be vectorized and transmitted to the retrieval system to construct an index. Based on historical text information one, the text can be predicted to be usable in certain scenarios, and based on the predicted scenarios, historical text information can be selected from the "scenario, question list". Based on the historical text information, segmented texts associated with the historical text information question can be retrieved from the retrieval system. For example, the top ten segmented texts with a correlation degree higher than the correlation degree threshold can be retrieved to obtain the "historical text information, segmented text list". The final historical reply content can be generated based on the "historical text information, segmented text list". Training data can be constructed based on the selected historical text information, segmented text list, and historical reply content.

[0122] As an optional implementation, training data is constructed based on multiple groups of segmented texts, including: in response to the type of the candidate question being a multi-document knowledge question-answering type, selecting a group of segmented texts from the multiple groups of segmented texts, and retrieving historical query information of the selected segmented texts, as well as historical reply content of the historical query information; from the multiple groups of segmented texts, retrieving segmented texts whose matching degree with the selected segmented texts is higher than a matching degree threshold; and determining the selected segmented texts, segmented texts whose matching degree is higher than a matching degree threshold, historical query information, and historical reply content as training data.

[0123] In this embodiment, in addition to using single-text data to construct training data, multi-text data can also be used to construct training data for partial knowledge-based training data. Furthermore, in response to the candidate question being a single-document knowledge question-answering type, multiple pieces of historical text information can be obtained, and the multiple pieces of historical text information can be segmented to obtain multiple groups of segmented texts. A group of segmented texts can be selected from the multiple groups of segmented texts, and the historical query information of the selected segmented texts and the historical reply content of the historical query information can be retrieved; among the multiple groups of segmented texts, the segmented texts whose matching degree with the selected segmented texts is higher than the matching degree threshold are retrieved; the selected segmented texts, the segmented texts whose matching degree is higher than the matching degree threshold, the historical query information, and the historical reply content are determined as training data.

[0124] Optionally, when the reply content of the inquiry information raised by the user comes from multiple texts, it is necessary to automatically aggregate and summarize the relevant information of the reply content. Therefore, in order to improve the accuracy of the large language model in processing data fragments from multiple articles, in this embodiment, a larger text library can be constructed. For example, the text library can include 1,000 texts. Obtain historical text information one, historical text information two, and historical text information three. Segment the three historical text information respectively to obtain multiple segmented texts. Transfer the segmented texts to the retrieval system to build an index (also known as an ES index). Randomly sample segmented text two from the multiple segmented texts, and query the retrieval system for segmented text three and segmented text five that have an associated relationship with segmented text two. Based on segmented text two, automatically generate multiple historical inquiry information. From the multiple historical inquiry information, select historical inquiry information two, and manually answer the historical reply content of historical inquiry information two. Segmented text two, segmented text three, segmented text five, historical inquiry information, and historical reply content can be used as training data.

[0125] In this embodiment, the training data constructed as described above can be used to complete the supervised fine tuning (SFT) training task based on Tongyi Qianwen, and ultimately obtain a large language model that can perform data prediction and processing.

[0126] In an embodiment of the present disclosure, when an inquiry message is received, the inquiry message is rewritten into a candidate question for executing a customized question, a data segment associated with the candidate question is determined from the multimedia information, and the data segment is processed using a large language model to obtain at least one reply content that matches the inquiry message, thereby achieving the technical effect of being able to provide personalized answers to the content of the multimedia information and solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0127] From the perspective of human-computer interaction, the present embodiment also provides another method for processing multimedia information, which can be applied to a generative question-answering system. FIG4 is a flowchart of another method for processing multimedia information according to an embodiment of the present invention. As shown in FIG4 , the method may include the following steps.

[0128] Step S402 : Responding to an input instruction on the operation interface, obtaining input query information.

[0129] In the technical solution provided in the above step S402 of the present disclosure, input query information may be obtained in response to an input instruction on the operation interface, wherein the input instruction may be triggered by the user and used to transmit the query information to the generative question-answering system.

[0130] Optionally, the user inputs query information in the operation interface. After the query information is input, the user clicks an input instruction on the operation interface. In response to the input instruction acting on the operation interface, the input query information can be obtained.

[0131] Step S404 , in response to a query instruction applied to the operation interface, rewrite the query information into candidate questions for performing customized questioning on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system.

[0132] In the technical solution provided in step S404 of the present disclosure, the user can trigger the query control to issue a query instruction. In response to the query instruction on the operation interface, the query information can be rewritten into at least one candidate question for performing a customized question on the multimedia information.

[0133] Step S406: query and obtain data segments associated with the candidate question from the question-answer database, wherein the question-answer database stores at least one set of data segments, and the data source of the data segments is multimedia information.

[0134] Step S408: Process the data segment using the large language model to generate at least one response content that matches the query information.

[0135] Step S410: Display the reply content on the operation interface.

[0136] In an embodiment of the present disclosure, in response to an input instruction acting on an operation interface, input query information is obtained; in response to a query instruction acting on the operation interface, the query information is rewritten into at least one candidate question for performing customized questions on multimedia information, wherein the multimedia information is content pre-uploaded to a generative question-answering system; data segments associated with the candidate questions are queried from a question-answering database, wherein the question-answering database stores multiple groups of data segments, and the data source of the data segments is the multimedia information; the data segments are processed using a large language model to generate at least one reply content that matches the query information; and the reply content is displayed on the operation interface, thereby achieving a technical effect of providing personalized answers to the content of the multimedia information and solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0137] Currently, generative question-answering systems typically use AI assistants for work and learning to answer inquiries. With the continuous development of generative question-answering systems, new requirements have been placed on them for audio and video content scenarios, such as meetings, classes, interviews, training, live broadcasts, watching videos, and listening to podcasts. Generative question-answering systems are required to automatically take notes on audio data and organize interviews. However, most current AI functions are invoked through built-in triggering opportunities, and the results of the AI ​​functions are displayed directly on the product page. Users cannot actively interact with the large model, making it impossible to provide users with personalized audio and video question-answering capabilities. This leads to the technical problem of being unable to provide personalized answers to multimedia data content.

[0138] To solve the above problems, the present disclosure proposes a system and method for building an intelligent question-and-answer assistant based on audio and video. This method uses artificial intelligence technologies such as large models to quickly extract and precipitate knowledge from uploaded audio and video for audio and video content scenarios, helping to efficiently complete the transcription, retrieval, summary and organization of audio and video content anytime and anywhere. For example, large models can be used to automatically take notes, organize interviews, extract PPTs, etc. Users can ask customized questions about single or multiple uploaded audio records, and can also search, review and summarize historical audio records, thereby helping users to better understand the current demand scenario in a flexible and convenient way, and complete quick questions and answers and searches through a more natural human-computer interaction method.

[0139] In this embodiment, users can ask multiple rounds of questions directly through natural language, and the large model can provide accurate and reliable answers based on the audio and video content, thereby helping users understand the audio and video content in a more flexible, in-depth, and more personalized way. The method includes system links (offline and online) and specific algorithm capability construction, covering a complete set of solutions from audio and video uploading, transcription, slicing, retrieval to understanding. In this embodiment, users can ask multiple rounds of questions through natural language, and the large model can provide accurate and reliable answers based on the audio and video content, thereby helping users understand the audio and video content in a more flexible, in-depth, and more personalized way, thereby achieving the technical effect of providing personalized answers to the content of multimedia information, and solving the technical problem of not being able to provide personalized answers to the content of multimedia data.

[0140] The following further introduces a personalized clothing size recommendation method based on user body shape and historical purchasing behavior proposed in an embodiment of the present disclosure.

[0141] Figure 5 is a schematic diagram of a multimedia information processing system according to an embodiment of the present disclosure. As shown in Figure 5, the multimedia information processing system 50 may include: an offline link and an online link. Among them, the offline link is mainly used to process the uploaded multimedia information. After the audio data is uploaded, the transcription, slicing, vectorization and storage of the audio data are completed through the offline link. The online link is mainly used to determine the reply content of the user's inquiry information, which may include rewriting the inquiry information, vectorizing the inquiry information, retrieving data segments related to the inquiry information, sorting the data segments, diverting the data segments, assembling the data segments, and determining the reply content through large model reasoning.

[0142] As an optional embodiment, the user's query information is obtained and rewritten to obtain candidate questions.

[0143] In this embodiment, as shown in FIG5 , a user query 501 is obtained. Based on multiple rounds of conversation samples stored by the generative question-answering system over a historical period, the query 501 is rewritten to obtain a candidate question 502. The candidate question can also be referred to as a standard query. The query information can also be referred to as a user query.

[0144] Optionally, since the user will omit the previously mentioned context information during the multi-round interaction with the LLM model (which can also be called the big model in this embodiment), and only give a simple instruction. For example, the first round of questions is: "Which buildings in Hangzhou are mentioned in the article?", the big model answers: "It mentions IFC and xxx in Hangzhou", and then the user asks: "When was this built?" It can be seen that the second round of user query itself does not bring in a subject. If the query information is directly used for ES retrieval, the recall result will be inaccurate. Therefore, in order to avoid the above problems, this embodiment proposes to use open source multi-round dialogue training data to add a multi-round rewriting task based on the big model, so that after the query information 501 and historical information are input, a more semantically complete and rich candidate question 502 can be obtained, so that more accurate related chunk fragments can be recalled. Among them, the above-mentioned big model can be an LLM model. Context information includes historical query information or historical reply content.

[0145] As an optional embodiment, the candidate questions are divided and the number of data segments required to be retrieved to answer the candidate questions is determined.

[0146] In this embodiment, the candidate questions 502 are divided, and based on the results of the division, the number of data segments that the retrieval module 504 needs to retrieve for the candidate questions 502 is determined.

[0147] Because the types of query information input by users vary, to improve data processing efficiency, this embodiment provides a question splitting module 503 to determine the number of data segments required when the retrieval module 504 inputs the LLM model. Different numbers of data segments are provided for different types of questions. In other words, in this embodiment, the number of retrieved data segments can be flexibly adjusted based on the type of candidate question. This ensures that the majority of the responses to the candidate questions are included in the multimedia information, avoiding the introduction of excessive noise information and thus achieving the goal of improving the accuracy of the response content.

[0148] Optionally, the retrieval module 504 may be a distributed search engine.

[0149] Optionally, the candidate questions 502 may be judged by the question distribution module 503 to determine the number of data segments that the retrieval module 504 needs to provide.

[0150] For example, candidate questions can be categorized as more knowledge-based, more summary-based, and unanswerable. For example, if the candidate question is: "What is the height of Mount Everest?", it can be determined that the candidate question type is more knowledge-based. This type of candidate question requires fewer data segments. For example, three data segments may need to be retrieved, and these can be the top three data segments with the semantic similarity to the candidate question after segmentation. If the candidate question is: "What are the main points of this article?", it can be determined that the candidate question type is more summary-based. This type of candidate question requires more data segments. For example, ten data segments may need to be retrieved, and these can be the top ten data segments with the semantic similarity to the candidate question after segmentation. If the candidate question is: "What is the temperature in Hangzhou today?", it can be determined that the candidate question type is unanswerable. This type of candidate question does not require data segments and can be answered based on a pre-set standard answer.

[0151] As an optional implementation, the data segments associated with the candidate questions are determined by the retrieval module, and the selected data segments are assembled.

[0152] In this embodiment, the retrieval module 504 determines multiple groups of data segments associated with candidate questions in the multimedia information, and constructs the input of the LLM model based on the selected multiple groups of data segments.

[0153] Since the input of the LLM model usually has a certain length limit, the recalled multiple related fragments cannot all be put into the input. For this reason, a ranking model can be used to score the relevance of the candidate questions and the data fragments (chunks), and the scores are arranged from high to low. The data fragments within the preset ranking are filtered out, and the data fragments within the preset ranking are put into the prompt information assembly unit 507 (prompt assembly) one by one according to the generation time for assembly until the assembly result exceeds the input length limit.

[0154] Optionally, as shown in Figure 5, the feature vector 519 of the candidate information 502 is determined, and the feature vector is transmitted to the retrieval module 504. The retrieval module 504 determines the data fragment 513, data fragment 514 and data fragment 515 associated with the candidate question through the feature vector 519, determines the similarity between the data fragment and the candidate question, and sorts the multiple data fragments based on the similarity to obtain data fragment 516, data fragment 517 and data fragment 518.

[0155] For example, the Transformer-based encoder BERT in ranking model 520 can be used to simultaneously encode and calculate candidate questions and data segments, output a relevance score, and determine the data segments associated with the candidate questions based on the relevance score. After determining the data segments associated with the candidate questions, the data segments can be assembled according to the time of their generation.

[0156] As an optional embodiment, the assembled data segments are transmitted to the LLM model for reasoning to obtain the reply content.

[0157] In this embodiment, the candidate questions and the assembled data segments may be used as input and transmitted to a large language model (also referred to as an LLM model) to process and obtain the answer content of the candidate questions.

[0158] For example, by calling the model-as-a-service application program interface, the response content can be returned in a streaming manner and displayed to the user through the front-end dialogue interface.

[0159] As an optional embodiment, after obtaining audio data (Audio), the audio data can be transcribed to obtain text data 506, and the text data 506 can be sliced ​​to obtain data segments 507, 508, and 509. The audio data can be uploaded by the user or recorded in real time. This is only an example and does not specifically limit the method of obtaining the audio data.

[0160] In this embodiment, since the input length of the LLM model is limited, the text data after each audio data transcription needs to be sliced ​​to obtain data segments that do not exceed a specific size.

[0161] Optionally, the text data may be segmented into slices based on length and / or semantic information. The size of the slices may not only consider the input length during LLM reasoning, but also the input length allowed by the model during subsequent vectorization and sorting. For example, the slice size may be 300 to 500 words.

[0162] As an optional embodiment, the segmented data segments may be vectorized and the vectorized data may be stored in a distributed search engine for retrieval.

[0163] 5 , data segments 507 , 508 , and 509 may be vectorized to obtain feature vectors 510 , 511 , and 512 . Feature vectors 510 , 511 , and 512 are stored in retrieval module 504 .

[0164] Optionally, the model can convert discrete natural language data fragments into continuous feature vectors to facilitate subsequent semantic-based retrieval. For example, the Transformer-based encoder BERT can be used to vectorize the data fragments. The vectorized data fragments are stored in Elastic Search. For online candidate information that has also been vectorized, multiple recall methods, such as BM25-based word matching and Proxima-based semantic similarity, can be used to find multiple relevant data fragments as reference input for the large model to answer the query information.

[0165] Considering that the types of user candidate questions can be divided into knowledge-based and summary-based types, in order to improve the accuracy of the LLM model in processing different types of candidate questions, in this embodiment, multiple types of training data are constructed to train the LLM model.

[0166] As an optional embodiment, training data containing historical query information of a partial knowledge question-answering type is constructed.

[0167] In this embodiment, considering that the generative question-answering system is primarily oriented towards the consumer side (C-side for short), its historical query information, which is biased towards knowledge-based question-answering, generally has the following characteristics: high accuracy requirements for the response content, a wide range of sources, possible cross-slice or cross-text situations, and a complex overall algorithm chain. Therefore, for historical query information that is biased towards knowledge-based question-answering, data production schemes can be designed with both single-text and multi-text sources, and all query information selection and response content writing are performed manually, with the entire offline chain simulating the construction of real online scenarios.

[0168] FIG6 is a schematic diagram of constructing training data according to an embodiment of the present disclosure. As shown in FIG6 , given a text 601, it can be segmented according to the length of the text to obtain multiple segmented texts of length 2000, such as segmented text 602, segmented text 603, and segmented text 604. The multiple segmented texts are randomly sampled, and one of the segmented texts 603 is randomly selected. Three lists of historical query information are generated based on segmented text 603 by GPT-4, wherein the historical query information lists may include historical query information 605, historical query information 606, and historical query information 607. The historical query information can be randomly selected manually and answered 606 to obtain the answer content 608, thereby ensuring the quality of the historical query information and the answer content. Furthermore, the selected segmented text can be segmented into sub-segmented texts of smaller length, such as sub-segmented text 609, sub-segmented text 610, and sub-segmented text 611. The sub-segmented text 609 , the sub-segmented text 610 , the sub-segmented text 611 , the selected historical inquiry information, and the reply content may be used as training data.

[0169] Alternatively, FIG7 is a schematic diagram of manual annotation according to an embodiment of the present disclosure. As shown in FIG7 , data annotation can be performed in the intelligent data annotation platform. For example, the name of the selected historical text information (i.e., the file name) and the automatically generated multiple historical query information can be given in the intelligent data annotation platform 70. The historical query information to be answered can be manually selected, and the data fragments in the historical text information can be selected to obtain the historical reply content.

[0170] In this embodiment, in addition to using single-text data to construct the training data, multi-text data can also be used to construct the training data for partial knowledge-based training data.

[0171] Optionally, when the response content of the user's inquiry information comes from multiple texts, it is necessary to automatically aggregate and summarize the relevant information of the response content. Therefore, in order to improve the accuracy of the LLM model in processing data fragments from multiple pieces, in this embodiment, a larger text library is obtained, for example, the text library may include 1,000 texts. Figure 8 is a schematic diagram of constructing training data according to an embodiment of the present disclosure. As shown in Figure 8, text 801, text 802, and text 803 are obtained. The three texts are segmented respectively. After text 801 is segmented, segmented text 804, segmented text 805, ... are obtained; after text 802 is segmented, segmented text 806, segmented text 807, ... are obtained; after text 803 is segmented, segmented text 808, segmented text 809, ... are obtained. The segmented texts are transmitted to the retrieval system 810 to construct an index (also known as an ES index). Segmented text 806 is randomly sampled, and segmented text 807 and segmented text 808 that have an associated relationship with segmented text 806 are queried in the retrieval system. Based on segmented text 806, historical query information 811, historical query information 812, and historical query information 813 are automatically generated. Historical query information 812 is selected, and historical response content 814 for historical query information 812 is manually answered. Segmented text 806, segmented text 807, segmented text 808, historical query information 812, and historical response information 813 can be used as training data.

[0172] As an optional embodiment, training data containing historical query information of a summary type is constructed.

[0173] In this embodiment, historical queries that tend to be summary-oriented refer to general questions about specific audio and video scenarios, such as those in meetings and lectures. These questions tend to summarize and generalize the key points of the article. Compared to knowledge-based questions, this presents an additional challenge: how to define general questions in audio and video scenarios.

[0174] FIG9 is a schematic diagram of another method of constructing training data according to an embodiment of the present disclosure. As shown in FIG9 , a seed scenario can be set first. For example, the seed scenario can include five scenarios such as meetings and speeches. Other scenarios can be expanded through GPT-4. For example, multiple fine-grained scenarios such as talk shows and news broadcasts can be further expanded. Historical query information is generated for each set scenario, where the historical query information can be a general question. For example, for a meeting scenario, multiple standard questions such as "What is the goal of this meeting?", "Who are the participants in the meeting?", and "What core agendas were discussed in the meeting?" can be derived. After generating the historical query information for each scenario, in order to control the quality of the historical query information, the generated historical query information list can be manually filtered and modified to eventually form a fixed set of "scenario, question list".

[0175] Considering that writing summaries is more difficult for humans than for machines, in this embodiment, the GPT model can be used to generate the final training data. The overall process is similar to the process of constructing training data containing knowledge-based historical query information. As shown in Figure 9, text 901 can be selected and sliced ​​to obtain segmented text 902, segmented text 903, ..., segmented text 90n. Segmented text 902, segmented text 903, ..., segmented text 90n can be transmitted to a retrieval system 91 to construct an index. Based on text 901, the text's usage scenario is predicted. Based on the predicted scenario, historical query information is selected from the "scenario, question list". Based on the historical query information, segmented texts associated with the historical query information are retrieved from the retrieval system. For example, the top 10 segmented texts with a degree of relevance above a relevance threshold can be retrieved to obtain a "historical query information, segmented text list". The final response content is generated based on this "historical query information, segmented text list". Training data can be constructed based on the selected historical query information, segmented text list, and response content.

[0176] Optionally, the LLM model can process historical query information in a meeting scenario. In this scenario, the LLM model can be used to complete meeting minutes, record to-do items in the meeting, and determine the conclusions of the meeting. It can also process historical query information in a speech scenario. In this scenario, the LMM model can be used to determine the golden sentences of the speech, speech cases, and the core ideas during the speech. It can also process historical query information in a learning scenario. In this scenario, the LLM model can be used to determine the core knowledge points in learning and reading suggestions. It can also process historical query information in a general scenario. In this scenario, the LLM model can be used to determine milestone events, key data, concept explanations, etc. Figure 10 is a schematic diagram of a meeting minutes according to an embodiment of the present disclosure. As shown in Figure 10, the LLM model can generate meeting minutes based on audio data during the meeting. The meeting minutes can include the meeting title, participants, main content and to-do items, and the referenced data segments can be represented in the meeting minutes in the form of labels.

[0177] In this embodiment, the training data constructed as described above can be used to complete the supervised fine tuning (SFT) training task based on Tongyi Qianwen, and finally obtain an LLM model that can perform data prediction and processing.

[0178] Figure 11 is a schematic diagram of an interactive scenario according to an embodiment of the present disclosure. As shown in Figure 11, a question-and-answer service can be provided in the interactive form of a chat bot in the operation interface 1101, and inspiration for asking questions can be provided to users. Figure 12 is a schematic diagram of a reply content according to an embodiment of the present disclosure. As shown in Figure 12, while the inquiry information 1205 and the reply content 1206 are displayed on the operation interface 1201, answer tracing is also supported. The multimedia information 1202, multimedia information 1203, and multimedia information 1204 used to determine the reply content can be displayed, as well as at least one timestamp corresponding to each multimedia. By clicking on the timestamp in the operation interface, the corresponding location can be located to help users better understand the detailed information. Figure 13 is a schematic diagram of a recommended question according to an embodiment of the present disclosure. As shown in Figure 13, in order to lower the user's usage threshold, the operation interface 1301 can also specifically list question templates for various scenarios in the question inspiration section to guide users to ask questions and think.

[0179] In this embodiment, considering the high degree of colloquialism of the text after audio and video transcription, a multi-way recall based on ES is introduced to retrieve candidate chunks at the word level and semantic level, covering as many fragment contents as possible, thereby achieving the purpose of improving the accuracy of the retrieved data fragments. In addition, multiple rounds of interaction are carried out through natural language and a large language model. Considering that it is easy to omit the context that has been mentioned during communication, a rewrite is introduced to take into account the current sentence and historical input and output, and restore it to a user query with rich context and complete semantic information, thereby improving the recall rate and accuracy of the results in the retrieval step. At the same time, by introducing more input information in a diversion manner, the coverage of the answer can be guaranteed while the link is reused, thereby achieving the technical effect of providing personalized answers to the content of multimedia information, and solving the technical problem of not being able to provide personalized answers to the content of multimedia data.

[0180] The method embodiment provided in Example 1 of the present disclosure can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 14 is a hardware structure block diagram of a computer terminal (or mobile device) according to a method for processing multimedia information in accordance with an embodiment of the present disclosure. As shown in Figure 14, the computer terminal 140 (or mobile device) may include one or more (1402a, 1402b, ..., 1402n are used in the figure to illustrate) processors 1402 (the processor 1402 may include but is not limited to a microprocessor (Microcontroller Unit, referred to as MCU) or a programmable logic device (Field Programmable Gate Array, referred to as FPGA) and other processing devices), a memory 1404 for storing data, and a transmission device 1406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 14 is only for illustration and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 140 may also include more or fewer components than shown in FIG. 14 , or have a configuration different from that shown in FIG. 14 .

[0181] The hardware structure block diagram shown in Figure 14 can not only serve as an exemplary block diagram of the above-mentioned computer terminal 140 (or mobile device), but also as an exemplary block diagram of the above-mentioned server. In an optional embodiment, Figure 2 shows in a block diagram an embodiment of using the computer terminal 140 (or mobile device) shown in Figure 14 as a computing node in the computing environment 201.

[0182] The memory 1404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method in the embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 1404, that is, realizing the above-mentioned data processing method. The memory 1404 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1404 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 140 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0183] Transmission device 1406 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of computer terminal 140. In one embodiment, transmission device 1406 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 1406 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0184] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 140 (or mobile device). In another optional embodiment, FIG15 shows a block diagram of an embodiment using the computer terminal 140 (or mobile device) shown in FIG1 as a service grid. FIG15 is a structural block diagram of a service grid according to an embodiment of the present disclosure. As shown in FIG15, the service grid 1500 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them to run on different clusters / machines.

[0185] As shown in Figure 15 , microservices may include application service instance A and application service instance B, which form the functional application layer of service grid 1500. In one embodiment, application service instance A runs as a container / process 1508 on a machine / workload container group 1514 (POD), and application service instance B runs as a container / process 1510 on a machine / workload container group 1516 (POD).

[0186] In one implementation, application service instance A may be an inquiry information query service, and application service instance B may be an audio conversion service.

[0187] As shown in Figure 15 , application service instance A and grid proxy (sidecar) 1503 coexist in machine workload container group 1514, while application service instance B and grid proxy 1505 coexist in machine workload container 1514. Grid proxy 1503 and grid proxy 1505 form the data plane layer of service grid 1500. Grid proxy 1503 and grid proxy 1505 each run as container / process 1504, which can receive requests 1512 for querying information, and grid proxy 1506. Bidirectional communication is possible between grid proxy 1503 and application service instance A, and between grid proxy 1505 and application service instance B. Furthermore, bidirectional communication is possible between grid proxy 1503 and grid proxy 1505.

[0188] In one embodiment, all traffic for application service instance A is routed to the appropriate destination via grid proxy 1503, and all network traffic for application service instance B is routed to the appropriate destination via grid proxy 1505. It should be noted that network traffic mentioned herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), a high-performance, general-purpose open source framework (g-RPC), and an open source in-memory data structure storage system (Redis).

[0189] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in service mesh 1500. Service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 1503 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.

[0190] As shown in Figure 15 , the service grid 1500 also includes a control plane layer. The control plane layer can be comprised of a set of services running in a dedicated namespace, hosted by a managed control plane component 1501 within a machine / workload container group (machine / pod) 1502. As shown in Figure 15 , managed control plane component 1501 communicates bidirectionally with grid agent 1503 and grid agent 1505. Managed control plane component 1501 is configured to perform certain control and management functions. For example, managed control plane component 1501 receives telemetry data transmitted by grid agent 1503 and grid agent 1505 and can further aggregate this telemetry data. Managed control plane component 1501 can also provide user-oriented application programming interfaces (APIs) for easily manipulating network behavior and providing configuration data to grid agent 1503 and grid agent 1505.

[0191] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for the object to choose to authorize or refuse.

[0192] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0193] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present disclosure.

[0194] According to an embodiment of the present disclosure, a multimedia information processing device for implementing the multimedia information processing method shown in FIG. 3 is also provided.

[0195] Figure 16 is a schematic diagram of a multimedia information processing device according to an embodiment of the present disclosure. As shown in Figure 16, the multimedia information processing device 1600 may include: a receiving component 1602, a first rewriting component 1604, a first query component 1606 and a first generation component 1608.

[0196] The receiving component 1602 is configured to receive query information.

[0197] The first rewriting component 1604 is configured to rewrite the query information into candidate questions for performing customized questioning on multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system.

[0198] The first query component 1606 is configured to query and obtain data segments associated with the candidate question from the question and answer database, wherein the question and answer database stores at least one set of data segments, and the data source of the data segments is multimedia information.

[0199] The first generating component 1608 is configured to process the data segment using a large language model to generate at least one response content matching the query information.

[0200] Here, the receiving component 1602, the first rewriting component 1604, the first query component 1606, and the first generating component 1608 correspond to steps S302 to S308 in Example 1. The four components and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above components can be hardware components or software components stored in a memory (e.g., memory 1404) and processed by one or more processors (e.g., processors 1402a, 1402b..., 1402n). The above components can also be part of the device and can be run in the computer terminal 140 provided in Example 2.

[0201] According to an embodiment of the present disclosure, there is also provided a multimedia information processing device configured to implement the multimedia information processing method shown in FIG. 4 .

[0202] Figure 17 is a schematic diagram of another multimedia information processing device according to an embodiment of the present disclosure. As shown in Figure 17, the multimedia information processing device 1700 may include: an acquisition component 1702, a second rewriting component 1704, a second query component 1706, a second generation component 1708 and a display component 1710.

[0203] The acquisition component 1702 is configured to respond to an input instruction acting on the operation interface and acquire input query information.

[0204] The second rewriting component 1704 is configured to respond to a query instruction on the operation interface and rewrite the query information into a candidate question configured to perform a custom question on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system.

[0205] The second query component 1706 is configured to query and obtain data segments associated with the candidate question from the question-answer database, wherein the question-answer database stores at least one set of data segments, and the data source of the data segments is multimedia information.

[0206] The second generating component 1708 is configured to process the data segment using the large language model to generate at least one response content matching the query information.

[0207] The display component 1710 is configured to display the reply content on the operation interface.

[0208] It should be noted that the five components, namely, the acquisition component 1702, the second rewriting component 1704, the second query component 1706, the second generation component 1708, and the display component 1710, are the same as the examples and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above-mentioned embodiment 1. It should be noted that the above-mentioned components can be hardware components or software components stored in a memory (e.g., the memory 1404) and processed by one or more processors (e.g., processors 1402a, 1402b, ..., 1402n). The above-mentioned components can also be part of the device and can be run in the computer terminal 140 provided in embodiment 2.

[0209] In the multimedia information processing device, when an inquiry message is received, the inquiry message is rewritten into a candidate question for executing a customized question, a data segment associated with the candidate question is determined from the multimedia information, and the data segment is processed using a large language model to obtain at least one reply content that matches the inquiry message, thereby achieving the technical effect of being able to provide personalized answers to the content of the multimedia information and solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0210] The embodiment of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.

[0211] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0212] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the multimedia information processing method: receiving query information; rewriting the query information into candidate questions for performing customized questions on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; querying from the question-answering database to obtain data segments associated with the candidate questions, wherein the question-answering database stores at least one set of data segments, and the data source of the data segments is the multimedia information; using the large language model to process the data segments to generate at least one reply content that matches the query information.

[0213] Optionally, Figure 18 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure. As shown in Figure 18, the computer terminal A may include: one or more (only one is shown in the figure) processors 1802, a memory 1804 and a transmission device 1806.

[0214] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the multimedia information processing method and device in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned multimedia information processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0215] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving query information; rewriting the query information into candidate questions for performing customized questions on multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; querying and obtaining data segments associated with the candidate questions from the question-answering database, wherein the question-answering database stores at least one group of data segments, and the data source of the data segments is the multimedia information; using the large language model to process the data segments to generate at least one reply content that matches the query information.

[0216] Optionally, the processor may also execute program code for the following steps: obtaining multi-round dialogue samples stored by the generative question-answering system within a historical time period, wherein the multi-round dialogue samples record at least a plurality of semantically related historical query information occurring within the historical time period; identifying at least one historical query information matching the query information from the multi-round dialogue samples; and rewriting the query information based on the at least one historical query information to generate at least one candidate question.

[0217] Optionally, the processor may also execute the program code of the following steps: obtaining uploaded multimedia information; performing text conversion on the multimedia information to generate text data in text format; segmenting the text data to obtain at least one group of data segments; constructing a question-and-answer database based on at least one group of data segments, wherein each group of data segments is stored in the question-and-answer database in the form of a feature vector.

[0218] Optionally, the processor may further execute program code for the following steps: identifying the length of text data and / or semantic information of the text data; and segmenting the text data based on the length and / or semantic information to obtain at least one set of data segments.

[0219] Optionally, the processor may also execute the program code of the following steps: reading at least one set of data segments from the question-and-answer database; obtaining the semantic similarity between each set of data segments and the candidate question; and determining the screened data segments whose semantic similarity exceeds a similarity threshold as data segments associated with the candidate question.

[0220] Optionally, the processor may also execute the following program code: obtaining the feature vectors of each group of data segments and the feature vectors of the candidate questions; performing similarity calculations on the feature vectors of each group of data segments and the feature vectors of the candidate questions to obtain the semantic similarity between each group of data segments and the candidate questions.

[0221] Optionally, the processor may also execute program code for the following steps: sorting all screened data segments associated with the candidate questions according to the semantic similarity between the data segments and the candidate questions; extracting at least one group of data segments within a preset ranking according to the sorting result; reading the generation time of each group of data segments within the preset ranking; assembling each group of data segments in sequence according to the generation time of each group of data segments within the preset ranking; and processing the assembled data segments using a large language model to generate at least one reply content that matches the query information.

[0222] Optionally, the processor may also execute the program code of the following steps: based on the type of the candidate question, determine whether to perform a query operation on the data segments recalled from the question-and-answer database; if a query operation is performed, determine the number of retrievals when the query operation is performed; based on the number of retrievals, determine the expected number of data segments obtained by the query, wherein the expected number of segments is used to determine the number of data segments that match the candidate question.

[0223] Optionally, the processor may further execute program code for the following steps: displaying the reply content, the data segments used when generating the reply content, and the timestamps corresponding to the used data segments in the operation interface.

[0224] Optionally, the processor may further execute program code of the following steps: in response to a touch operation on a timestamp in the operation interface, locating a data segment corresponding to the timestamp in the operation interface.

[0225] Optionally, the processor may also execute the following program code: based on the type of the candidate question, obtain at least one historical text information; segment the historical text information to obtain multiple groups of segmented texts; construct training data based on the multiple groups of segmented texts; and use the training data to train a large language model.

[0226] Optionally, the processor may also execute the following program code: in response to the candidate question being a single-document knowledge question-answering type, obtaining any piece of historical text information, and segmenting the historical text information to obtain multiple groups of segmented texts, and then retrieving the historical query information of the segmented text, as well as the historical response content of the historical query information; segmenting the segmented text to obtain multiple sub-segmented texts; and determining the historical query information, historical response results, and multiple sub-segmented texts as training data.

[0227] Optionally, the processor may also execute the program code of the following steps: in response to the type of the candidate question being a multi-document knowledge question-answering type, selecting a group of segmented texts from multiple groups of segmented texts, and retrieving the historical query information of the selected segmented texts, as well as the historical reply content of the historical query information; among the multiple groups of segmented texts, retrieving the segmented texts whose matching degree with the selected segmented texts is higher than the matching degree threshold; determining the selected segmented texts, the segmented texts whose matching degree is higher than the matching degree threshold, the historical query information, and the historical reply content as training data.

[0228] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: in response to the input instruction acting on the operation interface, obtain the input query information; in response to the query instruction acting on the operation interface, rewrite the query information into candidate questions for performing customized questions on multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; from the question-answering database, query and obtain data segments associated with the candidate questions, wherein the question-answering database stores at least one group of data segments, and the data source of the data segments is the multimedia information; use the large language model to process the data segments to generate at least one reply content that matches the query information; and display the reply content on the operation interface.

[0229] According to the embodiment of the present disclosure, when an inquiry message is received, the inquiry message is rewritten into a candidate question for executing a customized question, a data segment associated with the candidate question is determined from the multimedia message, and the data segment is processed using a large language model to obtain at least one reply content that matches the inquiry message, thereby achieving the technical effect of being able to provide personalized answers to the content of the multimedia message and solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0230] Those skilled in the art will appreciate that the structure shown in FIG18 is merely illustrative, and that computer terminal A may also be a smartphone (e.g., an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG18 does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may include more or fewer components (e.g., a network interface, a display device, etc.) than those shown in FIG18 , or may have a configuration different from that shown in FIG18 .

[0231] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0232] The embodiment of the present disclosure further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the multimedia information processing method provided in the first embodiment.

[0233] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0234] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: receiving an inquiry message; rewriting the inquiry message into a candidate question for performing a customized question on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question-answering system; querying and obtaining data segments associated with the candidate question from a question-answering database, wherein the question-answering database stores at least one set of data segments, and the data source of the data segments is the multimedia information; and processing the data segments using a large language model to generate at least one reply content that matches the inquiry message.

[0235] Optionally, the computer-readable storage medium may also execute program code for the following steps: obtaining multi-round dialogue samples stored by the generative question-answering system within a historical time period, wherein the multi-round dialogue samples record at least a plurality of semantically associated historical query information occurring within the historical time period; identifying at least one historical query information matching the query information from the multi-round dialogue samples; and rewriting the query information based on the at least one historical query information to generate at least one candidate question.

[0236] Optionally, the computer-readable storage medium can also execute the program code of the following steps: obtaining uploaded multimedia information; converting the multimedia information into text to generate text data in text format; segmenting the text data to obtain at least one group of data segments; constructing a question-and-answer database based on at least one group of data segments, wherein the data segments are stored in the question-and-answer database in the form of feature vectors.

[0237] Optionally, the computer-readable storage medium may also execute program code for the following steps: identifying the length of text data and / or semantic information of the text data; and segmenting the text data based on the length and / or semantic information to obtain at least one set of data segments.

[0238] Optionally, the computer-readable storage medium may also execute program code for the following steps: reading at least one set of data segments from a question-and-answer database; obtaining the semantic similarity between each set of data segments and the candidate question; and determining the screened data segments whose semantic similarity exceeds a similarity threshold as data segments associated with the candidate question.

[0239] Optionally, the computer-readable storage medium may also execute program code for the following steps: obtaining feature vectors of each group of data segments and feature vectors of candidate questions; performing similarity calculations on the feature vectors of each group of data segments and the feature vectors of the candidate questions to obtain semantic similarity between each group of data segments and the candidate questions.

[0240] Optionally, the computer-readable storage medium may also execute program code for the following steps: sorting all screened data segments associated with the candidate questions according to the semantic similarity between the data segments and the candidate questions; extracting at least one group of data segments within a preset ranking according to the sorting results; reading the generation time of each group of data segments within the preset ranking; assembling each group of data segments in sequence according to the generation time of each group of data segments within the preset ranking; and processing the assembled data segments using a large language model to generate at least one reply content that matches the query information.

[0241] Optionally, the computer-readable storage medium may also execute program code for the following steps: based on the type of candidate question, determine whether to perform a query operation on the data segments recalled from the question-and-answer database; if a query operation is performed, determine the number of retrievals when the query operation is performed; based on the number of retrievals, determine the expected number of data segments obtained by the query, wherein the expected number of segments is used to determine the number of data segments that match the candidate question.

[0242] Optionally, the computer-readable storage medium may also execute program code for the following steps: displaying the reply content, the data segments used when generating the reply content, and the timestamps corresponding to the used data segments in the operation interface.

[0243] Optionally, the computer-readable storage medium may further execute program code for the following steps: in response to a touch operation on a timestamp in an operation interface, locating a data segment corresponding to the timestamp in the operation interface.

[0244] Optionally, the computer-readable storage medium can also execute the program code of the following steps: based on the type of candidate question, obtain at least one historical text information; segment the historical text information to obtain multiple groups of segmented texts; construct training data based on the multiple groups of segmented texts; and use the training data to train a large language model.

[0245] Optionally, the computer-readable storage medium can also execute the program code of the following steps: in response to the type of the candidate question being a single-document knowledge question-answering type, retrieve the historical query information of the segmented text and the historical reply content of the historical query information; segment the segmented text to obtain multiple sub-segmented texts; determine the historical query information, historical reply results, and multiple sub-segmented texts as training data.

[0246] Optionally, the computer-readable storage medium can also execute the program code of the following steps: in response to the type of the candidate question being a multi-document knowledge question-answering type, selecting a group of segmented texts from multiple groups of segmented texts, and retrieving the historical query information of the selected segmented texts, as well as the historical reply content of the historical query information; among the multiple groups of segmented texts, retrieving the segmented texts whose matching degree with the selected segmented texts is higher than the matching degree threshold; determining the selected segmented texts, the segmented texts whose matching degree is higher than the matching degree threshold, the historical query information, and the historical reply content as training data.

[0247] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: in response to an input instruction acting on an operation interface, obtaining input query information; in response to a query instruction acting on the operation interface, rewriting the query information into candidate questions for performing custom questions on multimedia information, wherein the multimedia information is content pre-uploaded to a generative question-answering system; querying and obtaining data segments associated with the candidate questions from a question-answering database, wherein the question-answering database stores at least one set of data segments, and the data source of the data segments is the multimedia information; processing the data segments using a large language model to generate at least one reply content that matches the query information; and displaying the reply content on the operation interface.

[0248] In an embodiment of the present disclosure, when an inquiry message is received, the inquiry message is rewritten into a candidate question for executing a customized question, a data segment associated with the candidate question is determined from the multimedia information, and the data segment is processed using a large language model to obtain at least one reply content that matches the inquiry message, thereby achieving the technical effect of being able to provide personalized answers to the content of the multimedia information and solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0249] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.

[0250] The embodiments of the present application further provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program that, when executed by a processor, implements the method provided in the embodiments above.

[0251] An embodiment of the present disclosure may provide an electronic device, which may include a memory and a processor.

[0252] Figure 19 is a block diagram of an electronic device according to a method for processing multimedia information in accordance with an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0253] As shown in FIG19 , device 1900 includes a computing component 1901 that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage component 1908 into a random access memory (RAM) 1903. Various programs and data required for the operation of device 1900 may also be stored in RAM 1903. Computing component 1901, ROM 1902, and RAM 1903 are connected to each other via a bus 1904. An input / output (I / O) interface 1905 is also connected to bus 1904.

[0254] Various components in device 1900 are connected to I / O interface 1905, including: input component 1906, such as a keyboard, mouse, etc.; output component 1904, such as various types of displays, speakers, etc.; storage component 1908, such as a magnetic disk, optical disk, etc.; and communication component 1909, such as a network card, modem, wireless communication transceiver, etc. Communication component 1909 allows device 1900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0255] The computing component 1901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing component 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing components that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing component 1901 performs the various methods and processes described above, such as the data verification method. For example, in some embodiments, the data verification method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage component 1908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1900 via the ROM 1902 and / or the communication component 1909. When the computer program is loaded into the RAM 1903 and executed by the computing component 1901, one or more steps of the data verification method described above can be performed. Alternatively, in other embodiments, the computing component 1901 may be configured to execute the data verification method in any other appropriate manner (eg, by means of firmware).

[0256] According to an embodiment of the present disclosure, a method for processing multimedia information is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0257] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0258] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0259] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0260] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display, monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0261] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0262] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0263] It should be noted that the serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.

[0264] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0265] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of components is only a logical function division. In actual implementation, there may be other division methods, such as multiple components or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of components or modules, which can be electrical or other forms.

[0266] Components described as separate parts may or may not be physically separate, and components shown as components may or may not be physical components, that is, they may be located in one place or distributed across multiple network components. Some or all of these components may be selected based on actual needs to achieve the objectives of this embodiment.

[0267] In addition, the functional components in the various embodiments of the present disclosure may be integrated into a single processing component, each component may exist physically separately, or two or more components may be integrated into a single component. The aforementioned integrated components may be implemented in the form of hardware or software functional components.

[0268] If the integrated components are implemented in the form of software functional components and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage media include: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disk, etc., various media that can store program codes.

[0269] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure. Industrial Applicability

[0270] The solution provided by the embodiments of the present disclosure can be applied to the processing of multimedia information, where query information is received; the query information is rewritten into candidate questions for performing customized questions on the multimedia information, wherein the multimedia information is content pre-uploaded to a generative question-answering system; data fragments associated with the candidate questions are queried from a question-answering database, wherein at least one set of data fragments is stored in the question-answering database, and the data source of the data fragments is the multimedia information; the data fragments are processed using a large language model to generate at least one reply content that matches the query information, thereby solving the technical problem of being unable to provide personalized answers to the content of the multimedia data.

[0271] .

Claims

1. A method for processing multimedia information, applied to a generative question answering system, the method comprising: Receiving inquiry information; rewriting the query information into candidate questions for performing custom questions on multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system; Querying and obtaining data segments associated with the candidate question from a question-and-answer database, wherein the question-and-answer database stores at least one set of data segments, and a data source of the data segments is the multimedia information; The data segment is processed using a large language model to generate at least one reply content matching the query information.

2. The method according to claim 1, wherein: The step of rewriting the query information into the candidate question for performing a custom question on the multimedia information includes: Acquire a multi-round dialogue sample stored in a historical time period by the generative question answering system, wherein the multi-round dialogue sample at least records a plurality of semantically related historical query information occurring in the historical time period; Identifying at least one historical inquiry information matching the inquiry information from the multiple rounds of dialogue samples; The query information is rewritten based on the at least one historical query information to generate at least one candidate question.

3. The method according to claim 1, wherein: Before querying the question-answer database to obtain the data segment associated with the candidate question, the method further includes: Acquire the uploaded multimedia information; Converting the multimedia information into text to generate text data in a text format; Segmenting the text data to obtain at least one group of data segments; Based on at least one group of the data segments, the question-answer database is constructed, wherein each group of the data segments is stored in the question-answer database in the form of feature vectors.

4. The method according to claim 3, wherein: The step of segmenting the text data to obtain at least one set of data segments includes: Identifying the length of the text data and / or semantic information of the text data; Based on the length and / or the semantic information, the text data is segmented to obtain at least one group of the data segments.

5. The method according to claim 3 or 4, wherein: The querying and obtaining, from the question-answer database, data segments associated with the candidate question, includes: Reading at least one set of the data segments from the question-answer database; Obtaining semantic similarity between each group of the data segments and the candidate questions; The data segments whose semantic similarity exceeds the similarity threshold are determined as the data segments with the candidate The piece of data that the question is associated with.

6. The method according to claim 5, wherein: The obtaining of the semantic similarity between each group of the data segments and the candidate questions includes: Obtaining a feature vector of each group of the data segments and a feature vector of the candidate question; The feature vectors of each group of the data segments are respectively calculated to have similarities with the feature vectors of the candidate questions, so as to obtain the semantic similarity between each group of the data segments and the candidate questions.

7. The method according to claim 6, wherein: After determining the data segment associated with the candidate question, the method further includes: sorting all the screened data segments associated with the candidate questions according to the semantic similarity between the data segments and the candidate questions; Extracting at least one group of the data segments within a preset ranking according to the sorting result; Reading the generation time of each group of the data segments within the preset ranking; Assembling each group of the data segments in sequence according to the generation time of each group of the data segments within the preset ranking; Using a large language model to process the data segments to generate at least one reply content matching the query information includes: using a large language model to process the assembled data segments to generate the at least one reply content matching the query information.

8. The method according to claim 1, wherein: After rewriting the inquiry information into candidate questions for performing custom questioning on multimedia information, the method further includes: Based on the type of the candidate question, determining whether to perform a query operation on the data segment recalled from the question-and-answer database; If the query operation is performed, determining the number of retrievals when the query operation is performed; Based on the retrieval quantity, the expected number of data segments obtained by the query is determined, wherein the expected number of segments is used to determine the number of data segments matching the candidate question.

9. The method according to claim 1, wherein: The method further comprises: The reply content, the data segments used when generating the reply content, and the timestamps corresponding to the used data segments are displayed in the operation interface.

10. The method according to claim 9, wherein: The method further comprises: In response to a touch operation on the timestamp in the operation interface, a data segment corresponding to the timestamp is located in the operation interface.

11. The method according to claim 1, wherein: The method further comprises: Based on the type of the candidate question, obtaining at least one piece of historical text information; Segmenting the historical text information to obtain multiple groups of segmented texts; Constructing training data based on the multiple groups of segmented texts; The large language model is obtained by training using the training data.

12. The method according to claim 11, wherein: The step of constructing the training data based on the multiple groups of segmented texts includes: In response to the type of the candidate question being a single-document knowledge question-answering type, retrieving historical query information of the segmented text and historical answer content of the historical query information; Segmenting the segmented text to obtain a plurality of sub-segmented texts; The historical inquiry information, the historical answer results, and the plurality of sub-segmented texts are determined as the training data.

13. The method according to claim 11, wherein: The step of constructing the training data based on the multiple groups of segmented texts includes: In response to the type of the candidate question being a multi-document knowledge question-answering type, selecting a group of segmented texts from the multiple groups of segmented texts, and retrieving historical query information of the selected segmented texts and historical answer content of the historical query information; Retrieving, from the plurality of groups of segmented texts, segmented texts whose matching degree with the selected segmented text is higher than a matching degree threshold; The selected segmented text, the segmented text with a matching degree higher than a matching degree threshold, the historical inquiry information, and the historical reply content are determined as the training data.

14. A method for processing multimedia information, applied to a generative question answering system, comprising: Responding to an input instruction acting on the operation interface, obtaining input query information; In response to a query instruction acting on the operation interface, rewriting the query information into a candidate question for performing a custom question on the multimedia information, wherein the multimedia information is content pre-uploaded to the generative question answering system; Querying and obtaining data segments associated with the candidate question from the question-answer database, wherein the question-answer database stores at least one set of data segments, and the data source of the data segments is the multimedia information; Processing the data segment using a large language model to generate at least one reply content matching the query information; The reply content is displayed on the operation interface.

15. An electronic device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of any one of methods 1-14 are implemented.

16. A computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, wherein the computer program implements the method according to any one of claims 1 to 14 when executed by a processor.

Citation Information

Patent Citations

  • Multi-round question and answer method, device and equipment

    CN114547274A

  • Open domain natural language reasoning question-answering system and method driven by large language model

    CN116932708A

  • Knowledge base construction method and question and answer dialogue method and system based on generative large language model

    CN117056471A

  • Information processing method and device, electronic equipment and storage medium

    CN117112754A

  • Multimedia information processing method and system and electronic equipment

    CN117951318A

Cited By

  • Data error correction method, data error correction device and storage medium

    CN121434323A

  • Double-flow recording traceability and answer presentation method based on asynchronous triggering and semantic anchor points

    CN121708917A