Subject information extraction method based on audio and video data and computer equipment
Through the multimodal information fusion of audio and video data, the accuracy of the theme information refinement in online video conferences is solved, and more reliable theme information extraction is achieved.
Patent Information
- Application Number
- CN202410013950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-11
AI Technical Summary
In the online video conference of prior art, refining theme information with a single audio information is easily disturbed by user language expression and other factors, resulting in the extracted theme information being inaccurate and reliable enough.
By obtaining audio and video data, extracting audio data and converting it into text content, and fusing information with modal data such as speech features and image information, filtering fusion fragments that meet preset recognition conditions, and extracting theme information.
It improves the accuracy and reliability of the main information, eliminates the interference of user language expression and other factors, and ensures the fluency and readability of the extracted information.
Smart Images

Figure CN120296647A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of natural language processing, and in particular, to a method for extracting keynote information based on audio-visual data and a computer device. Background Art
[0002] Nowadays, online video conferencing has gradually become the main way for most enterprises to hold meetings, and people's demand for quickly browsing the meeting content of online video conferences is also increasing. Related technologies can convert the audio information of the meeting content into text content, slice the text content according to natural language processing algorithms to obtain each chapter segment of the meeting content, and analyze the chapter segment to obtain the keynote information of the meeting content. However, related technologies only rely on single audio information to extract keynote information, and are easily interfered by the language expressions of different users or other factors, resulting in the inability of related technologies to accurately and reliably slice, and thus the extracted keynote information is not accurate and reliable enough. Summary of the Invention
[0003] An object of the embodiments of the present application is to provide a method for presenting keynote information, a computer device, and a computer-readable storage medium, so as to solve the technical problem that the keynote information provided by related technologies is not accurate and reliable enough.
[0004] In a first aspect, the embodiments of the present application provide a method for extracting keynote information based on audio-visual data, including:
[0005] Obtain audio-visual data;
[0006] Extract the audio data from the audio-visual data and convert it into text content;
[0007] Extract the modality data matching the text content from the audio-visual data;
[0008] Fuse the text content and the modality data to obtain a plurality of fusion segments;
[0009] Screen the fusion segments that meet the preset recognition conditions for extracting keynote information;
[0010] Summarize all the extracted keynote information as the keynote information of the audio-visual data.
[0011] Optionally, the information fusion of the text content and the modality data to obtain multiple fusion segments includes: performing information fusion on the text content and the modality data to obtain the fused text content, and segmenting the fused text content to obtain multiple fusion segments. This embodiment can fuse the modality data with the text content before the segment segmentation stage, and subsequent vector dimension conversion is not required, which is conducive to quickly integrating the modality data into the text content and improving the recognition efficiency of chapter segments.
[0012] Optionally, the information fusion of the text content and the modality data to obtain multiple fusion segments includes: segmenting the text content to obtain multiple text segments, performing information fusion on the text segments and the modality data to obtain the fused text segments, and the multiple fused text segments form multiple fusion segments. This embodiment can fuse information different from the attributes of the text segments with the text segments during the segment processing stage, achieving the purpose of cross-modal fusion, and can assist the preset large language model to comprehensively understand the text content of the text segments, enabling this embodiment to reliably and accurately extract the main information in various scenarios.
[0013] Optionally, the modality data includes first-class modality information and second-class modality information. The information fusion of the text content and the modality data to obtain multiple fusion segments includes: performing information fusion on the text content and the first-class modality information to obtain the fused text content, segmenting the fused text content to obtain multiple text segments, performing information fusion on the text segments and the second-class modality information to obtain the fused text segments, and the multiple fused text segments form multiple fusion segments. This embodiment can perform information fusion on the text content and the first-class modality information before the segment segmentation stage, and after the segment segmentation, that is, during the segment processing stage, perform information fusion on the text segments and the second-class modality information. In this way, this embodiment can not only assist in enhancing the information content of the text segments from the speech feature dimension or the text description dimension, but also enhance the information content of the text segments from the image dimension, thus achieving multi-modal information fusion, which is conducive to accurately and reliably identifying chapter segments and accurately and reliably extracting the main information.
[0014] Optionally, the performing information fusion on the text content and the first-class modality information to obtain the fused text content includes: determining an insertion position in the text content and splicing the first-class modality information at the insertion position to obtain the fused text content.
[0015] Optionally, the text content includes multiple sentences, the first type of modality information includes speech feature information, and determining the insertion position in the text content includes: determining the sentence associated with the speech feature information as the target sentence, and determining the target position of the target sentence as the insertion position.
[0016] Optionally, the modality data further includes a third type of modality information. Splitting the fused text content to obtain multiple text segments includes: performing at least two splitting operations on the fused text content according to a preset window length to obtain at least two candidate segments, splicing the third type of modality information at a specified position of the candidate segments obtained in each splitting operation to obtain spliced candidate segments, and multiple spliced candidate segments form multiple text segments.
[0017] Optionally, the second type of modality information includes image information matching the text segment. Fusing the text segment with the second type of modality information to obtain a fused text segment includes: fusing the text segment with the image information to obtain a fused text segment, where the image information is an image with a shooting time consistent with the time of the text segment. This fused segment not only fuses speech feature information from the speech feature dimension, fuses text description information from the text description dimension, but also fuses image information from the image dimension. In this way, it can enhance the information content of the text segment in multiple dimensions and modalities, which is beneficial to accurately and reliably identifying chapter segments and accurately and reliably extracting the main idea information.
[0018] Optionally, screening the fused segments that meet the preset recognition conditions for extracting the main idea information includes: screening the fused segments that meet the preset recognition conditions as chapter segments, and extracting the main idea information of the chapter segments.
[0019] Optionally, screening the fused segments that meet the preset recognition conditions as chapter segments includes: determining whether a target fused segment has chapter features, where the target fused segment is one of the multiple fused segments. If the target fused segment has chapter features, it is determined that the target fused segment meets the preset recognition conditions, and the target fused segment is determined as a chapter segment. If the target fused segment does not have chapter features, it is determined that the target fused segment does not meet the preset recognition conditions, and the target fused segment is determined as a non-chapter segment. In this embodiment, by analyzing whether the target fused segment has chapter features, it can effectively and reliably find chapter segments among multiple fused segments, which is beneficial to improving the accuracy and reliability of generating the main idea information.
[0020] Optionally, determining whether the target fusion segment has a chapter feature includes: sequentially inputting the target fusion segment into a preset score model to obtain a chapter feature score corresponding to the target fusion segment, and determining whether the target fusion segment has a chapter feature according to the chapter feature score. In this embodiment, based on the preset score model, the chapter feature scores of each character in the target fusion segment can be quickly and reliably output, so that it can be quickly and effectively identified whether the target fusion segment is a chapter segment.
[0021] Optionally, the target fusion segment includes multiple characters, the chapter feature includes a chapter start point, the chapter feature score includes a start point score, and determining whether the target fusion segment has a chapter feature according to the chapter feature score includes: determining whether the start point score is greater than or equal to a preset start threshold. If so, it is determined that the target fusion segment has a chapter feature, and the character with the start point score greater than or equal to the preset start threshold is determined as the chapter start point. If the start point scores of all characters in the target fusion segment are less than the preset start threshold, it is determined that the target fusion segment does not have a chapter feature.
[0022] Optionally, extracting the main idea information of the chapter segment includes: determining a reference segment of the target chapter segment, where the reference segment is a text segment arranged after the target chapter segment, and the target chapter segment is one of the multiple chapter segments. Determine the content coverage range of the target chapter segment according to the attribute of the reference segment, and generate main idea information according to the content coverage range of the target chapter segment. This embodiment can maximize the search for the content coverage range of the target chapter segment, so as to enrich the information content carried by the target chapter segment, avoid fragmentation of the content of the target chapter segment, and is conducive to generating reliable and accurate main idea information.
[0023] Optionally, determining the content coverage range of the target chapter segment according to the attribute of the reference segment includes: determining whether the attribute of the reference segment is a chapter segment attribute. If the attribute of the reference segment is a chapter segment attribute, splice the target chapter segment and the text segment between the target chapter segment and the reference segment to obtain a final segment, and determine the content range carried by the final segment as the content coverage range. If the attribute of the reference segment is a non-chapter segment attribute, determine the text segment arranged after the reference segment as a new reference segment, and return to the step of determining whether the reference segment is a chapter segment. This embodiment adopts an iterative and recursive method, which can maximize the search for the content coverage range of the target chapter segment, and is conducive to improving the reliability and accuracy of the generation of main idea information.
[0024] Optionally, generating the main idea information according to the content coverage of the target chapter segment includes: obtaining custom indication information and the chapter content corresponding to the content coverage, encapsulating the indication information and the chapter content into model input data, and sending the model input data to a preset language model, so that the preset language model outputs the main idea information according to the model input data. In this embodiment, taking advantage of the fact that the preset language model is good at creating titles and abstracts for texts, by constructing the model input data and sending it to the preset language model, the preset language model can output titles and abstracts that are reliable, accurate, highly fluent, and highly readable.
[0025] Optionally, the main idea information includes the start time point, title, and abstract corresponding to each chapter segment, where the chapter segment is a fusion segment that meets the preset recognition conditions. The summary of all the main idea information extracted and obtained as the main idea information of the audio-visual data includes:
[0026] According to the order of the start time points of the respective chapter segments, sequentially concatenating the titles and abstracts of the respective chapter segments to obtain the main idea information of the audio-visual data.
[0027] In a second aspect, an embodiment of the present application provides a computer device, which is characterized by including a memory and a processor, the memory is connected to the processor, the processor is configured to execute one or more computer programs stored in the memory, and when the processor executes the one or more computer programs, the computer device implements the above method.
[0028] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the above method.
[0029] The embodiments of the present application can achieve the following technical effects: obtaining audio-visual data in a preset scenario, extracting audio data from the audio-visual data and converting it into text content, extracting modal data matching the text content from the audio-visual data, performing information fusion on the text content and the modal data to obtain multiple fusion segments, screening the fusion segments that meet the preset recognition conditions for extracting main idea information, and summarizing all the main idea information extracted and obtained as the main idea information of the audio-visual data. This embodiment can fuse the text content and the modal data to obtain multiple fusion segments. The fusion segments carry rich information content, enhance the expression accuracy of the text content, and are conducive to eliminating the interference brought by the user's language expression or other factors, so that the main idea information can be accurately and reliably extracted. Description of the Drawings
[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the accompanying drawings required to be used in the description of the embodiments of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0031] Figure 1 Schematic diagram of the system architecture of a meeting theme generation system provided by an embodiment of the present application;
[0032] Figure 2 Schematic diagram of the system architecture of a meeting theme generation system provided by another embodiment of the present application;
[0033] Figure 3 Schematic diagram of the process flow of a theme information extraction method based on audio-visual data provided by an embodiment of the present application;
[0034] Figure 4 Schematic diagram of the agenda information composed of the titles and abstracts of r chapter segments provided by an embodiment of the present application;
[0035] Figure 5 Schematic diagram of the text content provided by an embodiment of the present application being cut into m text segments, where there are k overlapping characters between the i-th text segment and the (i + 1)-th text segment;
[0036] Figure 6 Schematic diagram of the text content provided by an embodiment of the present application being cut into m text segments, where there are no overlapping characters between the j-th text segment and the (j + 1)-th text segment;
[0037] Figure 7 Schematic diagram of the sorting between chapter segments and non-chapter segments provided by an embodiment of the present application;
[0038] Figure 8 Schematic diagram of the structure of a theme information extraction device based on audio-visual data provided by an embodiment of the present application;
[0039] Figure 9 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0040] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0041] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the sequence in the flowchart. Furthermore, the terms "first", "second", "third", etc. adopted in the present application do not limit the data and the execution order, but only distinguish the same items or similar items with basically the same functions and effects.
[0042] In addition to the problems existing in the related technologies pointed out in the background art, the inventor also found a related technology in the process of implementing the embodiments of the present application. This related technology mainly uses the manual analysis method to analyze the meeting content to obtain the main idea information of the online video meeting. Among them, the manual analysis method is to manually observe and slice each frame of the meeting video, sort out each chapter segment of the meeting content, and manually analyze this chapter segment to obtain the main idea information of the meeting content. This method requires a lot of labor costs and has a low efficiency in generating the main idea information.
[0043] The inventor also found another related technology in the process of implementing the embodiments of the present application. This related technology sets up a collection device at the meeting site. The collection device includes a camera and a voice collector. The collection device collects user images and user audio data, and transmits the user images and user audio data to the background device. The background device converts the user audio data into text content, analyzes the facial expressions of the user according to the user images, identifies multiple chapter segments in the text content according to the facial expressions of multiple users, and determines the main idea information of the meeting content according to the chapter segments. The following problems exist in this related technology:
[0044] 1. It is necessary to increase the hardware cost. As mentioned above, the related technology needs to set up a collection device at the meeting site to collect the meeting site data, which will increase the hardware cost of the collection device.
[0045] 2. The scene applicability is poor. Since the related technology needs to identify the chapter segments according to the facial expressions of the users, however, in some meeting scenes, the facial expressions of some users remain unchanged throughout the meeting process, which easily leads to the failure of the related technology to successfully identify the chapter segments.
[0046] 3. The process of identifying the chapter segments is relatively complex. As mentioned above, the related technology needs to combine the facial expressions of multiple users to obtain unified chapter segments, which will increase the complexity of identifying the chapter segments.
[0047] Based on this, the embodiments of the present application can automatically identify chapter segments by machine equipment, without the need for manual slicing of the conference video and frame-by-frame analysis of the slices to identify the chapter segments. In addition, the embodiments of the present application can eliminate the influence of inconsistent main information of chapter segments caused by different understandings of different manual analysts in dividing chapter segments, can output the main information of each chapter segment with relatively high consistency, and at the same time save labor costs. In addition, the embodiments of the present application do not require additional acquisition equipment and can extract audio data from the audio-visual data of an online video conference, so that the cost of adding hardware can be saved.
[0048] Secondly, the embodiments of the present application adopt two processing stages to obtain the main information of the chapter segments. The first processing stage can fuse various modal information with the text content, greatly enriching the information content of the text content, which is beneficial to reliably and accurately identifying the chapter segments in the subsequent steps and eliminating the interference of noise data on the identification of chapter segments. In the process of identifying the chapter segments, the embodiments of the present application extract the chapter segments from the fused text content by means of span (phrase segment) extraction, and the second processing stage summarizes the chapter segments by means of text summary extraction, so as to obtain the main information of the chapter segments.
[0049] In the first processing stage, the embodiments of the present application extract audio data from the audio-visual data of the conference, and then use a preset speech recognition model to process and analyze the audio data to obtain the text content. The embodiments of the present application generate speech feature information according to the audio data and the text content, and insert the speech feature information into the text content, so as to obtain the text content after the first fusion. Since there are obvious features in the switching of chapter content, for example, there is an easy 3-second or 5-second speech interval between the previous chapter content and the next chapter content, this embodiment can integrate the speech feature information into the text content, so as to enhance the information content of the fused text content for identifying the chapter segments.
[0050] The embodiments of the present application segment the text content according to the improved chunking mechanism to obtain multiple text segments. Among them, the meeting description information is spliced in front of each text segment to obtain the fused text segment. In some meeting scenarios, people tend to chat and do not discuss the main topic. Incorporating the meeting description information into the text segment is beneficial for the subsequent preset language model to accurately and reliably extract the main information of the chapter segment.
[0051] In the embodiments of the present application, each fused text segment is encoded based on a preset language encoding model to obtain text vectors of each text segment. Then, the embodiments of the present application extract at least one frame of image information corresponding to the text segment from the audio-visual data, and encode the at least one frame of image information based on a preset image encoding model to obtain an image vector matching the text segment. In this embodiment, the text vector of each text segment is fused with the image vector to obtain a final segment.
[0052] The image information can describe the chapter segment from a dimension different from that of the text. The final segment can not only carry the information content of the text content, speech feature information, and meeting description information, but also carry the information content of the dynamic image information, which is beneficial to accurately and reliably identifying the chapter segment and accurately and reliably extracting the main idea information of the chapter segment.
[0053] Finally, the embodiments of the present application input the final segment into a preset score model to predict whether each text segment is a chapter segment, where the chapter segment includes the chapter start point of the chapter segment, the start time of the chapter start point, and the chapter end point.
[0054] In the second processing stage, each chapter segment obtained in the first processing stage is processed based on a preset large language model to obtain the main idea information of each chapter segment. Among them, the embodiments of the present application can sequentially concatenate the main idea information of each chapter segment according to the time sequence to obtain a meeting agenda with a clear time context and chapters. The preset large language model is obtained by fine-tuning training on the fine-tuning data using supervision instructions.
[0055] First, the embodiments of the present application can fuse the text content with the modal data to obtain multiple fused segments. The fused segments carry rich information content and enhance the expression accuracy of the text content, which is beneficial to eliminating the interference caused by the user's language expression or other factors, so as to accurately and reliably extract the main idea information.
[0056] Secondly, the embodiments of the present application simulate the method of artificial analysis method, and use machine equipment to simulate the recognition process of chapter segments and the extraction process of the main idea information of chapter segments. The technical effects achieved are as described above and will not be elaborated here.
[0057] Thirdly, the embodiments of the present application segment the text content according to the improved chunking mechanism, which can reduce the situation that the content belonging to the same chapter segment is unreasonably segmented into other chapter segments due to unreasonable segmentation, avoid content fragmentation, reduce the degree of content loss of the chapter segment, and improve the integrity of the content of the chapter segment.
[0058] Finally, the embodiment of the present application uses a preset large language model as the pre-training base and fine-tunes it with high-quality instruction data manually annotated, which can make full use of the generation ability and logical reasoning ability of the preset large language model for the main idea information, making the content of the main idea information relatively fluent and having relatively high readability.
[0059] The embodiment of the present application provides a system for generating the main idea of a meeting. Among them, the system 100 for generating the main idea of a meeting can adopt Figure 1 the communication architecture shown. Please refer to Figure 1 , the system 100 for generating the main idea of a meeting includes a plurality of terminal devices 11, a server 12, and a computer device 13. The server 12 is respectively communicatively connected to the plurality of terminal devices 11, and the computer device 13 can be communicatively connected to any one of the terminal devices 11 or the server 12. The communication connection includes a wired communication connection or a wireless communication connection. Among them, the wired communication connection includes various communication connections that transmit information using tangible media such as metal wires and optical fibers. The wireless communication connection includes 5G communication, 4G communication, 3G communication, 2G communication, CDMA communication, Zig-Bee communication, Bluetooth communication, or Wi-Fi communication, etc.
[0060] Users of the plurality of terminal devices 11 hold an online video meeting or an online voice meeting. Among them, each terminal device 11 sends the collected meeting data to the server 12, and the server then forwards the meeting data collected by each terminal device 11 to other terminal devices 11, and each terminal device 11 plays the corresponding meeting data locally. Among them, the meeting data includes audio-visual data and / or voice data.
[0061] The computer device 13 can obtain the meeting data from the terminal device 11 or the server 12 or an intermediate device and generate the main idea information according to the meeting data. Among them, the terminal device 11 or the server 12 can send the meeting data to the intermediate device for storage.
[0062] In some embodiments, the system 100 for generating the main idea of a meeting can also adopt Figure 2 the communication architecture shown. Please refer to Figure 2 , the system 100 for generating the main idea of a meeting includes a plurality of terminal devices 11, a server 12, and a meeting tablet 14. The server 12 is respectively communicatively connected to the plurality of terminal devices 11 and the meeting tablet 14, and the computer device 13 can be communicatively connected to any one of the terminal devices 11 or the server 12 or the meeting tablet 14.
[0063] Users of multiple terminal devices 11 and multiple users gathered in front of the conference tablet 14 hold an online video conference or an online voice conference. Among them, the conference data collected by each terminal device 11 and the conference data collected by the conference tablet 14 are both sent to the server 12, and the server 12 then forwards the conference data collected by each terminal device 11 and the conference data collected by the conference tablet 14 to other terminal devices 11 or the conference tablet 14, and each terminal device 11 and the conference tablet 14 play the corresponding conference data locally.
[0064] The computer device 13 can obtain conference data from the terminal device 11, the server 12, the conference tablet 14 or an intermediate device, and generate a keynote message based on the conference data.
[0065] It can be understood that the computer device 13 can be understood as a device that is relatively independent and parallel to the terminal device 11, the server 12, and the conference tablet 14. For example, the computer device 13 is a computer or the like.
[0066] It can also be understood that the computer device 13 can also be understood as a part of the components of the terminal device 11, the server 12, or the conference tablet 14. For example, the computer device 13 is a controller of the terminal device 11, the server 12, or the conference tablet 14.
[0067] It can also be understood that Figure 1 and Figure 2 This is only the scenario form of a conference keynote generation system provided by the embodiments of the present application. Those skilled in the art can also apply the computer device 13 to any scenario form according to their own needs, and no limitation is made here.
[0068] The embodiments of the present application provide a method for extracting keynote information based on audio-visual data. Among them, the execution subject of the method for extracting keynote information based on audio-visual data can be a computer device or a conference tablet with logical computing capabilities, and no specific limitation is made on the execution subject of the method for extracting keynote information based on audio-visual data here. It can be understood that the computer device can be understood as a part of the conference tablet or independent of the conference tablet. Please refer to Figure 3 , the method for extracting keynote information based on audio-visual data includes the following steps:
[0069] S31: Obtain audio-visual data.
[0070] In this step, the audio-visual data is data integrating audio data and image data. The audio data is the data corresponding to the voice collected in a preset scenario, and the image data is the data of the image captured in a preset scenario. The preset scenario is an online conversation scenario or an offline conversation scenario. The online conversation scenario includes an online video conference scenario, an online voice conference scenario, or other custom online scenarios. The offline conversation scenario includes an offline conference scenario or other custom offline scenarios.
[0071] When the preset scenario is an online video conference scenario, in some embodiments, obtaining the audio-visual data includes: obtaining the audio-visual data sent by a first target device in the online video conference scenario. The first target device can be a conference tablet or a terminal device in the online video conference scenario.
[0072] In some embodiments, obtaining the audio-visual data includes: obtaining the audio-visual data sent by a second target device. Among them, in the online video conference scenario, the conference tablet or terminal device communicates with the second target device. The second target device can store the audio-visual data corresponding to the online video conference scenario sent by the conference tablet or terminal device. When the second target device detects a first data acquisition request sent by a computer device, the second target device sends the audio-visual data to the computer device.
[0073] In this embodiment, there is no need to additionally set up a camera and a voice acquisition device. Only the voice data can be extracted from the original audio-visual data in the online video conference, or the original voice data in the online voice conference can be directly used. In this way, the hardware costs required for setting up a camera and a voice acquisition device can be saved.
[0074] S32: Extract the audio data from the audio-visual data and convert it into text content.
[0075] In this step, the audio data is the data belonging to the audio nature parsed from the audio-visual data, and the text content is the set of texts after the audio data is converted into characters.
[0076] Extracting the audio data from the audio-visual data and converting it into text content includes: extracting the audio data from the audio-visual data and converting the audio data into text content.
[0077] In some embodiments, extracting the audio data from the audio-visual data includes: extracting the audio data from the audio-visual data according to a preset audio separation algorithm. The preset audio separation algorithm can be based on a non-negative matrix factorization algorithm, etc.
[0078] In some embodiments, extracting the audio data from the audio-visual data includes: inputting the audio-visual data into a preset audio separation application program for processing to obtain the audio data. The preset audio separation application program can adopt the audio separation APP provided by related technologies and will not be elaborated here.
[0079] In some embodiments, converting the audio data into text content includes: converting the audio data into text content according to a preset speech recognition algorithm. The preset speech recognition algorithm includes an algorithm based on dynamic time warping, an algorithm of hidden Markov model based on parametric model, an algorithm of vector quantization based on non-parametric model, or a deep learning algorithm, etc.
[0080] In some embodiments, converting the audio data into text content includes: inputting the audio data into a preset speech recognition model for processing to obtain text content. The preset speech recognition model can be a speech recognition model provided by the related art.
[0081] S33: Extract modal data matching the text content from the audio-visual data.
[0082] In this step, the modal data is data for enhancing the information content of the text content. Among them, the modal data can enrich the information content of the text content from various dimensions. In some embodiments, the modal data may include a type of modal information. In this embodiment, a type of modal information can be used to characterize the characteristics of the text content to enrich the information content of the text content. Among them, this type of modal information can be speech feature information or text description information or image information. The speech feature information is used to enhance the expression characteristics of the text content from the audio dimension, the text description information is used to assist in understanding the text content, and the image information is used to enhance the expression characteristics of the text content from the image dimension.
[0083] In some embodiments, the modal data may include two types of modal information. The two types of modal information are any two of the speech feature information, text description information, and image information. In this way, this embodiment can fuse the modal information of two dimensions to further enhance the information content of the text content, which is beneficial to reliably and accurately identifying chapter segments and reliably and accurately extracting the main idea information of the chapter segments.
[0084] In some embodiments, the modal data may include three types of modal information, which are respectively speech feature information, image information, and text description information. In this way, this embodiment can fuse the modal information of three dimensions to greatly enhance the information content of the text content.
[0085] It can be understood that those skilled in the art can make other substitutions or deformations by combining various embodiments describing the modal data provided in the embodiments of the present application, and such substitutions or deformations should fall within the protection scope of the present application.
[0086] S34: Perform information fusion on the text content and the modal data to obtain a plurality of fusion segments.
[0087] In this step, information fusion is an operation of associating modal data with text content. The fusion fragment is a text fragment into which modal data is incorporated, and the text fragment is a fragment obtained by splitting the text content.
[0088] S35: Screen the fusion fragments that meet the preset recognition conditions to extract the main information.
[0089] In this step, screening the fusion fragments that meet the preset recognition conditions to extract the main information includes the following steps: Screening the fusion fragments that meet the preset recognition conditions as chapter fragments, and extracting the main information of the chapter fragments. Among them, the preset recognition conditions are used to screen out chapter fragments from multiple fusion fragments, and the fusion fragments that meet the preset recognition conditions are chapter fragments. A chapter fragment is a text fragment with chapter features, and a non-chapter fragment is a text fragment without chapter features. Among them, the chapter features include the chapter start point and / or the chapter end point. The chapter start point is the start point of the chapter content corresponding to the chapter fragment, and the chapter end point is the end point of the chapter content corresponding to the chapter fragment.
[0090] In some embodiments, the preset recognition conditions are configured as: the text fragment has chapter features. When the text fragment has chapter features, the text fragment meets the preset recognition conditions; when the text fragment does not have chapter features, the text fragment does not meet the preset recognition conditions.
[0091] S36: Summarize all the extracted main information as the main information of the audio-visual data.
[0092] In this step, the main information is the information obtained by summarizing the chapter content of the chapter fragment. Among them, the main information can be a title or an abstract or a combination of a title and an abstract. In some embodiments, after summarizing the main information of the audio-visual data, in this embodiment, the main information of each chapter fragment can be concatenated in sequence, and then the main information of each concatenated chapter fragment is presented.
[0093] Generally speaking, this embodiment can fuse the text content with the modal data to obtain multiple fusion fragments. The fusion fragments carry rich information content, enhance the expression accuracy of the text content, and are beneficial to eliminating the interference caused by the user's language expression or other factors, so as to accurately and reliably extract the main information.
[0094] In some embodiments, the main information includes the start time point, title, and abstract corresponding to each chapter fragment. The chapter fragment is a fusion fragment that meets the preset recognition conditions. Summarizing all the extracted main information as the main information of the audio-visual data includes the following steps: According to the order of the start time points of each chapter fragment, concatenate the titles and abstracts of each chapter fragment in sequence to obtain the main information of the audio-visual data.
[0095] The start time point of a chapter segment corresponds to the turning point of the chapter content. Generally, the start time point of the current chapter segment can be regarded as the end time point of the previous chapter segment. The chapter content of the previous chapter segment is different from that of the current chapter segment. Therefore, the start time point of the current chapter segment can be regarded as the turning point of the corresponding chapter content. This embodiment can encapsulate the start time point of the chapter segment in the main idea information, and later users can quickly locate the key chapter content they need to consult through the start time point of the chapter segment.
[0096] The title concisely summarizes the chapter content of the chapter segment, which can help users clearly understand the core content of the chapter segment at a glance.
[0097] The abstract expands the chapter content summarized by the title, which can lightly but more detailedly than the title show the main content of the chapter segment so that users can understand the important details of the chapter segment.
[0098] Please refer to Figure 4 , this embodiment has identified a total of r chapter segments. The a-th chapter segment includes the a-th title and the a-th abstract with the start time point of t a . Among them, a is a positive integer and less than r. As Figure 4 shown, the first chapter segment includes the 1st title and the 1st abstract with the start time point of t1, and the second chapter segment includes the 2nd title and the 2nd abstract with the start time point of t2.
[0099] As Figure 4 shown, this embodiment simulates the table of contents of a book, concatenates the titles and abstracts of each chapter segment according to the start time point, so as to form the main idea information of the audio-visual data with a time context, which is convenient for users to quickly master the context of the meeting and can intuitively and quickly understand the main content of the meeting.
[0100] In some embodiments, before splitting the text content, the method further includes the following steps: performing a sentence splitting operation on the text content to obtain multiple sentences, determining the time stamps of each sentence, and obtaining the text content including time stamps. Subsequently, this embodiment can split the text content including time stamps to obtain multiple text segments.
[0101] In this embodiment, the text content includes multiple characters. Performing a sentence splitting operation on the text content to obtain multiple sentences includes the following steps: sequentially traversing each character in the text content, determining whether the character is a preset end symbol. If the character is a preset end symbol, then the characters between the preset end symbol of the previous sentence and the current preset end symbol are used as a sentence. Among them, at the initial traversal, the preset end symbol of the previous sentence is the first character of the text content. If the character is not a preset end symbol, then continue to traverse the next character.
[0102] The preset end symbol may be a period “.”, a question mark “?”, an exclamation mark “!”, etc.
[0103] In this embodiment, determining the timestamp of each sentence includes: traversing the timestamp corresponding to the first character of each sentence in the time sequence corresponding to the voice data as the timestamp of the sentence.
[0104] For example, the text content is: Hi, everyone, welcome to watch this episode of constellations, I am the host, Mr. Gu. Today we will talk about the months of three constellations in the 12 constellations.
[0105] After the above text content is divided into sentences, "Hi, everyone, welcome to watch this episode of horoscopes. I am the host, Mr. Gu." is the first sentence, and "Today we will talk about the months of three zodiac signs among the 12 zodiac signs." is the second sentence.
[0106] The first character of the first sentence is "Hi". This embodiment performs a traversal operation on the time sequence of the voice data and obtains the timestamp of "Hi" as [timestamp 1].
[0107] The first character of the second sentence is "今". This embodiment performs a traversal operation on the time sequence of the voice data and obtains the timestamp of "今" as [timestamp 6].
[0108] After determining the timestamp of each sentence, the above text content can be changed to: [Timestamp 1] Hi, everyone, welcome to watch this episode of constellations, I am the host, Mr. Gu. [Timestamp 6] Today we will talk about the months of three constellations in the 12 constellations.
[0109] In other embodiments, before segmenting the text content, the method for presenting the main information may further include the following steps: processing the text content through ASR automatic speech recognition technology to obtain text content with original timestamps, performing sentence segmentation operations on the text content to obtain multiple sentences, determining a reference timestamp corresponding to each sentence, splicing the reference timestamp at the reference position of the sentence to determine the final timestamp of each sentence, and finally obtaining text content containing timestamps.
[0110] Determining the reference timestamp corresponding to each sentence includes the following steps: taking the original timestamp corresponding to the first character of the sentence as the reference timestamp of the sentence.
[0111] The reference position may be before the first character of a sentence or after the last character, etc.
[0112] For example, when the text content after ASR processing is: [Original timestamp 1] Hi, everyone, welcome to watch this episode of constellations, [Original timestamp 4] I am the host, Mr. Gu. [Original timestamp 6] Today we will talk about the three constellation months of the 12 constellations [Original timestamp 8].
[0113] After the above text content is divided into sentences, "Hi, everyone, welcome to watch this episode of horoscopes. I am the host, Mr. Gu." is the first sentence, and "Today we will talk about the months of three zodiac signs among the 12 zodiac signs." is the second sentence.
[0114] The first character of the first sentence is "Hi", and the original timestamp of "Hi" is [original timestamp 1], so the reference timestamp of the first sentence is [original timestamp 1].
[0115] The first character of the second sentence is "今", and the original timestamp of "今" is [original timestamp 6]. Therefore, the reference timestamp of the second sentence is [original timestamp 6].
[0116] After splicing the reference timestamp to the target position of each sentence, the above text content can become: [Original timestamp 1] Hi, everyone, welcome to watch this episode of constellations, I am the host, Mr. Gu. [Original timestamp 6] Today we will talk about the three constellation months in the 12 constellations. This embodiment configures a reference timestamp for each sentence, which is conducive to the subsequent generation of subject information with a starting time point.
[0117] In some embodiments, the modal data includes voice feature information, and extracting the modal data matching the text content from the audio and video data includes: extracting audio data from the audio and video data, performing a voice feature detection operation on the audio data, and obtaining the voice feature information.
[0118] The speech feature information includes speech interval information, intonation change information or timbre change information, etc.
[0119] The speech interval information is used to indicate the interval time between two adjacent sentences.
[0120] In some embodiments, the text content includes multiple sentences, the speech feature information includes speech interval information, and a speech feature detection operation is performed on the audio data to obtain the speech feature information, including the following steps: detecting whether the interval time between two adjacent sentences is greater than or equal to a preset time threshold; if so, generating speech interval information; if not, continuing to detect whether the interval time between the next pair of adjacent sentences is greater than or equal to the preset time threshold.
[0121] Two adjacent sentences are the first sentence and the second sentence. Detecting whether the interval time between two adjacent sentences is greater than or equal to a preset time threshold includes the following steps: determining the end time of the first sentence in the time sequence of the audio data, and the start time of the second sentence in the time sequence of the audio data, calculating the difference between the start time of the second sentence and the end time of the first sentence to obtain the interval time, and determining whether the interval time is greater than or equal to the preset time threshold. Among them, the preset time threshold is customized by the designer according to engineering experience. For example, the preset time threshold is 3 seconds or 5 seconds, etc.
[0122] Generally, during the switching process of different chapter contents, there is likely to be a transition time or a pause time between two adjacent chapter contents. The embodiment of the present application uses voice interval information to describe the interval time between two adjacent sentences, so as to enhance the expression characteristics of the text content and is conducive to accurately and reliably identifying chapter segments.
[0123] The intonation change information is used to represent the intonation difference between two adjacent sentences.
[0124] In some embodiments, the text content includes multiple sentences, and the voice feature information includes intonation change information. Performing a voice feature detection operation on the audio data to obtain the voice feature information includes the following steps: determining the intonation difference between two adjacent sentences, determining whether the intonation difference is greater than a preset intonation threshold. If so, generating intonation change information. If not, continuing to determine the intonation difference between the next pair of adjacent sentences.
[0125] Two adjacent sentences are the first sentence and the second sentence. Determining the intonation difference between two adjacent sentences includes: determining the first audio segment corresponding to the first sentence in the audio data, and the second audio segment corresponding to the second sentence in the audio data, calculating the average intonation of the first audio segment and the average intonation of the second audio segment, and calculating the absolute value of the difference between the average intonation of the second audio segment and the average intonation of the first audio segment to obtain the intonation difference.
[0126] Generally, during the switching process of different chapter contents, there is likely to be an intonation change between two adjacent chapter contents. For example, the intonation change can be a change from high to low or a change from low to high. The embodiment of the present application uses intonation change information to describe the intonation difference between two adjacent sentences, so as to enhance the expression characteristics of the text content and is conducive to accurately and reliably identifying chapter segments.
[0127] The timbre change information is used to represent the timbre change between two adjacent sentences.
[0128] In some embodiments, the text content includes multiple sentences, and the voice feature information includes timbre change information. Performing a voice feature detection operation on the audio data to obtain the voice feature information includes the following steps: respectively determining the timbres of two adjacent sentences, where the two adjacent sentences are the first sentence and the second sentence, and determining whether the timbre of the first sentence is consistent with that of the second sentence. If so, continue to determine the timbres of the next pair of adjacent sentences. If not, generate timbre change information.
[0129] Respectively determining the timbres of two adjacent sentences includes: determining a first audio segment corresponding to the first sentence in the audio data and a second audio segment corresponding to the second sentence in the audio data, and respectively processing the first audio segment and the second audio segment according to a preset timbre recognition algorithm to obtain the timbre of the first audio segment and the timbre of the second audio segment.
[0130] Generally, when the speaker of a meeting switches between speakers of different genders, it is often easy to have a switch in the chapter content at this time. For example, when switching from the speech of a male guest to the speech of a female guest, the chapter content expressed by the male guest is often different from the chapter content expressed by the female guest. The embodiments of the present application use timbre change information to describe the timbre difference between two adjacent sentences, so as to enhance the expression characteristics of the text content and facilitate the accurate and reliable identification of chapter segments.
[0131] In some embodiments, the modal data includes text description information, and the text description information can be used to describe the overall theme of the meeting. For example, the text description information is "create a harmonious and civilized society".
[0132] In some embodiments, extracting modal data matching the text content from the audio-visual data includes: obtaining text description information from the local area of the computer device.
[0133] In some embodiments, extracting modal data matching the text content from the audio-visual data includes: sending a description request to an external device so that the external device returns text description information according to the description request.
[0134] In some embodiments, the text content includes multiple text segments, and the modal data includes image information matching the text segments. Extracting modal data matching the text content from the audio-visual data includes: determining the time of the text segment and extracting at least one frame of image information from the audio-visual data whose time is the same as that of the text segment.
[0135] In terms of fusing the text content with the modal data, the embodiments of the present application provide the following multiple information fusion situations, specifically as follows:
[0136] 1. The first information fusion situation:
[0137] Fusing text content with modal data to obtain multiple fused segments includes the following steps: Fusing the text content with the modal data to obtain the fused text content, and segmenting the fused text content to obtain multiple fused segments.
[0138] In some embodiments, the modal data includes a first type of modal information. Fusing the text content with the modal data to obtain the fused text content includes the following steps: Determining an insertion position in the text content, and splicing the first type of modal information at the insertion position to obtain the fused text content.
[0139] The first type of modal information includes voice feature information. In this embodiment, the voice feature information can be spliced at the insertion position to obtain the fused text content. This embodiment can fuse the first type of modal information with the text content before the segment splitting stage, and there is no need for subsequent vector dimension conversion, which is beneficial to quickly integrating the first type of modal information into the text content and improving the recognition efficiency of chapter segments.
[0140] 2. The second case of information fusion:
[0141] Fusing text content with modal data to obtain multiple fused segments includes the following steps: Segmenting the text content to obtain multiple text segments, fusing the text segments with the modal data to obtain the fused text segments, and multiple fused text segments form multiple fused segments.
[0142] The modal data includes a second type of modal information. Fusing the text segment with the modal data to obtain the fused text segment includes the following steps: Fusing the text segment with the second type of modal information to obtain the fused text segment.
[0143] The second type of modal information includes image information. In this embodiment, the text segment is spliced with the image information to obtain the fused text segment. This embodiment can fuse image information with different attributes from the text segment during the segment processing stage, achieving the purpose of cross-modal fusion. The image information can assist the preset large language model to comprehensively understand the text content of the text segment, enabling this embodiment to reliably and accurately extract the main information in various scenarios.
[0144] Segmenting the text content to obtain multiple text segments includes the following steps: Performing at least two segmentation operations on the text content according to the target window length to obtain multiple text segments, where there are overlapping characters of a specified length between adjacent two text segments.
[0145] Perform at least two segmentation operations on the text content according to the target window length to obtain multiple text segments, including: determining the cutting end point of the previous cutting operation, sliding a preset step from the cutting end point to obtain the cutting start point of the current cutting operation, sliding the preset window length of the previous cutting operation from the cutting start point to obtain the text segment of the current cutting operation, until the text content is completely segmented to obtain multiple text segments, and the preset step is equal to the difference between the target window length and the specified length.
[0146] For example, the target window length is 1024, the specified length is 64, the preset step is 960, the length of each text segment can be 1024, and there will be 64 repeated characters between two adjacent text segments. Since there are overlapping characters of the specified length between two adjacent text segments, this embodiment can reduce the situation that the content belonging to the same chapter segment is unreasonably segmented into other chapter segments due to unreasonable segmentation, avoid content fragmentation, reduce the degree of content loss of the chapter segment, and improve the integrity of the content of the chapter segment.
[0147] 3. The third information fusion situation:
[0148] Perform information fusion on the text content and the modal data to obtain multiple fusion segments, including the following steps: perform at least two segmentation operations on the text content according to the preset window length to obtain at least one candidate segment, splice the modal data at the specified position of the candidate segment obtained in each segmentation operation to obtain the spliced candidate segment, and multiple spliced candidate segments form multiple fusion segments.
[0149] In some embodiments, the modal data includes the third type of modal information. Splicing the modal data at the specified position of the candidate segment obtained in each segmentation operation to obtain the spliced candidate segment includes the following steps: splicing the third type of modal information at the specified position of the candidate segment obtained in each segmentation operation to obtain the spliced candidate segment.
[0150] The third type of modal information includes text description information. In this embodiment, the text description information can be spliced at the specified position of the candidate segment to obtain the spliced candidate segment. This embodiment can fuse the first type of modal information with the candidate segment at the segment segmentation stage, so that each fusion segment carries the third type of modal information, which is beneficial to accurately and reliably identifying chapter segments and accurately and reliably extracting the main information.
[0151] 4. The fourth information fusion situation:
[0152] The modal data includes first - type modal information and second - type modal information. The first - type modal information can be one or both of speech feature information or text description information. The second - type modal information can be image information or other custom information. The other custom information is customized by the designer according to engineering experience. It can be understood that the other custom information does not exclude the first - type modal information. For example, the other custom information is the modal information after the vector dimension of the first - type modal information is converted, and the vector dimension of the first - type modal information will be consistent with the vector dimension of the text segment after conversion.
[0153] To perform information fusion on the text content and the modal data to obtain multiple fusion segments, the following steps are included: Perform information fusion on the text content and the first - type modal information to obtain the fused text content. Segment the fused text content to obtain multiple text segments. Perform information fusion on the text segments and the second - type modal information to obtain the fused text segments. The multiple fused text segments form multiple fusion segments.
[0154] This embodiment can perform information fusion on the text content and the first - type modal information before the segment - splitting stage, and after segment - splitting, that is, in the segment - processing stage, perform information fusion on the text segments and the second - type modal information. In this way, this embodiment can not only assist in enhancing the information content of the text segments from the speech feature dimension or text description dimension, but also enhance the information content of the text segments from the image dimension, thus realizing multi - modal information fusion, which is beneficial to accurately and reliably identifying chapter segments and accurately and reliably extracting the main information.
[0155] 5. The fifth information - fusion case:
[0156] The modal data includes first - type modal information and third - type modal information. To perform information fusion on the text content and the modal data to obtain multiple fusion segments, the following steps are included: Perform information fusion on the text content and the first - type modal information to obtain the fused text content. Perform at least two splitting operations on the fused text content according to a preset window length to obtain at least one candidate segment. Concatenate the third - type modal information at the specified position of the candidate segment obtained in each splitting operation to obtain the concatenated candidate segment. The multiple concatenated candidate segments form multiple fusion segments.
[0157] 6. The sixth information - fusion case:
[0158] The modal data includes second - type modal information and third - type modal information. To perform information fusion on the text content and the modal data to obtain multiple fusion segments, the following steps are included: Segment the text content to obtain multiple text segments, splice the third - type modal information at the specified positions of the text segments obtained by each segmentation operation to obtain spliced text segments, perform information fusion on the spliced text segments and the second - type modal information to obtain fused text segments, and the multiple fused text segments form multiple fusion segments.
[0159] 7. The seventh information fusion scenario:
[0160] The modal data includes first - type modal information, second - type modal information, and third - type modal information. To perform information fusion on the text content and the modal data to obtain multiple fusion segments, the following steps are included: Perform information fusion on the text content and the first - type modal information to obtain fused text content, perform at least two segmentation operations on the fused text content according to a preset window length to obtain at least one candidate segment, splice the third - type modal information at the specified positions of the candidate segments obtained by each segmentation operation to obtain spliced candidate segments, the multiple spliced candidate segments form multiple text segments, perform information fusion on the text segments and the second - type modal information to obtain fused text segments, and the multiple fused text segments form multiple fusion segments.
[0161] It can be understood that the above 7 information fusion scenarios may have the same or technically equivalent means in different literal forms. Therefore, the above 7 information fusion scenarios can mutually adopt the same or technically equivalent means in different literal forms, and the repeated technical means in the above 7 information fusion scenarios will not be elaborated in detail here.
[0162] In some embodiments, performing information fusion on the text content and the first - type modal information to obtain fused text content includes the following steps:
[0163] S3411: Determine the insertion position in the text content.
[0164] S3412: Splice the first - type modal information at the insertion position to obtain the fused text content.
[0165] In S341, the insertion position is the position where the first - type modal information is spliced into the text content. The text content includes multiple sentences, and the first - type modal information includes voice feature information. Determining the insertion position of the first - type modal information into the text content includes the following steps: Determine the sentence associated with the voice feature information as the target sentence, and determine the target position of the target sentence as the insertion position.
[0166] In some embodiments, when the speech feature information includes speech interval information, intonation change information, or timbre change information, determining the sentence associated with the speech feature information as the target sentence includes: detecting whether the interval time between two adjacent sentences is greater than or equal to a preset time threshold, or detecting whether the intonation difference between two adjacent sentences is greater than a preset intonation threshold, or detecting whether the timbres of two adjacent sentences are the same. If so, determining that both of the two adjacent sentences are target sentences, or determining the previous sentence of the two adjacent sentences as the target sentence, or determining the next sentence of the two adjacent sentences as the target sentence. If not, not determining the two adjacent sentences as target sentences.
[0167] The target position can be customized by the designer according to engineering experience. For example, the target position is the beginning, end, or middle of the target sentence.
[0168] In S342, the fused text content is the text content incorporating the first-modal information. In some embodiments, in this embodiment, the first type of modal information is directly inserted at the insertion position to obtain the fused text content. In some embodiments, in this embodiment, the first type of modal information can be converted into a special marker, and the special marker is inserted at the insertion position to obtain the fused text content. The first type of modal information can be speech feature information, etc.
[0169] This embodiment can fuse speech interval information on the text content with configured timestamps. This embodiment detects the text content with configured timestamps to detect whether the interval time between every two sentences is greater than or equal to a preset time threshold. If it is greater than or equal to, this embodiment selects both of the two sentences as target sentences.
[0170] For example, the text content includes sentence R1 and sentence R2, sentence R1 = r1, r2, r3, r4, r5, sentence R2 = r6, r7, r8, r9, r 10 . Since there is a speech interval between sentence R1 and sentence R2, therefore, in this embodiment, special markers [SEP] are inserted at the end of sentence R1 and the end of sentence R2, specifically as follows: r1, r2, r3, r4, r5, [SEP], r6, r7, r8, r9, r 10 , [SEP]. This embodiment uses the special marker [SEP] to replace the speech interval information and inserts the special marker [SEP] at the end of the target sentence, which is beneficial to enhancing the recognition ability of chapter segments.
[0171] In some embodiments, the modal data further includes a third type of modal information. The steps for splitting the fused text content to obtain multiple text segments include:
[0172] S3421: Perform at least two splitting operations on the fused text content according to a preset window length to obtain at least two candidate segments.
[0173] S3422: Concatenate the third type of modality information at the specified positions of the candidate segments obtained in each splitting operation to obtain the concatenated candidate segments, and multiple concatenated candidate segments form multiple text segments.
[0174] In S3421, the candidate segment is a part of the fused text content obtained by splitting according to the preset window length. Among them, the preset window length can be customized by the designer according to business requirements. For example, the preset window length is 512 or 1024 or 2048, etc.
[0175] It can be understood that there may be overlapping characters of a specified length between two adjacent candidate segments, and the characters can be marked as tokens. Among them, the characters include letters, punctuation marks, special symbols, Arabic numerals, etc. The specified length can be customized by the designer according to engineering experience. For example, the specified length is 32 or 64 or 128, etc.
[0176] Please refer to Figure 5 , the fused text content is split into m candidate segments. Among them, there are k overlapping characters between the i-th candidate segment and the (i + 1)-th candidate segment, and there are also k overlapping characters between the (i + 1)-th candidate segment and the (i + 2)-th candidate segment.
[0177] It can also be understood that there may be no overlapping characters of a specified length between two adjacent candidate segments. Among them, the last character of the previous candidate segment is adjacent to the first character of the next candidate segment.
[0178] Please refer to Figure 6 , the j-th candidate segment and the (j + 1)-th candidate segment do not overlap and the (j + 1)-th candidate segment is adjacent after the j-th candidate segment. The (j + 1)-th candidate segment and the (j + 2)-th candidate segment do not overlap and the (j + 2)-th candidate segment is adjacent after the (j + 1)-th candidate segment.
[0179] In some embodiments, performing at least two splitting operations on the fused text content according to the preset window length to obtain at least one candidate segment includes the following steps: performing at least two splitting operations on the fused text content according to the preset window length to obtain at least two candidate segments includes: determining the cutting end point of the previous cutting operation, sliding a specified step from the cutting end point to obtain the cutting start point of this cutting operation, sliding the preset window length of the previous cutting operation from the cutting start point to obtain the candidate segment of this cutting operation until the fused text content is completely split to obtain multiple candidate segments, and the specified step is equal to the difference between the preset window length and the specified length.
[0180] In S3422, the specified position can be the beginning, the end, or the middle of the candidate segment. The third type of modal information can be text description information, etc.
[0181] In some embodiments, the preset window length is equal to the length of the fusion segment minus the sum of the lengths of the first type of modal information and the third type of modal information. When the first type of modal information is represented by special markers, the preset window length is equal to the preset encoding length minus the sum of the length of the special markers and the length of the third type of modal information. It can be understood that the length of the fusion segment is equal to the preset encoding length, and the preset encoding length is the number of characters that the preset language encoding model can receive as input when encoding the fusion segment.
[0182] For example, the length of the fusion segment is 1024, the special marker of the first type of modal information is 24, the length of the third type of modal information is 48, the specified length is 64, and the preset window length is 1024 - 24 - 48 = 952, and the specified stride is 952 - 64 = 888.
[0183] For example, the preset window length is c, the specified stride is d, the text content consists of e characters, the length of the special marker of the first type of modal information is f, and the length of the third type of modal information is g. According to the above splitting method, the e characters of the text content can be split into the following text segments, as shown below:
[0184] Candidate segment 1: f1, f2, f3, …, [SEP], …, f c ;
[0185] Candidate segment 2: [SEP], f c-d+1 , f c-d+2 , f c-d+3 , ……, f c , f c+1 , f c+2 , ……, f 2c-d ;
[0186] Candidate segment 3: f 2c-2d+1 , f 2c-2d+2 , f 2c-2d+3 , …, [SEP], …, f 2c , f 2c+1 , f 2c+2 , ……, f 3c-2d ;
[0187] Next, in this embodiment, the third type of modal information is respectively spliced onto candidate segment 1, candidate segment 2, and candidate segment 3. It can be understood that candidate segment 1, candidate segment 2, and candidate segment 3 can be spliced with the same third type of modal information or different third type of modal information, as shown below:
[0188] Fusion fragment T1 corresponding to candidate fragment 1: Text description information, f1, f2, f3, …, [SEP], …, f c ;
[0189] Fusion fragment T2 corresponding to candidate fragment 2: Text description information, [SEP], f c-d+1 , f c-d+2 , f c-d+3 , ……, f c , f c+1 , f c+2 , ……, f 2c-d ;
[0190] Fusion fragment T3 corresponding to candidate fragment 3: Text description information, f 2c-2d+1 , f 2c-2d+2 , f 2c-2d+3 , …, [SEP], …, f 2c , f 2c+1 , f 2c+2 , ……, f 3c-2d ;
[0191] As described above, due to the existence of overlapping characters of a specified length between adjacent two text fragments, this embodiment can reduce the situation that the content belonging to the same chapter fragment is unreasonably segmented into other chapter fragments due to unreasonable segmentation, avoid content fragmentation, reduce the degree of content loss of the chapter fragment, and improve the integrity of the content of the chapter fragment.
[0192] In some embodiments, the second type of modal information includes image information matching the text fragment. Fusing the text fragment with the second type of modal information, the obtained fused text fragment S343 includes: fusing the text fragment with the image information to obtain the fused text fragment, and the image information is an image with the same time as the text fragment.
[0193] Multiple fused text fragments form multiple fusion fragments. Among them, the fusion fragment not only fuses the speech feature information from the speech feature dimension and the text description information from the text description dimension, but also fuses the image information from the image dimension. In this way, it can enhance the information content of the text fragment multi-dimensionally and multimodally, which is beneficial to accurately and reliably identifying the chapter fragment and accurately and reliably extracting the main information.
[0194] Fusing the text fragment with the image information to obtain the fused text fragment includes: performing encoding processing on the text fragment to obtain the encoded text fragment, performing encoding processing on the image information to obtain the encoded image information, and fusing the encoded text fragment with the encoded image information to obtain the fused text fragment.
[0195] Encode the text segment, and the encoded text segment includes: encoding the text segment based on a preset language encoding model to obtain the encoded text segment.
[0196] The preset language encoding model includes the BERT model or the LEBERT model, etc. The LEBERT model uses a Lexicon adaptation layer with Chinese sequence tags to enhance the BERT layer. Among them, the LEBERT model directly integrates external lexical information into the BERT layer through a Lexicon adaptation layer. Compared with existing methods, the LEBERT model helps to perform deep lexical knowledge integration in the lower layers of the BERT model.
[0197] When the preset language encoding model is the LEBERT model, in this embodiment, the base version or the large version of the LEBERT model can be selected. Among them, the base version of the LEBERT model has 12 layers of transformers, a hidden layer size of 768, and 12 heads in the multi-head attention mechanism. The large version of the LEBERT model has 24 layers of transformers, a hidden layer size of 1024, and 16 heads in the multi-head attention mechanism.
[0198] If this embodiment uses the large version of the LEBERT model to encode the text segment, the following output can be obtained:
[0199] E i = M1(T i )
[0200] Among them, E i is the i-th text segment after encoding, M1 is used to represent the large version of the LEBERT model, and T i is the i-th text segment before encoding.
[0201] Since the maximum input length of the LEBERT model is 1024 tokens, and this embodiment selects the large version of the LEBERT model, therefore, the dimension of the i-th fused segment Y i after encoding is 1024*1024. Among them, the first "1024" in "1024*1024" is the number of tokens of the i-th fused segment Yi after encoding, and the second "1024" is the dimension of each token in the i-th fused segment Yi after encoding.
[0202] Encode the image information, and the encoded image information includes: encoding the image information based on a preset image encoding model to obtain the encoded image information.
[0203] The preset image encoding model can be models such as the Taiyi model of The Investiture of the Gods. The Taiyi model is the first open-source Chinese CLIP model and has powerful visual-linguistic representation capabilities.
[0204] In this embodiment, models such as the Taiyi model are used to encode the image information to obtain the encoded image information P i , and the encoded image information P i has a vector dimension of 1*1024.
[0205] In this embodiment, the encoded fusion segment Y i is fused with the encoded image information P i to obtain the fused text segment Y i , and the fused text segment Y i has a vector dimension of 1024*2048.
[0206] In some embodiments, screening the fusion segments that meet the preset recognition conditions as chapter segments includes the following steps:
[0207] S3511: Determine whether the target fusion segment has chapter features. The target fusion segment is one of the multiple fusion segments.
[0208] S3512: If the target fusion segment has chapter features, determine that the target fusion segment meets the preset recognition conditions and determine that the target fusion segment is a chapter segment.
[0209] S3513: If the target fusion segment does not have chapter features, determine that the target fusion segment does not meet the preset recognition conditions and determine that the target fusion segment is a non-chapter segment.
[0210] In S3511, in this embodiment, one fusion segment is sequentially selected from the multiple fusion segments as the target fusion segment. For example, in this embodiment, the fusion segment T1 is first selected as the target fusion segment, and then the fusion segment T2 is selected as the target fusion segment, and so on.
[0211] In S3512, since the target fusion segment has chapter features, it indicates that the target fusion segment belongs to a chapter segment. For example, when the chapter start point and / or chapter end point are extracted from the target fusion segment in this embodiment, it indicates that the target fusion segment belongs to a chapter segment.
[0212] In S3513, since there are no chapter features in the target fusion segment, it indicates that the target fusion segment belongs to a non-chapter segment. For example, when the starting point and / or ending point of a chapter cannot be extracted from the target fusion segment in this embodiment, it indicates that the target fusion segment belongs to a non-chapter segment. For example, the content of the target fusion segment is chat content. By analyzing whether there are chapter features in the target fusion segment, this embodiment can effectively and reliably find chapter segments among multiple fusion segments, which is beneficial to improving the accuracy and reliability of generating the main idea information.
[0213] In some embodiments, determining whether there are chapter features in the target fusion segment includes the following steps:
[0214] S35111: Sequentially input the target fusion segment into a preset score model to obtain a chapter feature score corresponding to the target fusion segment.
[0215] S35112: Determine whether there are chapter features in the target fusion segment according to the chapter feature score.
[0216] In S35111, the preset score model is used to score whether the characters of the target fusion segment are chapter features. For example, the preset score model scores whether the characters of the target fusion segment are the starting point or ending point of a chapter. The preset score model includes a softmax model, a sigmoid model, a relu model, an SVD classification model, and so on.
[0217] In some embodiments, the preset score model is a single-layer pointer network layer. Sequentially inputting the target fusion segment into the preset score model to obtain a chapter feature score corresponding to the target fusion segment includes: Sequentially inputting the target fusion segment into the single-layer pointer network layer to obtain a chapter feature score corresponding to the target fusion segment.
[0218] In some embodiments, the preset score model is a multi-layer pointer network layer. Sequentially inputting the target fusion segment into the preset score model to obtain a chapter feature score corresponding to the target fusion segment includes: Sequentially inputting the target fusion segment into the multi-layer pointer network layer for dimensionality reduction processing, and outputting a chapter feature score corresponding to the target fusion segment.
[0219] The multi-layer pointer network layer includes a first pointer network layer and a second pointer network layer. Sequentially inputting the target fusion segment into the multi-layer pointer network layer for dimensionality reduction processing, and outputting a chapter feature score corresponding to the target fusion segment includes: Inputting the target fusion segment into the first pointer network layer for dimensionality reduction processing to obtain dimensionality reduction output data, and inputting the dimensionality reduction output data into the second pointer network layer to obtain a chapter feature score corresponding to the target fusion segment.
[0220] After the fusion segment fuses the second type of modal information that is image information, the vector dimension of the fusion segment will increase. For example, the vector dimension of the fusion segment that only fuses the first modal information and the second modal information is 1024*1024. When this fusion segment fuses the second type of modal information, where the vector dimension of the second type of modal information is 1*1024, the vector dimension of this fusion segment becomes: 1024*2048. In this embodiment, a first pointer network layer can be used to perform dimensionality reduction processing on this fusion segment. For example, the vector dimension of the fusion segment 1024*2048 is reduced to 1024*1024 or 1024*512, and then the fusion segment with a vector dimension of 1024*1024 or 1024*512 is input into the second pointer network layer to obtain a vector dimension of 1024*2. This output data of 1024*2 can represent the chapter feature score corresponding to the target fusion segment. Through the dimensionality reduction processing in this embodiment, the preset score model can be made relatively lightweight and convenient for data processing.
[0221] The chapter feature score is the score obtained by the preset score model for scoring whether the characters of the target fusion segment are chapter features. Among them, the preset score model represents the chapter feature score in a normalized manner. For example, the preset score model is a softmax model. The length of the target fusion segment is 1024. After each character is processed by the softmax model, the chapter feature score of this character is any value between [0,1].
[0222] In some embodiments, when the chapter feature score is less than 0.5, this embodiment sets the chapter feature score to 0. When the chapter feature score is greater than or equal to 0.5, this embodiment sets the chapter feature score to 1. Therefore, the chapter feature score corresponding to each character output by the softmax model is one of the values 0 and 1.
[0223] In S35112, in some embodiments, the target fusion segment includes multiple characters, the chapter feature includes a chapter start point, the chapter feature score includes a start point score, and determining whether the target fusion segment has a chapter feature according to the chapter feature score includes the following steps: determining whether the start point score is greater than or equal to a preset start threshold. If so, it is determined that the target fusion segment has a chapter feature, and the character whose start point score is greater than or equal to the preset start threshold is determined as the chapter start point. If not, it is determined that the target fusion segment does not have a chapter feature.
[0224] The start point score is used to represent the probability that the character of the target fusion segment belongs to the chapter start point. The preset start threshold can be customized by the designer according to the product situation. For example, when the preset score model is a softmax model, the preset start threshold is 0.5.
[0225] In some embodiments, the target fusion segment includes multiple characters, the chapter feature includes a chapter end point, the chapter feature score includes an end point score, and determining whether the target fusion segment has a chapter feature according to the chapter feature score includes the following steps: determining whether the end point score is greater than or equal to a preset end threshold, if so, determining that the target fusion segment has a chapter feature, and determining the character with the end point score greater than or equal to the preset end threshold as the chapter end point, if not, determining that the target fusion segment does not have a chapter feature.
[0226] The end point score is used to represent the probability that the character of the target fusion segment belongs to the chapter end point. The preset end threshold can be customized by the designer according to the product situation. For example, when the preset score model is the softmax model, the preset end threshold is 0.5.
[0227] When the chapter feature includes a chapter start point and a chapter end point, the preset score model can output a start point score and an end point score corresponding to each character of the target fusion segment.
[0228] For example, the dimension of the i-th target fusion segment Y i after encoding is 1024*2048. The i-th target fusion segment Y i after encoding is input into the preset score model, and the preset score model can output a vector of 1024*2 dimensions to obtain the start point score and end point score of each character, as follows:
[0229] D i = softmax(W e Y i + b e )
[0230] where D i is a vector of 1024*2 dimensions. Among them, D i includes the start point scores and end point scores of 1024 tokens. Among them, is the start point score of the j-th character in D i , and is the end point score of the j-th character in D i . If there is a character in D i whose start point score is greater than the preset start threshold and / or end point score is greater than the preset end threshold, it indicates that the target fusion segment T i corresponding to D i has a chapter feature, and the target fusion segment T i is a chapter segment. If there is no character in D i whose start point score is greater than the preset start threshold and end point score is greater than the preset end threshold, it indicates that the target fusion segment T i corresponding to Di Without chapter features, the target fusion segment T i is a non-chapter segment. Based on the preset score model, this embodiment can quickly and reliably output the chapter feature scores of each character in the target fusion segment, so as to quickly and effectively identify whether the target fusion segment is a chapter segment.
[0231] In some embodiments, determining the main idea information corresponding to each chapter segment includes the following steps:
[0232] S3521: Determine the reference segment of the target chapter segment, where the reference segment is the text segment arranged after the target chapter segment, and the target chapter segment is one of the multiple chapter segments.
[0233] S3522: Determine the content coverage range of the target chapter segment according to the attributes of the reference segment.
[0234] S3523: Generate the main idea information according to the content coverage range of the target chapter segment.
[0235] In S3521, for example, please refer to Figure 7 , the text content corresponds to G fusion segments, where the fusion segment Y h is a chapter segment, the fusion segment Y h+1 is a non-chapter segment, the fusion segment Y h+2 is a chapter segment, the fusion segment Y h+3 is a chapter segment, the fusion segment Y h+4 is a non-chapter segment, the fusion segment Y h+5 is a chapter segment.
[0236] When this embodiment selects the fusion segment Y h as the target chapter segment, since the fusion segment Y h+1 is arranged after the fusion segment Y h , therefore, the fusion segment Y h+1 is the reference segment.
[0237] When this embodiment selects the fusion segment Y h+2 as the target chapter segment, since the fusion segment Y h+3 is arranged after the fusion segment Y h+2 , therefore, the fusion segment Y h+3 is the reference segment.
[0238] And so on, which will not be elaborated here.
[0239] In S3522, the attributes of the reference fragment include chapter fragment attributes and non-chapter fragment attributes. The chapter fragment attributes are used to indicate that the reference fragment is a chapter fragment, and the non-chapter fragment attributes are used to indicate that the reference fragment is a non-chapter fragment. Among them, the attributes of the reference fragment can be represented by any identifier.
[0240] The content coverage range is the content limit that the target chapter fragment can express. It can be understood that in some embodiments, the content coverage range of the target chapter fragment is equal to the content range carried by the target chapter fragment. In some embodiments, the content coverage range of the target chapter fragment is equal to the sum of the content range carried by the target chapter fragment and the content range of the non-chapter fragments arranged after the target chapter fragment.
[0241] In some application scenarios, during the process of switching from chapter fragment A to chapter fragment B, sometimes there is a lot of chat content inserted between chapter fragment A and chapter fragment B, or although the text fragment between chapter fragment A and chapter fragment B does not have chapter characteristics, the content of this text fragment is still related to chapter fragment A. For example, the text fragment C between chapter fragment A and chapter fragment B does not have a chapter starting point, but the content of text fragment C is still related to chapter fragment A. Therefore, in this embodiment, it can be considered that the content coverage range of chapter fragment A is equal to the sum of the content range carried by chapter fragment A and the content range of text fragment C.
[0242] In S3523, this embodiment can generate the main idea information according to the content coverage range of the target chapter fragment. Therefore, this embodiment can find out the content coverage range of the target chapter fragment to the greatest extent, so as to enrich the information content carried by the target chapter fragment, avoid the fragmentation of the content of the target chapter fragment, and is conducive to generating reliable and accurate main idea information.
[0243] In some embodiments, determining the content coverage range of the target chapter fragment according to the attributes of the reference fragment includes the following steps:
[0244] S35221: Determine whether the attribute of the reference fragment is a chapter fragment attribute.
[0245] S35222: If the attribute of the reference fragment is a chapter fragment attribute, splice the target chapter fragment and the text fragment between the target chapter fragment and the reference fragment to obtain the final fragment, and determine the content range carried by the final fragment as the content coverage range.
[0246] S35223: If the attribute of the reference fragment is a non-chapter fragment attribute, determine the text fragment arranged after the reference fragment as the new reference fragment, and return to the step of determining whether the reference fragment is a chapter fragment.
[0247] In S35221, determining whether the attribute of the reference segment is a chapter segment attribute includes: determining whether the identifier of the reference segment is a chapter segment identifier. If the identifier of the reference segment is a chapter segment identifier, then determine that the reference segment is a chapter segment. If the identifier of the reference segment is a non-chapter segment identifier, then determine that the reference segment is a non-chapter segment.
[0248] In S35222, the final segment is an aggregate obtained by splicing the target chapter segment and the text segments between the target chapter segment and the reference segment. It can be understood that when the reference segment is a chapter segment and the reference segment is the first text segment arranged after the target chapter segment, then it is determined that the text segments between the target chapter segment and the reference segment are empty segments, and an empty segment is a segment without characters. Splicing the target chapter segment and the empty segment, the resulting final segment is equal to the target chapter segment.
[0249] Since the attribute of the reference segment is a chapter segment attribute, it indicates that the content coverage range of the target chapter segment has been found. The content coverage range of the target chapter segment is the sum of the content range of the target chapter segment and the content range of the text segments between the target chapter segment and the reference segment.
[0250] In S35223, since the attribute of the reference segment is a non-chapter segment attribute, it indicates that the content coverage range of the target chapter segment has not been found yet. It is necessary to use the text segment arranged after the reference segment as a new reference segment and return to step S35221 to continue searching for the content coverage range of the target chapter segment until a new reference segment is found to be a chapter segment, then stop the search operation and turn to the segment splicing operation.
[0251] For example, let the first reference segment arranged after the target chapter segment (the (i - 1)th text segment) be the ith text segment. The pseudo-code process for determining the content coverage range of the target chapter segment is as follows:
[0252] S41: Save the target chapter segment (the (i - 1)th text segment) in the open set.
[0253] S42: Select the ith text segment as the reference segment and determine whether the ith text segment is a chapter segment. If not, then execute step S43. If so, then execute step S44.
[0254] S43: Record the ith text segment in the open set, assign i = i + 1, and return to step S41.
[0255] S44: Splice the text segments recorded in the open set in sequence to obtain the final segment.
[0256] For example, please combine Figure 7 , when fusing segment Y hWhen it is the target chapter segment, this embodiment will fuse segment Y h and record it in the open set Q, and select the fused segment Y h+1 as the reference segment to determine whether the Y h+1 text segment is a chapter segment. Since the Y h+1 text segment is a non-chapter segment, therefore, this embodiment will record the fused segment Y h+1 in the open set Q. Then, the open set Q = {Y h , Y h+1}. Next, this embodiment selects the fused segment Y h+2 as the reference segment to determine whether the Y h+2 text segment is a chapter segment. Since the Y h+2 text segment is a chapter segment, therefore, this embodiment will splice the text segments recorded in the open set Q in sequence to obtain the final segment = Y h +Y h+1 .
[0257] For another example, when the fused segment Y h+2 is the target chapter segment, this embodiment will record the fused segment Y h+2 in the open set Q, and select the fused segment Y h+3 as the reference segment to determine whether the Y h+3 text segment is a chapter segment. Since the Y h+3 text segment is a chapter segment, therefore, this embodiment will splice the text segments recorded in the open set Q in sequence to obtain the final segment = Y h+2 .
[0258] This embodiment adopts an iterative and recursive method, which can find out the content coverage range of the target chapter segment to the greatest extent, and is beneficial to improving the reliability and accuracy of the generation of the main idea information.
[0259] In some embodiments, generating the main idea information according to the content coverage range of the target chapter segment includes the following steps: obtaining the custom indication information and the chapter content corresponding to the content coverage range, encapsulating the indication information and the chapter content into model input data, and sending the model input data to a preset language model so that the preset language model outputs the main idea information according to the model input data.
[0260] The indication information is used to represent various requirements of the designer for the preset language model to process the chapter content. In the indication information, the indication information declares what the preset language model needs to do, the output format of the output content, and the chapter content provided for the preset language model. What the indication information declares that the preset language model needs to do includes: extracting the title and abstract of the chapter segment. The output format includes the output formats of the title and abstract.
[0261] The model input data is the prompt terms of a pre - set language model. A prompt is a cue word for an artificial intelligence model and is a method of using natural language to guide or stimulate an artificial intelligence model to complete a specific task. Prompt terms can provide the context of the input information and the parameter information for inputting into the model to the artificial intelligence model. When training a supervised learning or unsupervised learning model, prompt terms can help the artificial intelligence model better understand the intention of the input and make corresponding responses.
[0262] In some embodiments, the expression of the model input data is: ins = "Please generate a text summary for the following text fragment and assign a title that can summarize the full text. The output format is: 'Title': 'Text'\n'Summary': 'Text'."
[0263] The form in which the model input data is input into the pre - set language model is the prompt information + the content coverage of the chapter fragment, as shown in the following formula:
[0264] O i = M2(ins + V i )
[0265] O i is the title and summary of the chapter fragment, M2 is the pre - set language model, ins is the mode prompt instruction, and V i is the content coverage of the chapter fragment.
[0266] In the embodiments of the present application, an open - source large - language model is used as the pre - training base, and fine - tuning training is carried out with tens of thousands of supervised instruction fine - tuning data manually annotated. The specific training method includes the method based on the pure - decoding model transformer structure, and the task of predicting the next token based on the known input is performed. Finally, a pre - set language model M2 with the ability to follow instructions is trained. Among them, the open - source large - language model can be various large - language models provided by the prior art.
[0267] In this embodiment, taking advantage of the characteristic that the pre - set language model is good at generating titles and summaries for texts, by constructing the model input data and sending the model input data to the pre - set language model, the pre - set language model can output titles and summaries that are reliable, accurate, highly fluent, and highly readable.
[0268] It should be noted that in the above - mentioned various embodiments, there is not necessarily a certain order between the above - mentioned steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present application that in different embodiments, the above - mentioned steps can have different execution orders, that is, they can be executed in parallel, or they can be executed alternately, etc.
[0269] As another aspect of the embodiments of the present application, the embodiments of the present application provide a device for extracting key information based on audio-visual data. Among them, the device for extracting key information based on audio-visual data may be a software module, and the software module includes several instructions stored in a memory. A processor can access this memory and call the instructions for execution to complete the method for extracting key information based on audio-visual data described in each of the above embodiments.
[0270] In some embodiments, the device for extracting key information based on audio-visual data can also be built by hardware devices. For example, the device for extracting key information based on audio-visual data can be built by one or more than two chips. Each chip can work in coordination with each other to complete the method for extracting key information based on audio-visual data described in each of the above embodiments. For another example, the device for extracting key information based on audio-visual data can also be built by various logic devices, such as being built by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.
[0271] Please refer to Figure 8 , the device 800 for extracting key information based on audio-visual data includes: a data acquisition module 81, a text conversion module 82, a modality extraction module 83, an information fusion module 84, a key extraction module 85, and a key summary module 86.
[0272] The data acquisition module 81 is used to acquire audio-visual data. The text conversion module 82 is used to extract audio data from the audio-visual data and convert it into text content. The modality extraction module 83 is used to extract modality data matching the text content from the audio-visual data. The information fusion module 84 is used to perform information fusion on the text content and the modality data to obtain multiple fusion segments. The key extraction module 85 is used to screen the fusion segments that meet the preset recognition conditions for key information extraction. The key summary module 86 is used to summarize all the key information extracted as the key information of the audio-visual data.
[0273] This embodiment can fuse the text content and the modality data to obtain multiple fusion segments. The fusion segments carry rich information content, enhance the expression accuracy of the text content, and are conducive to eliminating the interference brought by the user's language expression or other factors, so that key information can be accurately and reliably extracted.
[0274] In some embodiments, the information fusion module 84 is specifically used for: performing information fusion on the text content and the modality data to obtain the fused text content, and splitting the fused text content to obtain multiple fusion segments.
[0275] In some embodiments, the information fusion module 84 is specifically configured to: segment the text content to obtain a plurality of text segments, perform information fusion on the text segments and the modality data to obtain fused text segments, and the plurality of fused text segments form a plurality of fusion segments.
[0276] In some embodiments, the modality data includes first-class modality information and second-class modality information. The information fusion module 84 is specifically configured to: perform information fusion on the text content and the first-class modality information to obtain fused text content, segment the fused text content to obtain a plurality of text segments, perform information fusion on the text segments and the second-class modality information to obtain fused text segments, and the plurality of fused text segments form a plurality of fusion segments.
[0277] In some embodiments, the information fusion module 84 is specifically configured to: determine an insertion position in the text content, and splice the first-class modality information at the insertion position to obtain fused text content.
[0278] In some embodiments, the text content includes a plurality of sentences, and the first-class modality information includes speech feature information. The information fusion module 84 is specifically configured to: determine the sentence associated with the speech feature information as the target sentence, and determine the target position of the target sentence as the insertion position.
[0279] In some embodiments, the modality data further includes third-class modality information. The information fusion module 84 is specifically configured to: perform at least two segmentation operations on the fused text content according to a preset window length to obtain at least two candidate segments, splice the third-class modality information at a specified position of each candidate segment obtained by each segmentation operation to obtain spliced candidate segments, and the plurality of spliced candidate segments form a plurality of text segments.
[0280] In some embodiments, the second-class modality information includes image information matching the text segment. The information fusion module 84 is specifically configured to: perform information fusion on the text segment and the image information to obtain a fused text segment, and the image information is an image with a shooting time consistent with the time of the text segment.
[0281] In some embodiments, the theme extraction module 85 is specifically configured to: screen the fusion segments that meet the preset recognition conditions as chapter segments, and extract the theme information of the chapter segments.
[0282] In some embodiments, the theme extraction module 85 is specifically configured to: determine whether the target fusion segment has a chapter feature, where the target fusion segment is one of multiple fusion segments. If the target fusion segment has a chapter feature, it is determined that the target fusion segment meets the preset recognition condition, and the target fusion segment is determined to be a chapter segment. If the target fusion segment does not have a chapter feature, it is determined that the target fusion segment does not meet the preset recognition condition, and the target fusion segment is determined to be a non-chapter segment.
[0283] In some embodiments, the theme extraction module 85 is specifically configured to: sequentially input the target fusion segment into a preset score model to obtain a chapter feature score corresponding to the target fusion segment, and determine whether the target fusion segment has a chapter feature according to the chapter feature score.
[0284] In some embodiments, the target fusion segment includes multiple characters, the chapter feature includes a chapter start point, and the chapter feature score includes a start point score. The theme extraction module 85 is specifically configured to: determine whether the start point score is greater than or equal to a preset start threshold. If so, it is determined that the target fusion segment has a chapter feature, and the character with the start point score greater than or equal to the preset start threshold is determined to be the chapter start point. If the start point scores of all characters of the target fusion segment are less than the preset start threshold, it is determined that the target fusion segment does not have a chapter feature.
[0285] In some embodiments, the theme extraction module 85 is specifically configured to: determine a reference segment of the target chapter segment, where the reference segment is a text segment arranged after the target chapter segment, and the target chapter segment is one of multiple chapter segments. Determine the content coverage range of the target chapter segment according to the attribute of the reference segment, and generate theme information according to the content coverage range of the target chapter segment.
[0286] In some embodiments, the theme extraction module 85 is specifically configured to: determine whether the attribute of the reference segment is a chapter segment attribute. If the attribute of the reference segment is a chapter segment attribute, splice the target chapter segment and the text segment between the target chapter segment and the reference segment to obtain a final segment, and determine the content range carried by the final segment as the content coverage range. If the attribute of the reference segment is a non-chapter segment attribute, determine the text segment arranged after the reference segment as the new reference segment.
[0287] In some embodiments, the theme extraction module 85 is specifically configured to: obtain a custom indication information and the chapter content corresponding to the content coverage range, encapsulate the indication information and the chapter content into model input data, and send the model input data to a preset language model, so that the preset language model outputs theme information according to the model input data.
[0288] In some embodiments, the main idea information includes the start time point, title, and abstract corresponding to each chapter segment. The chapter segment is a fusion segment that meets the preset recognition conditions. The main idea extraction module 85 is specifically configured to: concatenate the titles and abstracts of each chapter segment in sequence according to the order of the start time points of each chapter segment to obtain the main idea information of the audio-visual data.
[0289] It should be noted that the above-mentioned main idea information extraction device based on audio-visual data can execute the main idea information extraction method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in the embodiments of the main idea information extraction device based on audio-visual data, reference can be made to the main idea information extraction method provided by the embodiments of the present application.
[0290] See Figure 9 , Figure 9 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 900 includes one or more processors 91 and a memory 92. The memory 92 is connected to one or more processors 91, for example, connected to the processor through a bus.
[0291] The processor 91 is configured to support the computer device to execute the corresponding functions in the method in the above method embodiments. The processor can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The above hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0292] The memory 92 is used to store program codes and the like. The memory may include volatile memory (VM), such as random access memory (RAM); the memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory may further include a combination of the above types of memories.
[0293] The memory 92 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for extracting the gist information based on audio-visual data in the embodiments of the present application. The processor executes various functional applications and data processing of the method for extracting the gist information based on audio-visual data and the device for extracting the gist information based on audio-visual data by running the non-volatile software programs, instructions, and modules stored in the memory, that is, realizes the functions of each module or unit of the method for extracting the gist information based on audio-visual data and the device for extracting the gist information based on audio-visual data provided in the above method embodiments.
[0294] The memory 92 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function. The data storage area can store data created according to the use of the device for extracting the gist information based on audio-visual data, etc. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the device for extracting the gist information based on audio-visual data through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0295] The one or more modules are stored in the memory and, when executed by the one or more processors, execute the method for extracting the gist information based on audio-visual data in any of the above method embodiments. For example, execute the method steps described in the above method embodiments and realize the functions of the modules described in the above device embodiments.
[0296] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the method as described in the foregoing embodiments.
[0297] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0298] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for extracting keynote information based on audio and video data, characterized in that, Including: Obtaining audio-visual data; Extracting audio data from the audio-visual data and converting it into text content; Extracting modality data matching the text content from the audio-visual data; Performing information fusion on the text content and the modality data to obtain multiple fusion segments; Screening the fusion segments that meet the preset recognition conditions for extracting the main information; Summarizing all the extracted main information as the main information of the audio-visual data.
2. The method according to claim 1, characterized in that, The performing information fusion on the text content and the modality data to obtain multiple fusion segments includes: Performing information fusion on the text content and the modality data to obtain the fused text content; Segmenting the fused text content to obtain multiple fusion segments.
3. The method according to claim 1, characterized in that, The performing information fusion on the text content and the modality data to obtain multiple fusion segments includes: Segmenting the text content to obtain multiple text segments; Performing information fusion on the text segments and the modality data to obtain the fused text segments, and multiple fused text segments form multiple fusion segments.
4. The method according to claim 1, characterized in that, The modality data includes the first type of modality information and the second type of modality information. The performing information fusion on the text content and the modality data to obtain multiple fusion segments includes: Performing information fusion on the text content and the first type of modality information to obtain the fused text content; Segmenting the fused text content to obtain multiple text segments; Performing information fusion on the text segments and the second type of modality information to obtain the fused text segments, and multiple fused text segments form multiple fusion segments.
5. The method according to claim 4, wherein The performing information fusion on the text content and the first type of modality information to obtain the fused text content includes: Determining the insertion position in the text content; Concatenating the first type of modality information at the insertion position to obtain the fused text content.
6. The method according to claim 5, wherein The text content includes multiple sentences, and the first type of modality information includes speech feature information. The determining the insertion position in the text content includes: Determining the sentence associated with the speech feature information as the target sentence; Determining the target position of the target sentence as the insertion position.
7. The method according to claim 4, wherein The modality data further includes the third type of modality information. The segmenting the fused text content to obtain multiple text segments includes: Performing at least two segmentation operations on the fused text content according to a preset window length to obtain at least two candidate segments; Concatenating the third type of modality information at the specified position of the candidate segments obtained by each segmentation operation to obtain the concatenated candidate segments, and multiple concatenated candidate segments form multiple text segments.
8. The method according to claim 4, characterized in that The second type of modality information includes image information matching the text segment. The performing information fusion on the text segment and the second type of modality information to obtain the fused text segment includes: Performing information fusion on the text segment and the image information to obtain the fused text segment, and the image information is an image with the same shooting time as the text segment.
9. The method according to any one of claims 1 to 8, characterized in that, The screening the fusion segments that meet the preset recognition conditions for extracting the main information includes: Screen the fusion fragments that meet the preset recognition conditions as chapter fragments; Extract the main information of the chapter fragments.
10. The method according to claim 9, wherein The screening of the fusion fragments that meet the preset recognition conditions as chapter fragments includes: Determine whether the target fusion fragment has chapter features, where the target fusion fragment is one of the multiple fusion fragments; If the target fusion fragment has chapter features, determine that the target fusion fragment meets the preset recognition conditions and determine that the target fusion fragment is a chapter fragment; If the target fusion fragment does not have chapter features, determine that the target fusion fragment does not meet the preset recognition conditions and determine that the target fusion fragment is a non-chapter fragment.
11. The method according to claim 10, wherein The determination of whether the target fusion fragment has chapter features includes: Sequentially input the target fusion fragment into a preset score model to obtain a chapter feature score corresponding to the target fusion fragment; Determine whether the target fusion fragment has chapter features according to the chapter feature score.
12. The method according to claim 11, wherein, The target fusion fragment includes multiple characters, the chapter features include chapter start points, and the chapter feature score includes a start point score. The determination of whether the target fusion fragment has chapter features according to the chapter feature score includes: Judge whether the start point score is greater than or equal to a preset start threshold; If so, determine that the target fusion fragment has chapter features and determine the character with the start point score greater than or equal to the preset start threshold as the chapter start point; If the start point scores of all characters of the target fusion fragment are less than the preset start threshold, determine that the target fusion fragment does not have chapter features.
13. The method according to claim 9, wherein The extraction of the main information of the chapter fragments includes: Determine the reference fragment of the target chapter fragment, where the reference fragment is a text fragment arranged after the target chapter fragment, and the target chapter fragment is one of the multiple chapter fragments; Determine the content coverage of the target chapter fragment according to the attribute of the reference fragment; Generate main information according to the content coverage of the target chapter fragment.
14. The method according to claim 13, wherein The determination of the content coverage of the target chapter fragment according to the attribute of the reference fragment includes: Judge whether the attribute of the reference fragment is a chapter fragment attribute; If the attribute of the reference fragment is a chapter fragment attribute, splice the target chapter fragment and the text fragment between the target chapter fragment and the reference fragment to obtain a final fragment, and determine the content range carried by the final fragment as the content coverage; If the attribute of the reference fragment is a non-chapter fragment attribute, determine the text fragment arranged after the reference fragment as a new reference fragment, and return to the step of judging whether the reference fragment is a chapter fragment.
15. The method according to claim 13, wherein The generation of the main information according to the content coverage of the target chapter fragment includes: Obtain the custom indication information and the chapter content corresponding to the content coverage; Package the indication information and the chapter content into model input data; Send the model input data to a preset language model so that the preset language model outputs main information according to the model input data.
16. The method according to any one of claims 1 to 15, characterized in that, The said key information includes the starting time point, title and abstract corresponding to each chapter segment, where the chapter segment is a fusion segment that meets the preset recognition conditions, and all the key information obtained by summary extraction as the key information of the audio-visual data includes: According to the order of the starting time points of each chapter segment, the titles and abstracts of each chapter segment are concatenated in sequence to obtain the key information of the audio-visual data.
17. A computer device, characterized in that, It includes a memory and a processor, the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device realizes the method according to any one of claims 1-16.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the processor executes the method according to any one of claims 1-16.
Citation Information
Cited By
Conference management method and system based on natural language processing and retrieval enhancement
CN120634500A