Information extraction method and apparatus, and computer program product and electronic device

By converting audio into text and performing entity recognition and classification, combined with phonetic features and task context information, the problem of low efficiency and insufficient accuracy in extracting meeting minutes information is solved, achieving efficient and accurate extraction of minutes information.

WO2026112973A1PCT designated stage Publication Date: 2026-06-04BOE TECHNOLOGY GROUP CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2024-11-29
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In existing technologies, the extraction efficiency of meeting minutes information is low and the accuracy is insufficient. Manual minutes processing is inefficient, while AI-generated minutes information is not accurate enough, which affects the rational application of minutes information.

Method used

By converting audio content into text content, entity recognition and classification are performed. A pre-trained entity classification network is used to extract entities of the target type. Entities are split and encoded according to phonetic features. Candidate entities are determined by combining the current task scenario information and then replaced to extract the minutes information.

Benefits of technology

It improves the efficiency and accuracy of extracting information from meeting minutes, accurately captures special entities in the text content, reduces reasoning time, and improves the precision of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135751_04062026_PF_FP_ABST
    Figure CN2024135751_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers. Provided are an information extraction method and apparatus, and a computer program product and an electronic device. The method comprises: converting collected audio content into text content; performing entity recognition on the text content, so as to obtain an entity recognition result and a feature vector of the entity recognition result, and on the basis of the entity recognition result and the feature vector, performing entity classification to extract an entity of a target type on the basis of a classification result, so as to obtain a target entity recognition result; splitting and encoding the target entity recognition result according to phonetic features, and matching an encoding result with candidate encoding results, so as to determine a target candidate entity on the basis of a matching result, wherein the candidate encoding results are obtained by means of splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined on the basis of current task scenario information; and replacing the target entity recognition result in the text content with the target candidate entity, and on the basis of an obtained target text, extracting summary information. The present disclosure improves the accuracy of summary information extraction. (FIG. 2)
Need to check novelty before this filing date? Find Prior Art

Description

Information extraction methods and devices, computer program products and electronic equipment Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to an information extraction method, an information extraction device, a computer program product, and an electronic device. Background Technology

[0002] In almost every industry, there is a need to record conversations, such as meeting minutes and phone records. In these scenarios, people engage in a large amount of multi-round dialogue. Simply archiving and tracing the audio of these conversations is insufficient for efficient information extraction and workflow support.

[0003] Currently, among the methods for extracting meeting minutes, manual meeting minutes processing is inefficient and cannot fully capture the core content. On the other hand, using AI (Artificial Intelligence) to process meeting minutes has the drawback of insufficient accuracy in the extracted information, which to some extent affects the rational application of the minutes information.

[0004] It should be noted that the information in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this disclosure is to provide an information extraction method, an information extraction device, a computer program product, and an electronic device, thereby improving the accuracy of extracting meeting minutes information to at least a certain extent.

[0006] According to one aspect of this disclosure, an information extraction method is provided, which converts collected audio content into text content; performs entity recognition processing on the text content to obtain entity recognition results and corresponding feature vectors, and performs entity classification based on the entity recognition results and feature vectors to extract entities of the target type according to the entity classification results, thereby obtaining target entity recognition results; splits and encodes the target entity recognition results according to phonetic features, and matches the encoded results with candidate encoded results to determine target candidate entities according to the matching results; wherein, the candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined according to the current task scenario information; replaces the target entity recognition results in the text content with the target candidate entities, and extracts the summary information corresponding to the text content based on the obtained target text.

[0007] In one exemplary embodiment of this disclosure, entity recognition processing is performed on text content to obtain entity recognition results and feature vectors corresponding to the entity recognition results. Entity classification is then performed based on the entity recognition results and feature vectors to extract entities of the target type according to the entity classification results, thereby obtaining target entity recognition results. This includes: extracting text features from the text content and performing entity recognition based on the text features to obtain entity recognition results; determining the feature vectors corresponding to the entity recognition results based on the text features; performing classification processing based on the entity recognition results and the corresponding feature vectors to obtain the entity type corresponding to the entity recognition results, thereby extracting target entity recognition results with the target type based on the entity type.

[0008] In one exemplary embodiment of this disclosure, classification processing is performed based on the entity recognition result and the corresponding feature vector to obtain the entity type corresponding to the entity recognition result, including: inputting the entity recognition result and the corresponding feature vector into a pre-trained entity classification network to obtain the entity type corresponding to the entity recognition result; wherein, the pre-trained entity classification network is obtained by model training based on entity samples, feature vectors of entity samples and entity type labels.

[0009] In one exemplary embodiment of this disclosure, the method further includes: obtaining current task scene information; extracting entities from the current task scene information to obtain scene entities; performing entity generation processing based on the scene entities to obtain reference entities with the same entity type as the scene entities; and determining candidate entities based on the reference entities and the scene entities.

[0010] In one exemplary embodiment of this disclosure, obtaining current task scenario information includes: responding to a user's touch operation and using the touch information corresponding to the touch operation as current task scenario information; and / or, performing topic recognition based on text content, determining the topic information of the text content, and using the topic information as current task scenario information.

[0011] In an exemplary embodiment of this disclosure, if topic recognition is performed based on text content to determine the topic information of the text content and the topic information is used as the current task scene information, the method further includes: if a first preset entity associated with the topic information exists in a preset entity library, then the first preset entity is simultaneously determined as a scene entity.

[0012] In one exemplary embodiment of this disclosure, the method further includes: performing speech recognition on the audio content, performing topic analysis based on the speech recognition results to obtain a conversation topic; and if a second preset entity associated with the conversation topic exists in a preset entity library, then the second preset entity is simultaneously determined as a scene entity.

[0013] In an exemplary embodiment of this disclosure, if in response to a user's touch operation, the touch information corresponding to the touch operation is used as the current task scene information, the method further includes: determining the target object in the current task scene based on the scene entity extracted from the current task scene information; obtaining the historical entity corresponding to the target object, and simultaneously determining the historical entity as the scene entity.

[0014] In one exemplary embodiment of this disclosure, the method further includes: displaying candidate entities; and updating candidate entities in response to touch operations on the candidate entities.

[0015] In one exemplary embodiment of this disclosure, the target entity recognition result is split and encoded according to phonetic features, and the encoded result is matched with candidate encoded results to determine the target candidate entity based on the matching result. This includes: splitting the pronunciation information of the target entity recognition result into multiple pronunciation units according to phonetic features; encoding each pronunciation unit to obtain multiple encoded results; determining the sub-similarity between each encoded result and the corresponding candidate encoded result, and fusion calculating the sub-similarity corresponding to each encoded result to obtain the fusion similarity corresponding to the candidate encoded result; and determining the target candidate entity based on the fusion similarity corresponding to the candidate encoded result.

[0016] In one exemplary embodiment of this disclosure, the pronunciation information of the target entity recognition result is divided into multiple pronunciation units according to the phonetic features, including: if the target entity recognition result is Chinese, the target entity recognition result is converted into Pinyin, and the Pinyin is divided into initial consonant units, final vowel units and tone units; or, if the target entity recognition result is English, the target entity recognition result is divided into at least one consonant phoneme unit and at least one vowel phoneme unit according to the phonetic symbols.

[0017] In one exemplary embodiment of this disclosure, the method further includes: the target candidate entity includes a primary entity and at least one secondary entity; wherein the at least one secondary entity is displayed in a manner different from that of the primary entity in the text content and / or in the minutes information.

[0018] In one exemplary embodiment of this disclosure, extracting summary information corresponding to the text content based on the obtained target text includes: obtaining processor information of the user device and determining a target summary extraction model based on the processor information; segmenting the target text based on the text length of the target text and performing summary extraction processing on the segmented text using the target summary extraction model to obtain segmented summary information corresponding to each segmented text; and concatenating the segmented summary information to obtain summary information corresponding to the text content.

[0019] In one exemplary embodiment of this disclosure, segmenting the target text based on its text length includes: performing topic identification on the target text, dividing the target text into multiple topic paragraphs based on the identification results; and for each topic paragraph, segmenting the topic paragraph based on its text length to obtain the segmented text corresponding to the topic paragraph.

[0020] In one exemplary embodiment of this disclosure, determining the target time summary extraction model based on processor information includes: if the processor information indicates that the user equipment is not configured with a graphics processing module, then determining a first quantization model as the target time summary extraction model; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information meets a preset model condition, then determining a second quantization model as the target time summary extraction model; wherein the second quantization model and the first quantization model are obtained by processing the initial time summary extraction model with different quantization schemes; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information does not meet the preset model condition, then determining the initial time summary extraction model as the target time summary extraction model.

[0021] In an exemplary embodiment of this disclosure, a target summary extraction model is used to extract summary information from segmented text to obtain segmented summary information corresponding to each segmented text. This includes: the target summary extraction model outputs segmented summary information of a preset length each time, performs content verification on the segmented summary information of the preset length using adjacent segmented summary information that was output earlier, and controls the extraction process of the target summary extraction model based on the verification results.

[0022] According to one aspect of this disclosure, an information extraction apparatus is provided, comprising: a speech-to-text module for converting acquired audio content into text content; an entity extraction module for performing entity recognition processing on the text content to obtain entity recognition results and feature vectors corresponding to the entity recognition results, and classifying entities based on the entity recognition results and feature vectors to extract entities of the target type according to the classification results, thereby obtaining target entity recognition results; an entity matching module for splitting and encoding the target entity recognition results according to phonetic features, and matching the encoded results with candidate encoded results to determine target candidate entities according to the matching results; wherein, the candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined according to the current task scenario; and an information extraction module for replacing the target entity recognition results in the text content with target candidate entities, and extracting the summary information corresponding to the text content based on the obtained target text.

[0023] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements any of the above methods.

[0024] According to one aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any of the above methods by executing the executable instructions.

[0025] This disclosure provides an information extraction method. On one hand, it converts collected audio content into text content, performs entity recognition processing on the text content to obtain entity recognition results and corresponding feature vectors, and then classifies entities based on the entity recognition results and feature vectors. Based on the classification results, it extracts entities of the target type, obtaining target entity recognition results. This method adds entity classification processing to the named entity recognition process, enabling the extraction of entities of the target type without outputting all entities. It accurately captures specific entities in the text content, providing an accurate entity basis for subsequent entity replacement, reducing inference time and improving information extraction efficiency. On the other hand, it splits and encodes the target entity recognition results according to phonetic features, and matches the encoded results with candidate encoded results to determine target candidate entities based on the matching results. The candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined based on the current task scenario information. The target candidate entities can then be used to replace the target entity recognition results in the text content, and the corresponding summary information can be extracted from the obtained target text. By considering the phonetic features of entities, entities are split and encoded to improve the comprehensiveness and accuracy of the target entity recognition results, so as to accurately distinguish different entities. Furthermore, by determining candidate entities based on the current task scenario information, without relying on historical records, the obtained target candidate entities are made to match the current task scenario, thereby improving the accuracy of entity replacement and thus improving the precision of extracting summary information.

[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0028] Figure 1 schematically illustrates an application environment diagram according to an example embodiment of the present disclosure.

[0029] Figure 2 schematically illustrates a flowchart of an information extraction method according to an exemplary embodiment of the present disclosure.

[0030] Figure 3 schematically illustrates an entity recognition model according to an exemplary embodiment of the present disclosure.

[0031] Figure 4 schematically illustrates a flowchart of an implementation of obtaining alternative entities according to an exemplary embodiment of this disclosure.

[0032] Figure 5 schematically illustrates an interactive interface for displaying alternative entities according to an exemplary embodiment of the present disclosure.

[0033] Figure 6 schematically illustrates a flowchart of an implementation of obtaining target candidate entities according to an exemplary embodiment of the present disclosure.

[0034] Figure 7 schematically illustrates a diagram of displaying minutes information according to an exemplary embodiment of the present disclosure.

[0035] Figure 8 schematically illustrates a flowchart of an implementation method for extracting summary information of target text according to an exemplary embodiment of the present disclosure.

[0036] Figure 9 schematically illustrates a block diagram of an information extraction apparatus according to an exemplary embodiment of the present disclosure.

[0037] Figure 10 schematically illustrates an electronic device for implementing the above-described information extraction method according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0038] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0039] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0040] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0041] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0042] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and the theory of algorithmic complexity. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0043] The technical solution provided by the exemplary embodiments of this disclosure relates to machine learning technology of artificial intelligence. Machine learning technology is used to convert the collected audio content into text content, perform entity recognition processing on the text content, extract the target entity recognition result with the target type, and match the target entity recognition result with the candidate entity determined based on the current task scenario, so as to replace the target entity recognition result in the text content with the obtained target candidate entity, thereby extracting the summary information from the obtained target text and improving the accuracy of extracting the summary information.

[0044] It should be noted that the information extraction method of the exemplary embodiments of this disclosure can be applied to various information extraction tasks, such as extracting meeting minutes, extracting audio recordings, and extracting video content summaries. The exemplary embodiments of this disclosure do not specifically limit the application field of the information extraction method.

[0045] Efficient information extraction facilitates operational support; for example, meeting minutes can identify corporate problems, guide planning, clarify responsibilities, and serve as crucial information for management in certain specific scenarios. However, manual minute-taking is inefficient and costly, while current AI-based methods suffer from insufficient accuracy in extracting information, hindering its proper application. Therefore, this exemplary embodiment provides an information extraction method that improves the accuracy of extracted minute information through multi-task entity extraction and target-type entity replacement.

[0046] It should be noted that the collection, updating, analysis, use, transmission, and storage of user personal information involved in the technical solution disclosed herein all comply with relevant laws and regulations, are used for legitimate and reasonable purposes, and are not shared, disclosed, or sold outside of these legitimate uses, and are subject to supervision and management by national regulatory authorities. Necessary measures should be taken to selectively block the use or access to personal information data to prevent unauthorized access to such personal information data, ensure that personnel authorized to access personal information data comply with relevant laws and regulations, and ensure the security of user personal information. Furthermore, once this user personal information data is no longer needed, the risk should be minimized by restricting or even prohibiting data collection and / or deleting the data.

[0047] The information extraction method provided by the exemplary embodiments of this disclosure can be applied to the application environment shown in FIG1. ​​In this environment, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be integrated onto server 102, or it can be located in the cloud or on another network server.

[0048] In one exemplary embodiment, the information extraction method provided by the exemplary embodiment of this disclosure can be executed by server 102, and the corresponding information extraction device is disposed in server 102. Correspondingly, in this method of execution by server 102, server 102 can start executing the steps in the technical solution of the exemplary embodiment of this disclosure in response to a triggering command, wherein the triggering command can be sent by a terminal used by a user, or can be triggered locally by the server in response to some automated events.

[0049] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Server 102 can execute background tasks.

[0050] Furthermore, in another exemplary embodiment, terminal 101 may also have similar functions to server 102, thereby performing the information extraction method provided by the exemplary embodiments of this disclosure.

[0051] The terminal 101 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, IoT device, or portable wearable device. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. It may also include a display terminal (e.g., an all-in-one conference machine, a smart city display screen), etc. The terminal 101 can also be referred to as a mobile terminal, terminal device, mobile device, etc. The exemplary embodiments of this disclosure do not limit the type of terminal 101.

[0052] Furthermore, the technical solutions of the exemplary embodiments of this disclosure can also be executed collaboratively by terminal 101 and server 102. In this manner, where terminal 101 and server 102 collaborate, some steps of the technical solutions provided in the exemplary embodiments of this disclosure are executed by terminal 101, while other steps are executed by server 102.

[0053] It should be noted that in this method where terminal 101 and server 102 work together, the steps executed by terminal 101 and server 102 can be dynamically adjusted according to the actual situation, and no special restrictions are imposed on this.

[0054] The terminal 101 and the server 102 can be connected directly or indirectly via wireless communication, and the exemplary embodiments of this disclosure are not particularly limited herein.

[0055] In an exemplary embodiment of this disclosure, an information extraction method is provided. Referring to FIG2, which is a flowchart of an exemplary embodiment of the information extraction method of this disclosure, the information extraction method includes steps S210 to S240, as follows:

[0056] In step S210, the collected audio content is converted into text content.

[0057] In exemplary embodiments of this disclosure, audio acquisition and audio-to-text conversion can be performed in parallel. That is, while acquiring audio content through the audio acquisition device, the process of converting the audio content into text content is executed simultaneously. Alternatively, all acquired audio content can be converted into text. Of course, the audio content can be received audio or audio content obtained from the cloud or a database. Exemplary embodiments of this disclosure include, but are not limited to, the methods of acquiring audio content described above. The audio acquisition device can be a recording program built into the terminal device or a dedicated recording device (such as a voice recorder); there is no specific limitation in this regard.

[0058] To improve audio processing efficiency, the acquired audio content can be broken down into multiple sub-audio segments, and each sub-audio segment can be converted into text in turn. For example, the audio content can be split into segments of 5 minutes each to reduce the processing pressure on subsequent models and improve processing speed.

[0059] In one exemplary embodiment, a finely tuned speech recognition model can be used to convert audio content into text content. Multiple audio segments can be input into the speech recognition model for audio-to-text conversion, and the converted text can be concatenated to obtain complete text content.

[0060] The speech recognition model can be constructed using lightweight neural networks, such as CNNs (Convolutional Neural Networks), RNNs (Recurrent Neural Networks), and DNNs (Deep Neural Networks). Alternatively, it can employ more complex network structures, such as heavyweight neural networks like Variational Autoencoders (VAEs), BERT (Bidirectional Encoder Representation from Transformers), Transformer (a language processing model), and any heavyweight neural network suitable for speech recognition. Furthermore, the speech recognition model can also be a large-scale model, referring to a deep learning model trained on massive amounts of data. The exemplary embodiments of this disclosure use a large language model as an example for illustration.

[0061] Optionally, to further improve the fine-tuning training speed of the speech recognition model, some parameters in the model can be converted to int8 integers. For example, the parameters of the self-attention part of each layer in the speech recognition model can be converted to int8 and fine-tuned, such as the fine-tuning of the weight matrices (e.g., the parameters of the four weight matrices Wq, Wk, Wv, and Wo) involved in the self-attention part. Converting from a typical 32-bit single-precision floating-point number to int8 can speed up computation and reduce training time. Of course, other model quantization methods can also be used.

[0062] The fine-tuning training of the speech recognition model can employ the LoRA (Low-Rank Adaptation of LLMs) fine-tuning technique. The essence of LoRA fine-tuning is to approximate the incremental parameters obtained from full-parameter fine-tuning of a large language model using fewer training parameters, thus achieving efficient fine-tuning with less GPU memory usage. Compared to full-parameter fine-tuning, LoRA fine-tuning requires less GPU memory and has a shorter training time. Of course, the exemplary embodiments of this disclosure can select the speech recognition model fine-tuning method according to the actual training task, and no special limitations are imposed in this regard.

[0063] In step S220, entity recognition processing is performed on the text content to obtain entity recognition results and feature vectors corresponding to the entity recognition results. Entity classification is performed based on the entity recognition results and feature vectors, and entities of the target type are extracted according to the classification results to obtain target entity recognition results.

[0064] In the exemplary embodiments of this disclosure, an entity refers to an objectively existing and distinguishable thing. A named entity is an entity identified by its name. Named entity recognition is a key technology in natural language processing, used to identify and label entities with specific meanings in text. Specifically, named entity recognition refers to identifying named entities in text, including names of people, places, organizations, dates and times, numbers, etc., and classifying them into corresponding named entity categories. Unlike traditional named entity recognition, the exemplary embodiments of this disclosure employ a multi-task learning framework, dividing named entity recognition into two tasks: extracting entity recognition results (i.e., named entities) and classifying the entity recognition results according to requirements. That is, in the exemplary embodiments of this disclosure, named entity recognition is used to extract entities and entity features from text content, and to classify entities using entities and entity features to extract target entity recognition results with target types.

[0065] The feature vector corresponding to the entity recognition result refers to the initial representation vector of the entity recognition result obtained by capturing the context information in the text. This initial representation vector predicts the named entity category of each word through sequence labeling.

[0066] A target type refers to a predefined type of entity. For example, an entity with a target type is an entity that is a proper noun. A proper noun represents a specific and unique person or thing (such as a person's name, a place name, etc.). The exemplary embodiments of this disclosure can set the target type according to the actual task requirements, and there are no special limitations on this. It should be understood that the exemplary embodiments of this disclosure are illustrated using the example of a proper noun as the target type.

[0067] For example, a pre-trained named entity recognition model can be used for entity recognition processing. For instance, BERT and CRF (conditional random field) models can be used to perform entity recognition processing on text content, obtaining entity recognition results and corresponding feature vectors. Specifically, the text content can be input into the BERT model to obtain a vector representation of the text content. Then, the vector representation of the text content is input into the CRF model for sequence labeling to obtain named entity categories. Based on the named entity categories and the vector representations of the text content, the feature vectors corresponding to each entity recognition result can be found. Here, the named entity categories are equivalent to the recognition results obtained from traditional named entity recognition. Compared to traditional named entity recognition, the exemplary embodiment of this disclosure adds an entity classification processing stage after the named entity recognition processing stage. It is necessary to first obtain the feature vectors corresponding to each entity recognition result (i.e., determine the feature vectors corresponding to each entity word based on the vector representation of the text content). Then, based on the entity recognition results and their feature vectors obtained in the named entity recognition processing stage, the entity recognition results are classified, thereby determining the target entity recognition result with the target type from all entity recognition results according to the classification results.

[0068] In the exemplary embodiments of this disclosure, by combining the entity recognition results and their corresponding feature vectors from the named entity recognition processing stage as the basic data for classification processing, no additional features need to be extracted, the reasoning time is low, and the processing efficiency is higher.

[0069] In step S230, the target entity recognition result is split and encoded according to the phonetic features, and the encoded result is matched with the candidate encoded result to determine the target candidate entity based on the matching result; wherein, the candidate encoded result is obtained by splitting and encoding the candidate entity according to the phonetic features, and the candidate entity is determined according to the current task scenario information.

[0070] In the exemplary embodiments of this disclosure, phonetic features refer to features describing the pronunciation of entities. For example, for Chinese phonetic features, these may include initials, finals, and tones. Each Chinese character in the target entity recognition result is segmented according to its phonetic features. The initial is the phoneme at the beginning of the Chinese character, the final is the part excluding the initial, and the tone is used to distinguish different syllables and meanings. A pre-trained entity segmentation network can be used to perform pinyin conversion on the Chinese characters, or existing third-party libraries (such as the Python py.pinyin library) can be used directly; no special limitations are imposed on this approach.

[0071] For example, Table 1 shows an example of splitting the target entity recognition result "Xiaoming" according to its phonetic features.

[0072] Table 1

[0073] The segmented features obtained by splitting according to phonetic features are encoded to obtain encoded features. The segmented features can be embedded to obtain encoded features. Embedding can map high-dimensional discrete data to a low-dimensional continuous space, making it easier for machine learning or deep learning models to process and understand, thus improving the computational efficiency and processing speed of the models. The exemplary embodiments of this disclosure do not specifically limit the method of feature encoding.

[0074] Similarly, for the phonetic features of English words, consonant phonemes and vowel phonemes can be included. Based on these phonetic features, the target entity recognition result of English can be segmented into at least one consonant phoneme unit and at least one vowel phoneme unit. The segmentation results are then encoded to obtain the encoding results corresponding to each phonetic unit. This can be done using a pre-trained entity segmentation network or by directly using existing third-party libraries (such as the Python nltk library), without any special limitations. For example, taking "cat" as an example, it is segmented into three phonetic units: the consonant phoneme " / k / ", the vowel phoneme " / ae / ", and the consonant phoneme " / t / ". Each of these three phonetic units is then encoded to obtain its corresponding encoding result. It should be understood that for other foreign languages, the segmentation into multiple phonetic units can also be performed based on phonetic features. The exemplary embodiments of this disclosure are described below using Chinese as an example of target entity recognition results.

[0075] Among them, an alternative entity refers to an entity with the same named entity type as the recognition result of the target entity, which is used to correct the recognition result of the target entity in the text content, so that the text content accurately reflects the true meaning of the audio content. For example, if the recognition result of the target entity is "ab", the alternative entities may include "little a", "little ab", "ab", "ah a", etc. The current task scenario information refers to the specific scenario information corresponding to the audio content, including the key content reflecting the information to be extracted in the summary. Taking the extraction of meeting minutes as an example, the current task scenario may include information such as participants and meeting topics. According to the current task scenario information, alternative entities for correcting the recognition result of the target entity can be extracted.

[0076] The method of splitting and encoding the alternative entity according to the phonetic features is the same as that of splitting and encoding the recognition result of the target entity, which will not be elaborated here.

[0077] Matching the encoding result of the target entity recognition result with the alternative encoding result is a process of obtaining a target alternative entity with phoneme features similar to those of the target entity recognition result from the alternative entities based on the encoding features, so as to use the target alternative entity to replace the target entity recognition result and improve the accuracy of the target text finally used to extract the summary information.

[0078] In step S240, the recognition result of the target entity in the text content is replaced with the target alternative entity, and the summary information corresponding to the text content is extracted according to the obtained target text.

[0079] In an exemplary embodiment of the present disclosure, the recognition result of the target entity in the text content can be replaced with the target alternative entity to obtain a target text that conforms to the original audio sentence, so as to extract the summary information according to the target text, making the summary information closer to the true expression of the original audio sentence.

[0080] This disclosure provides an information extraction method. On one hand, it converts collected audio content into text content, performs entity recognition processing on the text content to obtain entity recognition results and corresponding feature vectors, and then classifies entities based on the entity recognition results and feature vectors. Based on the entity classification results, it extracts entities of the target type, obtaining target entity recognition results. In the named entity recognition process, it can extract entities of the target type without outputting all entities, accurately capturing specific entities in the text content, providing an accurate entity basis for subsequent entity replacement, reducing inference time and improving information extraction efficiency. On the other hand, it splits and encodes the target entity recognition results according to phonetic features, and matches the encoded results with candidate encoded results to determine target candidate entities based on the matching results. The candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined based on the current task scenario. Then, the target candidate entities can be used to replace the target entity recognition results in the text content, and the corresponding summary information can be extracted from the obtained target text. By considering the phonetic features of entities, entities are split and encoded to improve the comprehensiveness and accuracy of the target entity recognition results, so as to accurately distinguish different entities. Furthermore, based on the current task scenario, candidate entities are determined, and without relying on historical records, the obtained target candidate entities are made to fit the current task scenario, thereby improving the accuracy of entity replacement and thus improving the precision of extracting summary information.

[0081] In one exemplary embodiment, an implementation method for obtaining target entity recognition results is provided. Entity recognition processing is performed on the text content to obtain entity recognition results and corresponding feature vectors. Entity classification is then performed based on the entity recognition results and feature vectors to extract entities of the target type according to the classification results. Obtaining target entity recognition results may include:

[0082] First, text features are extracted from the text content, and entity recognition is performed based on the text features to obtain entity recognition results. Then, based on the text features, the feature vector corresponding to the entity recognition results is determined. Finally, classification processing is performed based on the entity recognition results and the corresponding feature vectors to obtain the entity type corresponding to the entity recognition results, so as to extract the target entity recognition results with the target type based on the entity type.

[0083] Specifically, the entity type refers to a predefined type of entity, obtained by classifying the entity recognition results and corresponding feature vectors. The exemplary embodiment of this disclosure employs a multi-task learning framework, first extracting entities and then classifying them.

[0084] For example, Figure 3 shows a schematic diagram of an entity recognition model according to an exemplary embodiment of this disclosure. The entity recognition model includes a BERT model, a CRF model, and an entity classification network. The outputs of both the BERT model and the CRF model are input to the entity classification network so that the entity classification network can perform a classification task. The process of obtaining the target entity recognition result is described below with reference to Figure 3.

[0085] First, the BERT model is used to extract text features from the text content. These text features are then input into a CRF model to perform entity recognition based on the text features, yielding the entity recognition results. Next, the results are fed back to the BERT model to obtain the feature vector corresponding to each entity recognition result. Finally, both the entity recognition results and their corresponding feature vectors are used as input to an entity classification network for classification tasks.

[0086] The entity classification network may include at least one fully connected network layer to determine whether the entity type of the entity recognition result is the target entity type. That is, the entity classification network is a binary classification network, and the output results include belonging to the target entity type and not belonging to the target entity type.

[0087] It should be noted that the entity recognition model is pre-trained using a training dataset. The BERT model, CRF model, and entity classification network can be trained jointly, or the BERT+CRF model and the entity classification network can be trained separately. Finally, the training results are applied together. The exemplary embodiments of this disclosure do not impose special restrictions on the training process of the entity recognition model and can be flexibly selected according to the actual task requirements.

[0088] In some alternative embodiments, constraints can be added to the CRF model, which can be learned during CRF training to ensure the effectiveness of the training results.

[0089] For example, added constraints could include: the sentence must begin with "B-" or "O-", not "I-". Similarly, for "B-label1 I-label2 I-label3…", where categories 1 through 3 are the same entity category, "B-Person I-Person" is correct, while "B-Person I-Organization" is incorrect. Furthermore, "O I-label" is incorrect; the constraint indicates that the named entity should begin with "B-" instead of "I-".

[0090] In this specification, "B", "O", and "I" represent the ternary annotation method of the BIO annotation algorithm. "B" indicates the start of an entity, "O" indicates a non-entity, and "I" indicates a part of an entity, but not the start. It should be understood that quaternary or pentagonal annotations or other annotation methods of the BIO annotation algorithm can also be used. The exemplary embodiments disclosed herein are merely illustrative examples using the ternary annotation method of the BIO annotation algorithm.

[0091] After obtaining a dataset (such as the MSRA dataset), the training dataset can be labeled using the BIO annotation method to obtain the training dataset. When training the entity recognition model, all training data can be passed through a fully connected network, and the loss value can be calculated based on the obtained entity category prediction results and entity type labels, thereby adjusting the parameters of the fully connected network based on the loss value. Optionally, the loss corresponding to non-entity recognition results can be masked, that is, the weights of invalid positions can be set to a very small negative infinity to minimize their impact on model training. Here, passing all training data through a fully connected network means that for each entity sample, the entity recognition result (entity sample) and its corresponding feature vector (feature vector of the entity sample) are input into the fully connected network to adjust the network parameters.

[0092] After obtaining the trained entity recognition model, in practical applications, the entity recognition results and corresponding feature vectors can be input into a pre-trained entity classification network to obtain the entity type corresponding to the entity recognition result. Optionally, to improve the processing efficiency and accuracy of entity recognition, preset delimiters (such as ".", "!", "?") can be used to decompose the text content, and the decomposition results can be used as input to the entity recognition model.

[0093] The exemplary embodiments of this disclosure employ a multi-task framework entity recognition model, which sequentially performs entity extraction and entity classification tasks. Compared to current named entity recognition, which outputs all named entities, this model only outputs the target entity recognition results with the target type, without having to output all entities. It can accurately capture special entities in the text content, providing an accurate entity basis for subsequent entity replacement, thereby reducing reasoning time and improving the efficiency of information extraction.

[0094] In an exemplary embodiment, an implementation method for obtaining candidate entities is provided. As shown in FIG4, the implementation method for obtaining candidate entities may include:

[0095] Step S410: Obtain the current task scene information, extract entities from the current task scene information, and obtain scene entities.

[0096] The current task scenario information is related to the current task scenario. For example, in the meeting scenario mentioned above, the current task scenario information may include information such as the participants and the meeting topic. Scene entities are very likely to appear in the audio content of the meeting. By using scene entities to correct the target entity recognition results, it can be ensured that the content in the final minutes information is more in line with the actual scenario.

[0097] In this process, scene entities can be extracted from the current task scene information using a pre-trained entity extraction model. The exemplary embodiments disclosed herein do not impose specific restrictions on the method of extracting scene entities.

[0098] In some optional embodiments, the current task scenario information can be obtained in the following ways:

[0099] In response to user touch operations, the touch information corresponding to the touch operation is used as the current task scenario information; and / or, topic recognition is performed based on text content to determine the topic information of the text content, and the topic information is used as the current task scenario information.

[0100] In the first acquisition method, the user's touch operation can be an input operation, a selection operation, etc. For example, the user can input information such as attendees and meeting topics through a touch medium (such as a finger, stylus, etc.), or the user can select from a variety of provided function options (used to indicate relevant information about the meeting) to determine the current task scenario information.

[0101] In the second acquisition method, considering that users may not provide relevant information based on touch operations, topic recognition can be performed on the text content to determine topic information, which can then be used as the current task scenario information. The topic of the text content can be extracted using a pre-trained topic recognition model, and the training method for this model is not restricted. Of course, the second acquisition method can be used as a supplement to the first. If the current task scenario information can be obtained through the first method, then the first method should still be the preferred option.

[0102] Obtaining current scene information through user input and / or topic recognition can provide a source of alternative entities for subsequent correction of target entity recognition results, making the determined target alternative entities more suitable for the task scenario, thereby improving the accuracy of extracting minutes information.

[0103] Step S420: Perform entity generation processing based on scene entities to obtain reference entities with the same entity type as scene entities.

[0104] Entity generation processing refers to generating reference entities of the same type as the scene entities according to preset rules, based on scene entities. These reference entities reflect the actual situation of the current task scenario. The preset rules can be set according to actual task requirements without specific restrictions.

[0105] For example, if the scene entities include Zhang San, Li Si, and Zhao Wu, the reference entities generated according to the preset rules include: Xiao Zhang, Xiao Si, Xiao Wu, Xiao Li, Xiao Liu, Xiao Zhao, A San, A Si, A Wu, San San, Si Si, Wu Wu, etc.

[0106] Step S430: Determine the candidate entities based on the reference entity and the scene entity.

[0107] After obtaining the reference entities corresponding to each scene entity, the reference entities and scene entities are used together as candidate entities.

[0108] In addition, to avoid the candidate entities not conforming to normal language logic, the candidate entities can be validated in terms of language logic, such as by using a pre-trained logic validation model. There are no restrictions on this, so as to improve the accuracy of the candidate entities.

[0109] The exemplary embodiments of this disclosure determine candidate entities based on the current task scenario. Without relying on historical records, the obtained target candidate entities conform to the current task scenario, improving the accuracy of entity replacement and thus improving the precision of extracting minutes information.

[0110] In an exemplary embodiment, in a scenario where topic recognition is performed based on text content to determine topic information of the text content, and the topic information is used as current task scenario information, the method may further include:

[0111] If a first preset entity associated with the theme information exists in the preset entity library, then the first preset entity will also be identified as a scene entity.

[0112] Considering that the accuracy of scene entities is positively correlated with the accuracy of subsequent entity correction, increasing the comprehensiveness and accuracy of scene entities is beneficial to improving the accuracy of entity correction. The preset entity library is a pre-built entity library about different topics based on different task types. It is obtained by collecting entities from different task types and classifying them by topic, and is iteratively updated based on continuously collected entities to ensure that the preset entity library contains comprehensive entity results.

[0113] Specifically, after determining the topic information of the text content, a target topic associated with that topic information can be obtained from a preset entity library, and the first preset entity corresponding to the target topic is also determined as a scene entity. Since the scene entity, as a candidate entity, ultimately needs to be matched with the target entity recognition result to determine the target candidate entity, even if there are some potentially low-relevance entities in the scene entity, they will be eliminated during the matching process, thus avoiding the execution of erroneous entity replacement.

[0114] In one exemplary embodiment, a method for implementing supplementary scene entities is also provided. This implementation includes:

[0115] Perform speech recognition on the audio content, and analyze the topics based on the speech recognition results to obtain the conversation topics;

[0116] If a second preset entity related to the conversation topic exists in the preset entity library, then the second preset entity will also be identified as a scene entity.

[0117] In this context, considering that the audio content will involve topics related to the conversation as it progresses, speech recognition can be performed on the collected audio content. The conversation topic can then be obtained through topic analysis based on the speech recognition results. Speech recognition of the audio content can be performed using a pre-trained speech recognition model. The speech recognition model can be trained based on any neural network model, or it can be trained based on a pre-trained model; there are no specific limitations on this.

[0118] Topic analysis based on speech recognition results can extract keywords from the speech recognition results and determine the conversation topic based on the keywords. Alternatively, the topic analysis model can be used to directly process the speech recognition results to obtain the conversation topic. The topic analysis model can be trained based on any neural network model or a pre-trained model, and there are no specific limitations on this.

[0119] As mentioned above, the preset entity library can be a pre-built entity library for different task types, containing information on different topics. It is obtained by collecting entities from different task types and dividing them into topics, and is iteratively updated based on continuously collected entities to ensure the preset entity library contains comprehensive entity results. Similarly, entity libraries for different session topics can also be pre-built for different task types. The entity library based on session topics can be the same as or different from the session topics based on topic information, and can be pre-built based on relevant information from historical task types according to the actual task scenario.

[0120] An exemplary embodiment of this disclosure performs speech recognition on audio content, uses the conversation topic obtained from the speech recognition to obtain a second preset entity as a scene entity, increases the number of candidate entities that match the audio, thereby increasing the richness of scene entities and providing an entity basis for subsequent entity matching.

[0121] In an exemplary embodiment, when responding to a user's touch operation and using the touch information corresponding to the touch operation as the current task scene information, the scene entities can be enriched in the following ways:

[0122] Based on the scene entities extracted from the current task scene information, determine the target objects in the current task scene;

[0123] Obtain the historical entities corresponding to the target object, and simultaneously identify the historical entities as scene entities.

[0124] The target object can be a person or thing, such as attendees or the meeting topic. If the target object in the current task scenario is determined based on scene entities, and this target object has historical entities, and there is a possibility that historical entities will appear in this scenario, then the historical entities can also be determined as scene entities. For example, if the target object is "Zhang San", then the historical entity corresponding to "Zhang San" can be obtained and added as a scene entity.

[0125] It should be understood that historical entities of the target object can be pre-stored. To avoid situations where the target object is the same as a pre-stored historical object but is not actually the same object, tags such as gender, occupation, and domain tags can be added to the objects during pre-storage. This allows for further evaluation of whether the target object matches the tags of the historical objects, thus preventing the incorrect addition of historical entities as scene entities. It should also be understood that since scene entities, as candidate entities, ultimately need to be matched with the target entity recognition results to determine the target candidate entity, even if some potentially low-relevance entities exist within the scene entity pool, they will be removed during the matching process to avoid performing erroneous entity replacements.

[0126] By acquiring historical entities of the target object as scene entities, the number of alternative entities that match the target object is increased, thereby increasing the richness of scene entities.

[0127] In one exemplary embodiment, an implementation for updating a target candidate entity is also provided, including:

[0128] Display the candidate entities and update them in response to touch operations on the candidate entities.

[0129] The process of updating candidate entities can be performed by the user equipment in response to human entity update operations. In other words, updating candidate entities is based on human intervention; the user equipment can modify and update the provided candidate entities according to its own circumstances. Specifically, the user equipment can update the candidate entities in response to touch operations targeting them.

[0130] Figure 5 illustrates a schematic diagram of an interactive interface for displaying candidate entities, which allows for updating of candidate entities. It should be noted that all adjustable content can be configured within this interactive interface and viewed and modified using sliders. Furthermore, Figure 5 only exemplifies possible configurations of the interactive interface, and the exemplary embodiments of this disclosure do not limit this approach.

[0131] By showing users alternative entities and updating them based on touch input, the accuracy of subsequent entity replacement can be improved. It can also update and replace entities that have already been replaced, that is, use the updated alternative entities to replace the recognition results of the replaced target entities, thereby improving the overall accuracy of the final target text.

[0132] In one exemplary embodiment, an implementation method for obtaining target candidate entities is provided. As shown in Figure 6, the target entity recognition result is split and encoded according to phonetic features, and the encoded result is matched with candidate encoded results to determine the target candidate entity based on the matching result. This may include:

[0133] Step S610: According to the phonetic features, the pronunciation information of the target entity recognition result is divided into multiple pronunciation units.

[0134] As mentioned above, the pronunciation information of the target entity recognition results for different language types can be divided into multiple pronunciation units. For example, Chinese can be divided into initials, finals and tones according to pinyin, and English can be divided into consonant phonemes and vowel phonemes according to phonetic symbols. These will not be elaborated on here.

[0135] Step S620: Encode each of the obtained pronunciation units to obtain multiple encoding results; Step S630: For each encoding result, determine the sub-similarity between the encoding result and the corresponding candidate encoding result, and fuse the sub-similarity corresponding to each encoding result to obtain the fusion similarity corresponding to the candidate encoding result; Step S640: Determine the target candidate entity based on the fusion similarity corresponding to the candidate encoding result.

[0136] As an example, the recognition result of the Chinese target entity can be converted into pinyin, and then the pinyin is split into initial units, final units, and tone units, and the obtained split results are encoded respectively to obtain initial codes, final codes, and tone codes. For example, if the recognition result of the target entity is "小样", the converted pinyin is xiaoyang, and then xiaoyang is split into "x", "iao", "3", "y", "ang", "4", and each part is encoded respectively.

[0137] Among them, when encoding, first, based on the initial code, determine the initial similarity between the encoding result and the alternative encoding result. Then, according to the final code, determine the final similarity between the encoding result and the alternative encoding result. Next, based on the tone code, determine the tone similarity between the encoding result and the alternative encoding result. Finally, perform a fusion calculation through the initial similarity, final similarity, and tone similarity to obtain the fusion similarity, and match the encoding result with the alternative encoding result based on the fusion similarity to determine the target alternative entity according to the matching result.

[0138] Among them, calculating the similarity between the encoding result and the alternative encoding result based on each unit can calculate feature distances such as Manhattan distance and Euclidean distance between the corresponding units in the encoding result and the alternative encoding result, and there is no restriction on this.

[0139] Exemplarily, the process of calculating the fusion similarity can be represented by the following formula (1):

[0140] Among them, S(ci,ci’) is the fusion similarity, is the initial similarity, is the final similarity, is the tone similarity.

[0141] By converting the recognition result of the target entity into pinyin and encoding it, in addition to considering phonemes, the tone code is also used for similarity calculation, which can capture the characteristics of the recognition result of the target entity more comprehensively and provide an accurate entity basis for entity matching. For example, the two phonemes "繁" and "翻" are the same but have different tones. By encoding the initial unit, final unit, and tone unit respectively, it is possible to avoid entities that are difficult to distinguish in such situations, and based on the above formula, the fusion similarity between the encoding result of the recognition result of the target entity and each alternative entity can be calculated respectively. If the number of alternative encoding results is multiple, for any recognition result of the target entity, according to the fusion similarity corresponding to each different alternative encoding result, the target alternative entity can be determined from the alternative encoding results. For example, obtain the alternative encoding result corresponding to the highest fusion similarity, and determine the filing entity corresponding to this alternative encoding result as the target alternative entity.

[0142] An exemplary embodiment of this disclosure calculates the similarity between the encoding result of the target entity recognition result and the candidate encoding result of each candidate entity, and determines the target candidate entity based on the similarity. By first calculating the similarity between the pronunciation units separately, and then calculating the fusion similarity of the similarity corresponding to all pronunciation units, the feature correlation between the target entity recognition result and the candidate entity in each pronunciation part can be fully explored, improving the accuracy of feature matching, thereby improving the accuracy of the target candidate entity.

[0143] In one exemplary embodiment, to ensure that the final minutes information matches the needs of a real-world scenario, it may further include:

[0144] The target candidate entity includes one primary entity and at least one secondary entity; wherein, at least one secondary entity is displayed differently in the text content and / or in the minutes information than the primary entity.

[0145] Specifically, the fusion similarity of secondary entities is lower than that of primary entities. They can be sorted from high to low according to their fusion similarity, and the entities ranked at the top can be identified as primary and secondary entities, respectively.

[0146] Figure 7 illustrates a schematic diagram of displaying minutes information. As shown in Figure 7, the primary entities obtained from the extraction of minutes information can be displayed in one display mode, while secondary entities can be displayed in another. Optionally, secondary entities can be presented as thumbnails, so that when a user views the minutes information, the secondary entities can be opened by triggering the thumbnails.

[0147] By including a primary entity and at least one secondary entity in the target candidate entities, the possibility of error correction can be provided for the minutes information. That is, if the primary entity is considered inaccurate, the accuracy of the minutes information can be improved by looking at the candidate entities.

[0148] In one exemplary embodiment, a method for extracting summary information from target text is provided. As shown in Figure 8, extracting summary information corresponding to the text content from the obtained target text may include:

[0149] Step S810: Obtain the processor information of the user device and determine the target summary extraction model based on the processor information.

[0150] The processor information includes CPU (Central Processing Unit) information and / or GPU (Graphics Processing Unit) information. Of course, depending on the actual task requirements, the processor information may also include DSP (Digital Signal Processor) information. The exemplary embodiments of this disclosure are described with processor information including CPU information and GPU information.

[0151] Among them, the correspondence between processor information and minutes extraction model can be pre-configured, so as to dynamically select the appropriate minutes extraction model based on the processor information of the user device to meet the current needs and improve processing efficiency.

[0152] It should be noted that the summary extraction model of the exemplary embodiments of this disclosure can be a large model that has been fine-tuned and trained. The fine-tuning process is no different from that of traditional large models, and no specific restrictions are placed on the fine-tuning training process here.

[0153] Step S820: Segment the target text based on its length, and use the target summary extraction model to extract the summary information of each segment.

[0154] Considering that the long length of the target text increases the time required to extract the minutes, and that inputting the entire target text into the target minutes extraction model for inference would place excessive demands on GPU memory, the target text is segmented based on its length, and inference is performed by inputting the segments into the target minutes extraction model.

[0155] In some optional embodiments, the segment length can be selected based on GPU information (such as model). For example, a correspondence between GPU model and segment length can be pre-set to determine the segment length based on the correspondence. For instance, taking a segment length of 4096 (the number of characters in the text string) as an example, the overall text length of the target text is counted, the full text length (e.g., text_len) is divided by the segment length 4096, and the result is rounded up to obtain the number of segments, num. Then, the full text length is divided by the number of segments, num, to obtain the segmentation result. Based on this segmentation result, the target text is segmented and input into the target summary extraction model, thereby improving inference efficiency.

[0156] In some optional embodiments, segmenting the target text based on its length can also be done in the following way: first, the target text is identified by topic, and the target text is divided into multiple topic paragraphs according to the identification results; for each topic paragraph, the topic paragraph is segmented based on its length to obtain the segmented text corresponding to the topic paragraph.

[0157] Considering that segmenting the target text based on its length can lead to semantic truncation due to brute-force decomposition in certain scenarios, such as when segmenting according to the above method, two sentences within the same topic might be truncated into different segments, affecting the coherence and accuracy of the extracted minutes. However, if the target text is first segmented according to topic paragraphs, and then each topic paragraph is segmented using the method described above, it ensures that text content belonging to the same topic remains within a single topic paragraph, avoiding topic truncation caused by brute-force splitting and improving the accuracy of extracted minutes information.

[0158] The target text can be divided into topics using a pre-trained topic recognition model, which involves feature extraction, feature selection and topic classification. The topic recognition model is trained in a conventional way. The exemplary embodiments of this disclosure do not limit the type and training method of the topic recognition model.

[0159] Step S830: Combine the summary information of each segment to obtain the summary information corresponding to the text content.

[0160] After obtaining the minutes information for each segment, the minutes information for each segment is spliced ​​together to obtain the minutes information corresponding to the text content.

[0161] In some optional embodiments, a prompt can be set for the target minutes extraction model, requiring it to output segmented minutes information in the format of "content / title". By using the title-after method, the influence of the title on the content can be avoided, improving the accuracy of content extraction. When outputting segmented minutes information, to facilitate user reading habits, the title can be output in the first place. Finally, the titles and contents of each segmented minutes information are concatenated to obtain the final minutes information.

[0162] In the context of a large model, a prompt is an input method that guides the model to perform specific outputs or tasks through specific instructions or questions. Its main function is to provide context and parameter information to the large model, thereby guiding it to produce the expected output, such as answering questions, generating text, or translating languages. For example, a prompt could be set as: "You are a meeting secretary. Below is a transcript of an original meeting conversation. Please extract the meeting minutes. The output format of the meeting minutes should be 'Title\nContent' (the title should comprehensively summarize the content, and the events mentioned in the title must be reflected in the content). The meeting content is as follows:\n". Of course, this is only an example of a prompt, and the exemplary embodiments disclosed herein do not impose any special limitations on it.

[0163] In one exemplary embodiment, an implementation method for determining a target time summary extraction model is provided. Determining the target time summary extraction model based on processor information may include:

[0164] If the processor information indicates that the user equipment is not configured with a graphics processing module, then the first quantization model is determined as the target time summary extraction model; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information meets the preset model conditions, then the second quantization model is determined as the target time summary extraction model; wherein, the second quantization model and the first quantization model are obtained by processing the initial time summary extraction model with different quantization schemes; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information does not meet the preset model conditions, then the initial time summary extraction model is determined as the target time summary extraction model.

[0165] The initial summary extraction model is a pre-trained large model used to extract summary information, and there are no restrictions on the specific type of the initial summary extraction model.

[0166] Specifically, when the user device is configured with only a CPU module and not a GPU module, the initial summary extraction model can be quantized to obtain the first quantized model. Model quantization refers to the process of converting parameters or activation values ​​in the model from high precision (e.g., 32-bit floating-point numbers) to low precision (e.g., 8-bit integers), aiming to reduce the model size and computational complexity while maintaining performance as much as possible. For example, using the GGUF-INT4 quantization method to quantize the initial summary extraction model to obtain the first quantized model involves compressing the model weights and activation values ​​to 4-bit integers. By minimizing the mean squared error of the weights, the weights and activation values ​​are quantized to 4-bit integers, thereby significantly reducing memory usage and computational requirements while maintaining model performance.

[0167] The model information of the graphics processing module may include video memory and single / multi-card configuration. If the user device is configured with a CPU module and the GPU module is a single card with less than a preset video memory value (e.g., 24GB), the initial summary extraction model can be quantized to obtain a second quantized model. For example, the GGUF-INT8 quantization method can be used to quantize the initial summary extraction model to obtain the second quantized model. This involves converting the parameters in the model (such as weights and biases) from 32-bit floating-point numbers to 8-bit integers. This may include determining the scaling factor and zero point, and then mapping the floating-point numbers to the integer range, thereby significantly reducing the model's storage requirements and potentially accelerating the inference process while reducing power consumption.

[0168] Correspondingly, when the user device is configured with a CPU module and the GPU module is configured with multiple cards, the initial record extraction model can be directly used as the target record extraction model without model quantization. Inference can be performed through multi-card inference, which can also ensure the efficiency of record extraction.

[0169] It should be noted that the different quantization schemes and preset model conditions of the graphics processing module mentioned above are merely exemplary, and the exemplary embodiments of this disclosure can be flexibly set according to actual needs.

[0170] By selecting a target summary extraction model that is adapted to the processor information of the user device, the target summary extraction model that adapts to the processing needs of the user device can be dynamically selected, ensuring the efficiency of summary information extraction while maintaining device performance.

[0171] In one exemplary embodiment, a method for controlling the minutes extraction process is also provided, namely, using a target minutes extraction model to perform minutes extraction processing on segmented text, and obtaining segmented minutes information corresponding to each segmented text may further include:

[0172] The target summary extraction model outputs segmented summary information of a preset length each time. It then uses the adjacent segmented summary information that was output earlier to perform content verification on the segmented summary information of the preset length, and controls the extraction process of the target summary extraction model based on the verification results.

[0173] The preset length can be set according to actual needs. Taking a preset length of 100 characters as an example, the target minute extraction model outputs 100 characters and matches them with the 100 characters output previously. If the current 100 characters are duplicates of the previous 100 characters, or if the number of duplicates exceeds the preset number, the current 100 characters are deleted to control the target minute extraction model from repeatedly outputting the same minute information.

[0174] In other words, since the output of the current target summary extraction model is uncontrollable, there is a possibility that the output node cannot be obtained, which means that some content will be repeatedly output. The exemplary embodiment of this disclosure can avoid the problem of repeated output caused by the inability to obtain the end symbol in a large model by executing the model to extract summary information and verify the summary information in parallel, thereby improving the accuracy of the output summary information and avoiding the invalid work of the model.

[0175] This disclosure provides an information extraction method. On one hand, it converts collected audio content into text content, performs entity recognition processing on the text content to obtain entity recognition results and corresponding feature vectors, and then classifies entities based on the entity recognition results and feature vectors. Based on the classification results, it extracts entities of the target type, obtaining target entity recognition results. In the named entity recognition process, it can extract entities of the target type without outputting all entities, accurately capturing specific entities in the text content, providing an accurate entity basis for subsequent entity replacement, reducing reasoning time and improving information extraction efficiency. On the other hand, it splits and encodes the target entity recognition results according to phonetic features, and matches the encoded results with candidate encoded results to determine target candidate entities based on the matching results. The candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined based on the current task scenario. Then, the target candidate entities can be used to replace the target entity recognition results in the text content, and the corresponding summary information can be extracted from the obtained target text. By considering the phonetic features of entities, entities are segmented and encoded to improve the comprehensiveness and accuracy of target entity recognition results. This allows for accurate differentiation of different entities. Furthermore, by determining candidate entities based on the current task scenario without relying on historical records, the obtained candidate entities are made consistent with the current task scenario, improving the accuracy of entity replacement and thus enhancing the precision of extracting summary information. In addition, the segmented reasoning approach reduces memory usage, lowers hardware requirements, and is unaffected by the duration of the current task scenario.

[0176] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0177] The following are apparatus embodiments of this disclosure, which can be used to execute the method embodiments of this disclosure. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the method embodiments of this disclosure.

[0178] This disclosure also provides an information extraction device in an exemplary embodiment. Specifically, referring to FIG9, the information extraction device may include a speech-to-text module 910, an entity extraction module 920, an entity matching module 930, and an information extraction module 940. Wherein:

[0179] The speech-to-text module 910 converts the acquired audio content into text content; the entity extraction module 920 performs entity recognition processing on the text content to obtain the entity recognition result and the corresponding feature vector, and performs entity classification based on the entity recognition result and feature vector to extract the target type entity according to the classification result, thus obtaining the target entity recognition result; the entity matching module 930 splits and encodes the target entity recognition result according to the phonetic features, and matches the encoded result with the candidate encoded result to determine the target candidate entity according to the matching result; wherein, the candidate encoded result is obtained by splitting and encoding the candidate entity according to the phonetic features, and the candidate entity is determined according to the current task scenario information; the information extraction module 940 replaces the target entity recognition result in the text content with the target candidate entity, and extracts the summary information corresponding to the text content according to the obtained target text.

[0180] In one exemplary embodiment of this disclosure, the entity extraction module 920 is configured to perform: extracting text features from text content and performing entity recognition based on the text features to obtain entity recognition results; determining the feature vector corresponding to the entity recognition results based on the text features; performing classification processing based on the entity recognition results and the corresponding feature vectors to obtain the entity type corresponding to the entity recognition results, so as to extract target entity recognition results with target types based on entity types.

[0181] In one exemplary embodiment of this disclosure, the entity extraction module 920 is configured to perform: inputting the entity recognition result and the corresponding feature vector into a pre-trained entity classification network to obtain the entity type corresponding to the entity recognition result; wherein, the pre-trained entity classification network is obtained by model training based on entity samples, feature vectors of entity samples and entity type labels.

[0182] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to perform: obtaining current task scene information, extracting entities from the current task scene information to obtain scene entities; performing entity generation processing based on the scene entities to obtain reference entities of the same entity type as the scene entities; and determining candidate entities based on the reference entities and the scene entities.

[0183] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to perform: responding to a user's touch operation and using the touch information corresponding to the touch operation as the current task scene information; and / or, performing topic recognition based on text content, determining the topic information of the text content, and using the topic information as the current task scene information.

[0184] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to perform: if topic recognition is performed based on text content, the topic information of the text content is determined and the topic information is used as the current task scene information; if a first preset entity associated with the topic information exists in the preset entity library, the first preset entity is simultaneously determined as a scene entity.

[0185] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to perform: speech recognition on the audio content, topic analysis based on the speech recognition results to obtain the conversation topic; if a second preset entity associated with the conversation topic exists in the preset entity library, the second preset entity is simultaneously determined as a scene entity.

[0186] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to perform: if a user touch operation is responded to, the touch information corresponding to the touch operation is used as the current task scene information, and the scene entities extracted based on the current task scene information are used to determine the target object in the current task scene; the historical entity corresponding to the target object is obtained, and the historical entity is simultaneously determined as the scene entity.

[0187] In one exemplary embodiment of this disclosure, the entity extraction module 920 is further configured to: display candidate entities; and update candidate entities in response to touch operations on candidate entities.

[0188] In one exemplary embodiment of this disclosure, the entity matching module 930 is configured to perform the following: splitting the pronunciation information of the target entity recognition result into multiple pronunciation units according to the phonetic features; encoding each pronunciation unit to obtain multiple encoding results; determining the sub-similarity between each encoding result and the corresponding candidate encoding result, and performing a fusion calculation on the sub-similarity corresponding to each encoding result to obtain the fusion similarity corresponding to the candidate encoding result; and determining the target candidate entity based on the fusion similarity corresponding to the candidate encoding result.

[0189] In one exemplary embodiment of this disclosure, the entity matching module 930 is configured to perform: if the target entity recognition result is Chinese, convert the target entity recognition result into Pinyin and split the Pinyin into initial consonant units, final vowel units and tone units; and / or, if the target entity recognition result is English, split the target entity recognition result into at least one consonant phoneme unit and at least one vowel phoneme unit according to phonetic symbols.

[0190] In one exemplary embodiment of this disclosure, the entity matching module 930 is further configured to perform: the target candidate entity includes a primary entity and at least one secondary entity; wherein the at least one secondary entity is displayed in a manner different from that of the primary entity in the text content and / or in the minutes information.

[0191] In one exemplary embodiment of this disclosure, the information extraction module 940 is configured to perform: acquiring processor information of the user device and determining a target summary extraction model based on the processor information; segmenting the target text based on the text length of the target text and performing summary extraction processing on the segmented text using the target summary extraction model to obtain segmented summary information corresponding to each segmented text; and concatenating the segmented summary information to obtain summary information corresponding to the text content.

[0192] In one exemplary embodiment of this disclosure, the information extraction module 940 is configured to perform: topic recognition on the target text, dividing the target text into multiple topic paragraphs based on the recognition results; and for each topic paragraph, segmenting the topic paragraph based on the text length of the topic paragraph to obtain the segmented text corresponding to the topic paragraph.

[0193] In one exemplary embodiment of this disclosure, the information extraction module 940 is configured to perform the following: if the processor information indicates that the user equipment is not configured with a graphics processing module, then determine the first quantization model as the target summary extraction model; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information meets a preset model condition, then determine the second quantization model as the target summary extraction model; wherein the second quantization model and the first quantization model are obtained by processing the initial summary extraction model with different quantization schemes; or, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information does not meet the preset model condition, then determine the initial summary extraction model as the target summary extraction model.

[0194] In one exemplary embodiment of this disclosure, the information extraction module 940 is configured to perform the following: the target time summary extraction model outputs segmented time summary information of a preset length each time, performs content verification on the segmented time summary information of the preset length using adjacent previously output segmented time summary information, and controls the extraction process of the target time summary extraction model based on the verification results.

[0195] The specific details of each module in the above information extraction device have been described in detail in the corresponding information extraction methods, so they will not be repeated here.

[0196] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0197] Exemplary embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the information extraction method described above.

[0198] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.

[0199] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.

[0200] Computer program code can be written in one or more programming languages. Examples of programming languages ​​include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).

[0201] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic radiation, and infrared rays. Electronic devices can convert signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, the processor of the electronic device to execute) the method steps of various exemplary embodiments of this disclosure, such as the information extraction method described above, which includes the following steps: converting the acquired audio content into text content; performing entity recognition on the text content to obtain entity recognition results and feature vectors corresponding to the entity recognition results, and classifying entities based on the entity recognition results and feature vectors to extract entities of the target type according to the classification results, thereby obtaining target entity recognition results; splitting and encoding the target entity recognition results according to phonetic features, matching the encoded results with candidate encoded results, and determining target candidate entities according to the matching results; wherein the candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined according to the current task scenario; replacing the target entity recognition results in the text content with the target candidate entities, and extracting summary information based on the obtained target text.

[0202] Furthermore, in exemplary embodiments of this disclosure, an electronic device capable of implementing the above-described methods is also provided. Those skilled in the art will understand that various aspects of this disclosure can be implemented as systems, methods, or program products. Therefore, various aspects of this disclosure can be specifically implemented as entirely hardware embodiments, entirely software embodiments (including firmware, microcode, etc.), or embodiments combining hardware and software aspects, collectively referred to herein as "circuit," "module," or "system."

[0203] An electronic device 1000 according to such an embodiment of the present disclosure will now be described with reference to FIG10. The electronic device 1000 shown in FIG10 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present disclosure.

[0204] As shown in Figure 10, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one processing unit 1010, at least one storage unit 1020, a bus 1030 connecting different system components (including storage unit 1020 and processing unit 1010), and a display unit 1040.

[0205] The storage unit stores program code that can be executed by the processing unit 1010, causing the processing unit 1010 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1010 may perform the following steps: converting the acquired audio content into text content; performing entity recognition on the text content to obtain entity recognition results and corresponding feature vectors, and classifying entities based on the entity recognition results and feature vectors to extract entities of the target type according to the classification results, thereby obtaining target entity recognition results; splitting and encoding the target entity recognition results according to phonetic features, matching the encoded results with candidate encoded results, and determining target candidate entities according to the matching results; wherein the candidate encoded results are obtained by splitting and encoding candidate entities according to phonetic features, and the candidate entities are determined according to the current task scenario; replacing the target entity recognition results in the text content with the target candidate entities, and extracting summary information based on the obtained target text.

[0206] Storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1021 and / or a cache memory unit 1022, and may further include a read-only memory unit (ROM) 1023.

[0207] Storage unit 1020 may also include a program / utility 1024 having a set (at least one) program module 1025, such program module 1025 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0208] Bus 1030 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0209] Electronic device 1000 can also communicate with one or more external devices 1100 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1000, and / or any device that enables electronic device 1000 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1050. Furthermore, electronic device 1000 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1060. As shown, network adapter 1060 communicates with other modules of electronic device 1000 via bus 1030. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0210] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0211] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0212] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

Claims

1. An information extraction method, characterized in that, include: Convert the collected audio content into text content; The text content is subjected to entity recognition processing to obtain entity recognition results and feature vectors corresponding to the entity recognition results. Entity classification is performed based on the entity recognition results and the feature vectors to extract entities of the target type according to the classification results, thereby obtaining target entity recognition results. The target entity recognition result is split and encoded according to phonetic features, and the encoded result is matched with the candidate encoded result to determine the target candidate entity based on the matching result; wherein, the candidate encoded result is obtained by splitting and encoding the candidate entity according to phonetic features, and the candidate entity is determined according to the current task scenario information; The target entity identification results in the text content are replaced using the target candidate entities, and the summary information corresponding to the text content is extracted based on the obtained target text.

2. The method according to claim 1, characterized in that, The process of performing entity recognition processing on the text content to obtain entity recognition results and corresponding feature vectors, and then performing entity classification based on the entity recognition results and feature vectors to extract entities of the target type according to the classification results, thereby obtaining target entity recognition results, includes: Text features are extracted from the text content, and entity recognition is performed based on the text features to obtain the entity recognition result; Based on the text features, determine the feature vector corresponding to the entity recognition result; Based on the entity recognition result and the corresponding feature vector, classification processing is performed to obtain the entity type corresponding to the entity recognition result, so as to extract the target entity recognition result with the target type based on the entity type.

3. The method according to claim 2, characterized in that, The classification process based on the entity recognition result and the corresponding feature vector to obtain the entity type corresponding to the entity recognition result includes: The entity recognition result and the corresponding feature vector are input into a pre-trained entity classification network to obtain the entity type corresponding to the entity recognition result; The pre-trained entity classification network is obtained by training a model based on entity samples, the feature vectors of entity samples, and entity type labels.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the current task scene information, and extract entities from the current task scene information to obtain scene entities; Based on the scene entities, perform entity generation processing to obtain a reference entity with the same entity type as the scene entities; The candidate entity is determined based on the reference entity and the scene entity.

5. The method according to claim 4, characterized in that, The process of obtaining the current task scenario information includes: In response to the user's touch operation, the touch information corresponding to the touch operation is used as the current task scenario information; And / or, perform topic recognition based on the text content to determine the topic information of the text content, and use the topic information as the current task scenario information.

6. The method according to claim 5, characterized in that, If topic recognition is performed based on the text content to determine the topic information of the text content, and the topic information is used as the current task scenario information, the method further includes: If a first preset entity associated with the topic information exists in the preset entity library, then the first preset entity is simultaneously identified as the scene entity.

7. The method according to claim 5, characterized in that, The method further includes: The audio content is subjected to speech recognition, and the topic is analyzed based on the speech recognition results to obtain the conversation topic; If a second preset entity associated with the conversation topic exists in the preset entity library, then the second preset entity is simultaneously identified as the scene entity.

8. The method according to claim 5, characterized in that, If a user touch operation is responded to, and the touch information corresponding to the touch operation is used as the current task scenario information, the method further includes: Based on the scene entities extracted from the current task scene information, the target object in the current task scene is determined. Obtain the historical entity corresponding to the target object, and simultaneously identify the historical entity as the scene entity.

9. The method according to claim 4, characterized in that, The method further includes: The candidate entities are displayed; In response to a touch operation on a candidate entity, the candidate entity is updated.

10. The method according to claim 1, characterized in that, The step of splitting and encoding the target entity recognition result according to phonetic features, and matching the encoded result with candidate encoded results to determine the target candidate entity based on the matching result includes: Based on the phonetic features, the pronunciation information of the target entity recognition result is divided into multiple pronunciation units; Each of the obtained pronunciation units is encoded to obtain multiple encoding results; For each of the encoding results, the sub-similarity between the encoding result and the corresponding candidate encoding result is determined, and the sub-similarity corresponding to each encoding result is fused to obtain the fused similarity corresponding to the candidate encoding result; The target candidate entity is determined based on the fusion similarity corresponding to the candidate encoding results.

11. The method according to claim 10, characterized in that, The step of dividing the pronunciation information of the target entity recognition result into multiple pronunciation units according to the phonetic features includes: If the target entity recognition result is Chinese, then the target entity recognition result is converted into Pinyin, and the Pinyin is split into initial consonant units, final vowel units and tone units; And / or, if the target entity recognition result is in English, then the target entity recognition result is split into at least one consonant phoneme unit and at least one vowel phoneme unit according to phonetic symbols.

12. The method according to claim 10, characterized in that, The method further includes: The target candidate entity includes one primary entity and at least one secondary entity; The at least one secondary entity is displayed differently in the text content and / or in the minutes information than the primary entity.

13. The method according to claim 1, characterized in that, The step of extracting the summary information corresponding to the text content based on the obtained target text includes: Obtain the processor information of the user device and determine the target summary extraction model based on the processor information; The target text is segmented based on its length, and the target summary extraction model is used to extract the summary information of each segment. By splicing together the segmented summary information, the summary information corresponding to the text content is obtained.

14. The method according to claim 13, characterized in that, The segmentation of the target text based on its length includes: The target text is subjected to topic identification, and the target text is divided into multiple topic paragraphs based on the identification results; For each topic paragraph, the topic paragraph is segmented based on its text length to obtain the corresponding segmented text.

15. The method according to claim 13, characterized in that, The step of determining the target minutes extraction model based on processor information includes: If the processor information indicates that the user equipment is not configured with a graphics processing module, then the first quantization model is determined to be the target summary extraction model; Alternatively, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information meets the preset model conditions, then the second quantization model is determined to be the target summary extraction model; wherein, the second quantization model and the first quantization model are obtained by processing the initial summary extraction model with different quantization schemes; Alternatively, if the processor information indicates that the user equipment is configured with a graphics processing module and the model information does not meet the preset model conditions, then the initial summary extraction model is determined to be the target summary extraction model.

16. The method according to claim 13, characterized in that, The step of using the target summary extraction model to extract summary information from segmented texts to obtain segmented summary information corresponding to each segmented text includes: The target time summary extraction model outputs segmented time summary information of a preset length each time. It uses adjacent segmented time summary information that was output earlier to perform content verification on the segmented time summary information of the preset length, and controls the extraction process of the target time summary extraction model according to the verification results.

17. An information extraction device, characterized in that, The device includes: The speech-to-text module is used to convert the captured audio content into text content; The entity extraction module is used to perform entity recognition processing on the text content, obtain entity recognition results and feature vectors corresponding to the entity recognition results, and perform entity classification based on the entity recognition results and the feature vectors, so as to extract entities of the target type according to the classification results and obtain target entity recognition results; The entity matching module is used to split and encode the target entity recognition result according to phonetic features, and match the encoded result with the candidate encoded result to determine the target candidate entity based on the matching result; wherein, the candidate encoded result is obtained by splitting and encoding the candidate entity according to phonetic features, and the candidate entity is determined according to the current task scenario information; The information extraction module is used to replace the target entity recognition result in the text content with the target candidate entity, and extract the summary information corresponding to the text content based on the obtained target text.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 16.

19. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1-16 by executing the executable instructions.