Conference summary processing method and device, equipment and storage medium

By using multi-source matching data and topic identification technology, the problems of identity alignment and content reconstruction in meeting minutes generation were solved, improving the readability and review efficiency of meeting minutes, and achieving accurate reconstruction of meeting content and correction of contradictory information.

CN122019767APending Publication Date: 2026-05-12SHENZHEN TCL NEW-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TCL NEW-TECH CO LTD
Filing Date
2026-01-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing meeting minutes generation technologies suffer from inaccurate identity alignment, insufficient topication, inadequate implicit information mining, lack of self-correction and auditing mechanisms, and insufficient data integration, resulting in poor readability of meeting minutes and difficulty in quickly capturing core conclusions.

Method used

Multi-source matching data is introduced to match speaker identities, identify topics and generate summary information, update meeting minutes through topic identification and conflict verification, and reconstruct meeting content by combining intent identification and semantic analysis.

Benefits of technology

It improved the accuracy of speaker identification alignment, increased the readability of meeting minutes, shortened review time, and enabled retrospective correction and systematic auditing of conflicting information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019767A_ABST
    Figure CN122019767A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a conference summary processing method and device, equipment and a storage medium, which are used for performing content reconstruction on a conference summary so as to improve the readability of the conference summary. The method comprises the steps of obtaining a to-be-processed conference text and multi-source matching data, wherein the multi-source matching data comprises conference participant metadata, current conference information and historical conference data; spokesman matching is carried out on the to-be-processed conference text based on the multi-source matching data to obtain a first conference summary, and the first conference summary comprises text data of the to-be-processed conference text and spokesman information associated with the text data; performing topic identification on the first conference summary to obtain a plurality of topics corresponding to the first conference summary and summary information corresponding to each topic; and generating a target conference summary based on the plurality of topics and the summary information corresponding to each topics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, specifically to a method, apparatus, device, and storage medium for processing meeting minutes. Background Technology

[0002] Current technologies for automatically generating meeting minutes focus primarily on two areas: abstract extraction and keyword extraction. While they can extract basic information, they have several significant shortcomings in practical applications.

[0003] For example, inaccurate identity alignment makes it difficult to accurately match anonymous speakers with their semantic roles; insufficient topication results in generated content that simply lists information in chronological order without achieving topic-driven content reconstruction, making it difficult for users to quickly grasp the core conclusions; implied information mining is inadequate, resulting in poor identification of pain points, needs, and action items in meetings, and often omitting the responsible persons and deadlines for action items; there is a lack of self-correction and auditing mechanisms, lacking the ability to retrospectively correct contradictory information in meetings, and lacking systematic auditing processes and related evidence chains; and insufficient data integration fails to fully integrate multi-source enterprise data such as project data and system data, making it difficult to further improve the accuracy and intelligence level of minutes generation.

[0004] Therefore, there is an urgent need for a meeting minutes processing method that can improve the readability of meeting minutes. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for processing meeting minutes, used to reconstruct the content of meeting minutes to improve their readability.

[0006] The technical solution adopted by this invention to solve the problem is as follows: Firstly, this application provides a method for processing meeting minutes, including: Acquire the meeting text to be processed and multi-source matching data, which includes participant metadata, current meeting information, and historical meeting data; Based on the multi-source matching data, the speaker matching is performed on the meeting text to be processed to obtain the first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data; The agenda items of the first meeting minutes are identified to obtain multiple agenda items corresponding to the first meeting minutes and summary information for each agenda item. The target meeting minutes are generated based on the multiple topics and the summary information corresponding to each topic.

[0007] In some embodiments of this application, after generating the target meeting minutes based on the plurality of topics and the summary information corresponding to each topic, the method further includes: Perform conflict verification on the target meeting minutes to obtain the verification results; The meeting minutes for that target will be updated based on the verification result.

[0008] In some embodiments of this application, the conflict check on the target meeting minutes to obtain the check result includes: Detect the types of conflict in the meeting minutes of this target, which include conflicts of identity consistency, conflicts of issue positions, and conflicts of content integrity. Generate hypothesis branches for this type of conflict; Calculate the confidence level of this hypothesis branch; The verification result is determined based on this confidence level.

[0009] In some embodiments of this application, before identifying the topics of the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information corresponding to each topic, after performing speaker matching on the meeting text to be processed based on the multi-source matching data to obtain the first meeting minutes, the method further includes: Action item information is extracted and supplemented from the first meeting minutes to obtain structured data. This action item information is used to characterize the task content, task execution object, and task time contained in the meeting text to be processed. This structured data is used for topic identification.

[0010] In some implementations of this application, action item information is extracted and supplemented from the first meeting minutes to obtain structured data, including: Intent identification is performed on the first meeting minutes to obtain the intent classification of each text data in the first meeting minutes; Based on this intent, the text data to be processed is determined; Perform syntactic analysis on the text data to be processed to extract and supplement the task content, the task execution object, and the task time; This structured data is generated based on the task content, the task execution object, and the task time.

[0011] In some embodiments of this application, the first meeting minutes are subject to topic identification to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic, including: Semantic analysis was performed on the first meeting minutes to obtain the semantic vector of the text data in the first meeting minutes; The semantic vector is then subjected to topic clustering to obtain multiple initial topics; Topic switching detection is performed on the semantic vector to obtain the start time of the multiple initial topics; The semantic similarity of the semantic vector is calculated and aggregated to obtain the complete content of the multiple initial topics; Based on the start time and the complete content, construct multiple topics corresponding to the first meeting minutes; Extract the key nodes of the text data for each of the multiple topics; Based on this key node, the text data of each topic is compressed and recombined using summary generation technology to obtain summary information for each topic.

[0012] In some embodiments of this application, speaker matching is performed on the meeting text to be processed based on the multi-source matching data to obtain a first meeting summary, including: Based on at least one speaker matching operation, a set of candidate speakers corresponding to each segment of the meeting text to be processed is obtained; Based on the set of candidate speakers corresponding to each text data segment, the speaker corresponding to each text data segment is determined to obtain the first meeting minutes; The at least one speaker matching operation includes: A role vector library is constructed based on the participant's metadata and the historical meeting data; Obtain the semantic vector corresponding to each segment of the meeting text to be processed; Based on the semantic vector, similarity matching is performed from the role vector library to obtain the first candidate speaker corresponding to each segment of the meeting text to be processed; and / or; Named entity recognition is performed on the meeting text to be processed to obtain the title information in the meeting text; Based on the coreference resolution method and the title information, the second candidate speaker corresponding to each segment of text data in the meeting text to be processed is determined; and / or; Dialogue interaction recognition is performed on the text data in the meeting text to be processed in order to obtain the third candidate speaker corresponding to each segment of text data in the meeting text to be processed; and / or; Domain terminology is identified in the text data of the meeting text to be processed to obtain a set of domain terms; The fourth candidate speaker is obtained by matching each domain term in the domain terminology set with the domain terminology database, which is used to represent the relationship between domain terms and the roles of participants. and / or; Based on the temporal stickiness characteristics and the global optimal allocation method, speaker matching is performed on each segment of the conference text to be processed to obtain the fifth candidate speaker; The first candidate speaker, the second candidate speaker, the third candidate speaker, the fourth candidate speaker, and the fifth candidate speaker are included in the candidate speaker set.

[0013] Secondly, this application provides a meeting minutes processing device, comprising: The acquisition module is used to acquire the meeting text to be processed and multi-source matching data, which includes participant metadata, current meeting information, and historical meeting data. The processing module is used to perform speaker matching on the meeting text to be processed based on the multi-source matching data to obtain a first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data; to perform topic identification on the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information corresponding to each topic; and to generate a target meeting minutes based on the multiple topics and the summary information corresponding to each topic.

[0014] Thirdly, this application also provides a computer device, which includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the meeting minutes processing method of any of the first aspects.

[0015] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the meeting minutes processing method of any of the first aspects.

[0016] The beneficial effects of this invention are as follows: It introduces multi-source matching data for speaker identity matching, thereby improving the accuracy of speaker identity alignment. Simultaneously, it identifies topics in the meeting text to obtain multiple topics corresponding to the meeting text and their corresponding summary information. Then, based on these multiple topics and summary information, the meeting content is reconstructed to obtain the meeting minutes. This increases the readability of the meeting records while shortening the review time. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1This is a schematic diagram of the system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic flowchart of an embodiment of the meeting minutes processing method provided by the present invention; Figure 3 This is a schematic diagram of a speaker identity alignment process provided in an embodiment of the present invention; Figure 4 This is a flowchart illustrating the reflection-branching-backtracking process provided in an embodiment of the present invention; Figure 5 This is a flowchart illustrating a meeting minutes processing method provided in an embodiment of the present invention; Figure 6 This is a schematic block diagram of a specific embodiment of the meeting minutes processing device provided in this invention. Figure 7 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features.

[0021] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0022] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.

[0023] Current technologies for automatically generating meeting minutes primarily focus on two main areas: summary extraction and keyword extraction. While they can extract basic information, they suffer from several significant shortcomings in practical applications. For example, identity alignment is inaccurate, failing to precisely match anonymous speakers with their semantic roles; topic-based approach is insufficient, with generated content often simply listed in chronological order without achieving topic-driven content reconstruction, making it difficult for users to quickly grasp core conclusions; implicit information mining is inadequate, poorly identifying pain points, needs, and action items in meetings, often omitting responsible persons and deadlines for action items; self-correction and auditing mechanisms are lacking, failing to address contradictory information in meetings and lacking systematic auditing processes and supporting evidence chains; and data integration is insufficient, failing to fully integrate multi-source enterprise data such as project data and system data, hindering further improvements in the accuracy and intelligence of minute generation. Therefore, a meeting minutes processing method that can improve the readability of meeting minutes is urgently needed.

[0024] To address this technical problem, this application provides the following technical solution: Acquiring the meeting text to be processed and multi-source matching data, including participant metadata, current meeting information, and historical meeting data; performing speaker matching on the meeting text to be processed based on the multi-source matching data to obtain a first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data; identifying topics on the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic; generating a target meeting minutes based on the multiple topics and the summary information for each topic. This approach introduces multi-source matching data for speaker identity matching, thereby improving the accuracy of speaker identity alignment. Simultaneously, identifying topics on the meeting text to be processed to obtain multiple topics corresponding to the meeting text and the summary information for each topic, and then reconstructing the meeting content based on the multiple topics and summary information, yields the meeting minutes. This increases the readability of the meeting records while shortening the review time.

[0025] For ease of understanding, the relevant technologies involved in this application are explained below.

[0026] Automatic Speech Recognition (ASR) is the process of automatically converting speech signals into corresponding text using computer technology. It can also be called Speech-to-Text (STT) or Speech Recognition. Its applications include the following: Consumer electronics: including voice assistants on mobile phones (such as Siri and Xiao Ai), smart speakers, voice input methods, and in-vehicle voice control systems.

[0027] Enterprise office applications: automatic generation of meeting minutes, transcription of interview recordings, generation of online course subtitles, and analysis of customer service call content.

[0028] Media and Entertainment: Automatic video subtitle addition, live streaming subtitles, and content retrieval (video clip search via voice).

[0029] In the medical field: Automatic entry of doctors' dictated medical records (electronic medical records).

[0030] In the judicial field: real-time transcription of court hearing records improves the work efficiency of court clerks.

[0031] In the field of education: pronunciation assessment and classroom content recording in language learning applications.

[0032] Named Entity Recognition (NER) refers to the process of automatically identifying predefined entities with specific meanings from unstructured text and classifying them into appropriate categories. In practical applications, NER systems can recognize a wide variety of entity types, including general types and domain-specific types. General types can recognize information such as names of people, places, dates, proper nouns, numbers, and currencies. Domain-specific types can be used to identify information in the medical, financial, or legal fields.

[0033] Coreference resolution refers to the process of identifying all referents in a text that point to the same real-world entity and linking them together to form a "coreference chain" or "cluster." In essence, it can be understood as finding all words or phrases in a text that refer to the same thing. A referent is a word or phrase in the text used to refer to a particular entity. It can be a pronoun, a noun phrase, a proper noun, etc. An entity is a unique object in the real world, such as a person, a place, or a concept. Coreference occurs when two or more referents point to the same entity; they are in a coreference relationship.

[0034] Conditional Random Fields (CRFs) are discriminative undirected graphical probabilistic models. They are used to calculate a probability for all possible output sequences and then select the one with the highest probability as the final result.

[0035] Hidden Markov Models (HMMs) are generative directed graphical probabilistic models used to describe observation sequences generated by hidden state chains. In this application, they are used for topic switching / state transition modeling.

[0036] BERTopic: A topic modeling method based on transformers and topic representation clustering. Its core idea is that "semantically similar documents should belong to the same topic." It does not rely on the probability distribution of word frequencies, but instead utilizes powerful pre-trained language models (such as BERT) to transform documents into vectors that capture context and semantics. Then, it searches for document clusters in these vector spaces, with each cluster representing a topic. Therefore, in this application, it is used for topic clustering during issue identification.

[0037] Latent Dirichlet Allocation (LDA) topic model: a generative probabilistic model. Its core idea is that each document can be viewed as a mixture of multiple topics with different probabilities; and each topic can be viewed as a mixture of all words in the vocabulary with different probabilities.

[0038] Speaker Stickiness: The prior tendency for adjacent segments to be spoken by the same speaker. It can also be understood as the inertia of a conversation. For example, in a natural conversation, the two people usually speak for a short period of time before the next person speaks. Therefore, when analyzing an audio clip, we can first set a prior tendency: to keep the speaker the same for a period of time.

[0039] Evidence citation graph: This graph organizes the "evidence fragment - reasoning conclusion - confidence level - conflict relationship" into a graph structure. Its core objective is to formalize, visualize, and computably represent the reasoning process. Based on this evidence citation graph, information can be traced back along the edges of the graph for any conclusion. In this application, it is used to implement the backtracking of the processing of meeting minutes.

[0040] This application provides a method, apparatus, device, and storage medium for processing meeting minutes, used to reconstruct the content of meeting minutes to improve their readability. The electronic device provided in this application can be implemented as various types of user terminals or as a server.

[0041] Electronic devices can reconstruct the content of meeting minutes by running the meeting minutes processing method provided in the embodiments of this application, so as to improve the readability of the meeting minutes.

[0042] The above methods can be applied to many content processing fields, such as meeting minutes, experimental records, and so on.

[0043] In one exemplary solution, this meeting minutes processing method can be applied to online meetings. For example, in an online meeting scenario, it is necessary to process the audio data of the meeting recording into text. The specific process can be as follows: After the online meeting ends, the audio and video meeting records are obtained; the audio and video meeting records are subjected to speech recognition to obtain the meeting text to be processed; at the same time, multi-source matching data is obtained based on the information of the participants at that time. The multi-source matching data includes participant metadata (such as name, position and department), meeting information of the audio and video meeting records (such as meeting topic, meeting time, meeting check-in, project data involved in the meeting, etc.), and historical meeting data of the meeting participants; then, based on the above multi-source matching data, the speaker matching is performed on the text data of the meeting text to be processed to obtain the first meeting minutes; then, the topic identification is performed on the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and the summary information corresponding to each topic; finally, based on the multiple topics and the summary information corresponding to each topic, the meeting text to be processed is reconstructed to obtain the target meeting minutes. After this processing, the information from the current online meeting will be presented in a structured agenda, which will enhance the readability of the meeting minutes.

[0044] It should be understood that the above is only an exemplary application scenario of the meeting minutes processing method. There are many other possible application scenarios, which are not limited here.

[0045] The meeting minutes processing method provided in this application embodiment is applied to, for example, Figure 1 The system architecture diagram shown is for your reference. Figure 1To support a meeting minutes processing method, the terminal device 100 connects to the server 300 via a network 200, and the server 300 connects to the database 400. The network 200 can be a wide area network (WAN), a local area network (LAN), or a combination of both. The client for implementing the content operation plan is deployed on the terminal device 100, or it can run on the terminal device 100 as a standalone application. The specific presentation of the client is not limited here. The server 300 involved in this application can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The terminal device 100 can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication over a bidirectional communication link. Such a device can include cellular or other communication devices with single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays. The specific terminal device 100 can be a desktop terminal or a mobile terminal. Specifically, it can also be one of the following: Augmented Reality (AR) devices, mobile phones, tablets, laptops, in-vehicle devices, wearable devices, smart TVs, smart home appliances, aircraft, or intelligent voice interaction devices. The solution provided in this application can be completed by the terminal device 100 and the server 300 working together. The database 400, in short, can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of applications. A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. Database management systems (DBMS) can be categorized based on the database model they support, such as relational or Extensible Markup Language (XML); or based on the type of computer they support, such as server clusters or mobile devices; or based on the query language used, such as Structured Query Language (SQL) or XQuery; or based on performance priorities, such as maximum scale or highest operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, by simultaneously supporting multiple query languages.In this application, database 400 can be used to store data such as audio data, meeting text to be processed, participant metadata, historical meeting data, and meeting information, etc.

[0046] Those skilled in the art will understand that Figure 1 The system architecture diagram shown is merely one possible system architecture for this application and does not constitute a limitation on the system architecture of this application. Other system architectures may include more than one. Figure 1 The number of more or fewer terminal devices or servers shown, for example Figure 1 Only one server is shown in the diagram. It is understood that the system architecture may also include one or more other terminal devices or servers, which are not limited here.

[0047] It should be noted that, Figure 1 The system architecture shown is merely an example. The servers and scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of servers and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0048] like Figure 2 The diagram shown is a flowchart of an embodiment of the meeting minutes processing method in this application. The following description, using a server as the executing entity, details the meeting minutes processing method, which may include the following steps 201-204: 201. Obtain the meeting text to be processed and multi-source matching data, which includes participant metadata, current meeting information, and historical meeting data.

[0049] In this embodiment, the server can receive audio conference data recorded by a conference application, and then perform ASR transcription on the audio conference data to obtain the conference text to be processed.

[0050] Meanwhile, the participant metadata in this multi-source matching data can be understood as a structured set of information used to comprehensively describe, identify, classify, and associate individuals participating in meetings or activities. This can include the following information: Basic identity information, such as name, employee ID, job title, department, job level, and company. Meeting status information, such as meeting role (meeting organizer, speaker, participant, recorder, etc.), participation status (e.g., registered, checked in), registration time, and preferred meeting topics. Contact information, such as email address, phone number, and instant messaging account.

[0051] The current meeting information can be understood as a comprehensive description of the meeting. This includes information such as the online meeting identifier, number of participants, meeting start time, meeting end time, keynote speaker, and meeting topic.

[0052] Furthermore, the current meeting information can also include project data (such as project requirements, project defects, and project system logs related to the current meeting).

[0053] Historical meeting data can be understood as the meeting records of all participants in the current meeting. It can also refer to all historical meeting data archived by a company or organization. The specific definition is not limited here.

[0054] 202. Based on the multi-source matching data, perform speaker matching on the meeting text to be processed to obtain the first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data.

[0055] After receiving the meeting text to be processed and the multi-source matching data, the server can match the corresponding speaker for each segment of the meeting text to be processed based on the multi-source matching data. The specific processing flow can be as follows: the server obtains a set of candidate speakers corresponding to each segment of the meeting text to be processed based on at least one speaker matching operation; based on the set of candidate speakers corresponding to each segment of the text, the server determines the speaker corresponding to each segment of the text to obtain the first meeting minutes.

[0056] The following describes at least one speaker matching operation in this embodiment: In one exemplary solution, the server can match roles with text data based on a semantic matching scheme. Specifically, it generates a role vector library using the job descriptions and historical data of attendees, and calculates the similarity between the semantic vector of each text segment and this library, thereby achieving speaker matching. The process can be as follows: The server extracts semantic features from the participant's metadata and historical meeting data, and trains a role vector library to represent the semantic association between participants and text data. Then, the server obtains the semantic vector corresponding to each segment of the meeting text to be processed. Based on this semantic vector, it performs similarity matching with the semantic vectors of the text data in the role vector library, selecting the role corresponding to the text data with a similarity higher than a threshold as the first candidate speaker for each segment of the meeting text to be processed. This ensures semantic consistency between the meeting text to be processed and the speaker's role.

[0057] It should be understood that this similarity can be cosine similarity.

[0058] The following is a specific example: A transcript of a new product requirements review meeting. There were three key participants, but the ASR (Automatic Speech Recognition) system failed to accurately identify the speakers during transcription, resulting in some anonymous snippets of their speech.

[0059] Participants and their roles: Zhang San (Product Manager): Responsible for user research, requirement definition, feature planning, and market strategy.

[0060] Li Si (Backend Engineer): Responsible for server architecture, database design, API interfaces, and system performance.

[0061] Wang Wu (Front-end Engineer): Responsible for user interface (UI), user experience (UX), component development, and page performance.

[0062] When using the above-mentioned semantic matching method for role content, the process can be as follows: The server analyzes text data related to each participant's role and generates a unique "role semantic vector" for them.

[0063] During this process, the server can collect corpus data based on the historical meeting data of the aforementioned participants. For example, for Zhang San (product manager), his job description is: "Responsible for user needs research, writing PRD documents, planning product roadmaps, and monitoring user growth and market feedback...".

[0064] Its historical data includes: "Our DAU increased by 15% last quarter...", "The user profile for this feature is young white-collar workers...", "We must add this feature in the next version iteration to cope with competitive pressure...".

[0065] The server then calls a pre-trained language model (such as BERT, ERNIE, or other industry-specific models) to process the aforementioned historical corpus. The model encodes all relevant corpus data for the role, extracts its core semantics, and outputs a high-dimensional mathematical vector. For example, Vector_PM (Product Manager Role Vector): semantically highly related to words such as "user," "market," "demand," and "function."

[0066] During this analysis, the server uses a model to process the current speech and generate semantic vectors. These semantic vectors capture the core concepts in the current meeting text, such as "market," "demand," and "function."

[0067] The server can use a cosine similarity algorithm to calculate the similarity between the semantic vector of the current meeting speech and each vector in the role vector library. For example, it can calculate the cosine similarity between the semantic vector of the current meeting speech and Vector_PM (Product Manager). If the anonymous speech fragment in the current meeting frequently mentions information such as "user," "market," "demand," and "feature," then the similarity is considered high. Based on this similarity, the server concludes that the match with the Product Manager is 0.92. At this point, the server can determine that the semantic vector of this speech has the highest similarity to the role vector of the backend engineer, thus inferring that the speaker is Zhang San (Product Manager). In other words, the server can assign this anonymous speech fragment in the current meeting to Li Si.

[0068] In another exemplary scheme, the server identifies the title information in the meeting text to be processed using named entity recognition technology; then, based on the coreference resolution method and the title information, it determines the second candidate speaker corresponding to each segment of text data in the meeting text to be processed. In this way, in multilingual or colloquial scenarios, it can ensure that the title information matches the identity of the specific speaker, thereby improving the accuracy of speaker matching.

[0069] It should be understood that this appellation information may include names, job titles, aliases, and nicknames that appear in the text.

[0070] The following example illustrates this approach. Assume the current meeting is primarily conducted in Chinese, but includes English names and some colloquial expressions.

[0071] The attendee's metadata can be as follows: ID: 101, Name: Zhang Wei, Position: General Manager, Common Title: General Manager Zhang.

[0072] ID: 102, Name: Li Jing, Position: Marketing Director, Common Title: Director Li.

[0073] ID: 103, Name: Robert Chen, Position: Technical Lead, Chinese Name: Chen Yu, Common Titles: Robert, Engineer Chen.

[0074] ID: 104, Name: Wang Xiaoming, Position: Product Manager, Nickname: Mingzai.

[0075] The following is a text snippet from the current meeting: Speaker A: "First, I would like to ask Mr. Zhang to summarize the overall performance of the last quarter." Spokesperson B: “Okay. Our revenue last quarter exceeded expectations by 15%. This is mainly due to the marketing activities of Director Li’s team. Her online marketing strategy was very successful.” Spokesperson C: "To add to that, the technical team also deserves a lot of credit. Especially Luo Bo, who led the team to solve the performance bottleneck that had been troubling us for a long time. Engineer Chen can elaborate on this point." Spokesperson D: "Yes, we have restructured our core service modules. Also, regarding the new features for next quarter, I think Mingzai's proposal is very creative." Therefore, the execution flow of this server can be as follows: First, the NER model is called to extract title entities. This NER model is trained to recognize entities such as names, organizations, and positions / titles. For the above text data, the corresponding entity results can be as follows: Entity 1: General Manager Zhang; Entity 2: Director Li; Entity 3: She; Entity 4: Radish; Entity 5: He; Entity 6: Engineer Chen; Entity 7: Mingzai.

[0076] The server then attempts to directly and precisely match the entities identified by NER with the known attendee metadata. The matching process can be as follows: Entity 1 (General Manager Zhang) -> Search in metadata. Found that Zhang Wei (ID:101) is commonly addressed as "General Manager Zhang". Match successful: General Manager Zhang -> Zhang Wei (ID:101).

[0077] Entity 2 (Director Li) -> Search in metadata. It is found that Li Jing (ID:102)'s position is "Marketing Director," indicating that "Director Li" refers to her. Match successful: Director Li -> Li Jing (ID:102).

[0078] Entity 6 (Chen Gong) -> Search in metadata. Found "Chen Gong" in Robert Chen's (ID:103) common titles. Match successful: Chen Gong -> Robert Chen (ID:103).

[0079] Entity 7 (Mingzai) -> Search in metadata. Found that Wang Xiaoming (ID:104)'s nickname is "Mingzai". Match successful: Mingzai -> Wang Xiaoming (ID:104).

[0080] The current matching status: Mr. Zhang -> Zhang Wei (ID:101).

[0081] Director Li -> Li Jing (ID:102).

[0082] Chen Gong -> RobertChen (ID:103).

[0083] Mingzai -> Wang Xiaoming (ID:104).

[0084] "She", "radish", and "he" are still undetermined.

[0085] Therefore, the service can use coreference resolution to further identify the three entities "she", "radish", and "he". In this embodiment, the coreference resolution model analyzes syntactic structures and context to find the referent (antecedent) of the pronoun. For example, for the entity (she), according to the sentence: "This is mainly due to the marketing activities of Director Li's team. The online marketing strategy she developed was very successful." It can be seen that the pronoun "she" appears after "Director Li", and "Director Li" is female (Li Jing), which meets the gender reference. Semantically, the strategy developed by "she" immediately follows the description of the credit of Director Li's team. Therefore, the association relationship can be generated: she -> Director Li -> Li Jing (ID: 102).

[0086] Identify other entities in this way, and the final matching status can be obtained: Anaphora chain 1: {General Manager Zhang} -> Zhang Wei (ID: 101).

[0087] Anaphora chain 2: {Director Li, she} -> Li Jing (ID: 102).

[0088] Anaphora chain 3: {radish, he, Engineer Chen} -> Robert Chen (ID: 103).

[0089] Anaphora chain 4: {Mingzai} -> Wang Xiaoming (ID: 104).

[0090] In practical applications, the entity "radish" has not been identified through the metadata of the participants. At this time, the server can also call the phoneme approximate matching method for identification.

[0091] That is, the server extracts the phoneme representation of the unrecognized entity. For example, the server converts "radish" and all the names that have not been successfully matched in the metadata into phoneme representations. "radish" -> [luobo] (pinyin) or more standard international phonetic alphabet. "Robert" (from Robert Chen) -> phoneme representation of English pronunciation.

[0092] Then calculate the similarity of various phoneme representations. That is, the server uses a phoneme edit distance algorithm (such as a variant of Metaphone, Soundex) to calculate the similarity between phoneme sequences. When the phoneme similarity score exceeds the preset threshold, and "Robert Chen" is the only reasonable candidate in the current context (already linked to "Engineer Chen" and "he" through coreference resolution).

[0093] Matching success: radish -> Robert -> Robert Chen (ID: 103).

[0094] In this scheme, the final matching results of the speakers can be as follows: General Manager Zhang -> Zhang Wei (ID:101); Director Li, she -> Li Jing (ID:102); Radish, he, Engineer Chen -> Robert Chen (ID:103); Mingzai -> Wang Xiaoming (ID:104).

[0095] In another exemplary scheme, the server can perform dialogue interaction recognition on the text data in the meeting text to obtain the third candidate speaker corresponding to each segment of text data in the meeting text. That is, the server can analyze the interaction relationship between segments, such as the contextual association between "request" and "response". For example, if the "request" segment explicitly mentions a person, the "response" segment will prioritize matching that person. At the same time, the server may also analyze the references to other speeches in the segments (such as "as Mr. Zhang mentioned") to assist in identity alignment. Furthermore, the server can combine the consistency of the meeting topic, segment content and file name to enhance alignment accuracy. In this way, the logical rationality of speaker matching is enhanced through contextual interaction relationships.

[0096] In another exemplary scheme, the server can perform domain terminology recognition on the text data in the meeting text to be processed to obtain a domain terminology set; and match each domain term in the domain terminology set with a domain terminology database to obtain a fourth candidate speaker. This domain terminology database is used to represent the association between domain terms and participant roles. That is, the server can optimize the matching of people who frequently use specific terms with roles (for example, a technical leader is more likely to use technical terms). This combination of domain features further improves the accuracy of role matching.

[0097] In another exemplary approach, the server can perform speaker matching for each segment of the conference text to be processed based on temporal stickiness features and a globally optimal allocation method to obtain a fifth candidate speaker. That is, the server may construct a prior trend for the model based on temporal stickiness features, such as "speakers usually speak consecutively (i.e., adjacent segments often belong to the same person)". Then, the server constructs a conditional random field model, modeling the identity matching of each segment of the conference text as a sequence labeling problem; its input features can include the following information: semantic similarity; title recognition results; term weights; dialogue interaction graph scores; temporal stickiness (same-person transfer reward); and cross-topic switching penalty. Finally, the conditional random field is solved based on the above input features to obtain the identity alignment result.

[0098] Based on the above description, it can be seen that in the process of the server matching speakers for each segment of the conference text to be processed based on the temporal stickiness characteristics and the global optimal allocation method, its input data can be the output of the above four matching schemes. Therefore, in the execution process, this scheme can be used as a process of screening candidates for the above four matching schemes.

[0099] It should be understood that the solution algorithm can utilize the Viterbi algorithm or dynamic programming to calculate the globally optimal path. Specific limitations are not specified here.

[0100] During the speaker matching process, in addition to outputting the speaker's identity corresponding to each segment of text data in the meeting text to be processed, the server can also output the following information: Confidence score, which provides a confidence score for the identity matching of each statement, is used to measure the reliability of the match.

[0101] Candidate parallelism and evidence fragments, that is, for uncertain identity alignment, provide multiple candidate results and their corresponding evidence (such as semantic similarity, title matching results).

[0102] Among them, the candidate parallels and evidence fragments can serve as data sources for the subsequent "reflection-branching-backtracking mechanism".

[0103] Based on the above description, it can be said that... Figure 3 The flowchart shown illustrates the speaker identity matching process in this embodiment.

[0104] like Figure 3 The flowchart shown illustrates that after the server obtains the meeting text to be processed, it can perform role-semantic matching, title naming and common reference matching, dialogue interaction matching and domain terminology matching in parallel to obtain a candidate set; then, based on temporal stickiness and global allocation methods, it determines the final speaker identity matching result from the candidates and outputs the identity matching result and its corresponding confidence level.

[0105] The role-semantic matching, title naming and common reference matching, dialogue interaction matching and domain terminology matching schemes can be found in the above descriptions, and will not be elaborated here.

[0106] It should be understood that when the multi-source matching data is incomplete, it is not necessary to execute all speaker identity matching schemes.

[0107] 203. Identify the topics of the first meeting minutes to obtain the multiple topics corresponding to the first meeting minutes and the summary information of each topic.

[0108] In this embodiment, after obtaining the first meeting minutes, in order to enhance the readability of the first meeting minutes, the server can also identify the topics in the first meeting minutes to obtain multiple topics and the corresponding summary information of the topics.

[0109] Alternatively, the server may employ the following technical solutions when performing issue identification: The server performs semantic analysis on the first meeting minutes to obtain semantic vectors of the text data in the first meeting minutes; it then performs topic clustering on the semantic vectors to obtain multiple initial topics; it performs topic switching detection on the semantic vectors to obtain the start time of the multiple initial topics; it calculates and aggregates the semantic vectors to obtain the complete content of the multiple initial topics; based on the start time and the complete content, it constructs multiple topics corresponding to the first meeting minutes; it extracts key nodes from the text data of each topic; and based on the key nodes, it uses summarization generation technology to compress and reorganize the text data of each topic to obtain the summary information of each topic.

[0110] In this embodiment, the server can adopt the following scheme when performing topic clustering: In one exemplary scheme, the server extracts sentence vectors or paragraph vectors of the meeting text to be processed based on the BERTopic model, and then performs clustering on them.

[0111] In another exemplary scheme, the server uses a latent Dirichlet distribution model to classify the topics of the segments.

[0112] In this embodiment, the server can use a Hidden Markov Model (HMM) to model the state sequence of topics, where each topic state corresponds to an issue. Specifically, the semantic features of each text data in the meeting text to be processed are used as observations, and then the HMM is used to infer topic switching points, thereby outputting the start and end times of each issue.

[0113] The above solution utilizes advanced topic modeling and semantic analysis techniques to ensure efficient aggregation of relevant content.

[0114] In this embodiment, the server performs semantic similarity calculation and aggregation, which can be understood as aggregating related segments of the same topic across time periods to establish complete topic content. During this process, the server calculates the semantic similarity of segments within different time periods to determine if they belong to the same topic; it also identifies key statements within the topic, such as problem statements, analytical content, and concluding statements; and aggregates related segments of the same topic across time periods based on aggregation rules. These aggregation rules can be set as follows: "Segments of the same topic are rearranged in logical order (problem → analysis → conclusion → impact → action)." After the above processing, the server can obtain multiple topics of the meeting text to be processed and the complete content of each topic. This breaks time constraints and aggregates scattered topic content into a unified output.

[0115] In this embodiment, when the server extracts key nodes, these key nodes can be questions, conclusions, or action items. Then, based on these key nodes, key information is extracted, and finally, summarization and reorganization techniques (such as a Transformer-based generative model) are used to compress and reorganize the content, ultimately outputting a formatted summary text (i.e., summary information). This method of generating execution summaries significantly shortens review time.

[0116] Optionally, to more clearly display task information in the meeting, the server can also add a scheme for extracting action item information. The specific process can be as follows: the server extracts and supplements action item information from the first meeting minutes to obtain structured data. This action item information is used to characterize the task content, task execution object, and task time contained in the meeting text to be processed. This structured data is used for topic identification.

[0117] In one exemplary scheme, the specific process by which the server extracts and supplements action item information from the first meeting minutes to obtain structured data can be as follows: Intent recognition is performed on the first meeting minutes to obtain the intent classification of each text data segment in the first meeting minutes; the text data to be processed is determined based on the intent classification; syntactic analysis is performed on the text data to be processed to extract and supplement the task content, the task execution object, and the task time; the structured data is generated based on the task content, the task execution object, and the task time.

[0118] In this embodiment, the intent recognition is used to identify implicit intents in the meeting text to be processed, such as promises, requests, and suggestions. The server can use a classification model (such as the BERT model based on Transformer) to classify the intent of the segments. In an exemplary scheme, the server can use the classification model to extract features from verbs and verb-object structures (such as "needs to be completed", "suggests to be implemented") and tone markers (such as "should", "can", "must") in the meeting text to be processed, and then determine the intent type and corresponding confidence level of each segment of text data based on these features.

[0119] This action item information extraction is used to extract task content, responsible person, dependencies, and suggested deadlines from the meeting text to be processed. In this embodiment, the server can perform dependency parsing, such as using Dependency Parsing to analyze the syntactic structure of sentences. It extracts task-related predicates, objects, and subjects (e.g., "complete project report," "General Manager Zhang is responsible"). Event extraction techniques are used to identify task definitions (e.g., "complete task X") and conditions (e.g., "can only start after completion") within the sentences.

[0120] In task execution target identification, the current speaker is prioritized for matching. If the responsible person is not clearly identified, the responsibility is derived based on the job responsibility matrix.

[0121] In task time identification, if the date is not explicitly given, the task execution time can be inferred based on "meeting date + rule base" (such as common time rules like 2 days, 7 days, 14 days, etc.). If the enterprise system provides milestones or organizational calendars, the nearest feasible date is retrieved from them, and a "suggested deadline" label is added (such as "suggested completion time: 2023-12-01").

[0122] After performing the above operations, the server outputs a structured action item containing the task content, the task execution object, and the task time.

[0123] Simultaneously, the server can also mine pain points and needs from the meeting text to be processed, that is, identify the pain points or needs implied in the speeches, providing a basis for problem solving and follow-up actions. The server can utilize a predefined dictionary of pain points and needs (such as keywords like "problem," "difficulty," and "need") to assist in the mining. Alternatively, it can use semantic cue learning methods to learn and identify the expressions of pain points or needs through few-shot cues. Or it can use semantic similarity analysis, that is, calculating the similarity between the semantic vectors of a pre-trained language model (such as BERT) and the pain point / need template. Finally, the server outputs a structured list of pain points / needs, with each record containing a problem description and its context.

[0124] 204. Generate the target meeting minutes based on the multiple topics and the summary information corresponding to each topic.

[0125] After obtaining multiple topics from the first meeting minutes and the summary information corresponding to each topic, the server can generate the target meeting minutes based on the multiple topics and the summary information corresponding to each topic.

[0126] Optionally, to enhance the readability of the meeting minutes, after obtaining the complete content of each topic in the meeting minutes, the server can also perform structured processing on each topic. That is, the server decomposes the content of each topic into structured nodes, clearly defining the core elements of the content. These nodes can be divided as follows: Problem: The issues that need to be addressed or the core content of the discussion in the agenda.

[0127] Analysis: Analyze the background, causes, and current status of the problem.

[0128] Conclusion: The conclusions or decisions reached regarding the problem.

[0129] Impact: The potential impact of the agenda on projects, businesses, or attendees.

[0130] Actions: Specific follow-up tasks, including task details, responsible person, and deadline.

[0131] The server then generates topic nodes for each topic in the meeting to be processed based on the aforementioned node types and node generation rules. These node generation rules can be set as follows: using an intent classification model to identify problem descriptions, analytical statements, and concluding statements in the meeting text; leveraging syntactic analysis and event extraction techniques to extract key tasks, responsible persons, and other information from the text segments; and finally generating topic nodes in the logical format of "Problem → Analysis → Conclusion → Impact → Action." This topic node generation rule ensures that the elements of problem, analysis, conclusion, and action are presented completely.

[0132] After the above processing, the chronological meeting texts can be reconstructed into topic-driven structured content, making the target meeting minutes generated from the meeting texts easier to understand quickly.

[0133] Optionally, to ensure the accuracy of the target meeting minutes, the server can also perform conflict checking on the target meeting minutes to obtain the check result; and update the target meeting minutes based on the check result.

[0134] In one exemplary scheme, the server performs conflict verification on the target meeting minutes to obtain the verification result as follows: detect the target conflict type in the target meeting minutes, which includes identity consistency conflict, issue position conflict and content integrity conflict; generate a hypothesis branch for the target conflict type; calculate the confidence of the hypothesis branch; and determine the verification result based on the confidence.

[0135] It should be understood that the server can detect the following types of conflicts: identity consistency conflict (i.e., the speaker's identity does not match the semantics, or the title is inconsistent with the job role), issue position conflict (i.e., the content of the speech is contradictory in terms of issue position), and content integrity conflict (i.e., the action item is missing key elements).

[0136] During the detection process, for identity consistency conflicts, the consistency between the predicate evidence and the job semantics can be verified; for issue position conflicts, the logical consistency of the issue position can be checked; for content integrity conflicts, each action item can be reviewed to see if it contains the necessary elements such as task, person in charge, and deadline.

[0137] In this embodiment, the hypothesis branch is alternative information generated for the target conflict type. Its generation rules can be as follows: for identity consistency conflicts, a hypothesis branch for identity rematch can be generated; for issue position conflicts, a hypothesis branch for position adjustment can be generated; for content integrity conflicts, a hypothesis branch for supplementary information can be generated. Then, using Beam Search or Top-k path retention methods, each branch path is re-scored (i.e., the confidence level in this embodiment), where the branch scoring is based on evidence consistency, global logicality, and result integrity.

[0138] After obtaining the scores for each hypothetical branch, a validation result can be derived based on the scores. This validation result might be as follows: if the branch score is higher than the main path, the main path is replaced. If the branch score is lower than the main path, the original path is retained, and an "uncertainty" label is added. In other words, this validation result indicates whether the target meeting minutes need to be revised. In this embodiment, the server performs conflict detection on the target meeting minutes. If a conflict exists (i.e., the target conflict type), a hypothesis branch is generated for that conflict type. The confidence level of the hypothesis branch is calculated, and finally, the verification result is generated based on the confidence level. For example, in the target topics, a resource allocation conflict is found between the tasks of "user experience optimization" and "determining the release time".

[0139] At this point, the server can generate the following hypothetical branch: Assumption Branch 1: Prioritize user experience optimization and delay the release by 1 week.

[0140] Assumption Branch 2: Prioritize ensuring release time, and complete the user experience optimization part.

[0141] Branch rating: Calculate scores for each of the two branches, and select branch 2 based on the criteria of "minimizing the impact on the project" and "prioritizing the market plan".

[0142] Finally, the verification result allows you to choose hypothesis branch 2, and update the action items for the target issue according to hypothesis branch 2: prioritize ensuring the release time and partially complete user experience optimization; finally, output the revised issue diagram and action item list (i.e., output the revised target meeting minutes).

[0143] Based on the above description, it can be done by... Figure 4 The flowchart illustrates the reflection-branch-backtracking mechanism in this embodiment.

[0144] This involves performing speaker identity alignment, topic identification, and action item information extraction to obtain the target meeting minutes; then reflecting on (i.e., validating) the target meeting minutes. Specifically, speaker identity alignment undergoes identity consistency conflict detection; topic identification undergoes topic stance conflict detection; and action item information extraction undergoes content integrity conflict detection. If the validation passes, the target meeting minutes are output; otherwise, revisions are made until validation passes.

[0145] Optionally, to facilitate subsequent auditing, the server can also output verification records, evidence reference diagrams during the verification process, and support for human-machine collaborative modification when outputting the target meeting minutes.

[0146] This evidence citation diagram is used to record key evidence and its reasoning path during the generation of the target meeting minutes. The evidence citation diagram can be a graph structure of "evidence fragment—reasoning conclusion—confidence level—conflict relationship," where each edge represents the association between evidence and reasoning conclusion.

[0147] This verification log is used to record discovered conflicts and their correction status. Each verification log may include the conflict type (e.g., identity, stance, integrity); the assumed branch and its score; and the correction status (whether the main path was replaced).

[0148] Human-machine collaborative modification refers to allowing manual modification of the output results and writing the modification records back into the model. In other words, the module output is manually adjusted, and the model's learning data is updated by comparing the content before and after the modification.

[0149] Based on the above description, the following will be used as... Figure 5 The system architecture shown illustrates the meeting minutes processing method of this embodiment.

[0150] like Figure 5 As shown, the input layer of this system architecture receives ASR-transcribed meeting text, participant metadata, meeting information, and historical meeting data.

[0151] The input layer then passes the above information to the speaker alignment engine, which performs at least one of the following identity alignment operations based on the above information: role-semantic matching; title naming and common reference matching; dialogue interaction matching; domain terminology matching and temporal stickiness and global allocation.

[0152] After speaker alignment is completed, an issue graph is constructed. This issue graph construction includes the following processes: topic clustering, topic switch detection, and cross-time period aggregation.

[0153] After the issue map is constructed, implicit action item mining can be carried out. This implicit action item mining includes the following processes: pain point or need discovery and action item generation.

[0154] Then, the reflection-branch-backtracking operation is performed. This includes the following processes: evidence citation graph, verification, and hypothesis branching and backtracking.

[0155] Final output.

[0156] To better implement the meeting minutes processing method in the embodiments of this application, based on the meeting minutes processing method, the embodiments of this application also provide a meeting minutes processing device, such as... Figure 6 As shown, the meeting minutes processing device 600 includes: The acquisition module 601 is used to acquire the meeting text to be processed and multi-source matching data, which includes participant metadata, current meeting information, and historical meeting data. The processing module 602 is used to perform speaker matching on the meeting text to be processed based on the multi-source matching data to obtain a first meeting minutes, the first meeting minutes including the text data of the meeting text to be processed and the speaker information associated with the text data; to perform topic identification on the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information corresponding to each topic; and to generate a target meeting minutes based on the multiple topics and the summary information corresponding to each topic.

[0157] In this embodiment, multi-source matching data is introduced for speaker identity matching, thereby improving the accuracy of speaker identity alignment. Simultaneously, topic identification is performed on the meeting text to be processed to obtain multiple topics corresponding to the meeting text and their corresponding summary information. Then, the meeting content is reconstructed based on the multiple topics and summary information to obtain the meeting minutes. This increases the readability of the meeting records while shortening the review time.

[0158] In some embodiments of this application, the processing module 602 is specifically used for: Perform conflict verification on the target meeting minutes to obtain the verification results; The meeting minutes for that target will be updated based on the verification result.

[0159] In some embodiments of this application, the processing module 602 is specifically used for: Detect the types of conflict in the meeting minutes of this target, which include conflicts of identity consistency, conflicts of issue positions, and conflicts of content integrity. Generate hypothesis branches for this type of conflict; Calculate the confidence level of this hypothesis branch; The verification result is determined based on this confidence level.

[0160] In some embodiments, before identifying the topics of the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information corresponding to each topic, the processing module 602 performs speaker matching on the meeting text to be processed based on the multi-source matching data to obtain the first meeting minutes. Then, the processing module 602 is further used to extract and supplement action item information from the first meeting minutes to obtain structured data. The action item information is used to characterize the task content, task execution object, and task time contained in the meeting text to be processed. The structured data is used for topic identification.

[0161] In some embodiments, the processing module 602 is specifically used to perform intent recognition on the first meeting minutes to obtain the intent classification of each text data in the first meeting minutes; Based on this intent, the text data to be processed is determined; Perform syntactic analysis on the text data to be processed to extract and supplement the task content, the task execution object, and the task time; This structured data is generated based on the task content, the task execution object, and the task time.

[0162] In some embodiments, the processing module 602 is specifically used to perform semantic analysis on the first meeting minutes to obtain the semantic vector of the text data in the first meeting minutes; The semantic vector is then subjected to topic clustering to obtain multiple initial topics; Topic switching detection is performed on the semantic vector to obtain the start time of the multiple initial topics; The semantic similarity of the semantic vector is calculated and aggregated to obtain the complete content of the multiple initial topics; Based on the start time and the complete content, construct multiple topics corresponding to the first meeting minutes; Extract the key nodes of the text data for each of the multiple topics; Based on this key node, the text data of each topic is compressed and recombined using summary generation technology to obtain summary information for each topic.

[0163] In some embodiments, the processing module 602 is specifically used to obtain a set of candidate speakers corresponding to each segment of text data of the meeting text to be processed based on at least one speaker matching operation. Based on the set of candidate speakers corresponding to each text data segment, the speaker corresponding to each text data segment is determined to obtain the first meeting minutes; The at least one speaker matching operation includes: A role vector library is constructed based on the participant's metadata and the historical meeting data; Obtain the semantic vector corresponding to each segment of the meeting text to be processed; Based on the semantic vector, similarity matching is performed from the role vector library to obtain the first candidate speaker corresponding to each segment of the meeting text to be processed; and / or; Named entity recognition is performed on the meeting text to be processed to obtain the title information in the meeting text; Based on the coreference resolution method and the title information, the second candidate speaker corresponding to each segment of text data in the meeting text to be processed is determined; and / or; Dialogue interaction recognition is performed on the text data in the meeting text to be processed in order to obtain the third candidate speaker corresponding to each segment of text data in the meeting text to be processed; and / or; Domain terminology is identified in the text data of the meeting text to be processed to obtain a set of domain terms; The fourth candidate speaker is obtained by matching each domain term in the domain terminology set with the domain terminology database, which is used to represent the relationship between domain terms and the roles of participants. and / or; Based on the temporal stickiness characteristics and the global optimal allocation method, speaker matching is performed on each segment of the conference text to be processed to obtain the fifth candidate speaker; The first candidate speaker, the second candidate speaker, the third candidate speaker, the fourth candidate speaker, and the fifth candidate speaker are included in the candidate speaker set.

[0164] This application also provides a computer device that integrates any of the meeting minutes processing devices provided in this application. The computer device includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the meeting minutes processing method in any of the embodiments described above.

[0165] This application also provides a computer device that integrates any of the meeting minutes processing devices provided in this application. For example... Figure 7 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically: The computer device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 7 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 701 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 701.

[0166] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0167] The computer device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0168] The computer device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0169] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 702 according to the following instructions, and the processor 701 runs the application programs stored in the memory 702 to realize various functions, as follows: Acquire the meeting text to be processed and multi-source matching data, wherein the multi-source matching data includes participant metadata, current meeting information, and historical meeting data; Based on the multi-source matching data, the speaker matching is performed on the meeting text to be processed to obtain a first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data. The first meeting minutes are subject to topic identification to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic; The target meeting minutes are generated based on the multiple topics and the summary information corresponding to each topic.

[0170] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0171] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the meeting minutes processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps: Acquire the meeting text to be processed and multi-source matching data, wherein the multi-source matching data includes participant metadata, current meeting information, and historical meeting data; Based on the multi-source matching data, the speaker matching is performed on the meeting text to be processed to obtain a first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data. The first meeting minutes are subject to topic identification to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic; The target meeting minutes are generated based on the multiple topics and the summary information corresponding to each topic.

[0172] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0173] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0174] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0175] The foregoing has provided a detailed description of a meeting minutes processing method, apparatus, computer device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for processing meeting minutes, characterized in that, include: Acquire the meeting text to be processed and multi-source matching data, wherein the multi-source matching data includes participant metadata, current meeting information, and historical meeting data; Based on the multi-source matching data, the speaker matching is performed on the meeting text to be processed to obtain a first meeting minutes, which includes the text data of the meeting text to be processed and the speaker information associated with the text data. The first meeting minutes are subject to topic identification to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic; The target meeting minutes are generated based on the multiple topics and the summary information corresponding to each topic.

2. The method according to claim 1, characterized in that, After generating the target meeting minutes based on the multiple topics and the summary information corresponding to each topic, the method further includes: Perform conflict verification on the target meeting minutes to obtain the verification results; The target meeting minutes are updated based on the verification results.

3. The method according to claim 2, characterized in that, The conflict check on the target meeting minutes to obtain the check result includes: Detect the target conflict type in the target meeting minutes, which includes identity consistency conflict, issue position conflict, and content integrity conflict. Generate hypothesis branches for the target conflict type; Calculate the confidence level of the hypothetical branch; The verification result is determined based on the confidence level.

4. The method according to any one of claims 1 to 3, characterized in that, Before identifying topics in the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic, the method further includes performing speaker matching on the meeting text to be processed based on the multi-source matching data to obtain the first meeting minutes: Action item information is extracted and supplemented from the first meeting minutes to obtain structured data. The action item information is used to characterize the task content, task execution object, and task time contained in the meeting text to be processed. The structured data is used for topic identification.

5. The method according to claim 4, characterized in that, The action item information of the first meeting minutes is extracted and supplemented to obtain structured data including: Intent identification is performed on the first meeting minutes to obtain the intent classification of each text data segment in the first meeting minutes; The text data to be processed is determined based on the intent classification. Perform syntactic analysis on the text data to be processed to extract and supplement the task content, the task execution object, and the task time; The structured data is generated based on the task content, the task execution object, and the task time.

6. The method according to any one of claims 1 to 3, characterized in that, The step of identifying topics in the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information for each topic includes: Semantic analysis is performed on the first meeting minutes to obtain the semantic vector of the text data in the first meeting minutes; The semantic vectors are then subjected to topic clustering to obtain multiple initial topics; Topic switching detection is performed on the semantic vector to obtain the start time of the multiple initial topics; The semantic vectors are calculated for semantic similarity and then aggregated to obtain the complete content of the multiple initial topics; Based on the start time and the complete content, construct multiple topics corresponding to the first meeting minutes; Extract the key nodes of the text data for each of the multiple topics; Based on the key nodes, the text data of each topic is compressed and recombined using summary generation technology to obtain summary information for each topic.

7. The method according to any one of claims 1 to 3, characterized in that, Based on the multi-source matching data, speaker matching is performed on the meeting text to be processed to obtain the first meeting minutes, including: Based on at least one speaker matching operation, a set of candidate speakers corresponding to each segment of the meeting text to be processed is obtained; The speaker corresponding to each piece of text data is determined based on the set of candidate speakers corresponding to each piece of text data, so as to obtain the first meeting minutes; The at least one speaker matching operation includes: A role vector library is constructed based on the participant metadata and the historical meeting data; Obtain the semantic vector corresponding to each segment of the meeting text to be processed; Based on the semantic vector, similarity matching is performed from the role vector library to obtain the first candidate speaker corresponding to each segment of the meeting text to be processed; and / or; Named entity recognition is performed on the meeting text to be processed to obtain the title information in the meeting text to be processed; Based on the coreference resolution method and the title information, the second candidate speaker corresponding to each segment of text data in the meeting text to be processed is determined; and / or; Dialogue interaction recognition is performed on the text data in the meeting text to be processed to obtain the third candidate speaker corresponding to each segment of text data in the meeting text to be processed; and / or; Domain terminology recognition is performed on the text data in the meeting text to be processed to obtain a set of domain terms; The fourth candidate speaker is obtained by matching each domain term in the domain term set with the domain term database, which is used to represent the association between domain terms and the roles of participants. and / or; Based on temporal stickiness characteristics and a global optimal allocation method, speaker matching is performed on each segment of the meeting text to be processed to obtain a fifth candidate speaker; The first candidate speaker, the second candidate speaker, the third candidate speaker, the fourth candidate speaker, and the fifth candidate speaker are included in the candidate speaker set.

8. A meeting minutes processing device, characterized in that, include: The acquisition module is used to acquire the meeting text to be processed and multi-source matching data, including participant metadata, current meeting information, and historical meeting data. The processing module is configured to perform speaker matching on the meeting text to be processed based on the multi-source matching data to obtain a first meeting minutes, the first meeting minutes including the text data of the meeting text to be processed and the speaker information associated with the text data; to perform topic identification on the first meeting minutes to obtain multiple topics corresponding to the first meeting minutes and summary information corresponding to each topic; and to generate a target meeting minutes based on the multiple topics and the summary information corresponding to each topic.

9. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It contains a computer program that is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.