Conference data recording and annotation method and system for intelligent conferences
By collecting conference audio information for semantic extraction and clustering, and using text conversion and intelligent annotation technology, the problem of inefficiency in conference data recording and labeling is solved, and fast and accurate conference records and labeling is achieved.
Patent Information
- Application Number
- CN202510712934.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing conference systems are inefficient and error-prone in automatic recording and intelligent annotation of conference data.
The audio information of conferences is collected through the voice input component, semantic extraction and clustering is performed, text conversion model and natural language processing technology are used to generate text information, and key statement extraction and association analysis are performed through intelligent annotation channels to form conference records.
It realizes fast and accurate conference data recording and labeling, ensuring the completeness and accuracy of conference content.
Smart Images

Figure CN120235165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice information processing, and in particular to a conference data recording and annotation method and system for intelligent conferences. Background Art
[0002] In modern meetings, recording and annotating meeting data is crucial for summarizing, reviewing, and supporting decision-making. However, while existing conference systems offer a variety of functions (such as paperless meetings, simultaneous interpretation, and voting), they still lack the ability to automatically record and intelligently annotate meeting data. Meeting data, such as voice recordings, voting results, and simultaneous interpretation content, often requires manual compilation and annotation, resulting in inefficient and error-prone recording and annotation. Summary of the Invention
[0003] The present invention aims to solve the technical problems in the prior art of low efficiency and easy error in recording and annotating conference data, and provides a conference data recording and annotating method and system for intelligent conferences.
[0004] The technical solution of the present invention to solve the above technical problems is as follows:
[0005] In a first aspect, the present invention provides a conference data record annotation method for intelligent conferences, comprising: collecting conference audio information of a first topic time zone through a voice input component, performing semantic extraction, and obtaining a conference initial sentence data set; performing semantic clustering on the conference initial sentence data set to obtain multiple clusters of conference initial sentence data, extracting central semantics respectively, and obtaining a sentence list; traversing the sentence list, performing correlation analysis with the first topic time zone, and obtaining a related sentence list; receiving several sentence lists of non-first topic time zones, performing correlation analysis with the first topic time zone respectively, and obtaining an expanded related sentence list; setting the related sentence list and the expanded related sentence list as the first topic meeting record annotation.
[0006] Optionally, the conference audio information of the first topic time zone is collected through the voice input component, and semantic extraction is performed to obtain the conference initial sentence data set, including: performing text conversion on the conference audio information to obtain the first timbre text information up to the Mth timbre text information; combining natural language processing technology, traversing the first timbre text information up to the Mth timbre text information to perform preprocessing, obtaining the first timbre text basic features up to the Mth timbre text basic features, wherein the basic feature attributes include at least word segmentation features, part of speech features and syntactic features; sending the first timbre text information and the first timbre text basic features to the intelligent annotation channel for processing to obtain the first timbre text key sentence; until the Mth timbre text information and the Mth timbre text basic features are sent to the intelligent annotation channel for processing to obtain the Mth timbre text key sentence; adding the first timbre text key sentence up to the Mth timbre text key sentence into the conference initial sentence data set.
[0007] Among them, performing text conversion on the conference audio information to obtain the first timbre text information up to the Mth timbre text information includes: calling a text conversion model, wherein the text conversion model includes a timbre identification network and a text conversion network; processing the conference audio information through the timbre identification network to obtain the first timbre conference audio information up to the Mth timbre conference audio information, wherein the timbre identification network is constructed by machine learning training through multiple groups of data, and any group of the multiple groups of data includes audio recording data with at least two types of timbre and labels identifying several timbre audio recording information; processing the first timbre conference audio information up to the Mth timbre conference audio information through the text conversion network, and outputting the first timbre text information up to the Mth timbre text information, wherein the text conversion network is constructed by machine learning training through multiple groups of data, and any group of the multiple groups of data includes audio recording information and labels identifying text recording data.
[0008] Among them, the first timbre text information and the first timbre text basic features are sent to the intelligent annotation channel for processing to obtain the first timbre text key sentence, including: counting the set of named entities whose trigger frequency is greater than or equal to the trigger frequency threshold in the historical decision dependency data, and constructing a key sentence attribute library, wherein the historical decision dependency data is the key text selected by the user during historical decision-making; sending the first timbre text information and the first timbre text basic features to the text segmentation node of the intelligent annotation channel to extract a number of initial sentences and a number of initial sentence attributes; extracting the selected initial sentence attributes belonging to the key sentence attribute library from the several initial sentence attributes through the text matching node of the intelligent annotation channel; extracting the initial sentence corresponding to the selected initial sentence attribute from the selected initial sentence attribute, and adding it to the first timbre text key sentence.
[0009] Optionally, semantic clustering is performed on the conference initial statement data set to obtain multiple clusters of conference initial statement data, and central semantics are extracted respectively to obtain a statement list, including: performing pairwise semantic similarity analysis on the conference initial statement data set to obtain a number of semantic similarities; based on a predefined semantic similarity threshold, the conference initial statement data set is grouped in combination with the several semantic similarities to obtain the multiple clusters of conference initial statement data, wherein the multiple clusters of conference initial statement data have multiple clusters of timbre identification information; extracting the first conference initial statement data of the first cluster of conference initial statement data of the multiple clusters of conference initial statement data; based on the first conference The conference initial statement data is combined with the several semantic similarities, a preset number of adjacent conference initial statement data are selected from the first cluster conference initial statement data, and the semantic similarity average of the first conference initial statement data and the preset number of adjacent conference initial statement data is calculated, and set as the center coefficient of the first conference initial statement data; when the center coefficient of the first conference initial statement data is greater than or equal to the center coefficient threshold, the first conference initial statement data is set as the first cluster center semantics; until the Q-th cluster center semantics is obtained; the first cluster center semantics to the Q-th cluster center semantics, as well as the multi-cluster timbre identification information, are added into the statement list.
[0010] Optionally, the statement list is traversed, and a correlation analysis is performed with the first topic time zone to obtain a list of associated statements, including: extracting a first statement from the statement list, wherein the first statement has a timbre identification set; counting the triggering frequency characteristics of the first statement in the first topic time zone; configuring a first statement identity identification set based on the timbre identification set in combination with a timbre identity mapping library; performing professional-level mode identification on the first statement identity identification set through a predefined identity professional-level mapping library belonging to the first topic time zone to obtain a professional-level feature value; configuring a first weight for the trigger frequency feature normalization parameter, configuring a second weight for the professional-level feature value, performing weighted mean statistics, and obtaining a first statement associated feature value; when the first statement associated feature value is greater than or equal to the associated feature threshold, adding the first statement to the associated statement list.
[0011] Optionally, several statement lists of a time zone other than the first topic time zone are received, and correlation analysis is performed with the first topic time zone respectively to obtain an expanded associated statement list, including: extracting a first candidate statement based on the several statement lists, and calculating a candidate semantic similarity set between the first candidate statement and the associated statement list; when any one of the candidate semantic similarity sets is greater than or equal to an expandable semantic similarity threshold, adding the first candidate statement to a quasi-expanded statement set; for the quasi-expanded statement set, counting a quasi-expanded statement professional-level feature value set; based on the quasi-expanded statement professional-level feature value set, extracting quasi-expanded statements whose quasi-expanded statement professional-level feature values are greater than or equal to the professional-level feature threshold, and adding them to the expanded associated statement list.
[0012] In a second aspect, the present invention provides a conference data recording and annotation system for intelligent conferences, comprising:
[0013] The semantic extraction module is used to collect the conference audio information of the first topic time zone through the voice recording component, perform semantic extraction, and obtain the initial meeting sentence data set;
[0014] A semantic clustering module is used to perform semantic clustering on the initial conference statement data set to obtain multiple clusters of initial conference statement data, extract central semantics from each cluster, and obtain a statement list;
[0015] a correlation analysis module, configured to traverse the statement list, perform correlation analysis with the first topic time zone, and obtain a list of related statements;
[0016] an expansion association module, configured to receive a plurality of statement lists in a time zone other than the first topic, and perform correlation analysis on each of the statement lists with the first topic time zone to obtain an expansion association statement list;
[0017] The record annotation module is used to set the associated statement list and the expanded associated statement list as the first topic meeting record annotation.
[0018] By implementing the above technical solution, it is possible to collect the conference audio information of the first topic time zone through the voice input component, perform semantic extraction, obtain the initial conference sentence data set, completely record the audio information of the conference site, and convert the language information expressed by the participants into text information that can be recognized and analyzed by the computer, and identify the participants according to their timbre information to achieve accurate classification; through the above technical solution, it is possible to perform semantic clustering on the initial conference sentence data set, obtain multiple clusters of initial conference sentence data, extract the central semantics respectively, obtain a sentence list, and achieve rapid refinement of the conference content for the purpose of annotating the conference data;
[0019] By implementing the above technical solution, it is possible to traverse the statement list, perform correlation analysis with the first topic time zone, and obtain a list of related statements. This allows accurate extraction of related statements of the meeting topic based on the statement trigger frequency and statement source, and quickly locate the core focus of the meeting.
[0020] By implementing the above technical solution, it is possible to receive several statement lists in time zones other than the first topic time zone, perform correlation analysis with the first topic time zone respectively, obtain an expanded related statement list, and set the related statement list and the expanded related statement list as annotations for the first topic meeting record to expand the search scope of the related statements and ensure accurate and complete recording and annotation of the meeting content.
[0021] In summary, by implementing the present invention, the effect of quickly and accurately recording and marking conference data can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of a method for recording and annotating conference data for intelligent conferences provided by the present invention;
[0023] Figure 2 This is a structural diagram of a conference data recording and annotation system for intelligent conferences provided by the present invention.
[0024] In the accompanying drawings, the components represented by the reference numerals are as follows:
[0025] Semantic extraction module 11, semantic clustering module 12, association analysis module 13, expanded association module 14, record annotation module 15. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0028] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0029] Example 1, as Figure 1 As shown, an embodiment of the present invention provides a conference data recording and annotation method for intelligent conferences, including:
[0030] S100 collects the meeting audio information of the first topic time zone through the voice recording component, performs semantic extraction, and obtains the initial meeting sentence dataset;
[0031] S200 performs semantic clustering on the initial conference sentence data set to obtain multiple clusters of initial conference sentence data, extracts central semantics from each cluster, and obtains a sentence list;
[0032] S300 traverses the statement list, performs correlation analysis with the first topic time zone, and obtains a list of related statements;
[0033] S400: receiving a list of several statements in a time zone other than the first topic, and performing correlation analysis on each of the lists with the first topic time zone to obtain an expanded list of related statements;
[0034] S500 sets the associated sentence list and the expanded associated sentence list as the first topic meeting record annotation.
[0035] In step S100 of this embodiment, the audio information of the meeting in the first topic time zone is collected by the voice recording component, and semantic extraction is performed to obtain the initial meeting sentence dataset, including:
[0036] Performing text conversion on the conference audio information to obtain first timbre text information to Mth timbre text information;
[0037] In combination with natural language processing technology, traverse the first timbre text information until the Mth timbre text information to perform preprocessing, and obtain the first timbre text basic features until the Mth timbre text basic features, wherein the basic feature attributes include at least word segmentation features, part of speech features, and syntactic features;
[0038] Sending the first timbre text information and the first timbre text basic features to an intelligent annotation channel for processing to obtain key sentences of the first timbre text;
[0039] Until the Mth timbre text information and the Mth timbre text basic features are sent to the intelligent annotation channel for processing, and the Mth timbre text key sentence is obtained;
[0040] Add the first timbre text key sentence to the Mth timbre text key sentence into the conference initial sentence data set.
[0041] In this embodiment, the first topic time zone is the time range occupied by the discussion of the first topic (a single topic, such as personnel transfers or performance analysis) in the meeting, for example, from 2:00 PM to 3:00 PM. Meeting audio information can be collected using a paperless conferencing terminal (such as the DCS-2061 / 2062 series) and a simultaneous interpretation system (such as the GONSIN30000I).
[0042] The step of converting the conference audio information into text to obtain first to Mth timbre text information includes:
[0043] Retrieving a text conversion model, wherein the text conversion model includes a timbre identification network and a text conversion network;
[0044] Processing the conference audio information through the timbre identification network to obtain first timbre conference audio information through Mth timbre conference audio information, wherein the timbre identification network is constructed by machine learning training using multiple sets of data, each of which includes audio recording data having at least two types of timbre and labels identifying multiple timbre audio recording information;
[0045] Through the text conversion network, the first timbre conference audio information up to the Mth timbre conference audio information are processed respectively, and the first timbre text information up to the Mth timbre text information are output, wherein the text conversion network is constructed by using machine learning training through multiple groups of data, and any group of the multiple groups of data includes audio recording information and a label that identifies the text recording data.
[0046] In the embodiment of the present application, the text conversion model is mainly used to label the audio data collected from different sources at the conference site (mainly the speech audio of different people) according to the timbre (i.e., the aforementioned first timbre conference audio information to the said Mth timbre conference audio information) to distinguish the source of the audio, and then convert the audio information marked with different timbre labels into text, identify the language and text information contained therein, and output the aforementioned first timbre text information to the said Mth timbre text information (i.e., the information marked with the timbre label and converted into text).
[0047] In the construction of the text conversion model, the model contains two networks connected in series, namely the aforementioned timbre identification network (first part) and the text conversion network (second part). The conference audio information input into the model is first sent to the timbre identification network for timbre recognition, and then sent to the text conversion network after obtaining the timbre recognition label to identify the language and text information contained therein.
[0048] Among them, the input features of the timbre identification network of the first part are conference audio information, and the output features are the first timbre conference audio information to the Mth timbre conference audio information, that is, the conference audio information classified according to the timbre of the audio information.
[0049] The training data for the timbre identification network is any set of audio recordings containing at least two timbre categories and labels identifying information about the timbre audio recordings. The audio recordings containing at least two timbre categories can include audio recordings of speeches by different conference participants, audio recordings of a video played at the conference, or various types of noise at the conference, etc. The labels identifying the information about the timbre audio recordings are pre-labeled with the source types of the audio recordings of different timbre categories. For example, the label for audio recording A can be "007, Employee of Department A," the label for audio recording B can be "001, Director of Department B," and the label for audio recording C can be "Ambient background noise," and so on. Using the above method, multiple sets of data, each containing at least two timbre categories and labels identifying information about the timbre audio recordings, can be obtained as the data set for training the timbre identification network. The multiple sets of data containing the audio recordings containing at least two timbre categories and labels identifying information about the timbre audio recordings can be divided into a validation set and a training set in a ratio of 2:8 for training the timbre identification network.
[0050] Optionally, the timbre identification network can be constructed using a CNN-LSTM hybrid network, which is divided into three parts: an input layer, a feature extraction module, and an output layer.
[0051] The input layer of the timbre identification network is in audio format, specifically single-channel WAV format. It generates an 80-dimensional Mel spectrogram (with a 25ms window and a 10ms step size) and applies random band masking and temporal stretching for preprocessing. The feature extraction module of the timbre identification network consists of a convolution module and a time series modeling module. The convolution module uses a four-layer Conv2D with increasing filter counts from 64 to 256, kernel sizes from (3,3) to (5,5), and a ReLU activation function. The time series modeling module is a bidirectional LSTM layer with 256 hidden units, and the output states are concatenated to form a 512-dimensional feature vector. The output layer of the timbre identification network is a fully connected layer with an AM-Softmax activation function.
[0052] Furthermore, when training the timbre identification network, the initial learning rate of the CNN can be set to 0.0003, the dropout rate of the LSTM can be set to 0.3, the optimizer uses Lookahead+Rectified Adam, and the learning rate is fine-tuned using cosine annealing (with a 20-epoch cycle). The training sample size is at least 500 hours of clean speech (i.e., clear speech signals that have been processed to remove noise and interference), the training batch size is 128, and the number of training epochs is 100. Early stopping is implemented, meaning that training is stopped early if the validation set EER (Eerging Error Rate) shows no improvement over five consecutive epochs. The training data is the aforementioned multiple data sets, including audio recordings with at least two timbre categories and labels identifying information about the timbre of the audio recordings. The equal error rate (EER) can be used as the evaluation metric for the training of the timbre identification network. When the EER in the conference room scenario is ≤2.5%, the model is considered converged, and the timbre identification network is obtained.
[0053] After obtaining the timbre identification network, the original conference audio information is input, and the first timbre conference audio information to the Mth timbre conference audio information can be output, that is, the conference audio information classified according to the timbre of the audio information.
[0054] For the text conversion network in the second part of the text conversion model, its input features are the first to Mth timbre conference audio information output by the first timbre identification network, and its output features are the first to Mth timbre text information. This is the textual information contained in the audio information annotated with different timbre labels, such as specific text content, punctuation, etc. For example, if the timbre conference audio information output by the timbre identification network is the audio information generated by Department A employee 007 speaking at a meeting, then inputting this timbre-identified conference audio information (e.g., the timbre of Department A employee 007) into the text conversion network will output textual information with the same timbre identifier (e.g., the timbre of Department A employee 007). This textual information is the computer-recognizable human language information contained in the audio information, identified by the language recognition mechanism built into the timbre identification network. Therefore, it is necessary to prepare a large amount (e.g., 500 hours) of audio information with different timbre labels (e.g., at least 50 labels) in advance and manually annotate the textual information contained therein as training material for training the text conversion network. The training materials are divided into training set and validation set in a ratio of 8:2.
[0055] Furthermore, in building the text-to-text network, since the multi-timbre audio information output by the timbre identification network must be combined for text transcription, a hybrid Conformer-Transformer architecture can be employed to achieve high-precision multi-timbre text generation. This architecture comprises two main components: an audio encoder (Conformer-based) and a text decoder (Transformer-based). The text-to-text network takes as input the audio information from the first timbre conference, output by the timbre identification network, from the first to the Mth timbre conference. This information is then encoded using the audio encoder and decoded using the text decoder to obtain the text information contained in the audio information from the different timbre conferences.
[0056] Optionally, the audio encoder includes a six-layer Conformer module. Within each layer, three core components collaborate sequentially to extract features. These include a multi-head self-attention mechanism, which employs twelve parallel attention heads, each specifically analyzing a different aspect of the sound characteristics, such as pitch variation or articulation rhythm; a feedforward neural network, using a fully connected layer with a fourfold expansion ratio to temporarily expand the feature dimension to 3072 for deep processing, before compressing it back to 768 dimensions to maintain uniform dimensionality; and a depthwise separable convolutional layer, which uses a convolution kernel with a seven-times-step length to capture local continuity features in the sound signal. This design extracts detailed features while reducing computational complexity. This is accomplished by concatenating the 64-dimensional timbre feature vector with the original sound spectrum data and then performing a linear transformation to unify the vector to a 768-dimensional space. An intelligent adjustment module precedes each layer's processing. This module uses a sigmoid function to dynamically generate an adjustment coefficient between 0 and 1, automatically controlling the influence of the timbre feature on the current speech segment. For example, when ambient noise is detected, the system automatically reduces the weight of specific timbre features. The audio encoder input frame length can be 25ms per frame, the step length can be 10ms, and the depthwise separable convolution kernel size can be 7.
[0057] Optionally, the text decoder can include four layers of progressively processed decoding units, each layer equipped with eight independent attention heads. The decoding process adopts a cross-reference mechanism to simultaneously focus on two key information sources when generating each text - the semantics of the previous text generated and the sound feature map output by the encoder. The text decoder system has a built-in vocabulary of 50,000 commonly used words, covering the needs of mixed Chinese and English scenarios. A dual-task channel is set up in the final output stage, the main channel generates text content, and the auxiliary channel synchronously predicts punctuation. For example, when a question tone is recognized, the system will automatically add a question mark after the output text "project progress" to form a complete semantic expression. Among them, the maximum generation length of the text decoder can be selected as 512 tokens, and the temperature coefficient of the inference stage can be selected as 0.7.
[0058] For training the text conversion network, use the aforementioned training and validation set data. AdamW (with β1=0.9 and β2=0.98) is used as the optimizer. Cosine annealing with hot start is used as the learning rate scheduling strategy. The learning rate is reset after every 20 training rounds. The gradient clipping threshold is set to 1.0 to prevent gradient explosion.
[0059] The training evaluation standard of the text conversion network can be the character error rate. When the character error rate of the validation set stably drops to 5% or lower and meets the standard for three consecutive evaluations, the model is judged to have converged and the text conversion network is obtained.
[0060] By inputting the first timbre conference audio information to the Mth timbre conference audio information output by the first part of the timbre identification network into the text conversion network, the first timbre text information to the Mth timbre text information can be output, that is, the text information contained in the audio information marked with different timbre labels.
[0061] Next, natural language processing techniques are combined to traverse the first to the Mth timbre text information and perform preprocessing to obtain basic features of the first to the Mth timbre text. These basic features include at least word segmentation, part-of-speech, and syntactic features. Specifically, word segmentation breaks down a continuous text into individual words, for example, breaking "I love natural language processing" into "I / love / natural language processing." Part-of-speech tagging labels each word with its grammatical attributes, such as noun, verb, or adjective. Lexical analysis provides the foundation for subsequent syntactic and semantic analysis. Accurate word segmentation and part-of-speech tagging enable computers to better understand the relationships and grammatical structure between words. Syntactic analysis, the second stage of natural language processing, primarily analyzes the structural relationships between words in a sentence. Jieba's precise pattern segmentation can be used for word segmentation, combined with domain dictionaries (for example, by adding specialized terms such as "sales" and "cost reduction and efficiency improvement") to optimize the segmentation results. Part-of-speech tagging can be performed using the Penn Treebank annotation set from Stanford CoreNLP. The syntactic features can be extracted by constructing a dependency tree using an LTP dependency parser. Through the above method, the first timbre text basic features to the Mth timbre text basic features can be obtained.
[0062] In step S100 of this embodiment, the first timbre text information and the first timbre text basic features are sent to the intelligent annotation channel for processing to obtain the first timbre text key sentences, including:
[0063] Counting the named entity sets whose trigger frequency is greater than or equal to the trigger frequency threshold in the historical decision dependency data to build a key sentence attribute library, wherein the historical decision dependency data is the key text selected by the user in the historical decision;
[0064] Sending the first timbre text information and the first timbre text basic features to the text segmentation node of the intelligent annotation channel to extract a plurality of initial sentences and a plurality of initial sentence attributes;
[0065] extracting selected initial sentence attributes belonging to the key sentence attribute library from the plurality of initial sentence attributes through a text matching node of the intelligent annotation channel;
[0066] An initial sentence corresponding to the selected initial sentence attribute is extracted from the selected initial sentence attribute and added into the first timbre text key sentence.
[0067] In an embodiment of the present application, the intelligent annotation channel is used to count the set of named entities whose trigger frequency is greater than or equal to the trigger frequency threshold in the historical decision-dependent data in the timbre text information (i.e., the key text selected by the user during historical decision-making), and then use this named entity set to construct a key sentence attribute library.
[0068] The key texts selected during historical decision-making are key information in meeting texts automatically identified within a historical period (e.g., three months), such as topic names, speakers, voting results, etc.
[0069] Named entities refer to identifying the boundaries and categories of entities in Chinese text. For example, this includes identifying entity information such as names of people, places, and organizations. The trigger frequency threshold for named entities can be set manually. For example, the trigger frequency threshold can be set to 10 times. If an entity appears more than 10 times in the key text selected during historical decision-making, the entity will be included in the constructed key sentence attribute library.
[0070] The above steps can be implemented by using named entity recognition (NER) technology in natural language processing. This technology is an existing technology and will not be described in detail here.
[0071] Then, the first timbre text information (the language and text information contained in the first timbre audio) and the first timbre text basic features (including at least word segmentation features, part-of-speech features, and syntactic features) obtained in the above steps are sent to the text segmentation node of the intelligent annotation channel. The text segmentation node is used to segment the text sentences according to the first timbre text basic features. The segmented text is divided into sentences and segments according to normal language logic to obtain the aforementioned several initial sentences and the initial sentence attributes corresponding to the several initial sentences (mainly the named entity information contained in the initial sentences). Then, through the text matching node of the intelligent annotation channel, the selected initial sentence attributes belonging to the key sentence attribute library are extracted from the several initial sentence attributes; for example, if the initial sentence contains entity information such as a person's name, a place name, or an organization name that is included in the key sentence attribute library, the initial sentence containing the above entity information is added to the first timbre text key sentence.
[0072] Using the same method, multiple timbre text key sentences can be added until the Mth timbre text information and the Mth timbre text basic features are sent to the intelligent annotation channel for processing to obtain the Mth timbre text key sentence; then the first timbre text key sentence to the Mth timbre text key sentence are added to the meeting initial sentence data set to obtain the meeting initial sentence data set.
[0073] In step S200 of this embodiment, semantic clustering is performed on the initial meeting statement dataset to obtain multiple clusters of initial meeting statement data, and central semantics are extracted from each cluster to obtain a statement list, including:
[0074] Performing pairwise semantic similarity analysis on the initial meeting sentence dataset to obtain a number of semantic similarities;
[0075] Based on a predefined semantic similarity threshold and in combination with the plurality of semantic similarities, the conference initial sentence data set is grouped to obtain the multiple clusters of conference initial sentence data, wherein the multiple clusters of conference initial sentence data have multiple clusters of timbre identification information;
[0076] Extracting the first conference initial statement data of the first cluster of conference initial statement data from the plurality of clusters of conference initial statement data;
[0077] Based on the first conference initial statement data and in combination with the plurality of semantic similarities, a preset number of adjacent conference initial statement data are sorted from the first cluster of conference initial statement data, and an average of the semantic similarities between the first conference initial statement data and the preset number of adjacent conference initial statement data is calculated, and the average value is set as the first conference initial statement data center coefficient;
[0078] When the data center coefficient of the first conference initial statement is greater than or equal to the center coefficient threshold, setting the first conference initial statement data as the first cluster center semantics;
[0079] Until the Q-th cluster center semantics is obtained;
[0080] The first cluster central semantics to the Qth cluster central semantics, and the multiple clusters of timbre identification information are added to the sentence list.
[0081] In an embodiment of the present application, pairwise semantic similarity analysis is performed on the initial meeting sentence dataset in order to classify them according to the similarity between different initial sentences. Optionally, the similarity analysis method can be completed using a pre-trained language model such as Sentence-BERT. Specifically, Sentence-BERT (SBERT) is an improved model based on BERT. It encodes sentences through a twin network structure, converts text into high-dimensional semantic vectors, and then measures the consistency of the vector direction through cosine similarity to quantify semantic similarity. In this measurement method, semantic similarity is based on cosine similarity, with a value range of (0-1). The semantic similarity threshold can be set manually. For example, when the cosine similarity is (0.9-1.0), it is judged that the semantics are highly overlapped (such as "there will be a meeting tomorrow" and "there is a meeting tomorrow"). The method for obtaining the pre-trained language model is existing technology and will not be described in detail here.
[0082] Next, the meeting initial statement dataset needs to be grouped based on several semantic similarities to obtain the aforementioned multi-cluster meeting initial statement data. Specifically, the meeting initial statement dataset grouping method can be to randomly select several meeting initial statements as the meeting initial statements. Then, using the randomly selected meeting initial statements as a benchmark, the meeting initial statements with a cosine similarity greater than a certain value (e.g., 0.5) are grouped together with the randomly selected meeting initial statements to obtain the multi-cluster meeting initial statement data.
[0083] Furthermore, it is necessary to extract the first meeting initial statement data from the first cluster of multiple clusters of meeting initial statement data. The first meeting initial statement data can be the randomly selected meeting initial statement data mentioned above as a benchmark. The data center coefficient of the first meeting initial statement data can then be obtained by calculating the mean semantic similarity between the first meeting initial statement data and a preset number of adjacent meeting initial statement data. The preset number of adjacent meeting initial statement data can be a number of data (e.g., 10) with the highest similarity to the first meeting initial statement data. The mean semantic similarity between the first meeting initial statement data and the preset number of adjacent meeting initial statement data is calculated (e.g., the mean can be 0.9) and set as the first meeting initial statement data center coefficient. When the first meeting initial statement data center coefficient is greater than or equal to a center coefficient threshold, the first meeting initial statement data can be set as the center semantics of the first cluster. This threshold can be set based on actual needs. When high accuracy is required for statement data segmentation, the threshold can be appropriately increased, such as 0.9. When low accuracy is required for statement data segmentation, the threshold can be appropriately decreased, such as 0.7.
[0084] Using this method, the first cluster's central semantics, the second cluster's central semantics, and so on can be calculated until the Qth cluster's central semantics are obtained. Q is the total number of clusters of the aforementioned multi-cluster conference initial sentence data. Finally, the first cluster's central semantics through the Qth cluster's central semantics, along with the multi-cluster timbre identification information corresponding to the multi-cluster conference initial sentence data, are added to the sentence list to facilitate correlation analysis.
[0085] In step S300 of this embodiment, the statement list is traversed and correlation analysis is performed with the first topic time zone to obtain a list of related statements, including:
[0086] Extracting a first sentence from the sentence list, wherein the first sentence has a timbre identification set;
[0087] Counting the triggering frequency characteristics of the first statement in the first topic time zone;
[0088] Based on the timbre identification set and in combination with the timbre identity mapping library, a first sentence identity identification set is configured;
[0089] Performing professional-level mode identification on the first statement identity set using a predefined professional-level mapping library of identities belonging to the first topic time zone to obtain professional-level feature values;
[0090] Configuring a first weight for the trigger frequency feature normalization parameter, configuring a second weight for the professional-level feature value, performing weighted mean statistics, and obtaining a first sentence-related feature value;
[0091] When the first sentence association feature value is greater than or equal to an association feature threshold, the first sentence is added to the associated sentence list.
[0092] In an embodiment of the present application, the statement list is traversed to perform a correlation analysis for the first topic time zone. First, the relevant statement must be extracted from the statement list, namely the aforementioned first statement (i.e., the statement corresponding to the first topic time zone). Since the multi-cluster timbre identification information corresponding to the multi-cluster conference initial statement data was added to the statement list in step S200, the first statement has a corresponding timbre identification set. After extracting the first statement, the trigger frequency feature of the first statement in the first topic time zone must be counted. This trigger frequency feature is the frequency of the first statement appearing in the first topic time zone, such as 3 times, 6 times, etc. Next, since the first statement has a corresponding timbre identification set, the timbre identity mapping library (which can achieve a mapping relationship between timbre and identity through the timbre identification network established in step S100) can be combined to configure the first statement identity identification set. Specifically, the first statement is labeled with an identity based on its corresponding timbre identity, such as 007 for Department A employee, 001 for Department B director, etc. Because the reference value of speeches made by different individuals at a conference varies, it is necessary to use a predefined identity-level mapping library for the first topic's time zone to perform a professional-level mode identification on the first statement's identity set to obtain a professional-level feature value. This is used to associate the first statement with its professional level. For example, within the first topic's time zone, the professional-level rating of 007 from Department A is 1, and the professional-level rating of 001 from Department B is 2. Similarly, the professional levels of all attendees within the first topic's time zone are rated accordingly. Then, based on the attendees within the first topic's time zone (identified by the first statement's identity set), a professional-level mode identification is performed (i.e., the mode of the aforementioned professional-level ratings, e.g., the mode of the professional-level ratings can be 1, 2, etc.). The value of the professional-level mode identification corresponding to the first statement is used as the professional-level feature value (e.g., 1, 2, etc.).
[0093] Furthermore, it is necessary to comprehensively consider the trigger frequency feature and the professional level feature to obtain the first sentence association feature value as a basis for determining whether to add the first sentence to the associated sentence list.
[0094] Optionally, the trigger frequency feature needs to be normalized first to facilitate weighted calculation. For example, the weight of the highest frequency statement appearing in the first topic time zone can be defined as 1. The frequency of the first statement appearing in the first topic time zone divided by this highest frequency is the normalized parameter for the trigger frequency feature of the first statement. For example, if the highest frequency statement appearing in the first topic time zone is 10 times, and the frequency of the first statement appearing in the first topic time zone is 7 times, then the normalized parameter for the trigger frequency feature of the first statement is 7 / 10 = 0.7. Then, based on actual needs, weights are assigned to the trigger frequency feature normalization parameter value and the professional-level feature value, namely the first weight and the second weight. The sum of the first weight and the second weight can be 1, such as the first weight = 0.3 and the second weight = 0.7. The ratio of the first weight to the second weight can be adjusted based on the importance of the trigger frequency feature and the professional-level feature. Through weighted calculation, we can obtain the first sentence's associated feature value. For example, if the trigger frequency feature normalization parameter is 0.7, the professional-level feature value is 2, the first weight = 0.3, and the second weight = 0.7, then the corresponding first sentence associated feature value is 0.7*0.3+2*0.7=1.61. When the first sentence associated feature value is greater than or equal to the preset associated feature threshold (such as 1.5), it means that the first sentence meets the criteria for being an associated sentence and can be added to the associated statement list. By adding multiple first sentences in the same way, a list of associated statements can be obtained.
[0095] In step S400 of this embodiment, a plurality of statement lists in a time zone other than the first topic are received, and correlation analysis is performed with the first topic time zone to obtain an expanded correlation statement list, including:
[0096] Extracting a first candidate sentence based on the plurality of sentence lists, and calculating a candidate semantic similarity set between the first candidate sentence and the associated sentence list;
[0097] When any one of the candidate semantic similarity sets is greater than or equal to the expandable semantic similarity threshold, the first candidate sentence is added to the quasi-expanded sentence set;
[0098] For the set of quasi-expanded sentences, calculating a set of professional-level characteristic values of the quasi-expanded sentences;
[0099] Based on the professional-level feature value set of the quasi-expanded statements, quasi-expanded statements whose professional-level feature values are greater than or equal to a professional-level feature threshold are extracted and added to the expansion-related statement list.
[0100] In an embodiment of the present application, since meeting content related to the first topic may be generated outside the first topic time zone in reality, in order to avoid missing content related to the first topic, it is necessary to receive several statement lists in time zones other than the first topic time zone, perform correlation analysis with the first topic time zone respectively, and obtain an expanded related statement list.
[0101] First, it is necessary to extract the first candidate statement from several statement lists that are not in the first topic time zone. The first candidate statement can be a statement containing keywords related to the first topic. For example, if the content discussed in the first topic is sales performance, the first candidate statement can be a statement containing keywords such as "sales" and "performance" generated outside the first topic time zone. Then, the similarity between each first candidate statement and the statements in the associated statement list is calculated to obtain a candidate semantic similarity set. When any one of the candidate semantic similarity sets (such as the similarity between the candidate semantics and a statement in the associated statement list is 0.9) is greater than or equal to the expandable semantic similarity threshold (such as the threshold can be set to 0.7), the first candidate statement can be added to the quasi-expanded statement set as a quasi-expanded statement. The quasi-expanded statement set contains more than one first candidate statement. The calculation of the similarity can be completed by the method described in step S200, which will not be repeated here.
[0102] Next, the professional-level features of the quasi-expansion statements in the set of quasi-expansion statements need to be evaluated to determine whether they are trustworthy (i.e., whether they should be added to the list of expansion-related statements). Optionally, based on the set of professional-level feature values for the quasi-expansion statements, quasi-expansion statements with professional-level feature values greater than or equal to a professional-level feature threshold can be extracted and added to the list of expansion-related statements. For example, the professional-level feature threshold can be set to 2. If a quasi-expansion statement has a professional-level feature value greater than 2, the quasi-expansion statement is added to the list of expansion-related statements.
[0103] In step S500 of this embodiment, setting the associated sentence list and the expanded associated sentence list as the first-topic meeting record annotation includes setting the associated sentence list and the expanded associated sentence list as the first-topic meeting record annotation. Furthermore, for different topics, multiple topic meeting record annotations can be obtained using the same method to complete the meeting data record annotation.
[0104] Example 2, as Figure 2 As shown, based on the same inventive concept as the conference data recording and annotation method for smart conferences provided in Example 1, an embodiment of the present invention further provides a conference data recording and annotation system for smart conferences, including:
[0105] Semantic extraction module 11, used to collect the conference audio information of the first topic time zone through the voice recording component, perform semantic extraction, and obtain the initial meeting sentence data set;
[0106] The semantic clustering module 12 is used to perform semantic clustering on the initial conference sentence data set to obtain multiple clusters of initial conference sentence data, extract central semantics from each cluster, and obtain a sentence list;
[0107] A correlation analysis module 13 is configured to traverse the statement list, perform correlation analysis with the first topic time zone, and obtain a list of related statements;
[0108] An expansion association module 14 is configured to receive a plurality of statement lists in a time zone other than the first topic, and perform correlation analysis with the first topic time zone to obtain an expanded association statement list;
[0109] The record annotation module 15 is configured to set the associated statement list and the expanded associated statement list as first-topic meeting record annotations.
[0110] Furthermore, the semantic extraction module 11 is further configured to:
[0111] Performing text conversion on the conference audio information to obtain first timbre text information to Mth timbre text information;
[0112] In combination with natural language processing technology, traverse the first timbre text information until the Mth timbre text information to perform preprocessing, and obtain the first timbre text basic features until the Mth timbre text basic features, wherein the basic feature attributes include at least word segmentation features, part of speech features, and syntactic features;
[0113] Sending the first timbre text information and the first timbre text basic features to an intelligent annotation channel for processing to obtain key sentences of the first timbre text;
[0114] Until the Mth timbre text information and the Mth timbre text basic features are sent to the intelligent annotation channel for processing, and the Mth timbre text key sentence is obtained;
[0115] Add the first timbre text key sentence to the Mth timbre text key sentence into the conference initial sentence data set.
[0116] The step of converting the conference audio information into text to obtain first to Mth timbre text information includes:
[0117] Retrieving a text conversion model, wherein the text conversion model includes a timbre identification network and a text conversion network;
[0118] Processing the conference audio information through the timbre identification network to obtain first timbre conference audio information through Mth timbre conference audio information, wherein the timbre identification network is constructed by machine learning training using multiple sets of data, each of which includes audio recording data having at least two types of timbre and labels identifying multiple timbre audio recording information;
[0119] Through the text conversion network, the first timbre conference audio information up to the Mth timbre conference audio information are processed respectively, and the first timbre text information up to the Mth timbre text information are output, wherein the text conversion network is constructed by using machine learning training through multiple groups of data, and any group of the multiple groups of data includes audio recording information and a label that identifies the text recording data.
[0120] The first timbre text information and the first timbre text basic features are sent to the intelligent annotation channel for processing to obtain the first timbre text key sentences, including:
[0121] Counting the named entity sets whose trigger frequency is greater than or equal to the trigger frequency threshold in the historical decision dependency data to build a key sentence attribute library, wherein the historical decision dependency data is the key text selected by the user in the historical decision;
[0122] Sending the first timbre text information and the first timbre text basic features to the text segmentation node of the intelligent annotation channel to extract a plurality of initial sentences and a plurality of initial sentence attributes;
[0123] extracting selected initial sentence attributes belonging to the key sentence attribute library from the plurality of initial sentence attributes through a text matching node of the intelligent annotation channel;
[0124] An initial sentence corresponding to the selected initial sentence attribute is extracted from the selected initial sentence attribute and added into the first timbre text key sentence.
[0125] Furthermore, the semantic clustering module 12 is further configured to:
[0126] Performing pairwise semantic similarity analysis on the initial meeting sentence dataset to obtain a number of semantic similarities;
[0127] Based on a predefined semantic similarity threshold and in combination with the plurality of semantic similarities, the conference initial sentence data set is grouped to obtain the multiple clusters of conference initial sentence data, wherein the multiple clusters of conference initial sentence data have multiple clusters of timbre identification information;
[0128] Extracting the first conference initial statement data of the first cluster of conference initial statement data from the plurality of clusters of conference initial statement data;
[0129] Based on the first conference initial statement data and in combination with the plurality of semantic similarities, a preset number of adjacent conference initial statement data are sorted from the first cluster of conference initial statement data, and an average of the semantic similarities between the first conference initial statement data and the preset number of adjacent conference initial statement data is calculated, and the average value is set as the first conference initial statement data center coefficient;
[0130] When the data center coefficient of the first conference initial statement is greater than or equal to the center coefficient threshold, setting the first conference initial statement data as the first cluster center semantics;
[0131] Until the Q-th cluster center semantics is obtained;
[0132] The first cluster central semantics to the Qth cluster central semantics, and the multiple clusters of timbre identification information are added to the sentence list.
[0133] Furthermore, the association analysis module 13 is further configured to:
[0134] Extracting a first sentence from the sentence list, wherein the first sentence has a timbre identification set;
[0135] Counting the triggering frequency characteristics of the first statement in the first topic time zone;
[0136] Based on the timbre identification set and in combination with the timbre identity mapping library, a first sentence identity identification set is configured;
[0137] Performing professional-level mode identification on the first statement identity set using a predefined professional-level mapping library of identities belonging to the first topic time zone to obtain professional-level feature values;
[0138] Configuring a first weight for the trigger frequency feature normalization parameter, configuring a second weight for the professional-level feature value, performing weighted mean statistics, and obtaining a first sentence-related feature value;
[0139] When the first sentence association feature value is greater than or equal to an association feature threshold, the first sentence is added to the associated sentence list.
[0140] Furthermore, the expansion association module 14 is further configured to:
[0141] Extracting a first candidate sentence based on the plurality of sentence lists, and calculating a candidate semantic similarity set between the first candidate sentence and the associated sentence list;
[0142] When any one of the candidate semantic similarity sets is greater than or equal to the expandable semantic similarity threshold, the first candidate sentence is added to the quasi-expanded sentence set;
[0143] For the set of quasi-expanded sentences, calculating a set of professional-level characteristic values of the quasi-expanded sentences;
[0144] Based on the professional-level feature value set of the quasi-expanded statements, quasi-expanded statements whose professional-level feature values are greater than or equal to a professional-level feature threshold are extracted and added to the expansion-related statement list.
[0145] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0146] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0148] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0150] Although preferred embodiments of the present invention have been described, additional changes and modifications to these embodiments may occur to those skilled in the art once the basic inventive concepts become known.
[0151] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A conference data recording and annotation method for intelligent conferences, characterized in that: include: The voice recording component collects the meeting audio information in the first topic time zone, performs semantic extraction, and obtains the initial meeting sentence dataset; Performing semantic clustering on the initial conference sentence data set to obtain multiple clusters of initial conference sentence data, extracting central semantics from each cluster to obtain a sentence list; Traversing the statement list, performing correlation analysis with the first topic time zone, and obtaining a list of related statements; receiving a list of sentences in a time zone other than the first topic, and performing correlation analysis on each of the sentences with the first topic time zone to obtain an expanded list of related sentences; The associated sentence list and the expanded associated sentence list are set as the first topic meeting record mark; Perform semantic clustering on the initial conference statement dataset to obtain multiple clusters of initial conference statement data, extract the central semantics of each cluster, and obtain a statement list, including: Performing pairwise semantic similarity analysis on the initial meeting sentence dataset to obtain a number of semantic similarities; Based on a predefined semantic similarity threshold and in combination with the plurality of semantic similarities, the conference initial sentence data set is grouped to obtain the multiple clusters of conference initial sentence data, wherein the multiple clusters of conference initial sentence data have multiple clusters of timbre identification information; Extracting the first conference initial statement data of the first cluster of conference initial statement data from the plurality of clusters of conference initial statement data; Based on the first conference initial statement data and in combination with the plurality of semantic similarities, a preset number of adjacent conference initial statement data are sorted from the first cluster of conference initial statement data, and an average of the semantic similarities between the first conference initial statement data and the preset number of adjacent conference initial statement data is calculated, and the average value is set as the first conference initial statement data center coefficient; When the data center coefficient of the first conference initial statement is greater than or equal to the center coefficient threshold, setting the first conference initial statement data as the first cluster center semantics; Until the Q-th cluster center semantics are obtained, where Q is the total number of clusters of the aforementioned multi-cluster conference initial sentence data; The first cluster central semantics to the Qth cluster central semantics, and the multiple clusters of timbre identification information are added to the sentence list.
2. The method according to claim 1, wherein The voice recording component collects the meeting audio information of the first topic time zone, performs semantic extraction, and obtains the initial meeting sentence dataset, including: Performing text conversion on the conference audio information to obtain first timbre text information to Mth timbre text information; In combination with natural language processing technology, traverse the first timbre text information until the Mth timbre text information to perform preprocessing, and obtain the first timbre text basic features until the Mth timbre text basic features, wherein the basic feature attributes include at least word segmentation features, part of speech features, and syntactic features; Sending the first timbre text information and the first timbre text basic features to an intelligent annotation channel for processing to obtain key sentences of the first timbre text; Until the Mth timbre text information and the Mth timbre text basic features are sent to the intelligent annotation channel for processing, and the Mth timbre text key sentence is obtained; Add the first timbre text key sentence to the Mth timbre text key sentence into the conference initial sentence data set.
3. The method according to claim 2, wherein Converting the conference audio information into text to obtain first timbre text information to Mth timbre text information includes: Retrieving a text conversion model, wherein the text conversion model includes a timbre identification network and a text conversion network; Processing the conference audio information through the timbre identification network to obtain first timbre conference audio information through Mth timbre conference audio information, wherein the timbre identification network is constructed by machine learning training using multiple sets of data, each of which includes audio recording data having at least two types of timbre and labels identifying multiple timbre audio recording information; Through the text conversion network, the first timbre conference audio information up to the Mth timbre conference audio information are processed respectively, and the first timbre text information up to the Mth timbre text information are output, wherein the text conversion network is constructed by using machine learning training through multiple groups of data, and any group of the multiple groups of data includes audio recording information and a label that identifies the text recording data.
4. The method according to claim 2, wherein Sending the first timbre text information and the first timbre text basic features to the intelligent annotation channel for processing to obtain key sentences of the first timbre text includes: Counting the named entity sets whose trigger frequency is greater than or equal to the trigger frequency threshold in the historical decision dependency data to build a key sentence attribute library, wherein the historical decision dependency data is the key text selected by the user in the historical decision; Sending the first timbre text information and the first timbre text basic features to the text segmentation node of the intelligent annotation channel to extract a plurality of initial sentences and a plurality of initial sentence attributes; extracting selected initial sentence attributes belonging to the key sentence attribute library from the plurality of initial sentence attributes through a text matching node of the intelligent annotation channel; An initial sentence corresponding to the selected initial sentence attribute is extracted from the selected initial sentence attribute and added into the first timbre text key sentence.
5. The method according to claim 1, wherein Traverse the statement list, perform correlation analysis with the first topic time zone, and obtain a list of related statements, including: Extracting a first sentence from the sentence list, wherein the first sentence has a timbre identification set; Counting the triggering frequency characteristics of the first statement in the first topic time zone; Based on the timbre identification set and in combination with the timbre identity mapping library, a first sentence identity identification set is configured; Performing professional-level mode identification on the first statement identity set using a predefined professional-level mapping library of identities belonging to the first topic time zone to obtain professional-level feature values; Configuring a first weight for the trigger frequency feature normalization parameter, configuring a second weight for the professional-level feature value, performing weighted mean statistics, and obtaining a first sentence-related feature value; When the first sentence association feature value is greater than or equal to an association feature threshold, the first sentence is added to the associated sentence list.
6. The method according to claim 1, wherein Receive a list of several statements in a time zone other than the first topic, perform correlation analysis on each statement with the first topic time zone, and obtain an expanded list of related statements, including: Extracting a first candidate sentence based on the plurality of sentence lists, and calculating a candidate semantic similarity set between the first candidate sentence and the associated sentence list; When any one of the candidate semantic similarity sets is greater than or equal to the expandable semantic similarity threshold, adding the first candidate sentence to the quasi-expanded sentence set; For the set of quasi-expanded sentences, calculating a set of professional-level characteristic values of the quasi-expanded sentences; Based on the professional-level feature value set of the quasi-expanded statements, quasi-expanded statements whose professional-level feature values are greater than or equal to a professional-level feature threshold are extracted and added to the expansion-related statement list.
7. A conference data recording and annotation system for intelligent conferences, characterized in that: The system is used to perform the method according to any one of claims 1 to 6, comprising: The semantic extraction module is used to collect the conference audio information of the first topic time zone through the voice recording component, perform semantic extraction, and obtain the initial meeting sentence data set; A semantic clustering module is used to perform semantic clustering on the initial conference statement data set to obtain multiple clusters of initial conference statement data, extract central semantics from each cluster, and obtain a statement list; a correlation analysis module, configured to traverse the statement list, perform correlation analysis with the first topic time zone, and obtain a list of related statements; an expansion association module, configured to receive a plurality of statement lists in a time zone other than the first topic, and perform correlation analysis on each of the statement lists with the first topic time zone to obtain an expansion association statement list; The record annotation module is used to set the associated statement list and the expanded associated statement list as the first topic meeting record annotation.
Citation Information
Patent Citations
Court trial audio-based association degree mining method, apparatus and device, and storage medium
CN116756324A