Conference recording method based on Internet of Things
By using IoT technology to identify speakers and speech wheels in conference records, combining voice transcription and topic segmentation technology, the problem that traditional conference record methods are difficult to balance integrity and readability in high-density discussions among multiple people, and a higher quality and readability meeting record is achieved.
Patent Information
- Application Number
- CN202510257601.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Traditional conference record methods are difficult to take into account completeness and readability in high-density discussion scenarios for multiple people, and it is difficult to correctly identify the true topic of speech, resulting in reduced information relevance and affecting logical coherence and accuracy of meeting minutes.
Using an Internet of Things-based conference record method, the speaker's location is identified through a microphone array and direction sensor, and individual voice characteristics are acquired in combination with a wearable microphone, spokesperson recognition and speech wheel detection are performed, and speech timeline is generated. Then, based on speech transcription and topic segmentation techniques, the theme boundaries are dynamically adjusted, topics that are not discussed continuously are optimized, the theme structure is formed, and summary and task allocation information are extracted.
Improve the accuracy and readability of meeting records, enhance the relevance of topics across timelines, reduce misclassification and information loss, and significantly improve the quality, integrity and usability of meeting records.
Smart Images

Figure CN119993161A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of conference recording, and in particular to a conference recording method based on the Internet of Things. Background Art
[0002] Meeting minutes are extremely important in formal meeting scenarios such as business, academia, and government. Traditional meeting minute-taking methods include handwritten notes, typed text, or voice transcription. These methods are difficult to balance completeness and readability in scenarios with multiple people and high-density discussions. Manual recording relies on the recorder's understanding and speed, and is prone to missing information or filtering out key information. Although voice transcription can retain the complete content, it has too much redundant information and the text is messy and difficult to use directly.
[0003] In order to improve the structured level of records, automated methods based on topic segmentation have been introduced in recent years. Meeting content is converted into text through speech recognition and divided by topic to facilitate summary generation. However, this method assumes that meeting discussions proceed sequentially, which is not the case in real meetings.
[0004] In scenarios such as project review, participants often have multiple rounds of discussions around the same topic, and the discussions are scattered at different stages of the meeting. Traditional time series methods tend to split them into multiple independent topics, resulting in reduced information relevance and affecting logical coherence. In addition, cross-discussions are more complicated, and multiple participants may speak alternately on different topics at the same time. Traditional methods rely on a single voice stream and find it difficult to correctly identify the true topic of the speech. Misclassification often occurs, affecting the accuracy of meeting minutes. Therefore, there is an urgent need for a meeting recording method based on the Internet of Things to solve such problems. Summary of the invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a conference recording method based on the Internet of Things to solve the problems that traditional methods are difficult to cope with cross-discussion, topic fragmentation, serious misclassification, missing or redundant information, and affect the quality of conference records.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] The embodiment of the present invention provides a conference recording method based on the Internet of Things, which includes:
[0009] Step S1, collecting conference voice data;
[0010] Step S2, performing speaker identification based on the voice data, distinguishing the voice streams of different speakers in combination with turn detection technology, and generating a speech timeline to separate cross-speech;
[0011] Step S3, performing speech transcription based on the speech timeline generated in step S2 to generate a transcribed text;
[0012] Step S4, subject segmentation of the transcribed text generated in step S3, and dynamic adjustment based on the speaker position data and individual voice features obtained in step S1;
[0013] Step S5, optimizing and merging the topic segmentation results generated in step S4, integrating non-continuous discussions on the same topic to form a topic structure;
[0014] Step S6, based on the topic structure, perform summary extraction, combine decision point detection to identify key conclusions, and perform structured processing on the task assignment content, extract and generate task assignment information, form a to-do list, and finally form a meeting record.
[0015] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, in step S1, during the voice data collection process:
[0016] The microphone array and direction sensor are used to detect the speaker's position, and the wearable microphone is used to obtain individual voice features for identity binding.
[0017] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the direction sensor includes ultrasonic and infrared cameras.
[0018] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the steps of using a microphone array and a direction sensor to detect the speaker's position and combining a wearable microphone to obtain individual voice features and perform identity binding are as follows:
[0019] Set up multiple microphone arrays at fixed positions, use signal arrival time difference and beamforming technology to estimate the direction of the speech source in the environment.
[0020] Combine ultrasonic and infrared cameras to obtain the angle and position coordinates of the speaker and calculate its position relative to the microphone array.
[0021] Through data fusion technology, the acoustic information of the speech source is matched with the position information of the direction sensor to determine the speaker's position.
[0022] Collect high signal-to-noise ratio speech data from wearable microphones and extract individual features, including the speaker's timbre, voice texture, speaking speed, and pronunciation habits.
[0023] Calculate the similarity of speech features between the ambient microphone data and the wearable microphone data, bind the individual identity, and record the identity information.
[0024] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the speaker identification is performed based on the voice data, the voice streams of different speakers are distinguished by the turn detection technology, and the speech timeline is generated, and the step of separating the cross-speech is as follows:
[0025] The original speech signal x(t) is denoised and feature enhanced. The denoising formula is:
[0026] x′(t)=x(t)-n(t),
[0027] Among them, x'(t) represents the denoised speech signal, x(t) represents the original speech signal, and n(t) represents the estimated background noise.
[0028] Mel frequency cepstral coefficient MFCC is used for feature extraction, and the extraction formula is:
[0029] f i =MFCC(x′ i ),
[0030] Among them, f i represents the feature vector of the i-th frame, x' i represents the denoised i-th frame of speech data, MFCC(·) represents the Mel frequency cepstral coefficient calculation function,
[0031] Gaussian mixture model GMM or deep learning model is used for clustering classification. The classification formula is:
[0032]
[0033] Among them, P(s|f) represents the probability that the speech feature belongs to a certain speaker, M represents the number of Gaussian distributions, and w j is the weight of the j-th Gaussian distribution, Indicates that the mean is μ j , the covariance is Σ j Gaussian distribution,
[0034] The short-time energy E is used to calculate the voice activity. The calculation formula is:
[0035]
[0036] Among them, E t is the short-term energy at time t, N is the window size, x i is the speech sampling value in the time window,
[0037] By threshold θ E Perform turn-segmentation:
[0038] If E t >θE , then VAD(t)=1,
[0039] If E t ≤θ E , then VAD(t)=0,
[0040] Where VAD(t) is a binary function of voice activity detection, 1 indicates speech, and 0 indicates silence.
[0041] The Blind Source Separation (BSS) method is used to separate multiple voices. The separation method is expressed as:
[0042] Y=WX,
[0043] Among them, X is the observation signal matrix, W is the separation matrix, and Y is the separated independent speech signal.
[0044] As a preferred solution of the method for recording a conference based on the Internet of Things described in the present invention, in which: during the text transcription process of step S3, noise reduction processing is used to remove background noise;
[0045] Semantic compensation is used in combination with IoT devices to collect projection and whiteboard writing data for text correction.
[0046] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the dynamic adjustment of step S4 includes:
[0047] If individual speech features change and position data changes are detected, topic boundaries are automatically inserted;
[0048] If the projection switch in step S3 is detected, the theme switch is automatically inferred.
[0049] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the step of subject-segmenting the transcribed text generated in step S3 and dynamically adjusting the speaker position data and individual voice features obtained in step S1 is as follows:
[0050] First, the transcribed text generated in step S3 is semantically vectorized, and the text is converted into a high-dimensional vector v using the word embedding method. i :
[0051] v i =φ(T i ),
[0052] Among them, v i is the semantic vector of the i-th sentence, T i is the text content of the i-th sentence, φ(·) is the word embedding conversion function,
[0053] The cosine similarity is used to calculate the semantic similarity of adjacent sentences. The formula is:
[0054]
[0055] Among them, S i,j is the semantic similarity between the i-th sentence and the j-th sentence, v i and v j Represent the semantic vectors of the two sentences respectively, |·| is the Euclidean norm of the vector,
[0056] Use clustering algorithm to detect topic boundaries and define topic boundary function:
[0057] If S k,k+1 <θ S , then B k =1,
[0058] If S k,k+1 ≥θ S , then B k =0,
[0059] Among them, B k is the topic boundary marker, θ S is the similarity threshold. If the similarity between two sentences is lower than the threshold, it is marked as a new topic boundary.
[0060] Automatically insert topic boundaries if changes in individual speech characteristics are detected and speaker position data changes:
[0061] If (Δv k >θ v )∧(Δp k >θ p ), then B' k =1, otherwise B' k =B k ,
[0062] Among them, B' k is the adjusted topic boundary, Δv k is the individual speech feature change value, θ v is the speech feature change threshold, Δp k is the change value of position data, θ p is the position change threshold,
[0063] If the content of the projection device is detected to be switched, the subject boundary is automatically inserted:
[0064] If C k -C k-1 ≠0, then B″ k =1, otherwise B″ k =B' k
[0065] Among them, B″ k is the final adjusted subject boundary, C k is the content identifier of the projection device at time k.
[0066] As a preferred solution of the conference recording method based on the Internet of Things described in the present invention, the step of optimizing and merging the topic segmentation results generated in step S4, integrating non-continuous discussions on the same topic, and forming a topic structure is as follows:
[0067] For the topic fragments generated in step S4 and Calculate the semantic similarity between fragments:
[0068]
[0069] in, For the theme and The similarity between i ,v j For the subject and The sentence semantic vector of Respectively indicate the subject and is the number of inner sentences, |·| is the Euclidean norm,
[0070] If the similarity between two topics is higher than the threshold θ M , then merge the two topics:
[0071] like but
[0072] like but
[0073] in, is the topic merge tag, θ M is the similarity threshold for topic merging,
[0074] If the subject and If there is a long period of non-related content between them, the interval is calculated using the formula:
[0075]
[0076] in, Indicates the subject and The shortest interval between m,n is the set of sentence indexes between two topics,
[0077] like To allow topic merging:
[0078] like but otherwise
[0079] Among them, θ G is the maximum interval threshold,
[0080] After merging, the topics are renumbered and built into a tree structure:
[0081] H k =Tree(M),
[0082] Among them, H k is the reconstructed topic hierarchy, and Tree(·) is a function for hierarchical topic classification based on the merge matrix M.
[0083] As a preferred solution of the method for recording a meeting based on the Internet of Things described in the present invention, the steps of identifying key conclusions by combining decision point detection, and performing structured processing on task assignment content, extracting and generating task assignment information are as follows:
[0084] Based on the topic hierarchy generated in step S5, calculate the representative sentences of each topic.
[0085] Use term frequency-inverse document frequency (TF-IDF) or bidirectional transformer encoding (BERT) to extract the core sentences of the topic, and select the sentence with the highest importance score as the summary sentence.
[0086] For each topic, select the previous sentence with the largest importance score as the topic summary;
[0087] Identify decision-making sentences through keyword matching and dependency analysis, such as sentences containing keywords such as "decision", "plan", "goal", and "plan".
[0088] Combined with syntactic analysis, sentences containing active decision-making verbs are selected and matched with speaker identities;
[0089] Identify sentences containing task assignment information, focusing on sentences containing task-related terms such as "responsible", "executed", and "completion time".
[0090] Combined with the conference speaker information, the task executor is analyzed and matched to the specific speaker identity.
[0091] Extract key attributes of tasks, including: task executor, task content, task deadline, task priority, and task-related resources;
[0092] Format and store tasks according to their attributes to generate a structured to-do list: task ID, task description, person in charge, deadline, and task status;
[0093] Summarize the topic summary, decision points and task allocation information, and classify them according to chronological order and subject to form the final meeting record document.
[0094] Based on the speaker’s identity information and timeline data, organize each speaker’s views and present them in a structured manner.
[0095] The beneficial effects of the present invention are as follows: the present invention uses a microphone array and a direction sensor to locate the speaker, and combines a wearable microphone to obtain individual voice features, matches the voice with the speaker identity, and in the speaker identification process, combines the turn detection technology to distinguish the voice streams of different speakers, generates a speaking timeline, and for speech transcription, removes background noise through noise reduction technology, introduces semantic compensation, and corrects the text in combination with the projection and whiteboard writing data collected by the IoT device to improve recognition accuracy.
[0096] The present invention additionally performs topic segmentation, adopts a semantic vectorization method to convert text into a high-dimensional vector, and uses similarity calculation and clustering algorithm to determine the topic boundary. At the same time, it combines IoT location data with individual voice features for dynamic adjustment, accurately classifies non-continuous discussions on the same topic, and prevents cross-speech on different topics from being mistakenly classified as one. At the same time, it calculates the semantic similarity between topics, integrates non-continuous discussions on the same topic, and uses a maximum interval threshold to ensure that the topic merging is reasonable, and finally constructs a tree structure to improve the hierarchy of topic organization.
[0097] The present invention combines the TF-IDF or BERT model to extract meeting summaries, uses keyword matching and grammatical analysis to detect decision-making sentences, keeps the core content of the meeting, and effectively parses task assignment content, extracts key elements such as task executors, content, deadlines, etc., and finally presents them in the form of a structured to-do list. The meeting minutes have higher readability and traceability.
[0098] In summary, the present invention not only optimizes the accuracy of traditional meeting records in multi-speaker discussions and improves the relevance of topics across timelines, but also effectively reduces misclassification and information loss, thereby greatly improving the quality, completeness and usability of meeting records. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0100] Figure 1 It is a flow chart of the conference recording method based on the Internet of Things of the present invention. DETAILED DESCRIPTION
[0101] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0102] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0103] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0104] Example 1, reference Figure 1 , this embodiment provides a conference recording method based on the Internet of Things, comprising the following steps:
[0105] Step S1, collecting conference voice data;
[0106] In step S1, during the voice data collection process:
[0107] The speaker's position is detected using a microphone array and directional sensor, and individual voice features are acquired using a wearable microphone for identity binding.
[0108] Direction sensors include ultrasonic and infrared cameras;
[0109] The microphone array and direction sensor are used to detect the speaker's position, and the wearable microphone is used to obtain individual voice features. The steps for identity binding are:
[0110] Set up multiple microphone arrays at fixed positions, use signal arrival time difference and beamforming technology to estimate the direction of the speech source in the environment.
[0111] Combine ultrasonic and infrared cameras to obtain the angle and position coordinates of the speaker and calculate its position relative to the microphone array.
[0112] Through data fusion technology, the acoustic information of the speech source is matched with the position information of the direction sensor to determine the speaker's position.
[0113] Collect high signal-to-noise ratio speech data from wearable microphones and extract individual features, including the speaker's timbre, voice texture, speaking speed, and pronunciation habits.
[0114] Calculate the similarity of speech features between the ambient microphone data and the wearable microphone data, bind the individual identity, and record the identity information;
[0115] Step S2, performing speaker identification based on the voice data, distinguishing the voice streams of different speakers in combination with turn detection technology, and generating a speech timeline to separate cross-speech;
[0116] Based on the voice data, speaker identification is performed, and the speech flow of different speakers is distinguished by the turn detection technology, and the speech timeline is generated. The steps for separating the cross-speech are as follows:
[0117] The original speech signal x(t) is denoised and feature enhanced. The denoising formula is:
[0118] x′(t)=x(t)-n(t),
[0119] Among them, x'(t) represents the denoised speech signal, x(t) represents the original speech signal, and n(t) represents the estimated background noise.
[0120] Mel frequency cepstral coefficient MFCC is used for feature extraction, and the extraction formula is:
[0121] f i =MFCC(x′ i ),
[0122] Among them, f i represents the feature vector of the i-th frame, x' i represents the denoised i-th frame of speech data, MFCC(·) represents the Mel frequency cepstral coefficient calculation function,
[0123] Gaussian mixture model GMM or deep learning model is used for clustering classification. The classification formula is:
[0124]
[0125] Among them, P(s|f) represents the probability that the speech feature belongs to a certain speaker, M represents the number of Gaussian distributions, and w j is the weight of the j-th Gaussian distribution, Indicates that the mean is μ j , the covariance is Σ j Gaussian distribution,
[0126] The short-time energy E is used to calculate the voice activity. The calculation formula is:
[0127]
[0128] Among them, E t is the short-term energy at time t, N is the window size, x i is the speech sampling value in the time window,
[0129] By threshold θ E Perform turn-segmentation:
[0130] If E t >θ E , then VAD(t)=1,
[0131] If E t ≤θ E , then VAD(t)=0,
[0132] Where VAD(t) is a binary function of voice activity detection, 1 indicates speech, and 0 indicates silence.
[0133] The Blind Source Separation (BSS) method is used to separate multiple voices. The separation method is expressed as:
[0134] Y=WX,
[0135] Among them, X is the observation signal matrix, W is the separation matrix, and Y is the independent speech signal after separation.
[0136] Specifically, speech signal denoising, feature extraction, speaker clustering identification, turn detection and blind source separation techniques are used to distinguish the speech streams of different speakers and generate a speech timeline;
[0137] The background noise estimation method is used to denoise the signal and improve the speech clarity. Then, MFCC is used to extract the speaker features, and the speaker classification and identification is performed through GMM or deep learning model. Turn detection detects speech activity through short-time energy calculation to identify the time interval of each speaker's speech. Finally, blind source separation technology is used to separate the cross-speech and accurately construct the speech timeline.
[0138] Step S3, performing speech transcription based on the speech timeline generated in step S2 to generate a transcribed text;
[0139] During the text transcription process of step S3, noise reduction processing is used to remove background noise;
[0140] And semantic compensation is used in combination with IoT devices to collect projection and whiteboard writing data for text correction;
[0141] Step S4, subject segmentation of the transcribed text generated in step S3, and dynamic adjustment based on the speaker position data and individual voice features obtained in step S1;
[0142] The dynamic adjustment of step S4 includes:
[0143] If individual speech features change and position data changes are detected, topic boundaries are automatically inserted;
[0144] If the projection switch in step S3 is detected, the theme switch is automatically inferred;
[0145] The steps of subject segmenting the transcribed text generated in step S3 and dynamically adjusting it in combination with the speaker position data and individual voice features obtained in step S1 are as follows:
[0146] First, the transcribed text generated in step S3 is semantically vectorized, and the text is converted into a high-dimensional vector v using the word embedding method. i :
[0147] v i =φ(T i ),
[0148] Among them, v i is the semantic vector of the i-th sentence, T i is the text content of the i-th sentence, φ(·) is the word embedding conversion function,
[0149] The cosine similarity is used to calculate the semantic similarity of adjacent sentences. The formula is:
[0150]
[0151] Among them, S i,j is the semantic similarity between the i-th sentence and the j-th sentence, v i and v j Represent the semantic vectors of the two sentences respectively, |·| is the Euclidean norm of the vector,
[0152] Use clustering algorithm to detect topic boundaries and define topic boundary function:
[0153] If S k,k+1 <θ S , then B k =1,
[0154] If S k,k+1 ≥θ S , then B k =0,
[0155] Among them, B k is the topic boundary marker, θ S is the similarity threshold. If the similarity between two sentences is lower than the threshold, it is marked as a new topic boundary.
[0156] Automatically insert topic boundaries if changes in individual speech characteristics are detected and speaker position data changes:
[0157] If (Δv k >θ v )∧(Δp k >θ p , then B' k =1, otherwise B' k =B k ,
[0158] Among them, B' k is the adjusted topic boundary, Δv k is the individual speech feature change value, θ v is the speech feature change threshold, Δp k is the change value of position data, θ p is the position change threshold,
[0159] If the content of the projection device is detected to be switched, the subject boundary is automatically inserted:
[0160] If C k -C k-1 ≠0, then B″ k =1, otherwise B″ k =B' k
[0161] Among them, B″ k is the final adjusted subject boundary, C k is the content identifier of the projection device at time k;
[0162] Specifically, the text is processed by semantic vectorization method, the semantic similarity of adjacent sentences is calculated, and the threshold is used to judge the topic boundary. Dynamic adjustment is made based on the speaker position and individual voice feature changes to improve the accuracy of topic segmentation. In addition, the topic division is further optimized by combining the IoT device to detect the projection switching content.
[0163] Step S5, optimizing and merging the topic segmentation results generated in step S4, integrating non-continuous discussions on the same topic to form a topic structure;
[0164] The steps of optimizing and merging the topic segmentation results generated in step S4 and integrating the discontinuous discussions on the same topic to form a topic structure are as follows:
[0165] For the topic fragments generated in step S4 and Calculate the semantic similarity between fragments:
[0166]
[0167] in, For the theme and The similarity between i ,v j For the subject and The sentence semantic vector of Respectively indicate the subject and is the number of inner sentences, |·| is the Euclidean norm,
[0168] If the similarity between two topics is higher than the threshold θ M , then merge the two topics:
[0169] like but
[0170] like but
[0171] in, is the topic merge tag, θ M is the similarity threshold for topic merging,
[0172] If the subject and If there is a long period of non-related content between them, the interval is calculated using the formula:
[0173]
[0174] in, Indicates the subject and The shortest interval between m,n is the set of sentence indexes between two topics,
[0175] like To allow topic merging:
[0176] like but otherwise
[0177] Among them, θ G is the maximum interval threshold,
[0178] After merging, the topics are renumbered and built into a tree structure:
[0179] H k =Tree(M),
[0180] Among them, H k is the reconstructed topic hierarchy, Tree(·) is the function for hierarchical topic classification based on the merge matrix M;
[0181] Specifically, in step S5, similar topics are determined by calculating the semantic similarity between topics, and topic merging is performed based on the semantic similarity threshold. At the same time, in order to ensure that non-continuous discussion topics can be correctly merged, the shortest interval between topics is calculated, and the merging judgment is made in combination with the maximum interval threshold. Finally, the merged topic hierarchy is constructed through a tree structure, making the topic organization more structured.
[0182] Step S6: Based on the topic structure, perform summary extraction, combine decision point detection to identify key conclusions, and perform structured processing on the task assignment content, extract and generate task assignment information, form a to-do list, and finally form a meeting record;
[0183] Combined with decision point detection to identify key conclusions, and structurally process the task assignment content, the steps of extracting and generating task assignment information are as follows:
[0184] Based on the topic hierarchy H generated in step S5 k , calculate the representative sentences for each topic,
[0185] Use term frequency-inverse document frequency (TF-IDF) or bidirectional transformer encoding (BERT) to extract the core sentences of the topic, and select the sentence with the highest importance score as the summary sentence.
[0186] For each topic Select the first N sentences with the largest importance scores as the topic summary;
[0187] Identify decision-making sentences through keyword matching and dependency analysis, such as sentences containing keywords such as "decision", "plan", "goal", and "plan".
[0188] Combined with syntactic analysis, sentences containing active decision-making verbs are selected and matched with speaker identities;
[0189] Identify sentences containing task assignment information, focusing on sentences containing task-related terms such as "responsible", "executed", and "completion time".
[0190] Combined with the conference speaker information, the task executor is analyzed and matched to the specific speaker identity.
[0191] Extract key attributes of tasks, including: task executor, task content, task deadline, task priority, and task-related resources;
[0192] Format and store tasks according to their attributes to generate a structured to-do list: task ID, task description, person in charge, deadline, and task status;
[0193] Summarize the topic summary, decision points and task allocation information, and classify them according to chronological order and subject to form the final meeting record document.
[0194] Organize the opinions of each speaker based on the speaker's identity information and timeline data, and present them in a structured manner;
[0195] Specifically, in step S6, based on the topic structure generated in step S5, key content is further extracted to form a meeting summary and task assignment information. The core sentences are screened through TF-IDF or deep learning methods to achieve topic summary extraction. At the same time, decision point detection technology is used to extract key decision content in the meeting, and task assignment information is parsed in combination with grammatical analysis. After the task data is structured, a to-do list is generated, which summarizes all information to form a complete meeting record, making the archiving and tracking of meeting content more efficient.
[0196] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A conference recording method based on the Internet of Things, characterized in that: include, Step S1, collecting conference voice data; Step S2, performing speaker identification based on the voice data, distinguishing the voice streams of different speakers in combination with turn detection technology, and generating a speech timeline to separate cross-speech; Step S3, performing speech transcription based on the speech timeline generated in step S2 to generate a transcribed text; Step S4, subject segmentation of the transcribed text generated in step S3, and dynamic adjustment based on the speaker position data and individual voice features obtained in step S1; Step S5, optimizing and merging the topic segmentation results generated in step S4, integrating non-continuous discussions on the same topic to form a topic structure; Step S6, based on the topic structure, perform summary extraction, combine decision point detection to identify key conclusions, and perform structured processing on the task assignment content, extract and generate task assignment information, form a to-do list, and finally form a meeting record.
2. A conference recording method based on the Internet of Things as claimed in claim 1, characterized in that: In step S1, during the voice data collection process: The microphone array and direction sensor are used to detect the speaker's position, and the wearable microphone is used to obtain individual voice features for identity binding.
3. A conference recording method based on the Internet of Things as claimed in claim 2, characterized in that: The direction sensors include ultrasonic and infrared cameras.
4. A conference recording method based on the Internet of Things as claimed in claim 3, characterized in that: The steps of using a microphone array and a direction sensor to detect the speaker's position and combining a wearable microphone to obtain individual voice features and perform identity binding are: Set up multiple microphone arrays at fixed positions, use signal arrival time difference and beamforming technology to estimate the direction of the speech source in the environment. Combine ultrasonic and infrared cameras to obtain the angle and position coordinates of the speaker and calculate its position relative to the microphone array. Through data fusion technology, the acoustic information of the speech source is matched with the position information of the direction sensor to determine the speaker's position. Collect high signal-to-noise ratio speech data from wearable microphones and extract individual features, including the speaker's timbre, voice texture, speaking speed, and pronunciation habits. Calculate the similarity of speech features between the ambient microphone data and the wearable microphone data, bind the individual identity, and record the identity information.
5. A conference recording method based on the Internet of Things as claimed in claim 4, characterized in that: The steps of performing speaker identification based on voice data, distinguishing the voice streams of different speakers in combination with turn detection technology, and generating a speech timeline to separate cross-speech are as follows: The original speech signal x(t) is denoised and feature enhanced. The denoising formula is: x′(t)=x(t)-n(t), Among them, x'(t) represents the denoised speech signal, x(t) represents the original speech signal, and n(t) represents the estimated background noise. Mel frequency cepstral coefficient MFCC is used for feature extraction, and the extraction formula is: in i =MFCC(x′ i ), Among them, f i represents the feature vector of the i-th frame, x' i represents the denoised i-th frame of speech data, MFCC(·) represents the Mel frequency cepstral coefficient calculation function, Gaussian mixture model GMM or deep learning model is used for clustering classification. The classification formula is: Among them, P(s|f) represents the probability that the speech feature belongs to a certain speaker, M represents the number of Gaussian distributions, and w j is the weight of the j-th Gaussian distribution, Indicates that the mean is μ j , the covariance is Σ j Gaussian distribution, The short-time energy E is used to calculate the voice activity. The calculation formula is: Among them, E t is the short-term energy at time t, N is the window size, x i is the speech sampling value in the time window, By threshold θ E Perform turn-segmentation: If E t >θ E , then VAD(t)=1, If E t ≤θ E , then VAD(t)=0, Where VAD(t) is a binary function of voice activity detection, 1 indicates speech, and 0 indicates silence. The Blind Source Separation (BSS) method is used to separate multiple voices. The separation method is expressed as: Y=WX, Among them, X is the observation signal matrix, W is the separation matrix, and Y is the separated independent speech signal.
6. A conference recording method based on the Internet of Things as claimed in claim 5, characterized in that: During the text transcription process of step S3, noise reduction processing is used to remove background noise; Semantic compensation is used in combination with IoT devices to collect projection and whiteboard writing data for text correction.
7. A conference recording method based on the Internet of Things as claimed in claim 6, characterized in that: The dynamic adjustment of step S4 includes: If individual speech features change and position data changes are detected, topic boundaries are automatically inserted; If the projection switch in step S3 is detected, the theme switch is automatically inferred.
8. A conference recording method based on the Internet of Things as claimed in claim 7, characterized in that: The step of subject-segmenting the transcribed text generated in step S3 and dynamically adjusting it in combination with the speaker position data and individual voice features obtained in step S1 is as follows: First, the transcribed text generated in step S3 is semantically vectorized, and the text is converted into a high-dimensional vector v using the word embedding method. i : v i =φ(T i ), Among them, v i is the semantic vector of the i-th sentence, T i is the text content of the i-th sentence, φ(·) is the word embedding conversion function, The cosine similarity is used to calculate the semantic similarity of adjacent sentences. The formula is: Among them, S i,j is the semantic similarity between the i-th sentence and the j-th sentence, v i and v j Represent the semantic vectors of the two sentences respectively, |·| is the Euclidean norm of the vector, Use clustering algorithm to detect topic boundaries and define topic boundary function: If S k,k+1 <θ S , then B k =1, If S k,k+1 ≥θ S , then B k =0, Among them, B k is the topic boundary marker, θ S is the similarity threshold. If the similarity between two sentences is lower than the threshold, it is marked as a new topic boundary. Automatically insert topic boundaries if changes in individual speech characteristics are detected and speaker position data changes: If (Δv k >θ v )∧(Δp k >θ p ), then B' k =1, otherwise B' k =B k , Among them, B' k is the adjusted topic boundary, Δv k is the individual speech feature change value, θ v is the speech feature change threshold, Δp k is the change value of position data, θ p is the position change threshold, If the content of the projection device is detected to be switched, the subject boundary is automatically inserted: If C k -C k-1 ≠0, then B" k =1, otherwise B" k =B' k Among them, B k is the final adjusted subject boundary, C k is the content identifier of the projection device at time k.
9. A conference recording method based on the Internet of Things as claimed in claim 8, characterized in that: The step of optimizing and merging the topic segmentation results generated in step S4 and integrating the discontinuous discussions on the same topic to form a topic structure is as follows: For the topic fragments generated in step S4 and Calculate the semantic similarity between fragments: in, For the theme and The similarity between i ,v j For the subject and The sentence semantic vector of Respectively indicate the subject and is the number of inner sentences, |·| is the Euclidean norm, If the similarity between two topics is higher than the threshold θ M , then merge the two topics: like but like but in, is the topic merge tag, θ M is the similarity threshold for topic merging, If the subject and If there is a long period of non-related content between them, the interval is calculated using the formula: in, Indicates the subject and The shortest interval between m,n is the set of sentence indexes between two topics, like To allow topic merging: like but otherwise Among them, θ G is the maximum interval threshold, After merging, the topics are renumbered and built into a tree structure: H k =Tree(M), Among them, H k is the reconstructed topic hierarchy, and Tree(·) is a function for hierarchical topic classification based on the merge matrix M.
10. A conference recording method based on the Internet of Things as claimed in claim 9, characterized in that: The steps of combining decision point detection to identify key conclusions, and structurally processing the task assignment content, extracting and generating task assignment information are as follows: Based on the topic hierarchy H generated in step S5 k , calculate the representative sentences for each topic, Use term frequency-inverse document frequency (TF-IDF) or bidirectional transformer encoding (BERT) to extract the core sentences of the topic, and select the sentence with the highest importance score as the summary sentence. For each topic Select the first N sentences with the largest importance scores as the topic summary; Identify decision-making sentences through keyword matching and dependency analysis. Combined with syntactic analysis, sentences containing active decision-making verbs are selected and matched with speaker identities; Identify sentences that contain task assignment information, Combined with the conference speaker information, the task executor is analyzed and matched to the specific speaker identity. Extract key attributes of tasks, including: task executor, task content, task deadline, task priority, and task-related resources; Format and store tasks according to their attributes to generate a structured to-do list: task ID, task description, person in charge, deadline, and task status; Summarize the topic summary, decision points and task allocation information, and classify them according to chronological order and subject to form the final meeting record document. Based on the speaker’s identity information and timeline data, organize each speaker’s views and present them in a structured manner.
Citation Information
Patent Citations
Conference video shooting, recording, and propagating method, conference video shooting, recording, and propagating device, and server
CN108366216A
Meeting record generation method, apparatus, apparatus, and computer storage medium
CN109388701A
Conference record generation method based on voice recognition, device and storage medium
CN110335612A
Speech recognition text coherence processing method and device
CN111832308A
Paperless conference record archiving method and system
CN116911817A
Cited By
Automatic conference recording and abstract generating method for intelligent conference
CN120340497A
Voice separation and paragraph affiliation method and system for teleconference scene
CN120808757A
Speech separation and paragraph attribution method and system for teleconference scenarios
CN120808757B
Intelligent voice conference behavior analysis method and system based on multi-source perception
CN121462328A
Speech recognition method and device, electronic equipment and readable storage medium
CN122050363A