Internet online conference summary generation method and electronic equipment

By adjusting the encoding bit rate in real time and optimizing the server architecture, the problem of speech recognition errors caused by network fluctuations in online meetings was solved, and high-quality meeting minutes were generated.

CN120600028APending Publication Date: 2025-09-05CHONGQING CHENGXI DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510940579.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

During online meetings, network fluctuations can cause large deviations in the text content output by the speech recognition model, affecting the accuracy of meeting minutes.

Method used

By monitoring network bandwidth fluctuations in real time, dynamically adjusting the encoding bit rate, distinguishing between high-value and ordinary voice signal segments, and making data supplement requests on the server side to ensure the integrity and accuracy of voice signal data, combined with dynamic adjustment of the server architecture to optimize processing efficiency.

Benefits of technology

Under network fluctuations, the accuracy of the text output by the speech recognition model is improved, the errors in context and key content are reduced, the server load and cost are optimized, and the overall quality of meeting minutes is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600028A_ABST
    Figure CN120600028A_ABST
Patent Text Reader

Abstract

The invention relates to the field of online conference technology, in particular to an Internet online conference summary generation method, which comprises the following steps: acquiring and obtaining voice signal data; network bandwidth fluctuation is monitored in real time, and the coding bit rate is dynamically adjusted; transmitting the voice signal data to a voice recognition engine to generate a first text; correcting the first text according to a historical correction mapping relation and generating a second text; storing the second text in the shared document in real time to form a conference summary; in the embodiment of the invention, the shared link is generated for the online conference room, so that the multiple terminals join the online conference to consult and edit the conference summary, the transmission quality of the voice signal data is ensured, and the text content accuracy of the conference summary is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of online conference technology, and in particular to a method and electronic device for generating minutes of an Internet online conference. Background Art

[0002] Online meetings are not restricted by location or space, allowing participants from different regions to hold meetings in real time. With the rapid development of the internet, they are widely used by businesses. Furthermore, online meetings can automatically record participant speeches through voice recognition to form meeting minutes. However, this relies on the accuracy of voice recognition. If the speaker's accent is not standard or the voice recognition model misidentifies certain specialized terms, the resulting text cannot be used directly and requires manual adjustment.

[0003] To this end, some online conferences will combine the mapping relationships formed by manual correction records of some vocabulary during historical conferences to automatically correct the text obtained by speech recognition, so as to reduce the workload of repeatedly correcting the same vocabulary and improve the accuracy of speech-to-text and translation. For example, the online conference method, device, storage medium and electronic device with publication number CN116366800A.

[0004] Regarding the above-mentioned related technologies, once the network fluctuates during an online meeting, the voice signal obtained by the speech recognition engine will fluctuate, resulting in a large deviation in the text content converted and output by the speech recognition model. Summary of the Invention

[0005] In order to ensure that the text output by the speech recognition model can maintain a high accuracy when network fluctuations occur, the present application provides an Internet online meeting minutes generation method and electronic device.

[0006] In a first aspect, the present application provides a method for generating online conference minutes using the following technical solution.

[0007] A method for generating online conference minutes on the Internet specifically includes the following steps.

[0008] Collect and obtain voice signal data;

[0009] Monitor network bandwidth fluctuations in real time and dynamically adjust encoding bit rate;

[0010] transmitting the speech signal data to a speech recognition engine to generate a first text;

[0011] Correcting the first text with reference to the historical correction mapping relationship and generating a second text;

[0012] storing the second text in a shared document in real time to form meeting minutes;

[0013] Generate a shared link for the online meeting room so that multiple terminals can join the online meeting to view and edit the meeting minutes.

[0014] By adopting the above technical solution, the encoding bit rate of the voice signal data is adjusted according to real-time network fluctuations to reduce the impact on the voice signal data transmission quality, thereby making it less likely for the subsequent voice recognition engine to output text with large errors.

[0015] Optionally, the formula for dynamically adjusting the encoding bit rate is:

[0016]

[0017] in, is the encoding bit rate adjusted in real time, is the initial bit rate, is the packet loss rate threshold, is the current packet loss rate, ∈[0,1], To prevent zero minimum, for the Set the minimum threshold .

[0018] Optionally, the 64kbps or 128kbps, is 0.05 or 0.03, is 0.0001, 16kbps or 32kbps.

[0019] By adopting the above technical solution, the coding bit rate of the voice signal data is adjusted in real time according to the packet loss rate, and the coding bit rate is ensured not to be too small, so as to ensure that the voice signal data is not easily distorted during transmission.

[0020] Optionally, during the online conference, the number of access terminals is counted in real time and the load factor is calculated, and the server architecture is dynamically replaced according to the load factor.

[0021] By adopting the above technical solution, a suitable server architecture is selected according to the number of online conference access terminals, so that the server load is not too large, the processing efficiency of voice signal data in the server is effectively improved, and the server usage cost can be balanced.

[0022] Optionally, the load factor L is calculated as follows:

[0023]

[0024] Among them, N is the number of online conference access terminals, The maximum load capacity of a single server. is the average terminal network bandwidth, The baseline network bandwidth.

[0025] By adopting the above technical solution, the server architecture is adjusted in real time in combination with the network bandwidth of each terminal and the number of involved terminals, so as to make a more precise dynamic adjustment to the load factor, so that the adjustment and replacement of the server architecture is more accurate.

[0026] Optionally, the server architecture dynamic replacement condition is:

[0027] When L < 0.3, the single-server mode is adopted;

[0028] When 0.3≤L<0.7, the regional server cluster mode is adopted;

[0029] When L ≥ 0.7, the cloud service automatic expansion mode is adopted.

[0030] By adopting the above technical solution, three different server architectures are dynamically replaced according to the load factor, so as to meet the high-efficiency processing of online conference voice signal data while balancing server costs.

[0031] Optionally, after the voice signal data is acquired, the terminal performs local semantic prediction to distinguish high-value segments from ordinary segments, performs enhanced transmission on the high-value segments and basic transmission on the ordinary segments, and then the server receives the high-value segments and ordinary segments and performs complete voice recognition.

[0032] By adopting the above-mentioned technical solution, the terminal pre-processes and distinguishes the voice signal data, and transmits high-value segments and ordinary segments using different encoding bit rates, so as to ensure that quality problems are not likely to occur in the transmission of important content in the voice signal data, and make it less likely that there will be large errors in the key content of the first text output by the final voice recognition engine, thereby improving the accuracy of the final meeting minutes.

[0033] Optionally, when the timestamps of the high-value segments and the ordinary segments are discontinuous, the server sends a data supplement request to the terminal, and the request range includes the timestamps before and after the time gap. Duration.

[0034] By adopting the above technical solution, segmented voice signal data is received on the server. Once the timestamps of each segment of voice signal data are interrupted, the terminal will retransmit the corresponding lost voice signal data, so that the text output by the voice recognition engine in the server is less likely to have contextual interruptions.

[0035] Optionally, the front and back extension time The specific calculation formula is:

[0036]

[0037]

[0038] in, is the dynamic expansion coefficient, , The timestamp of the end of the previous voice segment. The timestamp of the end of the subsequent voice segment. is the base expansion coefficient and is 1, is the adjustment gain coefficient and is 0.2, To calculate the average speaking speed, It is the real-time speaking speed in the current speech signal data.

[0039] By adopting the above technical solution, the duration of the voice signal data that the terminal needs to send in addition is dynamically adjusted in combination with the real-time speaking speed in the voice data that needs to be recognized, so as to ensure that the context of the first text obtained after the final voice recognition is not prone to interruptions.

[0040] In a second aspect, the present application provides an electronic device that adopts the following technical solution.

[0041] An electronic device includes a processor and a memory, wherein the memory stores a computer program. When the computer program is executed by the processor, the electronic device executes the above-mentioned method for generating minutes of an Internet online conference.

[0042] In summary, this application has at least the following beneficial effects:

[0043] 1. Adjust the encoding bit rate of voice signal data based on real-time network fluctuations to minimize the impact on voice signal data transmission quality, thereby reducing the likelihood of large errors in the text output by the subsequent voice recognition engine.

[0044] 2. Select an appropriate server architecture based on the number of online conference access terminals to minimize server load, effectively improve the server's voice signal data processing efficiency, and balance server usage costs.

[0045] 3. The terminal pre-processes and differentiates the voice signal data, transmitting high-value segments and common segments at different encoding bit rates. This ensures that the transmission of important content in the voice signal data is less likely to have quality issues, making it less likely that there will be large errors in the key content of the first text output by the final speech recognition engine, thereby improving the accuracy of the final meeting minutes.

[0046] 4. When the server receives segmented voice signal data, if there is a gap in the timestamps of each segment of the voice signal data, the terminal will retransmit the corresponding lost voice signal data to prevent the text output by the voice recognition engine in the server from being interrupted in context. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flowchart of a method for generating online meeting minutes in the present application;

[0048] Figure 2 This is a logic block diagram of the dynamic replacement of the server architecture in the method for generating online conference minutes in the present application;

[0049] Figure 3 This is a structural diagram of an electronic device of the present application. DETAILED DESCRIPTION

[0050] The present application is further described in detail below with reference to the accompanying drawings.

[0051] The present application embodiment discloses a method for generating minutes of an online conference on the Internet. Figure 1 , specifically including the following steps.

[0052] First, the voice signal data is collected and obtained through the microphone of each participant's terminal, such as a mobile phone or laptop.

[0053] Then, network bandwidth fluctuations are monitored in real time and the encoding bit rate is dynamically adjusted. This can be monitored on the terminal using nload or iftop, and on the server using vnstat or Wireshark. The formula for dynamically adjusting the encoding bit rate is as follows.

[0054]

[0055] in, is the encoding bit rate adjusted in real time, is the initial bit rate, is the packet loss rate threshold, is the current packet loss rate, ∈[0,1], To prevent zero minima, is 0.0001, for the Set the minimum threshold .

[0056] After acquiring voice signal data, the terminal performs local semantic prejudgment to distinguish high-value segments from common segments. The terminal deploys a compressed bidirectional LSTM model, with model parameters controlled within 500kb, to reduce CPU and memory usage and ensure high-quality voice signal acquisition. High-value and low-value segments are distinguished through a comprehensive approach involving keyword detection, semantic analysis, and behavioral analysis.

[0057] Keyword detection is to preset a group of keywords or phrases related to the main content or directive content of the meeting, such as urgent, required, necessary, important and approved. Then, part of the content in each split voice signal data segment is converted into voice and text content is extracted. In combination with regular expression matching or a machine learning-based keyword detection model, it is used to detect whether the preset keywords exist. If so, the voice signal data segment containing the corresponding content is a high-value segment; otherwise, it is an ordinary segment.

[0058] Semantic analysis is to analyze the text after speech transcription through pre-trained semantic understanding models, such as BERT or RoBERTa, to extract semantic features and determine whether the segment involves important topics or emotional expressions, such as anger or joy. If so, the corresponding speech signal data is segmented into high-value segments; otherwise, it is an ordinary segment.

[0059] Behavioral analysis is to extract the volume and speaking rate information features within each speech signal data segment. The MFCC model can be used for volume feature extraction, and the DDBHMM model can be used for speaking rate feature extraction. If the volume and speaking rate exceed the corresponding thresholds, it indicates that the corresponding speech signal data segment is a high-value segment; otherwise, it is an ordinary segment.

[0060] After the high-value segments and common segments are distinguished, the high-value segments are enhanced and the common segments are basic. The encoding bit rates of both enhanced and basic transmissions are determined by the above dynamic adjustment formula. The difference between the two is that the relevant parameters in the dynamic adjustment formula of the enhanced transmission encoding bit rate are: 128kbps, is 0.03, is 32kbps; basic transmission related parameters are, 64kbps, is 0.05, is 16kbps.

[0061] The server then receives the high-value segments and the common segments and aligns and splices them according to the timestamps, so that the subsequent speech recognition engine can perform complete speech recognition. In addition, there may be discontinuous timestamps after splicing. In this case, the server sends a data supplement request to the terminal, and the request range includes the timestamps before and after the time gap. For example, when the server receives a segment with timestamps t0-t1 and a segment with timestamps t2-t3, there is a gap in the timestamps of the two segments and they cannot be aligned. At this time, the server sends a request signal to the terminal. After receiving the request signal, the terminal sends the timestamp t 1-Δ / 2 to t 2+Δ / 2 The fragments are sent to the server so that the server can obtain the complete voice signal data. The duration is designed to ensure smooth transition and contextual integrity of voice signal data while reducing data redundancy and transmission burden.

[0062] Before and after expansion time The specific calculation formula is as follows.

[0063]

[0064]

[0065] in, is the dynamic expansion coefficient, , The timestamp of the end of the previous voice segment. The timestamp of the end of the subsequent voice segment. is the base expansion coefficient and is 1, is the adjustment gain coefficient and is 0.2, To calculate the average speaking speed, It is the real-time speaking speed in the current speech signal data.

[0066] The speech signal data is transmitted to the speech recognition engine to generate a first text, and the first text is corrected with reference to the historical correction mapping relationship to generate a second text, and the second text is stored in a shared document in real time to form meeting minutes.

[0067] In addition, the terminal or server generates a shared link for the online conference room, allowing multiple terminals to join the online meeting and review and edit meeting minutes. Furthermore, during the online meeting, the number of connected terminals is counted in real time, and the load factor is calculated. The server architecture is dynamically adjusted based on the load factor. The load factor L is calculated as follows.

[0068]

[0069] Among them, N is the number of online conference access terminals, The maximum load capacity of a single server. is the average terminal network bandwidth, is the baseline network bandwidth and can be 128kbps.

[0070] The server periodically sends query requests to all access terminals, and the terminals then return the current bandwidth data. The server calculates the average value based on the bandwidth data returned by each terminal to obtain .

[0071] Reference Figure 2 The specific conditions for dynamic replacement of server architecture are: when L < 0.3, single server mode is adopted; when 0.3 ≤ L < 0.7, regional server cluster mode is adopted; when L ≥ 0.7, cloud service automatic expansion mode is adopted.

[0072] The present application also discloses an electronic device, referring to Figure 3 , including a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the electronic device executes the above-mentioned method for generating minutes of an Internet online meeting.

[0073] The memory may be any medium capable of storing program code, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk. The processor may be a central processing unit, an application-specific integrated circuit, a dedicated instruction set processor, a graphics processing unit, a physical processing unit, a digital signal processor, a field programmable gate array, a programmable logic device, a controller, a microcontroller unit, a reduced instruction set computer, or a microprocessor.

[0074] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A method for generating online conference minutes, characterized by: include: Collect and obtain voice signal data; Monitor network bandwidth fluctuations in real time and dynamically adjust encoding bit rate; transmitting the speech signal data to a speech recognition engine to generate a first text; Correcting the first text with reference to the historical correction mapping relationship and generating a second text; storing the second text in a shared document in real time to form meeting minutes; Generate a shared link for the online meeting room so that multiple terminals can join the online meeting to view and edit the meeting minutes.

2. The method for generating online conference minutes according to claim 1, wherein: The formula for dynamically adjusting the encoding bit rate is: , in, is the encoding bit rate adjusted in real time, is the initial bit rate, is the packet loss rate threshold, is the current packet loss rate, ∈[0,1], To prevent zero minimum, for the Set the minimum threshold .

3. The method for generating online conference minutes according to claim 2, wherein: described 64kbps or 128kbps, is 0.05 or 0.03, is 0.0001, 16kbps or 32kbps.

4. The method for generating online conference minutes according to claim 1, wherein: During the online conference, the number of access terminals is counted in real time and the load factor is calculated, and the server architecture is dynamically replaced according to the load factor.

5. The method for generating online conference minutes according to claim 4, wherein: The specific calculation formula of the load factor L is: , Among them, N is the number of online conference access terminals, The maximum load capacity of a single server. is the average terminal network bandwidth, The baseline network bandwidth.

6. The method for generating online conference minutes according to claim 4, wherein: The server architecture dynamic replacement conditions are: When L < 0.3, the single-server mode is adopted; When 0.3≤L<0.7, the regional server cluster mode is adopted; When L ≥ 0.7, the cloud service automatic expansion mode is adopted.

7. The method for generating online conference minutes according to claim 1, wherein: After the voice signal data is acquired, the terminal performs local semantic pre-judgment to distinguish high-value segments from ordinary segments, performs enhanced transmission on the high-value segments and basic transmission on the ordinary segments, and then performs complete voice recognition after receiving the high-value segments and ordinary segments.

8. The method for generating online conference minutes according to claim 7, wherein: When the timestamps of the high-value segments and the ordinary segments are interrupted, the server sends a data supplement request to the terminal, and the request range includes the timestamps before and after the time gap. Duration.

9. The method for generating online conference minutes according to claim 8, wherein: The before and after expansion time The specific calculation formula is: , , in, is the dynamic expansion coefficient, , The timestamp of the end of the previous speech segment. The timestamp of the end of the subsequent voice segment. is the base expansion coefficient and is 1, is the adjustment gain coefficient and is 0.2, To calculate the average speaking speed, It is the real-time speaking speed in the current speech signal data.

10. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, wherein: When the computer program is executed by a processor, the electronic device executes the method for generating Internet online conference minutes as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Online conference method and device, storage medium and electronic equipment

    CN116366800A