Conference real-time summary framework generation method and related equipment
Through the wireless microphone system, the real-time collection, transliteration, structure and visualization of conference voice processing is solved, and the real-time and accuracy of traditional conference minutes generation methods are realized, real-time summary and dynamic visualization in the conference process are improved, and collaborative efficiency and decision-making accuracy are improved.
Patent Information
- Application Number
- CN202510781138.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-15
AI Technical Summary
The traditional meeting minutes generation method has the problem of low real-time and accuracy, especially in multi-person meetings or high concurrency scenarios, the network bandwidth requirements are high, and it is prone to lag or interruption, affecting the accuracy and real-timeness of meeting records.
The wireless microphone system is adopted, including a transmitter and a receiver, and voice signals are collected in real time through the recording module, the ASR module is transcribed in real time, the LLM module is structured, and the graphics engine is visualized, and finally displays the visual data on the display device to realize the real-time summary framework generation of the conference.
It realizes the instant generation of visual structured discussion context and core conclusions during the meeting, without waiting for manual sorting after the meeting, and can correct errors or supplement omissions in real time, and ensure information synchronization through retrospective historical speeches, significantly improving coordination efficiency and decision-making accuracy.
Smart Images

Figure CN120496534A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of microphone systems, and in particular to a method for generating a real-time summary framework for a conference and related equipment. Background Art
[0002] Traditional meeting recording methods require manual compilation of meeting content to create minutes after the meeting. This approach suffers from information lags, omissions of key information, and low efficiency. Furthermore, in multi-person meetings or high-concurrency scenarios, the transmission of raw audio data requires high network bandwidth, making it prone to lags or interruptions, impacting the accuracy and real-time nature of meeting records.
[0003] It can be seen that the traditional method of generating meeting minutes has problems of low real-time performance and accuracy. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to propose a method for generating a real-time summary framework for a meeting and related equipment to solve the problems of low real-time performance and accuracy in traditional methods of generating meeting minutes.
[0005] To solve the above technical problems, an embodiment of the present application provides a method for generating a real-time conference summary framework. The method is applied to a wireless microphone, wherein the wireless microphone includes a transmitter and a receiver. The transmitter is composed of a recording module, an ASR module, and a transmitting end transceiver module. The receiver is composed of a receiving end transceiver module, an LLM module, a graphics engine, and a connection module. The method adopts the following technical solutions:
[0006] The recording module collects the current voice signal of the participant in real time;
[0007] Performing a real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data;
[0008] Sending the current voice and text data to the receiving end transceiver module according to the transmitting end transceiver module;
[0009] Performing a structured processing operation on the current voice and text data received by the receiving end transceiver module according to the LLM module to obtain a hierarchical logical framework;
[0010] Performing a visualization processing operation on the hierarchical logic framework according to the graphics engine to obtain visualization data;
[0011] The visualization data is transmitted to a display device for display according to the connection module.
[0012] Furthermore, the step of performing a real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data specifically includes the following steps:
[0013] The original audio segment corresponding to the current voice text data is cached in a sending end database, and an index identifier corresponding to the current voice text data is constructed.
[0014] Furthermore, after the step of caching the original audio segment corresponding to the current voice text data into the sending end database and constructing the index identifier corresponding to the current voice text data, the following step is also included:
[0015] An audio cleaning operation is performed on the original audio segments stored in the sending end database according to a preset cleaning strategy.
[0016] Furthermore, the step of performing a visualization processing operation on the hierarchical logical framework according to the graphics engine to obtain visualization data specifically includes the following steps:
[0017] Determine whether there is a historical hierarchical logical framework;
[0018] If a historical hierarchical logical framework exists, the hierarchical logical framework and the historical hierarchical logical framework are compared for different contents, and the different contents are rendered based on the historical visualization data to obtain the visualization data.
[0019] Furthermore, after the step of performing visualization processing on the hierarchical logical framework according to the graphic engine to obtain visualization data, the following steps are further included:
[0020] A backtracking trigger button is added to each title node of the visualization data, wherein the backtracking trigger button is used to obtain index identification information corresponding to the title node.
[0021] Furthermore, after the step of transmitting the visual data to a display device for display according to the connection module, the following step is further included:
[0022] When the user clicks the backtracking trigger button, the index identification information corresponding to the title node is obtained, and an audio request instruction carrying the index identification information is sent to the transmitter via the receiver;
[0023] The transmitter obtains the audio segment corresponding to the index identification information from the transmitting end database, and sends the audio segment to the receiver;
[0024] The receiver outputs the audio clip through a speaker of the display device.
[0025] In order to solve the above technical problems, the embodiment of the present application further provides a system for generating a real-time conference summary framework, which adopts the following technical solutions:
[0026] The transmitter and receiver are composed of a recording module, an ASR module, and a transmitting end transceiver module, and the receiver is composed of a receiving end transceiver module, an LLM module, a graphics engine, and a connection module, wherein:
[0027] The recording module is used to collect the current voice signals of the participants in real time;
[0028] The ASR module is used to perform a real-time transcription operation on the current voice signal to obtain current voice text data;
[0029] The transmitting end transceiver module is used to send the current voice text data to the receiving end transceiver module;
[0030] The LLM module is used to perform structured processing operations on the current voice and text data received by the receiving end transceiver module to obtain a hierarchical logical framework;
[0031] The graphics engine is used to perform visualization processing operations on the hierarchical logic framework to obtain visualization data;
[0032] The connection module is used to transmit the visual data to a display device for display.
[0033] Furthermore, the ASR module includes:
[0034] The voice text caching submodule is used to cache the original audio segment corresponding to the current voice text data into the sending end database and construct an index identifier corresponding to the current voice text data.
[0035] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0036] The system comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the method for generating a real-time conference summary framework as described above when executing the computer-readable instructions.
[0037] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0038] The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the method for generating a real-time conference summary framework as described above are implemented.
[0039] The present application provides a method for generating a real-time summary framework for a conference, comprising: collecting the current voice signal of the participant in real time according to the recording module; performing real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data; sending the current voice text data to the receiving end transceiver module according to the transmitting end transceiver module; performing structured processing operation on the current voice text data received by the receiving end transceiver module according to the LLM module to obtain a hierarchical logical framework; performing visualization processing operation on the hierarchical logical framework according to the graphics engine to obtain visualization data; transmitting the visualization data to a display device for display according to the connection module. Compared with the existing technology, the present application processes the content of the conference speech in real time and dynamically generates a visualization framework, so that the participants can intuitively see the structured discussion context and core conclusions during the discussion, without having to wait for manual sorting after the meeting. It can not only correct errors or supplement omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, it guides the meeting to focus on key issues with the help of dynamic visualization logic presentation, significantly improving collaboration efficiency and decision-making accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0042] Figure 2 This is a flowchart of the method for generating a real-time conference summary framework provided in an embodiment of the present application;
[0043] Figure 3 This is a structural diagram of a system for generating a real-time summary framework for a conference provided in an embodiment of the present application;
[0044] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0046] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0047] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0048] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0049] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0050] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0051] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0052] It should be noted that the method for generating a real-time conference summary framework provided in the embodiment of the present application is generally executed by a server / terminal device. Accordingly, the system for generating a real-time conference summary framework is generally set in the server / terminal device.
[0053] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0054] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for generating a real-time conference summary framework according to the present application. The method for generating a real-time conference summary framework includes: step S201, step S202, step S203, step S204, step S205 and step S206.
[0055] In step S201, the current voice signal of the participant is collected in real time according to the recording module;
[0056] In step S202, the ASR module performs a real-time transcription operation on the current voice signal to obtain current voice text data;
[0057] In step S203, the current voice text data is sent to the receiving end transceiver module according to the transmitting end transceiver module;
[0058] In step S204, the LLM module performs a structured processing operation on the current voice text data received by the receiving end transceiver module to obtain a hierarchical logical framework;
[0059] In step S205, a visualization processing operation is performed on the hierarchical logic framework according to the graphics engine to obtain visualization data;
[0060] In step S206 , the visualization data is transmitted to the display device for display according to the connection module.
[0061] In an embodiment of the present application, the wireless microphone includes a transmitter and a receiver. The ASR model is loaded on the processing module of the transmitter, which can be integrated into the original chip or an SOC-level chip can be added for special processing; the LLM model is integrated into the processing module of the receiver; when in use, the receiver is plugged into the display device, and the meeting starts. The voice signals of the participants are directly converted into text data on the transmitter chip through the end-side ASR model. The transmitter chip transmits the text data wirelessly to the receiver. The receiver converts the text data into visual structure data in real time through the LLM model and transmits it to the display device, thereby realizing the visual real-time meeting summary mentioned in the above scheme.
[0062] In the embodiment of the present application, the implementation process of the visual real-time meeting summary may be:
[0063] Step 1: Voice collection and preprocessing
[0064] The wireless microphone transmitter has a built-in high-sensitivity microphone array to collect the voice signals of participants in real time.
[0065] The hardware noise reduction module (such as DSP chip) is used to filter out environmental noise (keyboard sound, air conditioning sound) and enhance the clarity of human voice.
[0066] Step 2: Real-time ASR transcription on the device side
[0067] The ASR chip built into the transmitter (such as an SOC with integrated NPU) converts the pre-processed voice stream into text data in real time.
[0068] The text data is accompanied by a timestamp, speaker ID (to distinguish different users through voiceprint recognition), and is cached in segments (such as one segment every 5 seconds).
[0069] Step 3: Transfer text data
[0070] The transmitter sends the encrypted text data to the receiver via a low-power wireless protocol (such as Bluetooth 5.3 or Wi-Fi Direct) (wired connection can also be used).
[0071] The transmitted content only contains structured text and does not contain raw audio, which saves bandwidth and improves privacy.
[0072] Step 4: Receiver LLM Structuring Processing
[0073] The receiver has a built-in LLM model. After receiving text data, it combines the context cache (historical speech content) to generate a hierarchical logical framework.
[0074] Dynamically adjust the output structure based on preset rules (such as "meeting minutes template" or "mind map node type"), for example:
[0075] Identify “decision points” and automatically mark them as red nodes;
[0076] Extract "Task Assignment" to generate sub-branches and associate responsible persons.
[0077] Step 5: Visual data generation and rendering
[0078] The structured data (such as JSON tree) output by LLM is converted into a visualization format through the sink's graphics engine (built-in lightweight rendering library).
[0079] Data is transmitted to a display device via the HDMI / USB-C interface and rendered in real time as an interactive mind map or outline. The following functions are supported:
[0080] Node expansion / collapse: click on the parent node to view the child content;
[0081] Timeline backtracking: Slide to view historical discussion nodes.
[0082] In an embodiment of the present application, a method for generating a real-time summary framework for a conference is provided, including: collecting the current voice signal of the participant in real time according to the recording module; performing real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data; sending the current voice text data to the receiving end transceiver module according to the transmitting end transceiver module; performing structured processing operation on the current voice text data received by the receiving end transceiver module according to the LLM module to obtain a hierarchical logical framework; performing visualization processing operation on the hierarchical logical framework according to the graphics engine to obtain visualization data; transmitting the visualization data to the display device for display according to the connection module. Compared with the existing technology, the present application processes the content of the conference speech in real time and dynamically generates a visualization framework, so that the participants can intuitively see the structured discussion context and core conclusions during the discussion, without having to wait for manual sorting after the meeting. It can not only correct errors or supplement omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, with the help of dynamic visualization logic presentation, it guides the meeting to focus on key issues, significantly improving collaboration efficiency and decision-making accuracy.
[0083] In some optional implementations of the embodiments of the present application, the step of performing a real-time transcription operation on the current voice signal according to the ASR module to obtain the current voice text data specifically includes the following steps:
[0084] The original audio segment corresponding to the current voice text data is cached in the sending end database, and an index identifier corresponding to the current voice text data is constructed.
[0085] In this embodiment of the present application, while converting speech to text, the transmitter caches the original audio segments in local storage by timestamp (e.g., one segment every 5 seconds) and creates a unique index ID (e.g., timestamp + speaker ID) with the generated text data. For example, the text node "AI Assistant Module" is associated with the audio index 20231105_1430_A.
[0086] In some optional implementations of the embodiments of the present application, after the steps of caching the original audio segment corresponding to the current voice text data in the sending end database and constructing the index identifier corresponding to the current voice text data, the following steps are also included:
[0087] Perform audio cleaning operations on the original audio clips stored in the sending end database according to a preset cleaning strategy.
[0088] In the embodiment of the present application, the transmitter dynamically cleans up old audio clips according to the storage capacity (for example, only retaining the latest 2 hours of data) to avoid memory overflow. Specifically, the cleaning strategy can be configured to retain by time or by important node marks.
[0089] In some optional implementations of the embodiments of the present application, the step of performing a visualization operation on the hierarchical logical framework according to the graphics engine to obtain visualization data specifically includes the following steps:
[0090] Determine whether there is a historical hierarchical logical framework;
[0091] If there is a historical hierarchical logical framework, the difference contents between the hierarchical logical framework and the historical hierarchical logical framework are compared, and the difference contents are rendered based on the historical visualization data to obtain visualization data.
[0092] In this embodiment of the application, new speech content triggers an incremental update of the map, rendering only the changed parts (such as adding a branch or modifying the task status). Users can make real-time corrections (such as merging duplicate nodes and adjusting priorities) through the external touch screen or physical buttons on the receiver.
[0093] In some optional implementations of the embodiments of the present application, after the step of performing a visualization operation on the hierarchical logical framework according to the graphics engine to obtain visualization data, the following steps are further included:
[0094] A backtracking trigger button is added to each title node of the visualized data, wherein the backtracking trigger button is used to obtain index identification information corresponding to the title node.
[0095] In the embodiment of the present application, in the visualization structure data generated by the receiver, a backtracking icon (such as a speaker symbol) is automatically added after each title node. When the user clicks the icon, the front end triggers a callback event to obtain the audio index ID associated with the node.
[0096] In some optional implementations of the embodiments of the present application, after the step of transmitting the visual data to the display device for display according to the connection module, the following steps are further included:
[0097] When the user clicks the backtrack trigger button, the index identification information corresponding to the title node is obtained, and the audio request instruction carrying the index identification information is sent to the transmitter through the receiver;
[0098] The transmitter obtains the audio segment corresponding to the index identification information from the transmitting end database and sends the audio segment to the receiver;
[0099] The receiver outputs the audio clip through the display device's speakers.
[0100] In an embodiment of the present application, when the user clicks on the icon, the front end triggers a callback event and obtains the audio index ID associated with the node; the receiver sends an audio request instruction to the transmitter through a wireless link (such as Bluetooth), including the target index ID and the request type (such as "play"); the transmitter retrieves the corresponding audio clip from the local cache according to the index ID, and returns an error code if it has expired (such as exceeding the storage time); the audio stream is transmitted back to the receiver through a high-priority wireless channel (such as Wi-Fi QoS) in the format of compressed audio (such as OPUS encoding); the receiver has a built-in audio decoding module, which converts the compressed audio into PCM waveform data and outputs it through the speaker of the display device.
[0101] In some optional implementations of the embodiments of the present application, when an audio clip is played, the visual interface synchronously highlights the associated text node (such as the node border flashing) to enhance feedback.
[0102] In summary, this application processes the content of conference speeches in real time and dynamically generates a visualization framework, allowing participants to intuitively see the structured discussion context and core conclusions during the discussion without having to wait for manual sorting after the meeting. They can not only correct errors or supplement omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, with the help of dynamic visualization logic presentation, the meeting is guided to focus on key topics, significantly improving collaboration efficiency and decision-making accuracy, and solving the problems of repeated discussions and conclusion deviations caused by information lag in traditional solutions. In addition, through the three core designs of streaming processing pipeline, incremental framework update and interactive visualization, the system can achieve real-time summary and dynamic visualization of meeting content without human intervention. The key to the technology lies in balancing the depth of semantic understanding and real-time requirements. In the future, it can be combined with neural symbolic systems (Neuro-Symbolic AI) to further improve the accuracy of logical reasoning (based on content pre-reading). For example, when mentioning microphone product planning, the AI model can recognize keywords and display content related to microphone product market trends in recent years. Real-time dynamic summary realizes the instant structuring and visualization of meeting content through streaming data processing and incremental framework update. Its core advantage lies in transforming "lagging" information processing into a dynamic cognitive assistance tool that "accompanies" meetings, directly improving collaborative efficiency and decision-making quality. The end-side microphone only transmits text data in real time, significantly optimizing transmission efficiency and real-time performance; the end-side ASR model converts voice into text in real time, and only transmits text content. Compared with the original audio data (such as PCM / WAV format), the amount of data is reduced by dozens to hundreds of times, greatly reducing the pressure on wireless channel bandwidth and ensuring smooth transmission in high-concurrency scenarios (such as multi-person meetings); text transmission has a higher tolerance for network fluctuations, and even in environments with unstable signals or limited bandwidth (such as mobile meetings and remote areas), it can still maintain real-time updates to avoid loss of meeting information caused by audio stream freezes or interruptions; privacy and security enhancements: Original The original voice data is retained on the end-side device throughout the process, and only the desensitized text content is transmitted externally, avoiding the risk of sensitive information leakage from the source and meeting the compliance requirements of industries such as finance and healthcare; the on-demand backtracking function improves the convenience of interaction and the intelligence of the system; precise positioning and immediate response: users can trigger an audio backtracking request with one click by clicking the icon associated with the visual node. The system automatically locates and returns the original speech segment of the corresponding time period, eliminating the need to manually review lengthy recordings, significantly improving review efficiency; resource allocation on demand: audio data is only returned in segments when the user actively requests it, avoiding continuous occupation of the transmission link and optimizing system resource utilization. At the same time, the end-side audio cache is combined with an automatic cleanup mechanism (such as retaining the content of the last 2 hours) to maximize storage space savings while ensuring the availability of functions; multimodal interaction fusion: when the audio is playing, the visual interface simultaneously highlights the associated text nodes, and supports operations such as dragging the progress bar and playing at double speed, forming a closed-loop experience of "text positioning → audio verification → focus linkage", enhancing the reliability of meeting records and the convenience of review.
[0103] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0104] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0105] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0106] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0107] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a system for generating a real-time summary framework of a conference. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0108] like Figure 3As shown, the real-time conference summary framework generation system 200 of the embodiment of the present application includes:
[0109] Transmitter 210 and receiver 220. Transmitter 210 is composed of recording module 211, ASR module 212 and transmitting end transceiver module 213. Receiver 220 is composed of receiving end transceiver module 221, LLM module 222, graphics engine 223 and connection module 224.
[0110] The recording module 211 is used to collect the current voice signals of the participants in real time;
[0111] ASR module 212, used to perform real-time transcription operation on the current voice signal to obtain current voice text data;
[0112] The transmitting end transceiver module 213 is used to send the current voice text data to the receiving end transceiver module 221;
[0113] The LLM module 222 is used to perform structured processing operations on the current voice text data received by the receiving end transceiver module 221 to obtain a hierarchical logical framework;
[0114] The graphics engine 223 is used to perform visualization processing operations on the hierarchical logic framework to obtain visualization data;
[0115] The connection module 224 is used to transmit the visual data to the display device for display.
[0116] In an embodiment of the present application, a conference real-time summary framework generation system 200 is provided, including: a transmitter 210 and a receiver 220, the transmitter 210 is composed of a recording module 211, an ASR module 212 and a transmitting end transceiver module 213, and the receiver 220 is composed of a receiving end transceiver module 221, an LLM module 222, a graphics engine 223 and a connection module 224, wherein: the recording module 211 is used to collect the current voice signal of the participant in real time; the ASR module 212 is used to perform real-time transcription operations on the current voice signal to obtain current voice text data; the transmitting end transceiver module 213 is used to send the current voice text data to the receiving end transceiver module 221; the LLM module 222 is used to perform structured processing operations on the current voice text data received by the receiving end transceiver module 221 to obtain a hierarchical logical framework; the graphics engine 223 is used to perform visualization processing operations on the hierarchical logical framework to obtain visualization data; and the connection module 224 is used to transmit the visualization data to a display device for display. Compared with the existing technology, this application processes the content of conference speeches in real time and dynamically generates a visual framework, so that participants can intuitively see the structured discussion context and core conclusions during the discussion without having to wait for manual sorting after the meeting. It can not only correct errors or supplement omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, with the help of dynamic visual logic presentation, it guides the meeting to focus on key issues, significantly improving collaboration efficiency and decision-making accuracy.
[0117] In some optional implementations of the embodiments of the present application, the ASR module includes:
[0118] The voice text caching submodule is used to cache the original audio segment corresponding to the current voice text data into the sending end database and construct an index identifier corresponding to the current voice text data.
[0119] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device according to an embodiment of the present application.
[0120] The computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 300 having components 310-330, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0121] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0122] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as a hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk equipped on the computer device 300, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 310 may also include both the internal storage unit of the computer device 300 and its external storage device. In the embodiment of the present application, the memory 310 is generally used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for the method for generating a real-time conference summary framework. In addition, the memory 310 can also be used to temporarily store various data that has been output or is about to be output.
[0123] In some embodiments, the processor 320 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 320 is generally used to control the overall operation of the computer device 300. In the embodiment of the present application, the processor 320 is used to execute computer-readable instructions or process data stored in the memory 310, such as computer-readable instructions for executing the method for generating a real-time conference summary framework.
[0124] The network interface 330 may include a wireless network interface or a wired network interface. The network interface 330 is generally used to establish a communication connection between the computer device 300 and other electronic devices.
[0125] The computer equipment provided in this application processes the content of conference speeches in real time and dynamically generates a visual framework, so that participants can intuitively see the structured discussion context and core conclusions during the discussion without having to wait for manual sorting after the meeting. It can not only correct errors or supplement omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, with the help of dynamic visual logical presentation, it guides the meeting to focus on key issues, significantly improving collaboration efficiency and decision-making accuracy.
[0126] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions. The computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned method for generating a real-time summary framework of a meeting.
[0127] The computer-readable storage medium provided in this application processes the content of conference speeches in real time and dynamically generates a visualization framework, so that participants can intuitively see the structured discussion context and core conclusions during the discussion without having to wait for manual sorting after the meeting. It can not only correct errors or omissions immediately, but also ensure information synchronization by tracing back historical speeches at any time. At the same time, with the help of dynamic visualization logic presentation, it guides the meeting to focus on key issues, significantly improving collaboration efficiency and decision-making accuracy.
[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0129] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A method for generating a real-time summary framework for a meeting, characterized in that: The method is applied to a wireless microphone, wherein the wireless microphone includes a transmitter and a receiver, the transmitter is composed of a recording module, an ASR module, and a transmitting end transceiver module, and the receiver is composed of a receiving end transceiver module, an LLM module, a graphics engine, and a connection module. The method includes the following steps: The recording module collects the current voice signal of the participant in real time; Performing a real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data; Sending the current voice text data to the receiving end transceiver module according to the transmitting end transceiver module; Performing a structured processing operation on the current voice and text data received by the receiving end transceiver module according to the LLM module to obtain a hierarchical logical framework; Performing a visualization processing operation on the hierarchical logic framework according to the graphics engine to obtain visualization data; The visualization data is transmitted to a display device for display according to the connection module.
2. The method for generating a real-time conference summary framework according to claim 1, characterized in that: The step of performing a real-time transcription operation on the current voice signal according to the ASR module to obtain current voice text data specifically includes the following steps: The original audio segment corresponding to the current voice text data is cached in a sending end database, and an index identifier corresponding to the current voice text data is constructed.
3. The method for generating a real-time conference summary framework according to claim 2, characterized in that: After the steps of caching the original audio segment corresponding to the current voice text data in the sending end database and constructing the index identifier corresponding to the current voice text data, the method further includes the following steps: An audio cleaning operation is performed on the original audio segments stored in the sending end database according to a preset cleaning strategy.
4. The method for generating a real-time conference summary framework according to claim 1, characterized in that: The step of performing a visualization processing operation on the hierarchical logical framework according to the graphics engine to obtain visualization data specifically includes the following steps: Determine whether there is a historical hierarchical logical framework; If a historical hierarchical logical framework exists, the hierarchical logical framework and the historical hierarchical logical framework are compared for different contents, and the different contents are rendered based on the historical visualization data to obtain the visualization data.
5. The method for generating a real-time conference summary framework according to claim 1, characterized in that: After the step of performing a visualization processing operation on the hierarchical logical framework according to the graphic engine to obtain visualization data, the following step is also included: A backtracking trigger button is added to each title node of the visualization data, wherein the backtracking trigger button is used to obtain index identification information corresponding to the title node.
6. The method for generating a real-time conference summary framework according to claim 5, characterized in that: After the step of transmitting the visual data to a display device for display according to the connection module, the method further includes the following steps: When the user clicks the backtracking trigger button, the index identification information corresponding to the title node is obtained, and an audio request instruction carrying the index identification information is sent to the transmitter via the receiver; The transmitter obtains the audio segment corresponding to the index identification information from the transmitting end database, and sends the audio segment to the receiver; The receiver outputs the audio clip through a speaker of the display device.
7. A system for generating a real-time summary framework for a meeting, characterized in that: The system comprises: The transmitter and receiver are composed of a recording module, an ASR module, and a transmitting end transceiver module, and the receiver is composed of a receiving end transceiver module, an LLM module, a graphics engine, and a connection module, wherein: The recording module is used to collect the current voice signals of the participants in real time; The ASR module is used to perform a real-time transcription operation on the current voice signal to obtain current voice text data; The transmitting end transceiver module is used to send the current voice text data to the receiving end transceiver module; The LLM module is used to perform structured processing operations on the current voice and text data received by the receiving end transceiver module to obtain a hierarchical logical framework; The graphics engine is used to perform visualization processing operations on the hierarchical logic framework to obtain visualization data; The connection module is used to transmit the visual data to a display device for display.
8. The system for generating a real-time conference summary framework according to claim 7, characterized in that: The ASR module includes: The voice text caching submodule is used to cache the original audio segment corresponding to the current voice text data into the sending end database and construct an index identifier corresponding to the current voice text data.
9. A computer device comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the method for generating a real-time conference summary framework according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the method for generating a real-time conference summary framework according to any one of claims 1 to 6.
Citation Information
Cited By
Real-time voice interaction method based on large model and electronic equipment
CN121415784A