Conference summary generation method and device based on bracelet and storage medium
By using a wristband array microphone and voiceprint recognition technology, the problem of distinguishing speakers in meeting recordings and text conversions was solved, resulting in accurate meeting minutes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The text generated from the converted meeting recordings has difficulty automatically distinguishing the content of different speakers, resulting in low intelligibility of the recordings.
Audio data is collected by a wristband array microphone and the sound source is located. Combined with voiceprint recognition algorithm and speech recognition technology, text data with identity tags is generated. Voiceprint recognition and sound source localization dual matching technology are used to distinguish speakers, and sub-tags are used to mark the speech content at different locations.
It effectively distinguishes speakers when transcribing meeting recordings, improving the intelligibility of the recordings and the accuracy of speaker identification, and generating structured meeting minutes.
Smart Images

Figure CN121814486A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and in particular to a method, device and storage medium for generating meeting minutes based on a wristband. Background Technology
[0002] Smart bracelets, as a typical wearable smart device, have become an important category in the consumer electronics field in recent years. They are typically worn as wristbands, integrating multiple sensors and low-power processors to continuously monitor the user's physiological parameters, exercise status, and environmental information. Through wireless communication technology, they synchronize and interact with smartphones, cloud servers, and other terminals, providing users with functions such as health management, exercise guidance, and message reminders.
[0003] In today's fast-paced business and academic activities, meetings are a core component of information exchange and decision-making. Accurate and efficient meeting minutes are crucial for recording key points, tracking task assignments, and clarifying decisions. Traditional methods of generating meeting minutes primarily rely on manual recording, which is not only time-consuming and labor-intensive but also prone to missing key information due to the recorder's misunderstanding or distraction. Subsequently, methods emerged that utilize smartphones or dedicated voice recorders to record conversations, followed by manual review and transcription or the generation of transcripts using cloud-based speech-to-text services.
[0004] In recent years, some technologies have attempted to integrate voice recording and transcription functions into wearable devices. However, when faced with the specific and complex scenario of "meeting minutes generation," these existing solutions often suffer from low intelligibility in typical multi-participant meeting room environments. The resulting transcripts lack structure and struggle to automatically distinguish between different speakers. Therefore, a new technology is needed to address the problem of low intelligibility caused by the inability of existing meeting minutes generation technologies to automatically differentiate between speakers. Summary of the Invention
[0005] The main objective of this invention is to solve the technical problem that the text generated from the conversion of meeting recordings is difficult to automatically distinguish the content of different speakers, resulting in low intelligibility of the recordings.
[0006] The first aspect of this invention provides a method for generating meeting minutes based on a wristband, comprising the following steps: Real-time audio data acquisition based on wristband array microphone, and localization of the sound source corresponding to the audio data; According to a preset voiceprint recognition algorithm, the audio data is processed to obtain an identity tag; Determine whether the identity tag is in a preset tag positioning table, wherein the tag positioning table includes: tag data and positioning data corresponding to the tag data; If the identity tag and the sound source location mapping are not in the preset tag location table, then the identity tag and the sound source location mapping are written into the tag location table to obtain an updated tag location table; If the identity tag is in the preset tag positioning table, then query the tag positioning data corresponding to the tag data that matches the identity tag in the tag positioning table; When the sound source location matches the location data, the audio data is processed by a preset speech recognition algorithm to generate text data with identity tags. When the sound source localization does not match the mapping localization, the identity tag is marked with a sub-tag to obtain the identity sub-tag, and the audio data is processed by phoneto-text conversion according to the preset speech recognition algorithm to generate text data with identity sub-tags; Organize all text data to obtain real-time communication documents, and generate meeting minutes based on the real-time communication documents.
[0007] Optionally, in a first implementation of the first aspect of the present invention, the step of performing voiceprint recognition processing on the audio data according to a preset voiceprint recognition algorithm to obtain an identity tag includes: Based on a preset sliding window and a preset energy threshold, the audio data is subjected to sliding window recognition to obtain N high-energy sound ranges, where N is a positive integer. Filter out time-domain segments whose duration is less than a preset interval threshold between N high-energy sound domains to obtain M smooth sound domains, where M is a positive integer not greater than N; According to the preset voiceprint recognition algorithm, voiceprint recognition processing is performed on the M smooth voice domains to obtain the identity tags corresponding to the M smooth voice domains; When all the identity tags corresponding to the M smoothed sound ranges are consistent, the identity tags corresponding to the M smoothed sound ranges are determined as the identity tags corresponding to the audio data.
[0008] Optionally, in a second implementation of the first aspect of the present invention, the step of real-time acquisition of audio data based on the wristband array microphone and localization of the sound source corresponding to the audio data includes: Audio data is collected using the wristband array microphone, and the array delay data corresponding to the audio data is recorded. According to a preset spatial calculation algorithm, the array delay data is processed for positioning calculation to generate the sound source location corresponding to the audio data.
[0009] Optionally, in a third implementation of the first aspect of the present invention, after the step of querying the location data corresponding to the identity tag in the tag location table, the method further includes: Calculate the three-dimensional offset rate between the sound source localization and the localization data; When the three-dimensional offset rate is less than a preset offset threshold, it is determined that the sound source localization matches the localization data; When the three-dimensional offset rate is not less than a preset offset threshold, it is determined that the sound source localization does not match the localization data.
[0010] Optionally, in a fourth implementation of the first aspect of the present invention, the step of determining whether the identity tag is in a preset tag positioning table includes: Calculate the multidimensional vector distance between the identity tag and each tag data in the preset tag location table; If the multidimensional vector distance is less than a preset hit threshold, then the tag data corresponding to the smallest multidimensional vector distance is determined as the tag data that matches the identity tag. If the distances of the multidimensional vectors are not less than the preset hit threshold, then it is confirmed that the identity tag is not in the preset tag location table.
[0011] Optionally, in the fifth implementation of the first aspect of the present invention, the step of organizing all text data to obtain a real-time communication document includes: Based on the acquisition timestamp of the audio data, all text data are sorted and organized to obtain a real-time communication document.
[0012] Optionally, in a sixth implementation of the first aspect of the present invention, generating meeting minutes based on the meeting communication document includes: Input the real-time communication document and preset prompts into the preset large language model; Receive the meeting minutes output by the large language model.
[0013] Optionally, in the seventh implementation of the first aspect of the present invention, after the step of organizing all the text data to obtain the meeting communication document, the method further includes: Based on the Bluetooth protocol, the real-time communication document is transmitted to a preset display device.
[0014] A second aspect of the present invention provides a wristband-based meeting minutes generation device, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the instructions in the memory to cause the wristband-based meeting minutes generation device to execute the above-described wristband-based meeting minutes generation method.
[0015] A third aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described wristband-based meeting minutes generation method.
[0016] In this embodiment of the invention, by using a wristband array microphone to collect audio data and locate the sound source of the audio data, and by performing dual matching of voiceprint recognition and sound source localization on the audio data, the speech content of different speakers can be effectively distinguished. Audio data with the same identity tag but different positions are distinguished by sub-tags. With enhanced speaker classification, different speaking positions of the same speaker can be marked, effectively recording the meeting communication situation. This enables effective differentiation of speakers when converting meeting recordings to text, enhances the accuracy of speaker identification, and solves the technical problem that it is difficult to automatically distinguish the content of different speakers in the text generated by converting meeting recordings, resulting in low intelligibility of the recordings. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of an embodiment of the meeting minutes generation method based on a wristband according to the present invention; Figure 2 This is a schematic diagram of a specific embodiment of the 101 steps of the wristband-based meeting minutes generation method in this invention. Figure 3 This is a schematic diagram of a specific embodiment of the 102 steps of the wristband-based meeting minutes generation method in this invention. Figure 4 This is a schematic diagram of a specific embodiment of step 103 of the meeting minutes generation method based on a wristband in this invention. Figure 5 This is a schematic diagram of one embodiment of the meeting minutes generation device based on a wristband in this invention. Detailed Implementation
[0018] This invention provides a method, device, and storage medium for generating meeting minutes based on a wristband.
[0019] The embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0020] In the description of the embodiments disclosed in this invention, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0021] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 A schematic diagram of an embodiment of the meeting minutes generation method based on a wristband in this invention includes the following steps: 101. Real-time acquisition of audio data based on a wristband array microphone, and localization of the sound source corresponding to the audio data; In this embodiment, the wristband is worn on the user's wrist to activate the meeting data collection function. The wristband is equipped with an array microphone to collect the audio data of the meeting and simultaneously collect the sound source location corresponding to the audio data. The sound source location can be combined with subsequent voiceprint recognition to double confirm the speaker's identity.
[0022] Please see Figure 2 , Figure 2 This is a schematic diagram of a specific embodiment of step 101 of the meeting minutes generation method based on a wristband in this invention. Step 101 includes the following specific implementation methods: 1011. Acquire audio data based on the wristband array microphone, and record the array delay data corresponding to the audio data; 1012. According to the preset spatial calculation algorithm, the array delay data is processed for positioning calculation to generate the sound source positioning corresponding to the audio data.
[0023] In steps 1021-1022, the array microphones of the wristband are used to collect external audio data, and the array delay data of the audio data at different microphones are calculated at the same time.
[0024] The time difference between the arrival times of sound at two different microphones is collected. This time difference is a key quantity that can be measured and used for calculation, denoted as the TDOA measurement. Multiplying the TDOA measurement by the speed of sound *c* yields the distance difference between the sound source and the two microphones. This distance difference defines a hyperboloid, on which the sound source must lie. Each TDOA measurement defines a hyperboloid. The intersection of multiple hyperboloids in three-dimensional space determines the location of the sound source. Using a set of TDOA measurements, an overdetermined system of equations is solved to calculate the spatial coordinates of the audio data. Spatial solution algorithms can be combined with clustering algorithms or spatial filtering to classify different TDOA sets to different sound sources, generating the sound source localization corresponding to the audio data.
[0025] 102. Based on a preset voiceprint recognition algorithm, perform voiceprint recognition processing on the audio data to obtain an identity tag; In this embodiment, the voiceprint recognition algorithm can be a deep convolutional neural network, such as ResNet or ECAPA-TDNN models, which converts audio data into vectors, inputs them into the model data, and outputs the corresponding identity tags for the audio data.
[0026] Please see Figure 3 , Figure 3 This is a schematic diagram of a specific embodiment of step 102 of the meeting minutes generation method based on a wristband in this invention. Step 102 includes the following specific implementation methods: 1021. Based on a preset sliding window and a preset energy threshold, perform sliding window recognition on the audio data to obtain N high-energy sound ranges, where N is a positive integer; 1022. Filter out time-domain segments with a duration less than a preset interval threshold between N high-energy sound domains to obtain M smooth sound domains, where M is a positive integer not greater than N; 1023. According to the preset voiceprint recognition algorithm, perform voiceprint recognition processing on the M smooth voice ranges to obtain the identity tags corresponding to the M smooth voice ranges; 1024. When all the identity tags corresponding to the M smooth sound ranges are consistent, the identity tags corresponding to the M smooth sound ranges are determined as the identity tags corresponding to the audio data.
[0027] In steps 1021-1024, audio data is used for voiceprint recognition in 40-second cycles. An energy threshold of 0.02 is set to filter the audio data. Segments and frames with energy values below the threshold are ignored, i.e., silent recordings are ignored. Segments with energy values above the threshold are marked, and audio data with energy values above the threshold is marked as filtered audio data on the timeline.
[0028] First, set a sliding window, such as 10 seconds, and then perform energy threshold judgment processing on each audio data one by one to obtain multiple high-energy sound ranges that are spaced apart on the time axis.
[0029] Set an interval threshold of 0.2s, remove short pauses of less than 0.2s from the high-energy sound ranges of multiple intervals, and delete the timestamps of the short pauses at the same time. Then, splice the high-energy sound ranges at both ends of the deleted timestamps together to obtain multiple smooth sound ranges on the time axis.
[0030] Then, using a voiceprint recognition algorithm, voiceprint recognition processing is performed on multiple smooth vocal ranges to obtain identity tags corresponding to multiple smooth vocal ranges.
[0031] If the identity labels corresponding to multiple smooth vocal ranges are consistent, it means that the 40 seconds of audio data are all from the same speaker.
[0032] If the identity labels corresponding to multiple smooth audio ranges are inconsistent, then the different intervals of the audio data are marked according to the timestamps of the smooth audio range timelines of different identity labels. However, this processing actually divides a whole audio data into multiple audio sub-data with different identity labels. The processing logic for each audio sub-data with the same identity label is consistent with that for the whole audio data with the same identity label, which will not be elaborated here.
[0033] 103. Determine whether the identity tag is in a preset tag positioning table, wherein the tag positioning table includes: tag data and positioning data corresponding to the tag data; In this embodiment, the identity tag is actually a multi-dimensional vector, such as a 6-dimensional vector [0.55, 0.0766, 0.999, 0.101, 2, 0.77522]. In practice, it may have 40-70 dimensions. The system analyzes whether the identity tag is less than 0.7 in dimensional distance from the tag data in the tag location table. Each tag data corresponds to the speaker's voice source location data.
[0034] For details, please refer to Figure 4 , Figure 4 This is a schematic diagram of a specific embodiment of step 103 of the meeting minutes generation method based on a wristband in this invention. Step 103 includes the following specific implementation methods: 1031. Calculate the multidimensional vector distance between the identity tag and each tag data in the preset tag positioning table; 1032. If the multidimensional vector distance is less than a preset hit threshold, then the tag data corresponding to the smallest multidimensional vector distance is determined as the tag data that matches the identity tag. 1033. If the distances of the multidimensional vectors are not less than the preset hit threshold, then it is confirmed that the identity tag is not in the preset tag location table.
[0035] In steps 1031-1033, the multidimensional vector distance between each identity tag and each tag data in the tag positioning table is calculated one by one. The multidimensional vector distance is calculated using Euclidean distance. Then, it is determined whether the multidimensional vector distance is less than the hit threshold of 0.7. If there is a multidimensional vector distance less than the preset hit threshold of 0.7, the tag data corresponding to the smallest multidimensional vector distance among all the multidimensional vector distances is determined as the tag data that matches the identity tag. If there is no multidimensional vector distance less than the hit threshold, it is confirmed that the identity tag is not in the preset tag positioning table.
[0036] 104. If the identity tag and the sound source location mapping are not in the preset tag location table, then the identity tag and the sound source location mapping are written into the tag location table to obtain an updated tag location table; In this embodiment, the identity tag and sound source location are bound together as a set and written into the tag location table to obtain an updated tag location table. This can be achieved using the NumPy library, where {identity tag, sound source location} is written as a new record into the tag location table to generate the updated tag location table.
[0037] 105. If the location data is in the preset tag location table, then query the tag location data that matches the identity tag in the tag location table; In this embodiment, if the identity tag is in the preset tag location table, the tag location table is queried to find the location data corresponding to the matching tag data, so that the location data can be analyzed in the future to confirm whether the speakers with the same identity tag are speaking from the same location.
[0038] Furthermore, following step 105, the following specific implementation methods are also included: 1051. Calculate the three-dimensional offset rate between the sound source localization and the localization data; 1052. When the three-dimensional offset rate is less than a preset offset threshold, it is determined that the sound source localization matches the localization data; 1053. When the three-dimensional offset rate is not less than the preset offset threshold, it is determined that the sound source positioning does not match the positioning data.
[0039] In steps 1051-1053, assuming the sound source is located at (a, b, c) and the location data is (d, e, f), then the three-dimensional offset rate ε = √[(ad)² + (be)² + (cf)²]. It is then determined whether the three-dimensional offset rate ε is less than an offset threshold of 3%. If the three-dimensional offset rate is less than the preset offset threshold, it is determined that the sound source location matches the location data. If the three-dimensional offset rate is not less than the preset offset threshold, it is determined that the sound source location does not match the location data.
[0040] 106. When the sound source location matches the location data, the audio data is processed by audio-text conversion according to a preset speech recognition algorithm to generate text data with identity tags. In this embodiment, if the sound source localization matches the localization data, a speech recognition algorithm is invoked to perform audio-to-text conversion processing on the audio data, converting it into text data. The original identity tag [0.55, 0.0766, 0.999, 0.101, 2, 0.77522] is converted into a substitute tag m, which is then marked in the text data to indicate the speaker's identity. Based on sound source localization and voiceprint recognition, the accuracy of determining the speaker's identity is improved.
[0041] 107. When the sound source localization does not match the mapping localization, the identity tag is marked with a sub-tag to obtain the identity sub-tag, and the audio data is processed by a preset speech recognition algorithm to generate text data with identity sub-tags. In this embodiment, if the sound source localization does not match the localization data, the original identity label [0.55,0.0766,0.999,0.101,2,0.77522] is sub-labeled. For example, if the sub-labeling is performed by giving the index [0.55,0.0766,0.999,0.101,2,0.77522]1, the original identity sub-label [0.55,0.0766,0.999,0.101,2,0.77522] is converted into a substitute sub-label m1. The substitute sub-label m1 is marked in the text data, indicating that the speaker's identity in the text data is identified as m. Different identification positions can display different speaker's speaking status and mark important speeches. Users can manually modify the identity sub-label to p through the wristband, manually adjust the speaker of the identified text data, and facilitate manual final error correction of the content.
[0042] 108. Organize all text data to obtain real-time communication documents, and generate meeting minutes based on the real-time communication documents.
[0043] In this embodiment, all tagged text data is sorted and organized according to the collection time to generate a real-time communication document. Finally, AI is used to summarize and categorize the real-time communication document to generate meeting minutes.
[0044] Specifically, the step "organizing all text data to obtain real-time communication documents" in step 108 includes the following specific implementation methods: 1081. Based on the acquisition timestamp of the audio data, sort and organize all the text data to obtain a real-time communication document.
[0045] In step 1081, all tagged text data is sorted and organized using the timestamps of the audio data collection, with the identity tags written first and the text data written last, to generate the original transcript of the audio converted word by word, which is the real-time communication document.
[0046] Specifically, step 108, "Generate meeting minutes based on the meeting communication document," includes the following specific implementation methods: 1082. Input the real-time communication document and preset prompt words into the preset large language model; 1083. Receive the meeting minutes output by the large language model.
[0047] In steps 1082-1083, the real-time communication document and the preset prompt "Summarize and generate meeting minutes" are input into the preset large language model. After the large language model organizes the data based on the real-time communication document and the prompt, the meeting minutes output by the large language model are received.
[0048] Furthermore, after "organizing all text data to obtain real-time communication documents" in the 108 steps, it also includes: 1084. Based on the Bluetooth protocol, the real-time communication document is transmitted to a preset display device.
[0049] In step 1084, the wristband can connect to an external display device via Bluetooth to transmit the real-time communication document converted from meeting audio to a preset display device, so that users can view the communication document as subtitles on the display device.
[0050] In this embodiment of the invention, by using a wristband array microphone to collect audio data and locate the sound source of the audio data, and by performing dual matching of voiceprint recognition and sound source localization on the audio data, the speech content of different speakers can be effectively distinguished. Audio data with the same identity tag but different positions are distinguished by sub-tags. With enhanced speaker classification, different speaking positions of the same speaker can be marked, effectively recording the meeting communication situation. This enables effective differentiation of speakers when converting meeting recordings to text, enhances the accuracy of speaker identification, and solves the technical problem that it is difficult to automatically distinguish the content of different speakers in the text generated by converting meeting recordings, resulting in low intelligibility of the recordings.
[0051] Figure 5This is a schematic diagram of a wristband-based meeting minutes generation device 500 provided in an embodiment of the present invention. The wristband-based meeting minutes generation device 500 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 510 and memory 520, and one or more storage media 530 for storing application programs 533 or data 532. The memory 520 and storage media 530 can be temporary or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the wristband-based meeting minutes generation device 500. Furthermore, the processor 510 may be configured to communicate with the storage media 530 and execute the series of instruction operations in the storage media 530 on the wristband-based meeting minutes generation device 500.
[0052] The wristband-based meeting minutes generation device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, Free BSD, etc. Those skilled in the art will understand that... Figure 5 The illustrated wristband-based meeting minutes generation device structure does not constitute a limitation on wristband-based meeting minutes generation devices, which may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0053] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the wristband-based meeting minutes generation method.
[0054] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0055] Furthermore, although the operations are described in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0056] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for generating meeting minutes based on a wristband, characterized in that, Including the following steps: Real-time audio data acquisition based on wristband array microphone, and localization of the sound source corresponding to the audio data; According to a preset voiceprint recognition algorithm, the audio data is processed to obtain an identity tag; Determine whether the identity tag is in a preset tag positioning table, wherein the tag positioning table includes: tag data and positioning data corresponding to the tag data; If the identity tag and the sound source location mapping are not in the preset tag location table, then the identity tag and the sound source location mapping are written into the tag location table to obtain an updated tag location table; If the identity tag is in the preset tag positioning table, then query the tag positioning data corresponding to the tag data that matches the identity tag in the tag positioning table; When the sound source location matches the location data, the audio data is processed by a preset speech recognition algorithm to generate text data with identity tags. When the sound source localization does not match the mapping localization, the identity tag is marked with a sub-tag to obtain the identity sub-tag, and the audio data is processed by phoneto-text conversion according to the preset speech recognition algorithm to generate text data with identity sub-tags; Organize all text data to obtain real-time communication documents, and generate meeting minutes based on the real-time communication documents.
2. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, The step of performing voiceprint recognition processing on the audio data according to a preset voiceprint recognition algorithm to obtain an identity tag includes: Based on a preset sliding window and a preset energy threshold, the audio data is subjected to sliding window recognition to obtain N high-energy sound ranges, where N is a positive integer. Filter out time-domain segments whose duration is less than a preset interval threshold between N high-energy sound domains to obtain M smooth sound domains, where M is a positive integer not greater than N; According to the preset voiceprint recognition algorithm, voiceprint recognition processing is performed on the M smooth voice domains to obtain the identity tags corresponding to the M smooth voice domains; When all the identity tags corresponding to the M smoothed sound ranges are consistent, the identity tags corresponding to the M smoothed sound ranges are determined as the identity tags corresponding to the audio data.
3. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, The steps of real-time audio data acquisition based on the wristband array microphone and sound source localization corresponding to the audio data include: Audio data is collected using the wristband array microphone, and the array delay data corresponding to the audio data is recorded. According to a preset spatial calculation algorithm, the array delay data is processed for positioning calculation to generate the sound source location corresponding to the audio data.
4. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, After the step of querying the location data corresponding to the identity tag in the tag location table, the method further includes: Calculate the three-dimensional offset rate between the sound source localization and the localization data; When the three-dimensional offset rate is less than a preset offset threshold, it is determined that the sound source localization matches the localization data; When the three-dimensional offset rate is not less than a preset offset threshold, it is determined that the sound source localization does not match the localization data.
5. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, The step of determining whether the identity tag is in the preset tag location table includes: Calculate the multidimensional vector distance between the identity tag and each tag data in the preset tag location table; If the multidimensional vector distance is less than a preset hit threshold, then the tag data corresponding to the smallest multidimensional vector distance is determined as the tag data that matches the identity tag. If the distances of the multidimensional vectors are not less than the preset hit threshold, then it is confirmed that the identity tag is not in the preset tag location table.
6. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, The steps for organizing all text data to obtain real-time communication documents include: Based on the acquisition timestamp of the audio data, all text data are sorted and organized to obtain a real-time communication document.
7. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, The step of generating meeting minutes based on the meeting communication document includes: Input the real-time communication document and preset prompts into the preset large language model; Receive the meeting minutes output by the large language model.
8. The method for generating meeting minutes based on a wristband according to claim 1, characterized in that, After the step of organizing all the text data to obtain the meeting communication document, the following is also included: Based on the Bluetooth protocol, the real-time communication document is transmitted to a preset display device.
9. A wristband-based meeting minutes generation device, characterized in that, The wristband-based meeting minutes generation device includes: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a circuit; The at least one processor invokes the instructions in the memory to cause the wristband-based meeting minutes generation device to perform the wristband-based meeting minutes generation method as described in any one of claims 1-8.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the wristband-based meeting minutes generation method as described in any one of claims 1-8.