Earphone voice memo application method, device and computer readable storage medium

CN122738501APending Publication Date: 2026-09-11SHENZHEN ZOWEE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610845548.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0005]为了克服现有技术中的不足,本发明的目的在于提供一种耳机语音备忘应用方法、设备及计算机可读存储介质,以解决目前耳机的语音备忘应用方案存在的依赖用户主动触发、无法识别语音对象、需人工整理语音的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122738501A_ABST
    Figure CN122738501A_ABST
Patent Text Reader

Abstract

This invention discloses a method, device, and computer-readable storage medium for headphone voice memo applications, belonging to the field of headphone technology. The method includes: first, collecting voice data according to the current working mode while the headphones are worn; then, generating memo information corresponding to the voice data based on pre-stored feature identifiers. This invention achieves a more efficient and convenient headphone voice memo application solution. Based on the characteristic of wearable headphones being worn for extended periods, it provides users with a seamless and all-day memo recording experience, eliminating the need for manual user operation and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of headphone technology, and in particular to a headphone voice memo application method, device, and computer-readable storage medium. Background Technology

[0002] In the existing technology, with the continuous development of communication technology and wearable devices, the form and function of headphones are becoming more and more diverse; among them, wearable headphones (such as earbuds, ear clips, ear hooks, neckbands, eyeglass frames, etc.) are gradually becoming popular among users. The main feature of these wearable headphones is that they need to be worn for a long time.

[0003] Currently, headphone-based voice memo application solutions have the following problems: First, it requires users to actively trigger the recording, and cannot automatically capture everyday conversations; Secondly, it is unable to independently identify and label different speech objects; Third, relying on manual processing of recordings is time-consuming and labor-intensive.

[0004] Therefore, given the characteristics of wearable headphones being worn for extended periods, how to achieve a seamless, all-day, and adaptive voice memo application has become a pressing technical problem that needs to be solved. Summary of the Invention

[0005] In order to overcome the shortcomings of the prior art, the present invention aims to provide a method, device and computer-readable storage medium for headphone voice memo application, so as to solve the problems of current headphone voice memo application solutions that rely on user active triggering, cannot recognize voice objects and require manual transcription of voice.

[0006] This invention proposes a method for using headphone voice memos in wearable headphones. The method includes: Voice data is collected based on the current working mode while the device is being worn. Memo information corresponding to the voice data is generated based on pre-stored feature identifiers.

[0007] Optionally, the working mode is normal mode, and the feature identifier is identity identifier; Based on pre-stored feature identifiers, memo information corresponding to the voice data is generated, specifically including: In the voice data, extract and record the first voice information of the owner and / or contact person corresponding to the identity identifier, and the second voice information of the unknown identity; The first voice information is organized into a first memo containing dialogue information, and the second voice information is used as a memo to be processed.

[0008] Optionally, the working mode is conference mode, and the feature identifier is voiceprint identifier; Based on pre-stored feature identifiers, memo information corresponding to the voice data is generated, specifically including: Extract and record the second voice information of the meeting participants corresponding to the voiceprint identifier from the voice data; The second voice information is organized into a second memo containing a meeting summary, speaking percentage, and to-do items, and the second memo provides options for meeting participants to complete the information.

[0009] Optionally, the working mode is private mode, and the feature identifier is an instruction identifier; Based on pre-stored feature identifiers, memo information corresponding to the voice data is generated, further including: The triggering scenarios and actions for monitoring and private modes; When the generated instruction identifier is determined based on the triggering scenario and / or triggering operation, the voice data is sealed, and an authentication option corresponding to the private mode is set for the unsealing of the voice data.

[0010] Optionally, while wearing the device, voice data is collected according to the current working mode, specifically including: Determine the status information of the two paired earphones; The task mode for each earphone is determined based on the status information.

[0011] Optionally, the status information is the remaining battery power information; The task mode for each earphone is determined based on the status information, specifically including: Determine the primary relationship between the remaining battery levels of the two earphones; Based on the first size relationship, the earphone with more remaining battery power is selected as the earphone for voice data collection, and the earphone with less remaining battery power is selected as the earphone for voice data processing.

[0012] Optionally, the status information is the remaining storage information; The task mode for each earphone is determined based on the status information, specifically including: Determine the second size relationship of the remaining stored information of the two headphones; Based on the second size relationship, the earphone with more remaining storage is selected as the earphone for voice data acquisition, and the earphone with less remaining storage is selected as the earphone for voice data processing.

[0013] Optionally, the task mode of each earphone is determined based on the status information, further including: Determine the switching time between the task modes of the two headphones; The voice data collected by the two headphones are merged by combining the switching time and task mode.

[0014] The present invention also proposes an earphone voice memo application device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the earphone voice memo application method as described in any of the preceding claims.

[0015] The present invention also proposes a computer-readable storage medium storing an earphone voice memo application, wherein when the earphone voice memo application is executed by a processor, the steps of the earphone voice memo application method as described in any of the preceding claims are implemented.

[0016] The present invention provides a headphone voice memo application method, device, and computer-readable storage medium that collect voice data according to the current working mode while the headphones are worn; and generate memo information corresponding to the voice data based on pre-stored feature identifiers. This invention achieves a more efficient and convenient headphone voice memo application solution. Based on the characteristic of wearable headphones being worn for extended periods, it provides users with a seamless and all-day memo recording experience without manual operation, thus enhancing the user experience. Attached Figure Description

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the first embodiment of the headphone voice memo application method of the present invention; Figure 2 This is a flowchart of the second embodiment of the headphone voice memo application method of the present invention; Figure 3 This is a flowchart of the third embodiment of the headphone voice memo application method of the present invention; Figure 4 This is a flowchart of the fourth embodiment of the headphone voice memo application method of the present invention; Figure 5 This is a flowchart of the fifth embodiment of the headphone voice memo application method of the present invention; Figure 6 This is a flowchart of the sixth embodiment of the headphone voice memo application method of the present invention; Figure 7 This is a flowchart of the seventh embodiment of the headphone voice memo application method of the present invention; Figure 8 This is a flowchart of the eighth embodiment of the headphone voice memo application method of the present invention. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0019] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.

[0020] Example 1 Figure 1 This is a flowchart of the first embodiment of the headphone voice memo application method of the present invention. A headphone voice memo application method, applied to wearable headphones, includes: S1. Collect voice data according to the current working mode while wearing the device; S2. Generate memo information corresponding to the voice data based on the pre-stored feature identifiers.

[0021] In this embodiment, based on the wearable headphones' features of seamless wearing and long battery life, a headphone voice memo application solution is implemented throughout the entire wearing cycle. Specifically: automatic human voice detection and recording are employed, for example, using the VAD (Voice Activity Detection) algorithm to determine in real time whether someone is speaking in the environment on the device side (e.g., the headphone case or a mobile terminal paired with the headphone case), thereby triggering recording; voiceprint recognition and speaker separation are employed, for example, using the device side's VP (Voiceprint Recognition) technology to separate the user's speech from that of other people, recording the user's voiceprint features before use, while the user can actively mark the voiceprints of others; automatic recording organization and content generation are employed, for example, based on the device side's ASR (Automatic Speech Recognition) technology, converting the separated recordings into text for processing by the corresponding artificial intelligence model.

[0022] In this embodiment, the generation of memo information includes: first, segmenting the transcribed text labeled with the speaker according to time segments, and then generating corresponding memo information through semantic understanding using a large language model. Specifically: generating memo information according to a daily timeline (i.e., event summaries sorted by time); generating memo information according to a structured diary (i.e., a daily review organized in natural language); generating memo information according to a to-do list (i.e., commitments, agreements, and tasks extracted from dialogue); and generating memo information according to intelligent reminder suggestions (i.e., schedule reminders based on time semantics).

[0023] In this embodiment, the earphone collects voice data according to the current working mode while being worn, and performs all-weather wear perception, voice detection, voiceprint feature extraction and labeling, audio caching and transmission operations; the earphone case and / or the mobile terminal paired with the earphone case perform local processing of the voice data, and perform speech transcription, data management, user interaction, and act as a proxy for cloud services; the cloud service provides the reasoning capabilities of a large-scale language model to generate structured diaries, to-do items and reminder suggestions.

[0024] In this embodiment, the earphone is equipped with a microphone array module, which includes at least one feedforward microphone for picking up ambient sound and a call microphone for picking up human voice, thereby realizing multi-channel audio acquisition.

[0025] In this embodiment, the earphone is also equipped with a wear detection module, which uses a capacitive or infrared proximity sensor to determine in real time whether the earphone is in the ear.

[0026] In this embodiment, the VAD algorithm is also configured in the earphone and integrated into the earphone's DSP digital signal processor or low-power AI chip, so as to perform real-time human voice detection on the earphone side of the audio stream picked up by the microphone.

[0027] In this embodiment, the earphone is also equipped with a voiceprint feature extraction module, which is integrated into the aforementioned DSP digital signal processor or low-power AI chip. After the VAD detects a human voice, it extracts the speaker's voiceprint embedding vector from the audio frame, thereby achieving preliminary speaker identification.

[0028] In this embodiment, the earphone is also equipped with a voiceprint matching and tagging module, which compares the extracted voiceprint vector with the locally stored registered voiceprint database to tag the speaker's identity for the recorded segment.

[0029] In this embodiment, the headphones are also equipped with an audio caching and storage module, which includes solid-state flash memory for caching raw audio and pre-stored voiceprint library data.

[0030] In this embodiment, the earphone is also equipped with a Bluetooth control and transmission module, which is used to package and transmit audio data with speaker markers and related information (e.g., timestamps, durations, identification markers, etc.) to the earphone compartment or mobile terminal.

[0031] In this embodiment, the earphone is also equipped with a main control and low power management module, which is used to coordinate the work of other modules and execute dynamic left and right ear allocation modes, working modes and sleep strategies.

[0032] In this embodiment, the earphone is also equipped with a physical indicator light or buzzer to provide visual or auditory cues when recording is activated, thereby meeting the user's privacy and transparency requirements.

[0033] In this embodiment, the specific signal and data flow logic is as follows: sensing and rough processing are performed on the device side, followed by fine processing on the device side, or the earphone case, or the mobile terminal side, and finally deep processing is performed in the cloud.

[0034] Specifically, in this embodiment, the earphone's wear detection sensor outputs a "wearing / removing" signal to the main controller, which serves as a global enable switch for the memo application. When the mode is enabled, the VAD continuously monitors the audio signal after noise reduction via the microphone array. If the detected human voice probability exceeds a threshold and continues to exceed the shortest duration, a recording-triggered interrupt signal is output. The human voice probability is the proportion of the current human voice audio to all current audio, preferably 0.9, and the shortest duration is 200ms-500ms, preferably 300ms. Further, the interrupt signal triggers... The system uses a voiceprint feature extractor to extract voiceprint embeddings from the current audio frame and compares them with a local voiceprint library in real time. The comparison results are then time-aligned with the audio data stream to generate tagged audio segments. These tagged audio segments are temporarily stored in the headset's buffer with an appended header (e.g., containing start and end timestamps, speaker identity, and confidence level). Furthermore, the headset's Bluetooth control and transmission module transmits the completed audio files from the buffer to the headset charging case or mobile terminal according to a real-time or batch transmission strategy. This utilizes simultaneous recording and transmission, or delayed batch transmission, to save power.

[0035] Specifically, in this embodiment, the mobile terminal receives the audio file and calls the local ASR or cloud ASR to convert the speaker-separated audio into a text dialog box with speaker tags; then the text dialog box is sent to the large language model in the cloud, processed by PE (Prompt Engineering), to generate structured memo data, and sent back to the mobile terminal, so that the mobile terminal can present users with information such as diaries and to-do items corresponding to the voice data.

[0036] In this embodiment, the device initialization and voiceprint registration process specifically includes: after the user first wears the headphones and launches the relevant application on the mobile terminal, they enter the corresponding voiceprint management interface. Based on this interface, the user is guided to read a specified text (e.g., 30 seconds). The voiceprint feature extractor generates a unique voiceprint model for the user, which is marked as "me," i.e., the user themselves, and stored in the local voiceprint library. Furthermore, the user can also create voiceprint profiles for family members, colleagues, and other contacts. Similarly, after guiding the contact to read aloud for a short time or detecting a stable voiceprint from existing conversations, the user actively marks it as "Contact A," "Colleague B," etc. It should be noted that the above process only stores the anonymized voiceprint feature vector and does not save the original audio.

[0037] In this embodiment, the automatic operation process of the all-day memo includes: after the wearing detection module confirms that the earphone is in the ear, it enters a preset normal mode (or default mode, or smart mode, etc.). In this mode, in order to separate the human voice-triggered recording from the speaker, a VAD is used to continuously detect voice data or ambient sound data. When a valid human voice is detected, the voiceprint module is activated. Based on the voiceprint module, a matching judgment is performed. If the current speaker's voiceprint matches "my" registered voiceprint, it is marked as "user". If it matches an existing contact, it is marked as the corresponding identity. Otherwise, it is marked as a temporary identity such as "Stranger 1" or "Stranger 2". The user can then manually mark it in the above-mentioned interactive interface. Furthermore, when two or more speakers take turns speaking in the conversation, the SP (Speaker Diarization) technology on the device side is used to cut the voice segments of different people on the timeline, and each segment is compared and marked separately, thereby accurately generating the corresponding notes information of the person, time, and speech content.

[0038] In this embodiment, in order to achieve local caching and transmission, during the continuous recording process, an audio file segment is generated and cached in the headphone's flash memory every time a complete semantic pause is formed (e.g., when silence exceeds 1.5 seconds) or a fixed duration limit is reached (e.g., every 5 minutes).

[0039] In this embodiment, in order to achieve collaborative work between the left and right ears, one earphone is used as the main recording ear, while the other earphone is only used for VAD monitoring, thereby reducing power consumption; for example, when the storage space or power of one earphone is lower than the threshold, the other earphone is automatically switched to continue recording and buffering, and the audio streams are automatically merged in the future.

[0040] In this embodiment, during the idle window of Wi-Fi or Bluetooth Low Energy, the audio file is transmitted to the headphone compartment or mobile terminal in an encrypted manner.

[0041] Furthermore, in this embodiment, the speech transcription and content generation specifically include: after the mobile terminal receives the tagged audio, it calls the offline ASR to generate text with speaker tags, in the following format: [14:05] Me: Are the meeting materials for this afternoon ready? [14:06] Contact person A - Manager Zhang: I have sent them to your email, please check. Furthermore, in this embodiment, the mobile terminal concatenates text conversations from a time window (e.g., the past hour or a summary period manually triggered by the user) into contextual prompts and sends them to the cloud-based LLM (Large Language Model). The LLM then generates structured content according to the following instructions: a daily timeline, used to extract events mentioned in the conversation and list them in chronological order; a structured diary, used to organize today's core conversations and events into a concise diary entry in the first person; a to-do list, used to identify all commitments, agreements, and tasks, extract them as to-do items, and indicate their source and time; and intelligent reminder suggestions, used to automatically suggest creating calendar reminders based on the conversation content (e.g., a report needs to be completed by next Wednesday), allowing the user to add reminder items with a single click.

[0042] The beneficial effect of this embodiment is that it collects voice data based on the current working mode while the device is worn, and generates memo information corresponding to the voice data based on pre-stored feature identifiers. This achieves a more efficient and convenient headphone voice memo application solution. Based on the characteristic of wearable headphones being worn for extended periods, it can provide users with a seamless and all-day memo recording experience without requiring manual operation, thus enhancing the user experience.

[0043] Example 2 Figure 2 This is a flowchart of the second embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the working mode is normal mode, and the feature identifier is an identity identifier; generating memo information corresponding to the voice data according to the pre-stored feature identifier specifically includes: S21. Extract and record the first voice information of the owner and / or contact person corresponding to the identity identifier, and the second voice information of the unknown identity from the voice data; S22. Organize the first voice information into a first memo containing dialogue information, and use the second voice information as a memo to be processed.

[0044] In this embodiment, the working mode of the headset is set to normal mode and the feature identifier is the identity identifier, according to different scenario requirements. Based on this, the VAD is always on, and the headset is automatically marked according to the voiceprint library. The first voice information of the owner and / or contact person corresponding to the identity identifier is extracted and recorded, so as to realize the full recording of "my" speech and registered contacts and generate a diary. Similarly, the second voice information of unknown identity is extracted and recorded, so as to realize the processing of recording "unknown speaker".

[0045] In this embodiment, the processing priority of second voice information from an unknown identity is reduced, and user confirmation is required before it can be retained.

[0046] The beneficial effect of this embodiment is that, in normal mode, the first voice information with known identity is organized into the first memo information containing dialogue information, and the second voice information with unknown identity is used as the memo information to be processed, thereby improving the processing efficiency of voice information and eliminating the need for the user to process them one by one.

[0047] Example 3 Figure 3 This is a flowchart of the third embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the working mode is a conference mode, and the feature identifier is a voiceprint identifier; the memo information corresponding to the voice data is generated according to the pre-stored feature identifier, specifically including: S23. Extract and record the second voice information of the meeting participants corresponding to the voiceprint identifier from the voice data; S24. Organize the second voice information into a second memo containing a meeting summary, speaking percentage, and to-do items, and provide the meeting participants with options to complete the second memo.

[0048] In this embodiment, for multi-person scenarios, a speaker segmentation algorithm is applied. Unlike the embodiments described above, it no longer relies on pre-registration based on voiceprints, but instead uses real-time clustering based on sound features.

[0049] In this embodiment, the second voice information of the meeting participants corresponding to the voiceprint identifiers is directly extracted and recorded. This second voice information is then organized into a second memo containing a meeting summary, speaking percentage, and to-do items. For example, after the meeting ends, the meeting summary, speaking percentage, and to-do list are proactively pushed to the mobile terminal, and the user is provided with an option to quickly complete the speaker's name, thus creating a complete voice memo without complex editing operations.

[0050] The beneficial effect of this embodiment is that, in the meeting mode, the second voice information related to the meeting participants is organized into a second memo information that includes a meeting summary, speaking percentage and to-do items, and the second memo information provides the option to complete the meeting participants. Therefore, it is not necessary to perform voiceprint recognition one by one during the meeting, nor is it necessary for users to confirm unknown people one by one after the meeting. This improves the efficiency and automation of the generation of meeting voice memos.

[0051] Example 4 Figure 4 This is a flowchart of the fourth embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the working mode is private mode, and the feature identifier is a command identifier; generating memo information corresponding to voice data according to the pre-stored feature identifier further includes: S25. Triggering scenarios and operations for monitoring and privacy modes; S26. When the generated instruction identifier is determined based on the triggering scenario and / or triggering operation, the voice data is sealed, and an authentication option corresponding to the private mode is set for the unsealing of the voice data.

[0052] In this embodiment, unlike the normal mode or conference mode described above, when the working mode is private mode and the feature identifier is an instruction identifier, the corresponding instruction identifier is generated by physical button or voice command, thereby immediately pausing all VAD and recording activities.

[0053] In this embodiment, if the above-mentioned private mode is entered according to the preset trigger scenario and trigger operation, all the current cached data of the earphone is immediately sealed; furthermore, an application lock is added on the mobile terminal side, so that this private mode is more suitable for highly sensitive scenarios and gives the user full control.

[0054] In this embodiment, if the user exits the private mode according to a different triggering scenario and triggering operation than described above, an identity verification step is set for the headset or mobile terminal to further meet the user's privacy needs in this mode.

[0055] The beneficial effect of this embodiment is that by determining the generation of instruction identifiers based on the triggering scenario and / or triggering operation in privacy mode, and by sealing the voice data, and by setting authentication options corresponding to the privacy mode for the unsealing of the voice data, the security of seamless voice memos is improved, without requiring complicated operations by the user, and the user's privacy needs are fully met.

[0056] Example 5 Figure 5 This is a flowchart of the fifth embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, voice data is collected according to the current working mode while the headphones are being worn, specifically including: S01. Determine the status information of the two paired earphones; S02. Determine the task mode for each of the two earphones based on the status information.

[0057] In this embodiment, to implement a low-power design and a dynamic allocation mechanism between the left and right ears, enabling the wearable headphones to operate around the clock, multiple low-power optimization strategies are adopted: First, tiered wake-up. For example, in the wearing state, only low-power wear detection is implemented, while the voice activity detection module is always working, and other modules (e.g., voiceprint extraction, Bluetooth transmission, main processor, etc.) are in a deep sleep state. During this process, the aforementioned VAD interrupt triggers the other modules to wake up sequentially in a pipeline manner. Second, the left and right ears take turns controlling the main ear. The battery level, storage space, and microphone signal quality (e.g., signal-to-noise ratio) of the left and right ears are monitored in real time. Based on this, under normal circumstances, the side with higher battery level and better signal is selected as the main earphone, responsible for voice pickup, VAD judgment, voiceprint extraction, and audio storage, while the other side serves as the auxiliary earphone, retaining only wear detection and auxiliary noise reduction, thereby keeping power consumption at an extremely low level.

[0058] The beneficial effect of this embodiment is that by determining the status information of the two paired earphones and determining the task mode of the two earphones respectively based on the status information, the power consumption of the earphones is reduced, enabling the earphones to operate all day long and avoiding interruption of voice memos.

[0059] Example 6 Figure 6 This is a flowchart of the sixth embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the status information is the remaining battery level information; the task mode of each of the two headphones is determined according to the status information, specifically including: S11. Determine the primary relationship between the remaining battery levels of the two earphones; S12. Based on the first size relationship, select the earphone with more remaining battery power as the earphone for voice data acquisition, and select the earphone with less remaining battery power as the earphone for voice data processing.

[0060] In this embodiment, when the battery level of the main earphone drops to a certain threshold (e.g., below 20% of the battery level of the auxiliary earphone), the main control unit of the main earphone issues a smooth switching command. Based on this command, the main and auxiliary roles are swapped, the original auxiliary earphone is woken up and continues all recording tasks, and the original main earphone enters a sleep state. Optionally, in this embodiment, there is no sound reminder during the switching process to reduce user awareness, and a data merging method is used to avoid data loss.

[0061] Example 7 Figure 7This is a flowchart of the seventh embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the status information is the remaining storage information; the task mode of each of the two headphones is determined according to the status information, specifically including: S13. Determine the second size relationship of the remaining stored information of the two headphones; S14. Based on the second size relationship, select the earphone with more remaining storage as the earphone for voice data acquisition, and select the earphone with less remaining storage as the earphone for voice data processing.

[0062] In this embodiment, when the remaining storage of the main earphone drops to a certain threshold (e.g., below 30% of the remaining storage of the auxiliary earphone), the main control unit of the main earphone issues a smooth switching command, the main and auxiliary roles are swapped, the original auxiliary earphone is woken up and continues all recording tasks, and the original main earphone enters a sleep state; optionally, in this embodiment, there is no sound reminder during the switching process to reduce user awareness, and a data merging method is used to avoid data loss.

[0063] Example 8 Figure 8 This is a flowchart of the eighth embodiment of the headphone voice memo application method of the present invention. Based on the above embodiment, the task mode of each of the two headphones is determined according to the status information, and further includes: S15. Determine the switching time for the task modes of the two earphones; S16. Combine the switching time and task mode to merge the voice data collected by the two headphones.

[0064] In this embodiment, the switching time of the task modes of the two earphones is first determined, and then the voice data collected by the two earphones is merged by combining the switching time and the task mode. Specifically, in order to achieve balanced storage and collaborative caching, this embodiment configures both the left and right earphones to be in a mode that can independently cache recordings. Based on this, when Bluetooth data transmission is implemented, the mobile terminal establishes a file index for the received data, thereby merging the audio segments from the left and right earphones and sorting them by timestamp, thereby ensuring the accuracy and precision of the merging.

[0065] Furthermore, in this embodiment, a transmission window strategy is set. For example, the high-speed link of Bluetooth Low Energy (BLE) is used to transmit recording files in batches during idle periods when the mobile terminal screen is off and no audio is playing, thereby avoiding competition for bandwidth with media audio and further reducing the total power consumption.

[0066] The beneficial effect of this embodiment is that by determining the second size relationship of the remaining storage information of the two earphones, the earphone with more remaining storage is selected as the voice data acquisition earphone, and the earphone with less remaining storage is selected as the voice data processing earphone, thereby dynamically allocating the working state of the left and right earphones and evenly allocating processing resources and storage resources, which can save data transmission time and evenly extend the battery life of the left and right earphones.

[0067] Example 9 Based on the above embodiments, the present invention also proposes an earphone voice memo application device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the earphone voice memo application method as described in any of the above embodiments.

[0068] It should be noted that the above-described device embodiments and method embodiments belong to the same concept. The specific implementation process can be found in the method embodiments, and the technical features in the method embodiments are also applicable to the device embodiments, which will not be repeated here.

[0069] The beneficial effect of this embodiment is that it collects voice data based on the current working mode while the device is worn, and generates memo information corresponding to the voice data based on pre-stored feature identifiers. This achieves a more efficient and convenient headphone voice memo application solution. Based on the characteristic of wearable headphones being worn for extended periods, it can provide users with a seamless and all-day memo recording experience without requiring manual operation, thus enhancing the user experience.

[0070] Example 10 Based on the above embodiments, the present invention also proposes a computer-readable storage medium storing an earphone voice memo application, wherein when the earphone voice memo application is executed by a processor, the steps of the earphone voice memo application method as described in any of the above claims are implemented.

[0071] It should be noted that the above-described medium embodiments and method embodiments belong to the same concept. The specific implementation process can be found in the method embodiments, and the technical features in the method embodiments are also applicable to the medium embodiments, which will not be repeated here.

[0072] The headphone voice memo application method, device, and computer-readable storage medium of the present invention collect voice data according to the current working mode while wearing the headphones; generate memo information corresponding to the voice data based on pre-stored feature identifiers; and realize a more efficient and convenient headphone voice memo application solution. Based on the characteristics of wearable headphones being worn for a long time, it can bring users a seamless and all-day memo recording experience without the need for manual operation by the user, thus enhancing the user experience.

[0073] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0074] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0075] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0076] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for earphone voice memo application, applied to a wearable earphone, characterized in that, The method includes: Voice data is collected based on the current working mode while the device is being worn. Memo information corresponding to the voice data is generated based on pre-stored feature identifiers.

2. The earphone voice memo application method of claim 1, wherein, The working mode is the normal mode, and the feature identifier is the identity identifier; The step of generating memo information corresponding to the voice data based on pre-stored feature identifiers specifically includes: From the voice data and / or the environmental voice data, extract and record the first voice information of the owner and / or contact person corresponding to the identity identifier, and the second voice information of the unknown identity. The first voice information is organized into a first memo containing dialogue information, and the second voice information is used as a memo to be processed.

3. The earphone voice memo application method of claim 1, wherein, The working mode is a conference mode, and the feature identifier is a voiceprint identifier; The step of generating memo information corresponding to the voice data based on pre-stored feature identifiers specifically includes: Extract and record the second voice information of the meeting participants corresponding to the voiceprint identifier from the voice data and / or the environmental voice data; The second voice information is organized into a second memo containing a meeting summary, speaking percentage, and to-do items, and the second memo provides options for the meeting participants to complete the information.

4. The earphone voice memo application method of claim 1, wherein, The operating mode is a private mode, and the feature identifier is an instruction identifier; The step of generating memo information corresponding to the voice data based on pre-stored feature identifiers further includes: Monitor the triggering scenarios and triggering operations related to the aforementioned private mode; When the instruction identifier is generated based on the trigger scenario and / or the trigger operation, the voice data and / or the environmental voice data are sealed, and an authentication option corresponding to the private mode is set for the unsealing process of the voice data and / or the environmental voice data.

5. The earphone voice memo application method of claim 1, wherein, The process of collecting voice data based on the current working mode while wearing the device specifically includes: Determine the status information of the two paired earphones; The task mode of each of the two headphones is determined based on the status information.

6. The earphone voice memo application method of claim 5, wherein, The status information is the remaining battery power information; The step of determining the task mode of the two earphones based on the status information specifically includes: Determine the first relationship between the remaining battery information of the two earphones; Based on the first size relationship, the earphone with more remaining battery power is selected as the earphone for collecting the voice data and / or the environmental voice data, and the earphone with less remaining battery power is selected as the earphone for processing the voice data and / or the environmental voice data.

7. The earphone voice memo application method of claim 6, wherein, The status information refers to the remaining storage information; The step of determining the task mode of the two earphones based on the status information specifically includes: Determine the second size relationship of the remaining stored information of the two earphones; Based on the second size relationship, the earphone with more remaining storage is selected as the earphone for collecting the voice data and / or the environmental voice data, and the earphone with less remaining storage is selected as the earphone for processing the voice data and / or the environmental voice data.

8. The method of claim 6 or 7, wherein, The step of determining the task mode of the two earphones based on the status information further includes: Determine the switching time of the task mode of the two headphones; The voice data and / or environmental voice data collected by the two earphones are combined based on the switching time and the task mode.

9. A headset voice memo application device, characterized by comprising: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the headphone voice memo application method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an application for using a headset voice memo, which, when executed by a processor, implements the steps of the headset voice memo application method as described in any one of claims 1 to 8.