Earphone, earphone system and earphone interaction method

By employing a dual-audio module design and voice processing technology in the headphones, the problem of headphones lacking interactivity and immersion has been solved, achieving a stereo audio and intelligent interactive headphone experience.

CN121751040APending Publication Date: 2026-03-27GUANGDONG XIAOTIANCAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing headphones lack interactivity and immersion, and smart speakers and smart headphones have limited audio output methods, making it difficult to provide a deep interactive experience.

Method used

It adopts a dual audio module design, each corresponding to a different role or language model. The voice acquisition module collects user commands, the processing module analyzes and adjusts the audio data to enhance sound effects and spatial sense, and the storage module stores the audio data in categories.

Benefits of technology

It enables an interactive experience between the headphones and the user, enhancing personalization and immersion, and providing stereo audio effects and intelligent interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121751040A_ABST
    Figure CN121751040A_ABST
Patent Text Reader

Abstract

The invention provides an earphone, an earphone system and an earphone interaction method. The earphone comprises: a processing module for obtaining audio data to be played, the audio data to be played comprising first audio data and second audio data; the first audio data comprises audio data corresponding to a first role or a first language model, and the second audio data comprises audio data corresponding to a second role or a second language model; the first audio module is used for receiving and playing the first audio data; and the second audio module is used for receiving and playing the second audio data. According to the earphone provided by the invention, the two audio modules respectively and correspondingly play audio data of different sound production roles or intelligent assistants, so that the interactive experience and auditory experience of a user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio technology, and more particularly to a pair of headphones, a headphone system, and a method for interacting with the headphones. Background Technology

[0002] Currently, headphones on the market are mainly used for playing music, answering calls, and watching videos, typically providing only a single auditory experience and lacking interactivity and immersion. Existing audio interactive devices, such as smart speakers, can provide some interactive functions, but due to their fixed usage locations and single audio output methods, the user experience is limited, making it difficult to provide a truly immersive interactive experience.

[0003] Some smart headphones have voice assistant functions, which can control playback and answer calls through voice commands, but these functions are mainly focused on audio control and simple task execution, lacking a deep interactive experience. Summary of the Invention

[0004] The purpose of this invention is to provide a pair of headphones, a headphone system, and a headphone interaction method, which enables the playback of audio data corresponding to different roles through two audio modules, with each audio module providing different viewpoints and suggestions, thereby enhancing the interactive and auditory experience of using the headphones.

[0005] The technical solution provided by this invention is as follows:

[0006] This invention provides an earphone, comprising:

[0007] The processing module acquires audio data to be played, which includes first audio data and second audio data. The first audio data includes audio data corresponding to a first character or a first language model, and the second audio data includes audio data corresponding to a second character or a second language model.

[0008] The first audio module receives and plays the first audio data.

[0009] The second audio module receives and plays the second audio data.

[0010] Specifically, the headphones proposed in this invention can not only distribute the audio data of different characters to different audio modules for playback, but the language model in the headphones can also generate audio data to be played based on the user's interactive content, thereby realizing an interactive experience between the headphones and the user.

[0011] Furthermore, the headphones also include:

[0012] The voice acquisition module is used to collect user voice data.

[0013] The processing module identifies valid commands from the user's voice. Valid commands are interactive commands from the user in response to the first audio data or the second audio data. The module analyzes the valid commands and adjusts the first audio data or the second audio data accordingly.

[0014] Specifically, the earphone proposed in this invention has a voice acquisition module that collects the user's voice, a processing module that extracts valid commands for controlling audio playback, and a playback of audio data based on the valid commands.

[0015] Furthermore, it also includes:

[0016] The processing module creates audio data to be played based on the image of different characters; or, it enhances the sound effects of the first audio data and the second audio data based on the image of different characters.

[0017] Specifically, the headphones proposed in this invention can enhance the personalization and recognizability of characters by customizing specific sound effects for different characters, and can also enhance the spatial sense and immersiveness of the sound playback.

[0018] Furthermore, the headphones also include:

[0019] The storage module stores the audio data to be played; or, it stores the first audio data or the second audio data in categories; or, it stores the first language model and the second language model in categories.

[0020] Specifically, the headphones proposed in this invention store audio data to be played through a storage module, enabling the headphones to quickly access and play audio data, thus improving playback efficiency. When processing multi-track audio or multilingual content, the headphones provide more diverse content and services by classifying and storing different audio data and different language models.

[0021] The present invention also proposes an earphone system, comprising:

[0022] The smart device includes a first communication module that acquires and sends audio data to be played. The audio data to be played includes first audio data, second audio data, and narration audio data. The first audio data includes audio data corresponding to a first character or a first language model, and the second audio data includes audio data corresponding to a second character or a second language model.

[0023] The headphones include a second communication module, a first audio module, and a second audio module. The second communication module receives first audio data and second audio data, plays the first audio data through the first audio module, and plays the second audio data through the second audio module.

[0024] Specifically, through the headphone system proposed in this invention, smart devices can play content according to user needs or context, and distribute audio data of different characters to the headphones for playback. The language model in the smart device can also generate audio data to be played based on the user's interactive content, thereby realizing an interactive experience between the headphone system and the user.

[0025] Furthermore, the headphones include:

[0026] The voice acquisition module is used to collect user voice data.

[0027] The second processing module identifies valid commands from the user's voice. Valid commands are interactive commands from the user in response to the first audio data or the second audio data.

[0028] Specifically, through the headphone system proposed in this invention, the headphones can collect the user's voice and identify valid commands in the voice that affect the playback of audio data.

[0029] Furthermore, smart devices also include:

[0030] The first processing module receives and analyzes valid instructions, and adjusts the first audio data or the second audio data according to the valid instructions.

[0031] The first processing module also creates audio data to be played based on the image of different characters; or, it performs sound effect enhancement processing on the first audio data and the second audio data based on the image of different characters.

[0032] The storage module stores the audio data to be played; or, it stores the first audio data or the second audio data in categories; or, it stores the first language model and the second language model in categories.

[0033] Specifically, through the headphone system proposed in this invention, smart devices can not only adjust the playback of subsequent audio data according to valid instructions, but also customize specific sound effects for different characters, enhancing the personalization and recognizability of the voice-speaking characters.

[0034] Furthermore, smart devices also include:

[0035] The third audio module is used to play the narration audio data.

[0036] The present invention also proposes an interaction method for headphones, using the above-described headphones or headphone system, comprising the following steps:

[0037] Obtain the audio data to be played, which includes first audio data, second audio data, and narration audio data; the first audio data includes the audio data corresponding to the first character or the first voice model, and the second audio data includes the audio data corresponding to the second character or the second language model.

[0038] Play the first audio data or the second audio data.

[0039] Specifically, through the headphone interaction method proposed in this invention, different audio data correspond to different voice actors, and the language model generates audio data to be played based on the user's interactive content, thereby realizing an interactive experience between the headphone and the user.

[0040] Furthermore, before playing the first audio data or the second audio data, the following is included:

[0041] Collect user voice data and identify valid commands from it; valid commands are interactive commands from which the user responds to the first audio data or the second audio data.

[0042] Analyze valid instructions and adjust the first or second audio data accordingly.

[0043] Specifically, the headphone interaction method proposed in this invention realizes a cyclical process of audio playback involving voice interaction and dynamic adjustment of audio data, which improves user participation and sense of control, making the audio experience more personalized and interactive.

[0044] The present invention provides an earphone, an earphone system, and an earphone interaction method. By assigning each audio module in the earphone to a different role, and having each audio module provide different suggestions and play different audio data, users can enjoy a more intelligent and personalized interactive experience. Attached Figure Description

[0045] The preferred embodiments will now be described in a clear and easy-to-understand manner, with reference to the accompanying drawings, to further explain the above-mentioned characteristics, technical features, advantages, and implementation methods of an earphone, an earphone system, and an earphone interaction method.

[0046] Figure 1 This is a schematic diagram of the structure of an embodiment of an earphone according to the present invention;

[0047] Figure 2 This is a schematic diagram of another embodiment of an earphone according to the present invention;

[0048] Figure 3 This is a schematic diagram of the structure of an embodiment of a headphone system according to the present invention;

[0049] Figure 4 This is a schematic diagram of another embodiment of the headphone system of the present invention;

[0050] Figure 5 This is a flowchart of an embodiment of an interactive method for headphones according to the present invention;

[0051] Figure 6This is a flowchart of another embodiment of an interactive method for headphones according to the present invention. Detailed Implementation

[0052] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0053] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or sets.

[0054] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0055] Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the specific implementation methods of the present invention will be described below with reference to the accompanying drawings. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings and other implementation methods can be obtained based on these drawings without any creative effort.

[0057] Current smart headphones primarily focus on sound quality, comfort, and noise cancellation, but have limitations in interactivity and immersion. The audio output of smart headphones is typically unidirectional, only providing functions such as music playback, answering calls, and watching videos. While some smart headphones integrate voice assistants, these functions are mainly concentrated on audio playback control and basic task execution, such as playing music and answering calls, lacking more complex conversational capabilities and personalized interactions.

[0058] For example, when experiencing an immersive interactive story through headphones, such as an interactive novel or audiobook, users cannot obtain different perspectives or decisions from the headphones, nor can they interact with the characters in the headphones based on the interactive information played in the audio.

[0059] Therefore, this invention proposes an earphone in which two audio modules are assigned different roles, with each audio module providing different suggestions and playing different audio data, enabling users to enjoy a more intelligent and personalized interactive experience.

[0060] Based on the headphones proposed in this invention, the left and right earpieces correspond to different characters or language models, adjust the audio content being played according to voice interaction, and provide different viewpoints and suggestions based on the user's needs or context, thereby enhancing the user's immersive interactive experience when using the headphones.

[0061] The following description is in conjunction with the accompanying drawings:

[0062] In one embodiment of the present invention, please refer to Figure 1 This illustration shows a structural schematic of an earphone provided in some embodiments of the present disclosure. The earphone 100 includes:

[0063] Processing module 110 acquires audio data to be played, which includes first audio data and second audio data; the first audio data includes audio data corresponding to the first character or the first language model, and the second audio data includes audio data corresponding to the second character or the second language model.

[0064] The first audio module 121 receives and plays the first audio data.

[0065] The second audio module 122 receives and plays the second audio data.

[0066] Specifically, when the audio data to be played involves multiple characters, the processing module 110, according to the user's selection, splits the audio data corresponding to the two characters selected by the user into two independent audio data sets: the first audio data set and the second audio data set. For example, if the audio data to be played is a dialogue involving characters A, B, and C, and the user selects characters A and B, the processing module 110 assigns the dialogue content of character A to the first audio data set and the dialogue content of character B to the second audio data set.

[0067] While splitting the audio data to be played, the processing module 110 also adds timestamps and synchronization information to the first and second audio data to ensure that the audio of each character can be correctly synchronized during playback. When the audio data to be played is not split, the processing module 110 can identify the audio tracks corresponding to different characters in the audio data to be played and allocate the audio tracks of different characters to the first audio module 121 and the second audio module 122.

[0068] When the earphone 100 is used as a smart assistant, the left and right ears of the earphone 100 are similar to the left and right hemispheres of the human brain, which process different types of information respectively. The smart assistant needs to provide different opinions and suggestions based on the user's questions or needs. The processing module 110 generates the first audio data and the second audio data of the answer opinion according to the first language model and the second language model (such as natural language processing model, dialogue management model, speech synthesis model, etc.).

[0069] For example, the natural language processing model in processing module 110 understands the user's intent and question, the dialogue management model provides appropriate text response data based on the user's question, and the speech synthesis model converts the text response data into audio data to be played so that the user can hear opinions and suggestions.

[0070] In this embodiment, the split first and second audio data still need to be dubbed according to the character's image. The processing module 110 can also:

[0071] Create audio data to be played based on the image of different characters; or, enhance the sound effects of the first audio data and the second audio data based on the image of different characters.

[0072] Specifically, firstly, the processing module 110 analyzes the audio of different characters to ensure that the quality and expressiveness of the sound match the character's image, and adjusts the audio volume to avoid problems of being too loud or too soft.

[0073] Secondly, the processing module 110 also needs to adjust the frequency response of the audio to highlight the character's vocal characteristics. For example, if character A is a deep villain, the processing module 110 will enhance the low-frequency part to make the voice sound deeper and more powerful; if character B is a lively child character, the processing module 110 will enhance the high-frequency part to make the voice sound crisper.

[0074] Finally, the processing module 110 also involves compressing the dynamic range of the audio data to ensure a consistent listening experience at different volume levels; adding appropriate reverb effects to simulate the environment in which the character is located, such as indoors, outdoors, or special scenes, thereby enhancing the sense of space and immersion when the headphones 100 play audio.

[0075] The audio data to be played (first audio data and second audio data) processed by the processing module 110 will be transmitted to the first audio module 121 and the second audio module 122 for playback.

[0076] When storing the audio data to be played, the headphones 100 further include:

[0077] Storage module 130 stores audio data to be played; or, stores first audio data or second audio data in categories; or, stores first language model and second language model in categories.

[0078] Specifically, the storage module 130 is responsible for saving the audio data to be played, ensuring that these audio files can be quickly and accurately retrieved during playback, guaranteeing the smoothness and stability of the audio playback process. The audio data to be played also includes various audio formats, such as MP3, WAV, and AAC, ensuring that users can choose the appropriate audio format for playback according to their needs.

[0079] The storage module 130 can intelligently identify and distinguish the first audio data and the second audio data, and classify and store the split first audio data and second audio data to ensure that these audio data can be stored according to the classification rules, making the management of audio data more organized and systematic.

[0080] For example, the storage module 130 in the headphones 100 stores audio data about interactive courses. By selecting an interactive history course, students can role-play historical figures through the headphones 100, have conversations with the headphones 100, and gain a deeper understanding of historical events.

[0081] When the earphone 100 acts as a smart assistant, in order to provide users with different perspectives and suggestions, the storage module 130 will also store language models related to different smart assistants, such as a first language model and a second language model, so as to generate different question and answer information.

[0082] In this embodiment, the headphones 100 proposed by the present invention can not only distribute the audio data of different characters to different audio modules for playback, but the language model in the headphones 100 can also generate audio data to be played based on the user's interactive content, thereby realizing an immersive and interactive audio experience between the headphones 100 and the user.

[0083] In one embodiment of the present invention, please refer to Figure 2 The headphones 100 also include:

[0084] The voice acquisition module 140 is used to acquire user voice.

[0085] The processing module 110 identifies valid commands from the user's voice. Valid commands are interactive commands from the user in response to the first audio data or the second audio data. The module analyzes the valid commands and adjusts the first audio data or the second audio data according to the valid commands.

[0086] Specifically, the language acquisition module 140 is dedicated to capturing the user's voice input, and the acquired user language is sent to the processing module 110. The processing module 110 performs a series of preprocessing steps on the user's voice, identifying valid instructions from the user's interaction commands. These valid instructions are the user's response to the first audio data or the second audio data, such as confirmation, rejection, inquiry, or request for more information.

[0087] When analyzing valid instructions, the processing module 110 needs to perform semantic recognition, contextual association, and possible intent inference on the valid instructions to ensure that the valid instructions are successfully recognized and understood, and then adjust the first audio data or the second audio data, including but not limited to increasing or decreasing the volume, adjusting the playback speed, and selecting the playback content.

[0088] For example, when an audiobook is selected to be played in the storage module 130 of the headset 100, the first audio module 121 issues a voice prompt, "Do you want character A to go to the castle?" The voice acquisition module 140 captures the user's voice, and after the processing module 110 recognizes a valid "yes" instruction, it adjusts the playback of the first audio data for "after character A goes to the castle." Alternatively, in a puzzle game, the user interacts with the headset 100 through voice prompts to solve the puzzle.

[0089] In the earphone 100, the voice acquisition module 140 can be shared by the first audio module 121 and the second audio module 122, or each audio module can have its own dedicated voice acquisition module 140.

[0090] When the headset 100 has two built-in voice acquisition modules 140, the user issues interactive commands based on the playing audio, and both voice acquisition modules 140 will collect the user's voice. For example, when character A in the first audio module 121 interacts with the user, both voice acquisition modules 140 will collect the user's voice, and the processing module 110 will only process the user's voice collected by the voice acquisition module 140 related to the first audio module 121. The other voice acquisition module 140 only collects the user's voice and does not send the voice to the processing module 110 for recognition processing.

[0091] In this embodiment, the earphone 100 proposed in this invention uses a voice acquisition module 140 to acquire the voice emitted by the user, and a processing module 110 to extract valid instructions for controlling the playback of audio. Based on the valid instructions, the playback of audio data is dynamically adjusted to ensure that the playback of audio data meets the user's interactive expectations.

[0092] In one embodiment of the present invention, please refer to Figure 3 This illustration shows a schematic diagram of the structure of a headphone system provided in some embodiments of the present disclosure. The headphone system 200 includes:

[0093] The smart device 210 includes a first communication module 211, which acquires and sends audio data to be played. The audio data to be played includes first audio data, second audio data, and narration audio data. The first audio data includes audio data corresponding to a first character or a first language model, and the second audio data includes audio data corresponding to a second character or a second language model.

[0094] The earphone 220 includes a second communication module 221, a first audio module 222, and a second audio module 223. The second communication module 221 receives first audio data and second audio data, plays the first audio data through the first audio module 222, and plays the second audio data through the second audio module 223.

[0095] Specifically, when the earphone 220 is used in conjunction with the smart device 210 via a wireless or wired connection, it constitutes an earphone system 200. The smart device 210 can be a smartphone, smart tablet, smartwatch, smart TV, smart speaker, etc.

[0096] After the earphone 220 establishes a communication connection (Bluetooth / Wi-Fi / wired) with the smart device 210, the audio data to be played can be directly obtained from the application installed on the smart device 210. The application is responsible for downloading, managing, and updating the audio data to be played. The first audio data and the second audio data can be sent to the second communication module 221 in the earphone 220 through the first communication module 211 of the smart device 210, and then the second communication module 221 sends the first audio data and the second audio data to the first audio module 222 and the second audio module 223, respectively.

[0097] The smart device 210 can process the audio data to be played. It can acquire the audio data to be played and split the data into two or more audio data streams according to the different characters contained in the audio data. It can identify the voice features of each character in the audio and split the audio into first audio data and second audio data based on the character selected when the user starts the headphone system 200. The audio data corresponding to other characters is used as narration data.

[0098] The smart device 210 can also function as a smart assistant. In response to user questions, the multiple intelligent language models within the smart device 210 understand the user's questions and provide personalized suggestions and interactive content. For example, when the smart device 210 provides two smart assistants (representing different logical thinking styles, such as left and right brain), one smart assistant can be assigned to the left earpiece, and another to the right earpiece.

[0099] To ensure effective playback of the audio data to be played, the smart device 210 is typically equipped with multiple audio playback channels, which can be physical speakers, headphone jacks, or wireless audio transmission devices. Through these audio playback channels, the first and second audio data can be sent to the headphones 220 simultaneously or sequentially to meet the user's listening experience in different scenarios.

[0100] In this embodiment, the headphone system 200 proposed in this invention can intelligently process the audio data to be played, providing an innovative audio experience that enhances the user's auditory enjoyment and immersive interaction.

[0101] In one embodiment of the present invention, please refer to Figure 4 The diagram illustrates the structure of a headphone system provided in some embodiments of this disclosure. The headphone 220 of the headphone system 200 includes: a voice acquisition module 224 and a second processing module 225.

[0102] The voice acquisition module 224 is used to acquire user voice.

[0103] The second processing module 225 identifies valid commands from the user's voice. Valid commands are interactive commands from the user in response to the first audio data or the second audio data.

[0104] The smart device 210 includes: a first processing module 212, a storage module 213, and a third audio module 214.

[0105] The first processing module 212 receives and analyzes valid instructions, and adjusts the first audio data or the second audio data according to the valid instructions; the first processing module 212 also creates audio data to be played according to the image of different characters; or, performs sound effect enhancement processing on the first audio data and the second audio data according to the image of different characters.

[0106] Storage module 213 stores audio data to be played; or, stores first audio data or second audio data in categories; or, stores first language model and second language model in categories.

[0107] The third audio module 214 is used to play narration audio data.

[0108] Specifically, the first processing module 212 in the smart device 210 analyzes and adjusts the frequency response of the audio based on the audio of different characters to ensure that the quality and expressiveness of the sound match the character's image and highlight the character's vocal characteristics. The storage module 213 of the smart device can intelligently identify and classify the storage of the first audio data, the second audio data, and the narration audio data, making the management of audio data more organized and systematic.

[0109] For example, in a multi-character dialogue scenario, the first audio data may be played by the dialogue of character A in the first audio module 222 of the headphones 220, while the second audio data may be played by the dialogue of character B in the second audio module 223 of the headphones 220, and the narration audio data may be played by the third audio module 214 of the smart device 210, thereby creating a stereo and immersive audio effect.

[0110] When a user interacts, the voice acquisition module 224 of the earphone 220 acquires the user's voice. The second processing module 225 identifies valid commands for the first or second audio data and sends these commands to the first processing module 212 in the smart device 210 for processing. The first processing module 212 performs semantic recognition, contextual analysis, and possible intent inference on the valid commands to ensure that the commands are successfully recognized and understood. Adjustments to the first or second audio data are then made, including but not limited to volume adjustments, playback speed adjustments, and content selection.

[0111] After the first processing module 212 processes the valid instruction, it sends the adjusted first / second audio data to the second communication module 221 through the first communication module 211. Then, the first audio module 222 and the second audio module 223 play the audio data respectively.

[0112] It should be noted that the specific execution process of each module and the relationship between each module in this embodiment can be found in the corresponding content of the aforementioned headphone embodiment, and will not be repeated here.

[0113] In one embodiment of the present invention, please refer to Figure 5 The document illustrates a flowchart of an interaction method for headphones provided in some embodiments of this disclosure. The method includes the following steps:

[0114] S100 acquires audio data to be played, which includes the first audio data and the second audio data; the first audio data includes audio data corresponding to the first character or the first language model, and the second audio data includes audio data corresponding to the second character or the second language model.

[0115] The S200 controls the playback of either the first or second audio data.

[0116] Specifically, upon activating the headphones or the headphone system (establishing a communication connection between the headphones and the smart device), the user can select audio data (interactive stories or lessons) to be played from the audio data stored in the headphones or from the smart device's application. Based on the user's selection, the audio data to be played is split into first audio data and second audio data. The audio data other than the first and second audio data serves as narration. The left and right audio modules of the headphones play the voices of different characters respectively. For example, the left ear audio module plays the first audio data, and the right ear audio module plays the second audio data. Narration can be played simultaneously in both ears or by other devices.

[0117] When no audio data to be played is selected and the headphones are used as a smart assistant, the left and right audio modules of the headphones correspond to two smart assistants with different logical thinking. The two smart assistants correspond to two language models, and the left and right headphones provide two different answers or suggestions based on the user's questions.

[0118] In this embodiment, through the headphone interaction method proposed in this invention, users can not only experience stereo audio effects, but also interact with the intelligent assistant through question and answer to enhance their sense of participation and bring an immersive interactive experience.

[0119] In one embodiment of the present invention, before playing the first audio data or the second audio data, please refer to... Figure 6 The above step S200 includes:

[0120] S210 collects user voice and identifies valid commands from the user's voice; valid commands are interactive commands from which the user responds to the first audio data or the second audio data.

[0121] S220 analyzes valid commands and adjusts the first audio data or the second audio data according to the valid commands.

[0122] Specifically, before playing the first or second audio data, the audio data will emit interactive information. The user will issue interactive instructions based on the interactive information, and the valid instructions (e.g., which audio data the instruction controls) will be identified from the user's response.

[0123] Based on the selection made in the valid instructions, the audio data to be played is dynamically adjusted. The audio data to be played develops differently according to the user's selection until the audio data playback ends.

[0124] In this embodiment, through the headphone interaction method proposed in this invention, the user does not need to operate manually and can directly interact with the headphone through language to achieve interactive response, thereby enhancing the convenience of the user in controlling audio data playback.

[0125] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0126] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0128] Furthermore, the functional units in the various embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit.

[0129] It should be noted that the above embodiments can be freely combined as needed. The above are merely preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An earphone, characterized in that, include: The processing module acquires audio data to be played, which includes first audio data and second audio data; the first audio data includes audio data corresponding to a first character or a first language model, and the second audio data includes audio data corresponding to a second character or a second language model. The first audio module receives and plays the first audio data; The second audio module receives and plays the second audio data.

2. The earphone according to claim 1, characterized in that, Also includes: The voice acquisition module is used to collect user voice data. The processing module identifies valid commands from the user's voice, whereby the valid commands are interactive commands from the user in response to the first audio data or the second audio data; analyzes the valid commands, and adjusts the first audio data or the second audio data according to the valid commands.

3. The earphone according to claim 1, characterized in that, Also includes: The processing module generates the audio data to be played based on the image of different characters; Alternatively, sound effects enhancement processing can be applied to the first audio data and the second audio data based on the image of different characters.

4. The headphones according to any one of claims 1 to 3, characterized in that, Also includes: Storage module, for storing the audio data to be played; Alternatively, the first audio data or the second audio data can be stored separately. Alternatively, the first language model and the second language model can be stored separately.

5. A headphone system, characterized in that, include: The smart device includes a first communication module that acquires and sends audio data to be played, the audio data to be played including first audio data, second audio data, and narration audio data; the first audio data includes audio data corresponding to a first character or a first language model, and the second audio data includes audio data corresponding to a second character or a second language model; The earphone includes a second communication module, a first audio module, and a second audio module. The second communication module receives the first audio data and the second audio data, plays the first audio data through the first audio module, and plays the second audio data through the second audio module.

6. The headphone system according to claim 5, characterized in that, The headphones also include: The voice acquisition module is used to collect user voice data. The second processing module identifies valid instructions from the user's voice, which are interactive instructions from the user in response to the first audio data or the second audio data.

7. The headphone system according to claim 5, characterized in that, The smart device also includes: The first processing module receives and analyzes valid instructions, and adjusts the first audio data or the second audio data according to the valid instructions. The first processing module also generates the audio data to be played based on the image of different characters; or, performs sound effect enhancement processing on the first audio data and the second audio data based on the image of different characters. The storage module stores the audio data to be played; or, it stores the first audio data or the second audio data in categories; or, it stores the first language model and the second language model in categories.

8. The headphone system according to claim 5, characterized in that, The smart device also includes: The third audio module is used to play the narration audio data.

9. An interaction method for headphones, characterized in that, Using the headphones or headphone system according to any one of claims 1-8 includes the following steps: Acquire audio data to be played, wherein the audio data to be played includes the first audio data and the second audio data; the first audio data includes audio data corresponding to the first character or the first language model, and the second audio data includes audio data corresponding to the second character or the second language model. Play the first audio data or the second audio data.

10. The interaction method for headphones according to claim 9, characterized in that, Before playing the first audio data or the second audio data, the following is included: Collect user voice data and identify valid commands from the user voice data; the valid commands are interactive commands from the user in response to the first audio data or the second audio data. Analyze the valid instructions and adjust the first audio data or the second audio data according to the valid instructions.