A voice interaction method, a smart earphone and a storage medium

By collecting and parsing user voice request data in the charging case, the compatibility between the earbuds and the charging case is determined, and an appropriate translation package is configured for each smart earbud. This solves the problem of translation response delay in existing technologies and achieves efficient real-time translation.

CN121053970BActive Publication Date: 2026-07-21SHENZHEN TIDE COMMUNICATIONS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN TIDE COMMUNICATIONS CO LTD
Filing Date
2025-09-20
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies cannot configure appropriate translation packages or models for the left and right earbuds based on the actual scenario in which the user initiates a voice request, resulting in translation response delays and low efficiency.

Method used

The charging case collects voice request data initiated by the user, performs semantic parsing and language information extraction, determines the compliance relationship between the left and right earbuds and the charging case, and configures target translation packages for the left and right earbuds respectively for real-time translation based on language text information and compliance relationship.

Benefits of technology

It significantly reduces translation response latency, improves translation efficiency, and ensures the stability and security of the translation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053970B_ABST
    Figure CN121053970B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of speech recognition, and discloses a speech interaction method, an intelligent earphone and a storage medium. The method first collects speech request data initiated by a user by a charging bin, performs semantic analysis and language information extraction on the speech request data, obtains semantic text information representing a user operation intention and language text information representing a target language demand, determines that the operation intention represented by the semantic text information is a user request to enter a translation mode, sends interconnection instructions to a left ear earphone and a right ear earphone, then determines a following relationship of the left ear earphone, the right ear earphone and the charging bin according to acoustic parameters in a speech request process, and then configures target translation packages for the left ear earphone and the right ear earphone according to the language text information and the following relationship to perform real-time translation. Compared with a translation mode of uniformly loading general resources, the application greatly reduces translation response delay and improves translation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, specifically to a speech interaction method, a smart headset, and a storage medium. Background Technology

[0002] With the widespread adoption of smart wearable devices, smart headphones have evolved from traditional audio playback devices into "voice interaction terminals," and their application scenarios have expanded from simple music listening to diverse fields such as real-time translation, voice control, and social communication. Among these, real-time translation functionality based on smart headphones has become one of the key areas of current technological research and development due to its ability to meet the convenience needs of cross-language communication (such as cross-border travel, international conferences, and multilingual daily communication).

[0003] To achieve translation functionality in smart earphones, various solutions have been proposed in the existing technology. For example, Korean patent application KR102393112B1 discloses a method and apparatus for implementing translation functionality using earphones. This technology achieves bidirectional translation between a first and a second language through the cooperation of a first and a second earphone, thus reducing reliance on external terminals to some extent. However, this solution still has significant technical limitations: it cannot configure adapted translation packages or models for the left and right earphones respectively based on the actual scenario of the user initiating a voice request. It can only load general translation resources uniformly after translation is initiated, and then have the two earphones process the translation task synchronously, resulting in translation response delays and low efficiency. Summary of the Invention

[0004] The main objective of this invention is to provide a voice interaction method, a smart earphone, and a storage medium, aiming to solve the technical problem in the prior art that it is impossible to configure appropriate translation packages or translation models for the left and right earphones respectively according to the actual scenario of the user initiating a voice request, resulting in translation response delay and low efficiency.

[0005] To achieve the above objectives, in a first aspect, this application provides a voice interaction method applied to a smart earphone, the smart earphone including a charging case, a left earphone, and a right earphone, the method including:

[0006] The charging case collects voice request data initiated by the user, and performs semantic parsing and language information extraction on the voice request data to obtain semantic text information representing the user's operation intention and language text information representing the target language requirement.

[0007] If the operational intent represented by the semantic text information is determined to be a user request to enter translation mode, an interconnection command is sent to the left and right earpieces.

[0008] The compliance relationship between the left earbud, the right earbud, and the charging case is determined based on the acoustic parameters during the voice request process. The compliance relationship includes a first compliance relationship and a second sequential relationship. The first compliance relationship indicates that the left earbud and the charging case jointly comply with the voice request user, and the second compliance relationship indicates that the right earbud and the charging case jointly comply with the voice request user.

[0009] Based on the language text information and the dependency relationship, target translation packages are configured for the left and right earbuds respectively for real-time translation.

[0010] In one possible implementation, the acoustic parameters include a first time difference and a second time difference, wherein the first time difference is the time difference between the arrival of voice request data in the left earbud and the charging case, and the second time difference is the time difference between the arrival of voice request data in the right earbud and the charging case.

[0011] The determination of the compliance relationship between the left earbud, the right earbud, and the charging case based on acoustic parameters during the voice request process includes:

[0012] If the first time difference is determined to be less than the second time difference, the compliance relationship is determined to be a first compliance relationship.

[0013] If the first time difference is determined to be greater than the second time difference, the compliance relationship is determined to be the second compliance relationship.

[0014] In one possible implementation, after obtaining the language text information representing the target language requirements, the following is also included:

[0015] The language in the language text information that matches the voice request is marked as the first target language, and other languages ​​in the language text information that are not the first target language are marked as the second target language.

[0016] In one possible implementation, configuring target translation packages for real-time translation for the left and right earbuds respectively based on the language text information and dependency relationships includes:

[0017] The dependency relationship is determined to be a first dependency relationship. A first target translation package and a second target translation package are obtained from the cloud server. The first target translation package is a translation package that translates the second target language into the first target language, and the second target translation package is a translation package that translates the first target language into the second target language.

[0018] Configure the right earbud with the first target translation packet so that the left earbud outputs the first target language; and,

[0019] Configure the second target translation packet for the left earphone so that the right earphone outputs the second target language.

[0020] In one possible implementation, the method further includes:

[0021] During real-time translation, the speech data corresponding to the first target language collected by the right earphone is filtered, and the speech data corresponding to the second target language collected by the left earphone is also filtered.

[0022] In one possible implementation, configuring target language translation packages for real-time translation on the left and right earbuds respectively, based on the language text information and dependency relationships, includes:

[0023] The dependency relationship is determined to be a second dependency relationship. A first target translation package and a second target translation package are obtained from the cloud server. The first target translation package is a translation package that translates the second target language into the first target language, and the second target translation package is a translation package that translates the first target language into the second target language.

[0024] Configure the right earphone with the second target translation packet so that the left earphone outputs the second target language; and,

[0025] Configure the first target translation packet for the left earphone so that the right earphone outputs the first target language.

[0026] In one possible implementation, the method further includes:

[0027] During real-time translation, the speech data corresponding to the second target language collected by the right earphone is filtered, and the speech data corresponding to the first target language collected by the left earphone is also filtered.

[0028] In one possible implementation, after obtaining the semantic text information representing the user's operational intent, the following is also included:

[0029] Determine the operational intent represented by the semantic text information to enter translation mode without requesting it, and determine whether the left and right earphones are in an interconnected state.

[0030] If the left and right earbuds are in an interconnected state, send an interconnection release command to the left and right earbuds; and / or,

[0031] After determining the compliance relationship between the left earbud, the right earbud, and the charging case based on acoustic parameters during the voice request process, the method further includes:

[0032] Real-time monitoring of the exchange characteristics of the left and right earphones, the exchange characteristics including earphone motion trajectory parameters or real-time distance change parameters between the left and right earphones and the user, wherein the earphone motion trajectory parameters include the degree of overlap between the motion trajectories of the left and right earphones;

[0033] When headphone switching characteristics are detected, it is determined to be a headphone switching event;

[0034] In response to the headphone exchange event, the compliance relationship reset mechanism is automatically triggered, the acoustic parameters during the voice request process are re-acquired for secondary judgment, and the compliance relationship is synchronously updated based on the secondary judgment result.

[0035] Secondly, this application also provides a smart headset, including: a memory and a processor, wherein the memory is used to store program code; and the processor is used to call the program code to execute the method as described in the first aspect.

[0036] Thirdly, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the first aspect.

[0037] Unlike existing technologies, this application provides a voice interaction method, a smart earphone, and a storage medium. The method is applied to a smart earphone, which includes a charging case, a left earphone, and a right earphone. First, the charging case collects voice request data initiated by the user and performs semantic parsing and language information extraction on the data to obtain semantic text information representing the user's operational intent and language text information representing the target language requirement. The operational intent represented by the semantic text information is determined to be a user request to enter translation mode, and an interconnection command is sent to the left and right earphones. Then, the compliance relationship between the left and right earphones and the charging case is determined based on acoustic parameters during the voice request process. Next, based on the language text information and the compliance relationship, target translation packages are configured for the left and right earphones respectively for real-time translation. In other words, this application's solution can clarify the compliance relationship between the two earphones and the charging case based on the actual scenario of the user's voice request (quantifying scenario features through acoustic parameters), thereby pre-configuring appropriate translation packages for the left and right earphones according to this compliance relationship. Compared to a translation mode that uniformly loads general resources, this significantly reduces translation response latency and improves translation efficiency. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0039] Figure 1 These are interactive schematic diagrams of the smart earphones in some embodiments of this application;

[0040] Figure 2 This is a flowchart illustrating the voice interaction method in some embodiments of this application;

[0041] Figure 3 This is a flowchart illustrating step S300 of the voice interaction method in some embodiments of this application;

[0042] Figure 4 This is a schematic diagram of the hardware structure of the smart earphone in some embodiments of this application.

[0043] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0045] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0046] Furthermore, the use of terms such as "first" and "second" in this invention is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the term "and / or" throughout the text includes three solutions; taking A and / or B as an example, it includes technical solution A, technical solution B, and a technical solution that simultaneously satisfies A and B. Furthermore, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of a person skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0047] With the widespread adoption of smart wearable devices, smart headphones have evolved from traditional audio playback devices into "voice interaction terminals," and their application scenarios have expanded from simple music listening to diverse fields such as real-time translation, voice control, and social communication. Among these, real-time translation functionality based on smart headphones has become one of the key areas of current technological research and development due to its ability to meet the convenience needs of cross-language communication (such as cross-border travel, international conferences, and multilingual daily communication).

[0048] To achieve translation functionality in smart earphones, various solutions have been proposed in the existing technology. For example, Korean patent application KR102393112B1 discloses a method and apparatus for implementing translation functionality using earphones. This technology achieves bidirectional translation between a first and a second language through the cooperation of a first and a second earphone, thus reducing reliance on external terminals to some extent. However, this solution still has significant technical limitations: it cannot configure adapted translation packages or models for the left and right earphones respectively based on the actual scenario of the user initiating a voice request. It can only load general translation resources uniformly after translation is initiated, and then have the two earphones process the translation task synchronously, resulting in translation response delays and low efficiency.

[0049] To address the aforementioned technical problems, this application provides a voice interaction method that can be applied to electronic devices such as smart headsets, or to service systems that include smart headsets and servers. Figure 1As shown in this embodiment, the smart earphones include a charging case 100, a left earphone 200, and a right earphone 300. The charging case 100 not only provides basic charging functionality but also serves as a processing terminal for the left earphone 200 and right earphone 300. For example, it can control the left earphone 200 and right earphone 300 to interact and connect through the charging case 100, forming an earphone pair with translation service capabilities. The specific workflow is as follows: The charging case 100 pre-configures the corresponding translation package or translation model from the cloud server 400. When the left earphone 200 receives speech data in language A, it transmits it to the charging case 100. The charging case 100 uses the configured local translation package or translation model to translate the speech data, generating speech data in language B, which is then sent to the right earphone 300 for playback. Furthermore, during real-time translation, the charging case 100 can also record and save the received speech data, providing a data foundation for subsequent iterative optimization of the translation package or translation model.

[0050] The following explanation uses a smart headset as an example to illustrate this voice interaction method. It should be noted that although the flowchart shows the logical order, in some cases, the steps shown or described may be performed in a different order. Please refer to the appendix. Figure 2 The method includes the following steps S100-S400:

[0051] Step S100: The charging case collects voice request data initiated by the user, and performs semantic parsing and language information extraction on the voice request data to obtain semantic text information representing the user's operation intention and language text information representing the target language requirement.

[0052] The charging case contains an audio pickup and microprocessor, which can pick up and recognize voice request data initiated by the user, and analyze the voice data.

[0053] During the analysis, semantic parsing and language information extraction can be performed on the voice request data. Semantic parsing can specifically obtain semantic text information that represents the user's operational intent, such as translation needs, dialogue needs, or music playback needs.

[0054] Language information extraction specifically involves analyzing and extracting the language corresponding to the voice request data itself, as well as the languages ​​explicitly mentioned in the voice request data. The target language requirement refers to the target language the user wishes to interact with. For example, when a user's voice request enters translation mode, if the language corresponding to the voice request data itself is language A, and the voice request data explicitly mentions languages ​​A and B, then the user needs a translation service between languages ​​A and B. In this case, languages ​​A and B are identified as language text information. Similarly, if the language corresponding to the voice request data itself is language A, and the voice request data explicitly mentions language B, then the user also needs a translation service between languages ​​A and B, and languages ​​A and B are also identified as language text information.

[0055] After obtaining the language text information representing the target language requirement, the method further includes: marking the language in the language text information that is consistent with the voice request as the first target language, and marking other languages ​​in the language text information other than the first target language as the second target language.

[0056] For example, when a user issues the voice command "Please enter Chinese-English translation mode" in Chinese, the charging case, after collecting and analyzing the voice request data, obtains the semantic text information representing the user's intention as "Enter translation mode," and the language text information representing the target language requirement as "Chinese, English." In this case, Chinese is marked as the first target language, and English is marked as the second target language.

[0057] For example, when a user issues the voice command "Switch to Chinese-English bilingual dialogue mode" in English, the charging case collects and analyzes the voice request data, obtaining the semantic text information representing the user's intention as "Switch to bilingual dialogue mode"; and the language text information representing the target language requirement as "Chinese, English". In this case, English is marked as the first target language, and Chinese is marked as the second target language.

[0058] Step S200: Determine that the operational intent represented by the semantic text information is a user request to enter translation mode, and send an interconnection command to the left earphone and the right earphone;

[0059] After obtaining the semantic text information in step S100, if it is determined that the operation intent represented by the information is a user request to enter translation mode, the charging case will send an interconnection command to the left and right earbuds. At this time, the left and right earbuds establish a dedicated communication channel through the charging case, which can use communication methods such as Bluetooth or a local area network.

[0060] Understandably, once a dedicated communication channel is established, it possesses independent connection isolation characteristics, unaffected by connection requests from other external devices. This effectively avoids communication interruptions, data transmission delays, or information leaks caused by external connection requests, ensuring the stability and security of voice data transmission during translation. Furthermore, by optimizing data forwarding paths, the dedicated communication channel ensures that voice data received by the left earpiece, after translation processing, can be directly transmitted to the right earpiece without delay; conversely, voice data received by the right earpiece, after translation, can also reach the left earpiece in real time, thus achieving instant response in two-way translation.

[0061] Step S300: Determine the compliance relationship between the left earphone, the right earphone and the charging case based on the acoustic parameters during the voice request process. The compliance relationship includes a first compliance relationship and a second sequential relationship. The first compliance relationship indicates that the left earphone and the charging case jointly comply with the voice request user. The second compliance relationship indicates that the right earphone and the charging case jointly comply with the voice request user.

[0062] It should be noted that in translation mode, the user initiating the voice request needs to hold the charging case and hand one of the earbuds to the other party, maintaining a certain distance between the left and right earbuds. This operation method helps the system determine the relationship between the left and right earbuds and the charging case, thus laying the foundation for subsequent automatic matching of appropriate language translation packages.

[0063] Therefore, it's understandable that the distance between the target user initiating the voice request and the left and right earbuds usually differs when issuing the command. The system can determine whether the target user is holding the left or right earbud by using acoustic parameters collected during the voice request process. Alternatively, it can determine whether the left and right earbuds, along with the charging case, are both subordinate to the voice requesting user, or vice versa. "Both subordinate to the voice requesting user" means that both the earbuds and the charging case are close to the target user, or that both are controlled and held by the target user.

[0064] For example, if the distance between the left earbud and the charging case and the user making the voice request is less than 1 unit, while the distance between the right earbud and the user making the voice request is greater than 1 unit, then the left earbud and the charging case both conform to the user making the voice request; this is a first conformity relationship. Conversely, if the distance determines that the right earbud and the charging case both conform to the user making the voice request, then it is a second conformity relationship.

[0065] In other words, when the compliance relationship is the first compliance relationship, the left earbud, charging case, and voice request user are linked together; when the compliance relationship is the second compliance relationship, the right earbud, charging case, and voice request user are linked together. Through this identity linking, the charging case can accurately identify the "user-side device group," thereby configuring translation packages suitable for each role for both the left and right earbuds.

[0066] It is understandable that, due to limitations in size, power consumption, cost, and other factors, headphones cannot integrate a distance sensor within the body to determine sequential relationships. Therefore, in one embodiment, acoustic parameters may include a first time difference and a second time difference. The first time difference is the time difference between the arrival of voice request data in the left earbud and the charging case, and the second time difference is the time difference between the arrival of voice request data in the right earbud and the charging case.

[0067] Step S300: Determining the compliance relationship between the left earphone, the right earphone, and the charging case based on the acoustic parameters during the voice request process, including:

[0068] S310. Determine that the first time difference is less than the second time difference, and determine that the compliance relationship is the first compliance relationship;

[0069] S320. Determine that the first time difference is greater than the second time difference, and determine that the compliance relationship is the second compliance relationship.

[0070] Specifically, when a voice request asks a user to wear the left earbud while holding the charging case, and the translator is wearing the right earbud, the time difference between the voice request data reaching the left earbud and the charging case is relatively small, while the time difference between the voice request data reaching the right earbud and the charging case is relatively large. Therefore, when the first time difference is less than the second time difference, the compliance relationship is determined to be a first compliance relationship. Conversely, when the first time difference is greater than the second time difference, the compliance relationship is determined to be a second compliance relationship. When the first time difference equals the second time difference, other acoustic parameters (such as sound intensity, phase difference, etc.) can be introduced to further determine the compliance relationship, which will not be discussed in detail here.

[0071] For example, the left earbud, right earbud, and charging case all have their voice reception functions enabled simultaneously, and each earbud records the precise timestamp (in milliseconds) of receiving a voice request from user A (such as "Please enter Chinese-English translation mode"), denoted as T. 左 T 右 T 仓 For example: T 左 = 1620000001000ms (left ear reception time), T 仓 = 1620000001502ms (charging case receiving time), T 右= 1620000002010ms (Right ear reception time). Time difference calculation: First time difference ΔT1 = |T 左 - T 仓 |=|1620000001000 - 1620000001502|=502ms; Second time difference ΔT2=|T 右 - T 仓 |=|1620000002010 - 1620000001502|=508ms. Since ΔT1 < ΔT2 (e.g., 502ms < 508ms), it is determined to be the first compliance relationship, indicating that the left earphone is closer to the sound source than the right earphone. Therefore, the left earphone is bound to the charging case, and the left ear corresponds to the first target language.

[0072] In other embodiments, when the difference between the first time difference and the second time difference is less than a preset threshold, or when the two are equal, the signal strength parameters of the voice request data at the left earphone, the right earphone, and the charging case can be obtained first; then, the compliant device can be determined according to the signal strength parameters. If the signal strength of the left earphone is greater than that of the right earphone, it is determined to be a first compliant relationship, otherwise it is determined to be a second compliant relationship. The signal strength parameters include the decibel value and signal-to-noise ratio of the voice signal.

[0073] To avoid translation errors caused by the switching of headphones by both parties after the system has established a compliance relationship, and to improve the system's adaptability to dynamic scenarios, in one embodiment, after determining the compliance relationship based on acoustic parameters, the method further includes: real-time monitoring of the switching characteristics of the left and right headphones, wherein the switching characteristics include headphone motion trajectory parameters or real-time distance change parameters between the left and right headphones and the user; when headphone switching characteristics are detected, it is determined as a headphone switching event; in response to the headphone switching event, a compliance relationship reset mechanism is automatically triggered, the acoustic parameters during the voice request process are re-acquired for secondary determination, and the compliance relationship is synchronously updated based on the secondary determination result.

[0074] Specifically, the headphone motion trajectory parameters can include the degree of overlap between the left and right headphone motion trajectories. For example, by using the headphone's built-in six-axis inertial sensors (gyroscope and accelerometer) to collect displacement vector sequences in three-dimensional space in real time, when it is detected that within a preset time window (e.g., 3-5 seconds), the motion trajectories of the left and right headphones form intersecting paths in the three-dimensional coordinate system, and the spatial overlap area of ​​the two trajectories accounts for more than 70% of the total length of their respective trajectories, the trajectory overlap feature is determined to be satisfied. The real-time distance change parameters between the left and right headphones and the user can be determined by the duration of receiving the corresponding voice data. When the system simultaneously detects the overlap feature of the left and right headphone motion trajectories and the distance change feature, it determines that a headphone exchange event has occurred. At this time, the compliance relationship reset mechanism is automatically triggered, and the acoustic parameters during the voice request process are re-collected for secondary judgment (the system issues a voice prompt to re-request the voice), and the compliance relationship is synchronously updated based on the secondary judgment result.

[0075] Step S400: Based on the language text information and the dependency relationship, configure the target translation package for the left earphone and the right earphone respectively for real-time translation.

[0076] In one embodiment, if the dependency relationship obtained in step S300 is a first dependency relationship, a first target translation package and a second target translation package are first obtained from the cloud server and temporarily stored locally. The first target translation package is a translation package that translates the second target language into the first target language, and the second target translation package is a translation package that translates the first target language into the second target language. Then, the first target translation package is configured for the right earphone so that the left earphone outputs the first target language. The second target translation package is configured for the left earphone so that the right earphone outputs the second target language.

[0077] Specifically, if the compliance relationship is the first compliance relationship, it indicates that the voice requesting user has completed the identity binding with the charging case (i.e., the left earbud, the charging case, and the voice requesting user form an identity binding relationship). At this time, the language marked by the voice requesting user in step S100 is the first target language, which means that the input and output languages ​​of the left earbud are both the first target language, while the input and output languages ​​of the right earbud correspond to the second target language.

[0078] Thus, when the compliance relationship is the first compliance relationship, the system will configure the first target translation package (used to convert the second target language into the first target language) for the right earphone, that is, call the locally temporarily stored first target translation package to translate the second target language content received by the right earphone into the first target language, and output it through the left earphone; at the same time, configure the second target translation package (used to convert the first target language into the second target language) for the left earphone, that is, call the locally temporarily stored second target translation package to translate the first target language content received by the left earphone into the second target language.

[0079] To avoid misinterpretation of the translated content due to external noise interference (such as information mixing caused by user B speaking at the same time as user A), this embodiment also sets up a voice filtering mechanism: filtering the first target language voice data collected by the right earphone and filtering the second target language voice data collected by the left earphone, thereby ensuring the accuracy of the translated content.

[0080] In another embodiment, if the relationship is determined to be a second compliance relationship, it indicates that the voice request user has completed identity binding with the charging case (i.e., the right earbud, the charging case, and the voice request user form an identity binding relationship). At this time, the language marked by the voice request user in step S100 is the first target language, which means that the input and output languages ​​of the right earbud are both the first target language, while the input and output languages ​​of the left earbud correspond to the second target language.

[0081] Thus, when the compliance relationship is the second compliance relationship, the system will configure a second target translation packet (used to convert the first target language into the second target language) for the right earphone to translate the first target language content received by the right earphone into the second target language and output it through the left earphone; at the same time, it will configure a first target translation packet (used to convert the second target language into the first target language) for the left earphone to translate the second target language content received by the left earphone into the first target language.

[0082] Similarly, to avoid external noise interference that could lead to mistranslation (e.g., information mixing caused by user B speaking while user A is speaking), this embodiment also includes a voice filtering mechanism: filtering the second target language voice data collected by the right earphone and filtering the first target language voice data collected by the left earphone, thereby ensuring the accuracy of the translation.

[0083] It should be noted that in this application, the charging case establishes a connection with the cloud server via a wireless network, enabling it to retrieve the corresponding target translation package from the cloud server within seconds, thereby completing the localization configuration of the translation package in a very short time. Compared to the translation mode that directly loads general resources from the cloud for each translation, this method significantly reduces translation response latency and effectively improves translation efficiency.

[0084] Additionally, after each translation task is completed (e.g., if no voice data is received within 20 seconds, the translation task is considered complete), the charging case's built-in memory automatically clears the local translation package, reserving storage space for configuring a new translation package for the next charging case cycle. Alternatively, an overwrite mechanism can be used: the next newly acquired translation package will automatically replace the previous one, eliminating the need for manual clearing to update the local translation package configuration. Under the overwrite mechanism, after completing the current translation, when performing the next translation, it can be checked whether the target translation package to be used next time is consistent with the currently configured target translation package. If the target translation package to be used next time is inconsistent with the currently configured target translation package, the currently configured target translation package will be overwritten and deleted; if the target translation package to be used next time is consistent with the currently configured target translation package, the currently configured target translation package will be retained.

[0085] In this way, by detecting the consistency between the next translation package and the current translation package, overwriting and deletion are only performed when there is a discrepancy, avoiding the unconditional deletion and re-downloading operation after each translation. When users repeatedly use the same language combination (such as continuously performing Chinese-English translation), existing translation packages can be reused directly, reducing the unnecessary occupation of limited storage space on the earphones / charging case and reducing resource waste caused by duplicate storage. Furthermore, for repetitive translation scenarios (such as the same user A and user B), since there is no need to download the same translation package again, the time from initiating a translation request to completing the configuration is significantly shortened, making real-time translation start up faster, reducing user waiting delays, and improving the smoothness of interaction. At the same time, when the translation package does not need to be updated, the process of repeatedly obtaining resources from the cloud is avoided, reducing dependence on wireless networks and data traffic consumption, especially in scenarios with unstable network environments or limited traffic, ensuring the stable availability of the translation function.

[0086] In other embodiments, after obtaining the semantic text information representing the user's operation intention in step S100, the method further includes: determining that the operation intention represented by the semantic text information is not a request to enter the translation mode, and determining whether the left earphone and the right earphone are in an interconnected state; if the left earphone and the right earphone are in an interconnected state, sending an interconnection release command to the left earphone and the right earphone to release connection resources and provide the possibility for other devices (such as mobile terminals) to connect to the earphones (e.g., for music playback scenarios).

[0087] Specifically, if the semantic text information represents an operation intent to enter recording mode, the system first determines whether the left and right earbuds are interconnected. If they are interconnected, the charging case immediately sends an interconnection release command to both earbuds, releasing collaborative communication resources. Subsequently, the charging case establishes independent encrypted communication links with each earbud to ensure the independence and stability of data transmission. After the links are established, the system defaults to setting the left earbud as the primary recording device and the right earbud as the secondary recording device, and plays a 0.3-second prompt tone (1200Hz single-frequency signal) through the left earbud to indicate that the initial configuration is complete.

[0088] Furthermore, the system can actively collect user voice commands through the charging case microphone (supports a 3-second command input window). For example, users can select the following recording methods by voice: (1) Main device single recording, only the left earphone microphone is enabled, and the right earphone enters a low-power listening state (microphone is turned off, touch response is retained), which is suitable for close-range one-way recording scenarios; (2) Auxiliary device single recording, the right earphone is switched to the main recording channel (the main device parameter configuration is loaded synchronously), and the left earphone is put into standby mode, which is suitable for users to temporarily switch earphones; (3) Binocular stereo recording, the left and right earphone microphones are activated at the same time, and the charging case synthesizes dual-channel audio through timestamp alignment technology (synchronization error ≤10ms). The left channel retains the original sound pickup characteristics, and the right channel enhances the sense of environmental sound layering, which is suitable for multi-sound source scenarios such as meetings and interviews.

[0089] After the user makes a selection, the system will confirm the selection through the corresponding headphone playback mode (left ear for main device recording, right ear for auxiliary device recording, and simultaneous dual-ear recording for stereo recording). The system supports switching between the two modes during recording, and users can adjust the mode in real time using voice commands such as "switch main record" and "turn on stereo".

[0090] In this way, the adaptability of the recording function to different scenarios is ensured, and the certainty of user operation is improved through standardized interactive feedback, solving the problems of mode switching efficiency and audio continuity when recording with multiple devices.

[0091] To facilitate a further understanding of the technical solution of this application, the steps of the voice interaction method are described below through more specific embodiments:

[0092] S510, User A wears the left earbud and holds the charging case, while User B wears the right earbud. User A initiates a Chinese-English translation request, for example, by saying the voice command "Please enter Chinese-English translation mode" in Chinese;

[0093] After receiving the Chinese voice message "Please enter Chinese-English translation mode", the S520 and charging case determine the operation intention as "enter translation mode" through semantic analysis, and at the same time obtain the language text information of "Chinese and English" through language information extraction.

[0094] S530. Mark the language (Chinese) corresponding to user A (the voice request initiator) as the first target language, and correspondingly mark English as the second target language;

[0095] Based on the time difference analysis of the received voice request data, the S540, left earphone, right earphone and charging case determine that the left earphone and charging case are closer to user A, and then establish the identity binding relationship between the left earphone, charging case and the user making the voice request, forming the first compliance relationship;

[0096] After confirming that the left earbud has been successfully linked to the user making the voice request and that the input / output language of the left earbud is Chinese, the S550 and charging case automatically associate the input / output language of the right earbud with English. At this point, the charging case retrieves two types of translation packages from the cloud server: a first target translation package for converting English to Chinese and a second target translation package for converting Chinese to English; and configures the second target translation package locally for the left earbud and the first target translation package locally for the right earbud.

[0097] S560. When user A inputs Chinese voice data through the left earbud, since the left earbud is configured with a second target translation package (Chinese to English), the charging case receives the voice data, calls the local second target translation package to translate it into English, and then plays it through the right earbud, so user B can obtain the English translation content through the right earbud; when user B inputs English voice data through the right earbud, since the right earbud is configured with a first target translation package (English to Chinese), the charging case receives the data, calls the local first target translation package to translate it into Chinese, and then plays it through the left earbud, so user A can obtain the Chinese translation content.

[0098] S570. After the translation task is completed, the charging case will automatically clear the translation package stored in the local configuration to reserve storage space for receiving new corresponding translation packages for the next translation.

[0099] Based on this, the technical solution of this application can clarify the compliance relationship between the dual earphones, the charging case and the user making the voice request based on the actual scenario of the user initiating a voice request (by quantifying the scenario characteristics through acoustic parameters). Thus, according to this compliance relationship, the corresponding localized translation packages are configured in advance for the left and right earphones respectively. Compared with the translation mode of uniformly loading general resources, this greatly reduces the translation response latency and improves the translation efficiency.

[0100] like Figure 4 As shown, Figure 4The diagram below shows the hardware structure of a smart headset in some embodiments of this application. The smart headset provided in the embodiments of this application also includes a memory 1000 and a processor 2000. The memory 1000 is used to store computer-readable instructions, and the processor 2000 is used to call the computer-readable instructions to execute the voice interaction method as described above.

[0101] The processor 2000 provides computing and control capabilities to control the smart earphones to perform corresponding tasks, such as controlling the smart earphones to perform the voice interaction method in any of the above method embodiments. The method includes: the charging case collecting voice request data initiated by the user, and performing semantic parsing and language information extraction on the voice request data to obtain semantic text information representing the user's operation intention and language text information representing the target language requirement; determining that the operation intention represented by the semantic text information is a user request to enter translation mode, and sending an interconnection command to the left and right earphones; determining the compliance relationship between the left and right earphones and the charging case based on the acoustic parameters during the voice request process, wherein the compliance relationship includes a first compliance relationship and a second sequential relationship, the first compliance relationship representing that the left earphone and the charging case jointly comply with the voice request user, and the second compliance relationship representing that the right earphone and the charging case jointly comply with the voice request user; and configuring target translation packages for real-time translation for the left and right earphones respectively based on the language text information and the compliance relationship.

[0102] The processor 2000 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0103] The memory 1000, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the voice interaction method in the embodiments of this application. The processor 2000 can implement the voice interaction method in any of the above method embodiments by running the non-transitory software programs, instructions, and modules stored in the memory 1000.

[0104] Specifically, memory 1000 may include volatile memory (VM), such as random access memory (RAM); memory 1000 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or other non-transitory solid-state storage devices; memory 1000 may also include combinations of the above types of memory.

[0105] In summary, the smart earphone of this application adopts the technical solution of any of the above-described voice interaction method embodiments, and therefore has at least the beneficial effects brought about by the technical solutions of the above embodiments, which will not be elaborated further here.

[0106] This application also provides a computer-readable storage medium, such as a memory including program code, which can be executed by a processor to complete the voice interaction method described in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0107] This application also provides a computer program product comprising one or more lines of program code stored in a computer-readable storage medium. The processor of the early warning system reads the program code from the computer-readable storage medium and executes the program code to complete the voice interaction method steps provided in the above embodiments.

[0108] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program or program code related to hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0109] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software and a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0111] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A voice interaction method, characterized in that, Applied to smart earphones, the smart earphones including a charging case, a left earphone, and a right earphone, the method includes: The charging case collects voice request data initiated by the user, and performs semantic parsing and language information extraction on the voice request data to obtain semantic text information representing the user's operation intention and language text information representing the target language requirement. If the operational intent represented by the semantic text information is determined to be a user request to enter translation mode, an interconnection command is sent to the left and right earpieces. The compliance relationship between the left earbud, the right earbud, and the charging case is determined based on the acoustic parameters during the voice request process. The compliance relationship includes a first compliance relationship and a second sequential relationship. The first compliance relationship indicates that the left earbud and the charging case jointly comply with the voice request user, and the second compliance relationship indicates that the right earbud and the charging case jointly comply with the voice request user. Based on the language text information and the compliance relationship, target translation packages are configured for the left and right earphones respectively for real-time translation, and the target translation packages are deleted or overwritten after the translation is completed. After obtaining the semantic text information representing the user's operational intent, the following is also included: Determine the operational intent represented by the semantic text information to enter translation mode without requesting it, and determine whether the left and right earphones are in an interconnected state. If the left and right earbuds are in an interconnected state, send an interconnection release command to the left and right earbuds; and / or, After determining the compliance relationship between the left earbud, the right earbud, and the charging case based on acoustic parameters during the voice request process, the method further includes: Real-time monitoring of the exchange characteristics of the left and right earphones, the exchange characteristics including earphone motion trajectory parameters or real-time distance change parameters between the left and right earphones and the user, wherein the earphone motion trajectory parameters include the degree of overlap between the motion trajectories of the left and right earphones; When headphone switching characteristics are detected, it is determined to be a headphone switching event; In response to the headphone exchange event, the compliance relationship reset mechanism is automatically triggered, the acoustic parameters during the voice request process are re-acquired for secondary judgment, and the compliance relationship is synchronously updated based on the secondary judgment result.

2. The voice interaction method as described in claim 1, characterized in that, The acoustic parameters include a first time difference and a second time difference. The first time difference is the time difference between the arrival of voice request data in the left earbud and the charging case, and the second time difference is the time difference between the arrival of voice request data in the right earbud and the charging case. The determination of the compliance relationship between the left earbud, the right earbud, and the charging case based on acoustic parameters during the voice request process includes: If the first time difference is determined to be less than the second time difference, the compliance relationship is determined to be a first compliance relationship. If the first time difference is determined to be greater than the second time difference, the compliance relationship is determined to be the second compliance relationship.

3. The voice interaction method as described in claim 1, characterized in that, After obtaining the language text information representing the target language requirements, the following is also included: The language in the language text information that matches the voice request is marked as the first target language, and other languages ​​in the language text information that are not the first target language are marked as the second target language.

4. The voice interaction method as described in claim 3, characterized in that, The step of configuring target translation packages for real-time translation for the left and right earbuds based on the language text information and dependency relationships includes: The dependency relationship is determined to be a first dependency relationship. A first target translation package and a second target translation package are obtained from the cloud server. The first target translation package is a translation package that translates the second target language into the first target language, and the second target translation package is a translation package that translates the first target language into the second target language. Configure the right earbud with the first target translation packet so that the left earbud outputs the first target language; and, Configure the second target translation packet for the left earphone so that the right earphone outputs the second target language.

5. The voice interaction method as described in claim 4, characterized in that, The method further includes: During real-time translation, the speech data corresponding to the first target language collected by the right earphone is filtered, and the speech data corresponding to the second target language collected by the left earphone is also filtered.

6. The voice interaction method as described in claim 3, characterized in that, The step of configuring target language translation packages for the left and right earbuds respectively for real-time translation based on the language text information and the dependency relationship includes: The dependency relationship is determined to be a second dependency relationship. A first target translation package and a second target translation package are obtained from the cloud server. The first target translation package is a translation package that translates the second target language into the first target language, and the second target translation package is a translation package that translates the first target language into the second target language. Configure the right earphone with the second target translation packet so that the left earphone outputs the second target language; and, Configure the first target translation packet for the left earphone so that the right earphone outputs the first target language.

7. The voice interaction method as described in claim 6, characterized in that, The method further includes: During real-time translation, the speech data corresponding to the second target language collected by the right earphone is filtered, and the speech data corresponding to the first target language collected by the left earphone is also filtered.

8. A smart earphone, characterized in that, include: Memory and processor, wherein the memory is used to store program code; The processor is used to call the program code to perform the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.