Charging box and audio playing method and system
By incorporating a microphone and audio recognition module within the charging case, the device achieves automatic recognition and synchronized playback of external audio, resolving issues of complex operation and poor synchronization, and improving user experience and playback accuracy.
Patent Information
- Application Number
- CN202511126248.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, users need to manually search or identify software to make the head-mounted playback device play the same audio as the outside world, which is complicated and not always accurate, affecting the user experience.
The charging case is equipped with a microphone and an audio recognition module to collect ambient audio signals and identify target audio files. The audio is then transmitted wirelessly to the head-mounted playback device for synchronized playback. Playback synchronization is achieved by using a buffer and clock adjustment.
It simplifies user operation, improves audio synchronization and playback accuracy, enhances the user listening experience, and reduces power consumption and latency.
Smart Images

Figure CN120916093A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio equipment, in particular, a charging box, an audio playing method and system are provided. BACKGROUND
[0002] Audio playing is performed in various scenarios, for example, music is played in restaurants, cafes and the like. After hearing the audio played by the outside world, the user may wish to play it on a head-mounted playing device such as an earphone due to various reasons, such as liking the music played, wanting to hear more clearly due to the outside world being noisy, and the like.
[0003] At present, if the head-mounted playing device is required to play the same audio as in the scenario, the user needs to determine the name of the audio played by searching for software or identifying software, and then searches and plays through the corresponding playing software. This method is relatively complex to operate, and the user may not be able to obtain the accurate audio and play it. SUMMARY
[0004] Therefore, the present application aims to provide a charging box, an audio playing method and system to simplify the operation complexity of the user listening to the same audio as the outside world using a head-mounted playing device and improve the user experience.
[0005] In a first aspect, a charging box is provided, which is used to charge a head-mounted playing device; the charging box comprises a processor, an audio acquisition module, a wireless module and an audio recognition module; the processor is connected with the audio acquisition module, the wireless module and the audio recognition module respectively; the audio recognition module is connected with the audio acquisition module and the wireless module respectively; the wireless module is used to be communicatively connected with the head-mounted playing device and a sound source device; the audio acquisition module is used to acquire an environmental audio signal; the environmental audio signal comprises an audio signal of audio played by an outside device; the audio recognition module is used to obtain an identification result of the environmental audio signal; the identification result is used to represent a target audio file comprising the audio played by the outside device; the processor is further used to obtain the target audio file from the sound source device through the identification result and the wireless module; and the processor is further used to send the target audio file to the head-mounted playing device for playing through the wireless module.
[0006] In the present application, a microphone and an audio recognition module are configured in the charging box, the microphone can acquire an environmental audio signal comprising audio played by an outside device, and the corresponding audio is identified through the audio recognition module. Thus, the user can determine the audio played by the outside device without manually searching or using an identification software to identify and play the audio on a head-mounted playing device, which reduces the operation complexity of the user and improves the listening experience of the user.
[0007] In an embodiment, the audio recognition module is further configured to determine target audio file information and a target position based on the environmental audio signal, wherein the target position represents a position of the environmental audio signal in the target audio file; and the processor is configured to continuously acquire an audio data stream of the target audio file after the target position in the target audio file based on the target audio file information and the target position, so that the head-mounted playback device starts playing the target audio file from the target position.
[0008] In the embodiments of the present application, the target audio file information and the target position corresponding to the environmental audio signal are identified, and the head-mounted playback device acquires an audio data stream starting from the target position, so that the head-mounted playback device receives and plays the audio data stream starting from the target position, thereby making the audio heard by the user consistent with the sound currently played by the external device, synchronizing the sound heard by the user, and improving the user experience.
[0009] In an embodiment, the processor is further configured to determine a target time difference between the acquisition of the environmental audio signal by the audio acquisition module and the acquisition of the target audio file from the sound source device, and control the playback of the target audio file by the head-mounted playback device to be synchronized with the external device based on the target time difference.
[0010] The head-mounted playback device and the external device are different devices, and may have different playback rates and playback times due to different reasons, thereby causing the feeling of asynchronous audio playback and affecting the listening experience of the user. In many cases, the user cannot directly control the external device. Therefore, in the embodiments of the present application, the charging box can determine a target time difference between the acquisition of the environmental audio signal by the audio acquisition module and the acquisition of the target audio file from the sound source device, and control the playback of the head-mounted playback device to be synchronized with the external device based on the target time difference, thereby improving the listening experience of the user. In addition, the sound played by the head-mounted playback device may be leaked. The audio acquisition module is arranged in the charging box instead of using the microphone of the head-mounted playback device to acquire the environmental audio signal, which can effectively avoid the case that the audio played by the head-mounted playback device is acquired in the environmental audio signal, thereby improving the accuracy of the synchronous playback control.
[0011] In an embodiment, the charging box comprises a first buffer and a second buffer; the first buffer is connected with the second buffer and the wireless module; the second buffer is equal to or greater than the audio buffer size of the head-mounted playback device, so that the data buffered by the second buffer is more than or equal to the data buffered by the head-mounted playback device; the first buffer is used to temporarily store the target audio file before the target audio file is sent to the head-mounted playback device; the first buffer is also used to output the target audio file to the second buffer when the target audio file is output to the head-mounted playback device; the processor is used to perform audio correlation processing on the environmental audio signal and the target audio file in the second buffer to determine the target time difference; the audio correlation processing is used to determine the time difference between the same features of the input data.
[0012] The head-mounted playback device buffers the audio to be played next through the audio buffer, so that when the data of the audio file cannot be stably obtained, the data in the audio buffer can be output to stably play, thereby reducing the playback lag, abnormality and the like. In the embodiment of the application, the first buffer can be arranged in the charging box to pre-buffer the data to be sent to the head-mounted playback device for playing, so as to reduce the audio playback lag, abnormality and the like caused by poor communication between the charging box and the audio source device. Even in the case of poor communication, part of the data can be transmitted to the head-mounted playback device for playing. At the same time, the second buffer is also arranged to simulate the audio buffer of the head-mounted playback device. Thus, the charging box itself can also determine the playing time of the target audio file information without the feedback of the head-mounted playback device to the charging box. Further, the charging box can quickly determine the target time difference to control the head-mounted playback device and the external device to play synchronously, reduce the power consumption of the head-mounted playback device, reduce the delay caused by transmission, and improve the timeliness and reliability of the synchronous playback control.
[0013] In an embodiment, the size of the second buffer is equal to the sum of the first buffer size and the second buffer size, the first buffer size is the audio buffer size of the head-mounted playback device, and the second buffer size is the buffer size corresponding to the preset maximum time delay of the data sent by the charging box to the head-mounted playback device.
[0014] The data transmission between the head-mounted playback device and the charging box also needs a certain time, if the second buffer and the head-mounted playback device audio buffer have the same cache, the data cached in the second buffer and the data cached in the head-mounted playback device can be different, after the data in the second buffer is replaced by new data, the data cached in the head-mounted playback device can still cache the data in the second buffer which has been replaced, thereby it can cause that the target time difference cannot be calculated. Therefore, in the embodiment of the present application, the size of the second buffer can be equal to the sum of the first cache size and the second cache size, so that the second buffer caches more data, thereby the same data as the head-mounted playback device can be cached.
[0015] In an embodiment, the processor is configured to adjust an acquisition rate of the wireless module for acquiring an audio data stream of the target audio file from the sound source device based on the target time difference and a preset nominal frequency, so that the acquisition rate of the wireless module for acquiring the target audio file matches a playing rate of the target audio file by the external device.
[0016] The playback synchronization includes consistent playing speed. Due to the hardware performance, environment and other reasons, the clock frequency of different playback devices and the nominal frequency exist differences, and then the playing speed of different playback devices for the same audio exists differences. Therefore, in the embodiment of the present application, the acquisition rate of the target audio file audio data stream can be adjusted based on the target time difference and the nominal frequency, so as to control the playing rate of the target audio file audio data stream output to the head-mounted playback device for playing, thereby making the playing rate of the head-mounted playback device and the external device consistent, so as to realize the playback synchronization.
[0017] In an embodiment, a clock adjustment instruction is generated based on the target time difference, wherein the clock adjustment instruction is used to instruct the head-mounted playback device to adjust a playing clock, so that a second playing time of the playing clock matches a sum of a first acquisition time and the target time difference, the first acquisition time is an acquisition time of the audio acquisition module for the environmental audio signal, and the second playing time is a playing time of the head-mounted playback device for the target audio file.
[0018] The playback synchronization also includes playing the same data at the same time. However, due to various reasons, such as identification of the collected environmental audio signal, obtaining the target audio file, audio data stream transmission, and the like, a certain delay is required, so that the head-mounted playback device and the external device do not play the audio at the same time. Therefore, in the embodiment of the present application, a target time difference generation clock adjustment instruction can be used to adjust the playback clock of the head-mounted playback device, so that the second playback time of the target audio file by the head-mounted playback device matches the first collection time of the environmental audio signal by the audio collection module, thereby synchronizing the audio playback of the head-mounted playback device and the external device.
[0019] In an embodiment, the processor is further configured to determine that the reliability of the environmental audio signal is greater than a preset threshold before determining the target time difference between the collection of the environmental audio signal by the audio collection module and the acquisition of the target audio file from the sound source device.
[0020] In the embodiment of the present application, when the reliability of the environmental audio signal is low, the sound of the environmental audio signal is also relatively small. After the user wears the head-mounted playback device, the physical structure of the head-mounted playback device blocks or even prevents the user from hearing the audio played by the external environment. Therefore, when the reliability of the environmental audio signal is low, synchronization is not required. Thus, the processor needs to perform less work, thereby reducing resource occupation and power consumption and improving the endurance.
[0021] In an embodiment, the audio identification module is configured with an audio identification model, and the audio identification module is configured to identify the environmental audio signal by using the audio identification model to obtain the identification result.
[0022] In the embodiment of the present application, the audio identification model is locally configured in the charging box to perform identification locally, which helps to reduce the delay caused by data transmission in the identification process and improve the identification efficiency.
[0023] In an embodiment, the sound source device includes an audio identification model; and the audio identification module is configured to send the environmental audio signal to the sound source device through the wireless module and obtain the identification result of the environmental audio signal by the audio identification model from the sound source device.
[0024] In the embodiment of the present application, the audio identification model is deployed in the sound source device, and the audio identification module is configured to control the sending of the environmental audio signal and the acquisition of the identification result. On the one hand, this can reduce the storage space requirement of the model stored locally in the charging box and reduce the power consumption of the charging box locally. On the other hand, the sound source device can deploy a larger and more accurate identification model to obtain more accurate identification results and improve the identification accuracy.
[0025] In an embodiment, the processor is further configured to receive an identification instruction, and in response to the identification instruction, control the audio acquisition module to acquire the ambient audio signal to enable identification of the ambient audio signal.
[0026] In an embodiment of the present application, the user can issue an instruction when he wants to listen to the same audio outside, and the charging box acquires and identifies the ambient audio signal only after receiving the identification instruction, thereby meeting the user's needs while avoiding power consumption caused by long-time opening of the acquisition and identification function.
[0027] In a second aspect, an embodiment of the present application provides an audio playback system, which includes the charging box and the head-mounted playback device provided by any of the embodiments of the first aspect, and wherein the charging box is configured to charge the head-mounted playback device.
[0028] In a third aspect, an embodiment of the present application provides an audio playback method applied to a processor of a charging box, the charging box being configured to charge a head-mounted playback device, and the charging box including a processor, an audio acquisition module, a wireless module, and an audio identification module. The audio playback method includes: acquiring an ambient audio signal by the audio acquisition module; the ambient audio signal including an audio signal of audio played by an external device; obtaining an identification result of the ambient audio signal by the audio identification module; the identification result being configured to represent a target audio file including the audio signal played by the external device; obtaining the target audio file from the audio source device based on the identification result and the wireless module; and sending the target audio file to the head-mounted playback device for playback by the wireless module. BRIEF DESCRIPTION OF DRAWINGS
[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0030] Figure 1 Structure diagram of the charging box provided by an embodiment of the present application; Figure 2 Relationship diagram of the first buffer and the second buffer provided by an embodiment of the present application; Figure 3 Flowchart of an audio playback method provided by an embodiment of the present application.
[0031] Figure legend: charging box 100; processor 110; audio acquisition module 120; wireless module 130; audio identification module 140. DETAILED DESCRIPTION
[0032] Firstly, the embodiment of the present application provides a charging box 100 for charging a head-mounted playback device. Wherein, the head-mounted playback device can be a truly wireless earphone, smart glasses, etc., accordingly, the charging box 100 can be a charging box of the truly wireless earphone, used for storing and charging the truly wireless earphone. Similarly, the charging box 100 can be a storage box of the smart glasses, and the storage box is built-in with a charging module to charge the smart glasses.
[0033] In the embodiment of the present application, the head-mounted playback device can include a left component and a right component. For example, if the head-mounted playback device is a truly wireless earphone, the left component can be a left earphone and the right component can be a right earphone. Similarly, the smart glasses can also include a left component and a right component.
[0034] Please refer to Figure 1 , Figure 1 A structural schematic diagram of a charging box 100 provided by an embodiment of the present application. In the embodiment of the present application, the charging box 100 includes a processor 110, a memory, an audio acquisition module 120, a wireless module 130 and an audio recognition module 140. Wherein, the processor 110 is connected with the memory, the audio acquisition module 120, the wireless module 130 and the audio recognition module 140 respectively, and the audio recognition module 140 is connected with the audio acquisition module 120 and the wireless module 130 respectively.
[0035] The wireless module 130 is used for communication connection with other devices. In the embodiment of the present application, the wireless module 130 can include a Bluetooth module, a WiFi module, a cellular module, etc., which is not limited here.
[0036] In the embodiment of the present application, the charging box 100 can be in communication connection with a sound source device and a head-mounted playback device through the wireless module 130. Wherein, the sound source device is a device that can provide an audio file to the charging box 100 and the head-mounted playback device. For example, the sound source device can be a mobile phone, a tablet computer, a computer, etc., and the sound source device can also be a cloud server, which is not limited here.
[0037] In the embodiment of the present application, the wireless module 130 can support WiFi direct, wireless access point, DLNA (Digital Living Network Alliance), etc. In some embodiments of the present application, the wireless module 130 can also support low-power Bluetooth, HDT (Higher Data Throughput) Bluetooth technology, etc., which is not limited here.
[0038] In the embodiments of the present application, the audio acquisition module 120 can include a microphone and a corresponding supporting circuit to acquire the sound of the environment in which the charging box 100 is located, that is, in the embodiments of the present application, the audio acquisition module 120 can be used to acquire the environmental audio signal.
[0039] The environmental audio signal is the sound in the environment in which the charging box 100 is located, including but not limited to human voice and sound played by other devices. That is, the environmental audio signal can include the audio signal of the audio played by the external device.
[0040] For example, when the charging box 100 is placed in a restaurant, there can be voices of people talking, sounds of various objects colliding, and music played by the restaurant sound box in the restaurant. The restaurant sound box is an external device, and the music sound played by the sound box included in the environmental audio signal is the audio signal of the audio played by the external device.
[0041] In the embodiments of the present application, the audio recognition module 140 can be used to obtain the recognition result of the environmental audio signal.
[0042] When the user hears some interesting audio, he or she may wish to play it on his or her own head-mounted playback device. For example, he or she may hear good music, interesting audio novels, storytelling, stand-up comedy, etc. due to too much noise in the external environment, too small sound of the audio played by the external device, or wish to listen to the audio more immersed, etc. Therefore, it is necessary to play it on his or her own head-mounted playback device such as earphones. At this time, the audio acquisition module 120 can acquire the audio played by the external device, identify it through the charging box 100, and then obtain the audio file of the audio from the audio source device and play it through the head-mounted playback device. The recognition result is used to represent the target audio file including the audio signal played by the external device. For example, the recognition result can be the name, code, version number, keyword, signal feature, etc. of the audio played by the external device. Through the recognition result, the audio source device can find the corresponding audio to obtain the target audio file.
[0043] In the embodiments of the present application, an audio recognition model can be configured in the audio recognition module 140. The audio recognition module 140 is used to identify the environmental audio signal through the audio recognition model to obtain the recognition result. Configuring the audio recognition model locally in the audio recognition module 140 can reduce the time caused by data transmission and improve the recognition efficiency.
[0044] In some other embodiments of the present application, the audio recognition model is configured on a recognition device such as a mobile phone, a tablet computer, a cloud, etc. The audio recognition module 140 is used to send the environmental audio signal to the recognition device through the wireless module 130, and obtain the recognition result of the environmental audio signal by the audio recognition model from the recognition device.
[0045] The recognition device can be a mobile phone, a cloud, etc., which has lower requirements for power consumption and size, and can be configured with an audio recognition model with better performance. Thus, sending the environmental sound signal to the recognition device for recognition can improve the accuracy of recognition.
[0046] In some embodiments of the present application, the audio recognition module 140 and the recognition device can both be configured with an audio recognition model, where the recognition device is configured with a teacher model of the audio recognition model, and the audio recognition module 140 is configured with a student model of the audio recognition model. When the recognition result confidence of the audio recognition module 140 is lower than a preset value, the environmental audio signal can be sent to the recognition device to obtain the recognition result of the recognition device. Thus, the recognition efficiency and the recognition accuracy can be considered at the same time.
[0047] In an embodiment of the present application, the processor 110 can be configured to obtain the target audio file from the sound source device through the recognition result and the wireless module 130.
[0048] In some embodiments, the recognition device and the sound source device can be the same device or the same server.
[0049] In this embodiment, the processor 110 can send the recognition result to the sound source device through the wireless module 130, so that the sound source device can find the corresponding target audio file through the recognition result. For example, the sound source device can find the corresponding audio file in the audio database through the audio name, the target audio file name, the version number, the audio keyword, etc.
[0050] Then, the processor 110 can also send the target audio file to the head-mounted playback device for playback through the wireless module 130 after obtaining the target file.
[0051] Thus, the user can play the same audio as the external device through the head-mounted playback device, meet the user's listening demand for audio, simplify the user's operation, and improve the user's experience.
[0052] Wherein, the audio file found due to various reasons can not be completely the same as the audio file played by the external playback device, for example, the audio file can be damaged due to various reasons, there is partial distortion in the process of playing by the external device, the audio played by the external device is not collected, etc. Therefore, in an embodiment of the present application, playing the same audio can refer to playing an audio file with a similarity greater than a preset value. For example, for the same song of the same singer, the external device may play the earlier version, and the sound source device can select the later version of the audio file for playing when the earlier version of the audio file is not searched.
[0053] In some embodiments of the present application, the processor 110 is further configured to receive an identification instruction, and in response to the identification instruction, control the audio acquisition module 120 to acquire the ambient audio signal to enable identification of the ambient audio signal.
[0054] The type of identification instruction can be various, including but not limited to a voice instruction, an instruction from a button on the charging box 100, or an instruction from a smart device such as a mobile phone. For example, the charging box 100 can be configured with a voice recognition function, and the user can issue a voice instruction to control the charging box 100 to acquire the ambient audio signal for identification. For another example, the charging box 100 is configured with a button, and after the user clicks the button, the charging box 100 acquires the ambient audio signal for identification. For another example, the user can issue an instruction to the charging box 100 through the mobile phone to acquire the ambient audio signal for identification by the charging box 100. The above are only examples, and the implementation manner can be various, which is not limited herein.
[0055] By issuing an instruction to control the acquisition of the ambient audio signal for identification, the identification is only performed when the user has a demand, which can avoid the power consumption caused by the identification function of the charging box 100 being always on, and avoid the interference to the audio currently played by the head-mounted playback device. Generally, the user will only search when an interesting audio segment is heard, and if the head-mounted playback device starts playing from the beginning of the target audio file, the user can not be interested and needs to manually adjust the progress to the interesting position. In addition, if the external device continuously plays audio, the content played by the user's head-mounted playback device and the external device is inconsistent and out of sync, which can also affect the user's listening experience.
[0056] In an embodiment of the present application, the audio identification module 140 can be further configured to determine the target audio file information and the target position based on the ambient audio signal. That is, in this embodiment of the present application, the identification result can further include the target audio file information and the target position.
[0057] In addition, the processor 110 can be further configured to continuously acquire the audio data stream after the target position in the target audio file based on the target audio file information and the target position, so that the head-mounted playback device starts playing the target audio file from the target position.
[0058] In an embodiment of the present application, the target audio file information is information used to represent the target audio file, for example, the target audio information can include the file name, the version number, the serial number, the number, the song name, the keyword, the author, etc., and for another example, the target audio file information can also be a sentence of lyrics, a section of melody, etc. The target audio file information can be various, which is not limited herein.
[0059] Then, the target position represents a position of the environmental audio signal in the target audio file. The target audio file can be determined by the target audio file information. Then, it is determined that the audio played by the external device in the environmental audio signal corresponds to a position in the target audio file, so as to determine the target position.
[0060] In the embodiments of the present application, the target position can be a file offset or a byte offset. For example, the byte offset can be a byte position in the audio file played by the external device or the target audio file, which corresponds to the audio signal played by the external device in the environmental audio signal collected by the microphone. For example, the audio signal played by the external device in the environmental audio signal corresponds to the Xth byte of the target audio file, and X is a positive integer. The file offset is a time position or a playing progress in the target audio file, which corresponds to the audio signal played by the external device in the environmental audio signal collected by the microphone. For example, the target audio file information is at the Yth moment, and Y is a positive integer.
[0061] In the embodiments of the present application, the recognition result can include one or more target audio file information, which is not limited herein.
[0062] In the embodiments of the present application, after the target audio file is obtained, the processor 110 can determine the position of the environmental audio signal in the target audio file by the environmental audio signal. In other embodiments, the processor 110 can send the environmental audio signal to the recognition device or the sound source device through the wireless module 130, so as to obtain the target position fed back by the recognition device or the sound source device. The recognition device and the sound source device can be the same device, for example, a mobile phone, a tablet computer, a cloud server, etc.
[0063] When the sound source device provides the audio file to various playing devices, the audio file is usually not stored locally in the playing device, but the audio data stream of the audio file is continuously transmitted to the playing device through communication. Therefore, in the embodiments of the present application, when the charging box 100 obtains the target audio file, the charging box 100 also continuously obtains the audio data stream of the target audio file. Similarly, the charging box 100 also continuously transmits the obtained audio data stream to the head-mounted playing device.
[0064] In the embodiments of the present application, the charging box 100 obtains the audio data stream of the target audio file after the target position. Correspondingly, the playing of the head-mounted playing device also starts from the target position. Therefore, the user can start listening from the same position of the audio currently played by the external device, so as to avoid listening to the content that is not interested in, and at the same time, the user can be consistent with the content played by the external device to a certain extent, so as to achieve a certain synchronization effect, and reduce the bad listening experience brought to the user when the sound played by the external device and the sound played by the head-mounted playing device are inconsistent.
[0065] In fact, the head-mounted playback device and the external device can not be able to achieve playback synchronization due to various reasons, for example, communication delay caused by data transmission, different types of devices, different environments, and thus different clock and playback frequencies. Therefore, in the embodiments of the present application, the processor 110 can also be configured to determine a target time difference between the environmental audio signal collected by the audio collection module 120 and the audio data stream corresponding to the target audio file obtained from the audio source device, and control the playback of the target audio file by the head-mounted playback device to be synchronized with the external device based on the target time difference.
[0066] In the embodiments of the present application, the target time difference is the time difference between the environmental audio signal collected by the audio collection module 120 and the audio data stream corresponding to the target audio file obtained from the audio source device. The audio data stream corresponding to the target audio file refers to the audio data stream corresponding to the audio signal played by the external device in the target audio file in the environmental audio signal.
[0067] In some embodiments, the target time difference can be determined by audio correlation processing between the environmental audio signal collected by the audio collection module 120 and the audio data stream corresponding to the target audio file obtained from the audio source device. The audio correlation processing can be time domain correlation processing or frequency domain correlation processing. There are various ways to determine the target time difference, which will not be described here.
[0068] In the embodiments of the present application, the user is usually close to the charging box 100, for example, carried by the user, and thus the time at which the audio collection module 120 collects the environmental audio signal can be used to represent the time at which the user hears the sound of the audio played by the external device.
[0069] Generally, compared with the position change between the external device and the user, the position relationship between the user and the charging box 100 does not change much, for example, the external device is a sound box, which can be placed tens of meters away, and the position change between the user and the sound box can be several meters to tens of meters, while the user usually carries the charging box 100, and the position change is usually within one meter. In contrast, the position relationship between the user and the charging box 100 does not change much and can be considered relatively fixed.
[0070] In the case that the position relationship between the user and the charging box 100 is relatively fixed, the time delay of the charging box 100 sending the target audio file to the head-mounted playback device is relatively fixed during the user wearing the head-mounted playback device. Therefore, the time at which the charging box 100 obtains the target audio file from the audio source device can be used to represent the time at which the head-mounted playback device plays the audio data stream in the target audio file.
[0071] In some embodiments, the time difference can also be referred to as a time delay. In the calculation of the target time difference, the target audio file information can have been played, but does not affect the calculation of the target time difference.
[0072] When playing, the audio playing device uses the buffer to cache part of the audio data stream, and plays the cached audio data stream according to the control of the clock signal. Therefore, during communication, even if the data transmission is suspended due to communication anomalies or other reasons, the audio playing device can still play normally through the cached data.
[0073] Correspondingly, in the embodiments of the present application, a first buffer can also be arranged in the charging box 100 to temporarily store the audio data stream of the target audio file obtained from the sound source device, and the data cached by the first buffer is output to the audio playing device through the wireless module 130.
[0074] In the audio playing device such as a head-mounted playing device, through the cached data, the playing time of the cached data can be determined, for example, a certain frame of data will be played after how many seconds. Therefore, in the embodiments of the present application, a buffer equivalent to the audio buffer of the head-mounted playing device can also be arranged in the charging box 100, so that the charging box 100 can cache the same size of data in the head-mounted playing device to predict the playing time of the target audio file information by the head-mounted playing device.
[0075] Therefore, in some embodiments of the present application, the charging box 100 can include a first buffer and a second buffer. The first buffer is connected with the second buffer and the wireless module 130. The second buffer is greater than or equal to the size of the audio buffer of the head-mounted playing device, so that the data cached by the second buffer is more than or equal to the data cached by the head-mounted playing device.
[0076] In other embodiments, the size of the second buffer can also be equal to the sum of the first buffer size and the second buffer size, the first buffer size is the size of the audio buffer of the head-mounted playing device, and the second buffer size is the cache size corresponding to the preset maximum time delay of the data sent by the charging box 100 to the head-mounted playing device.
[0077] Please refer to Figure 2 , Figure 2 The relationship diagram of the first buffer and the second buffer provided by an embodiment of the present application.
[0078] In this embodiment, the first buffer is used to temporarily store the target audio file before sending the target audio file to the head-mounted playing device. At the same time, the first buffer is also used to output the target audio file to the second buffer when the target audio file is output to the head-mounted playing device.
[0079] Thus, the charging box 100 can determine the playing time of the audio data stream through the audio data stream in the second buffer, and further can determine the playing time of the audio played by the external device in the ambient audio signal on the head-mounted playback device. In the determination of the playing time, there is enough data in the first buffer and the second buffer to determine the playing time.
[0080] The charging box 100 needs to forward the audio data stream of the target audio file to the head-mounted playback device. If the second buffer outputs data according to the output speed of the audio buffer in the head-mounted playback device, and the size of the second buffer is the same as that of the audio buffer of the head-mounted playback device, the data in the second buffer may be output while the head-mounted playback device has not been output. In this case, the data corresponding to the ambient audio signal is output, and the audio played by the external device is not stored in the second buffer, so that the same data cannot be determined, and the target time difference cannot be determined.
[0081] Therefore, in the embodiments of the present application, the size of the second buffer can be greater than the size of the audio buffer of the head-mounted playback device.
[0082] In some embodiments, the size of the second buffer can also be equal to the sum of the first cache size and the second cache size, the first cache size is the size of the audio buffer of the head-mounted playback device, and the second cache size is the cache size corresponding to the preset maximum delay of the charging box 100 sending data to the head-mounted playback device. The preset maximum delay includes the time delay of the charging box wireless module sending the audio data stream and the time delay of the head-mounted playback device receiving the audio data stream.
[0083] In some embodiments of the present application, the processor 110 is configured to perform audio correlation processing on the ambient audio signal and the target audio file in the second buffer to determine the target time difference. The audio correlation processing can be frequency domain correlation processing or time domain correlation processing, which is not limited herein.
[0084] The correlation processing of the signal is a technology of extracting information or synchronizing signals by analyzing the similarity between signals, mainly including autocorrelation and cross-correlation. For details, please refer to the prior art, which will not be expanded here.
[0085] In the embodiments of the present application, the audio correlation processing can be used to determine the time difference between the same features in the input data. That is, the input data are respectively the audio data stream in the target audio file obtained from the sound source device and the environmental sound signal collected by the audio collection module 120. By performing the audio correlation processing on the audio data stream in the target audio file and the environmental sound signal, the same features of the two can be determined. The same features are the features corresponding to the audio played by the external device in the environmental audio signal. Then, the target time difference is determined by the same features, for example, the positions of the features in the two are determined, and the difference between the positions is calculated to determine the target time difference.
[0086] After the target time difference is determined, the data played by the head-mounted playback device can be controlled to be synchronized with the data played by the external device based on the target time difference.
[0087] The synchronization includes the synchronization of the playback rate and the synchronization of the playback time. The head-mounted playback device and the external device belong to different devices, and the clocks used by the different devices are different. There can be a deviation between the clocks, and thus the playback rates are different. Therefore, in the embodiments of the present application, the playback rate can be adjusted to make the head-mounted playback device and the external device play synchronously.
[0088] In an embodiment, the processor 110 can be configured to adjust the acquisition rate of the wireless module 130 for acquiring the audio data stream of the target audio file from the sound source device based on the target time difference and a preset nominal frequency, so that the acquisition rate of the wireless module 130 for acquiring the target audio file matches the playback rate of the target audio file by the external device.
[0089] The nominal frequency is the working frequency corresponding to the nominal clock. Due to different reasons, the clock for controlling the audio playback of each device has a deviation from the nominal clock, and thus the frequency for controlling the audio playback of each device has a deviation from the nominal frequency, for example, the deviation is 100 ppm (parts per million, relative deviation value per million unit frequency), 50 ppm, 20 ppm, 10 ppm, 2 ppm, etc. This will cause the playback rates of the devices to be different.
[0090] The working frequency of the clock source is difficult to adjust. Therefore, in the embodiments of the present application, if the target time difference does not meet the preset value, the acquisition rate of the audio data stream of the target audio file from the sound source device can be adjusted, for example, the acquisition rate is accelerated or slowed down, and thus the rate of the data output from the first buffer to the head-mounted playback device for playback is accelerated or slowed down, so as to adjust the playback rate of the head-mounted playback device and reduce the difference between the playback rates of the head-mounted playback device and the external device.
[0091] In the embodiments of the present application, because the head-mounted playback device and the charging box 100 are different devices, the charging box 100 cannot directly control the playback of the head-mounted playback device, and therefore, the charging box 100 can control the rate at which the target audio file data is obtained from the sound source device, and further control the speed at which the target audio file data is output to the head-mounted playback device, so as to achieve the playback control of the head-mounted playback device.
[0092] In the embodiments of the present application, the rate at which the target audio file data is obtained from the sound source device can be adjusted so that the target time difference tends to a predetermined value, for example, 2 sampling points, or is within a predetermined interval, for example, the target time difference is within the range of 1 to 3 sampling points.
[0093] In the embodiments of the present application, adjusting the playback rate can be adjusting the sampling rate or the bit rate.
[0094] Correspondingly, in addition to the difference in playback rate, there is also a difference in playback time, and therefore, in some embodiments of the present application, the processor 110 can also be configured to generate a clock adjustment instruction based on the target time difference; wherein the clock adjustment instruction is used to instruct the head-mounted playback device to adjust the playback clock, so that the second playback time of the playback clock matches the sum of the first collection time and the target time difference, the first collection time being the collection time of the environmental audio signal by the audio collection module 120, and the second playback time being the playback time of the target audio file by the head-mounted playback device.
[0095] Therefore, in the embodiments of the present application, the playback of the target audio file by the head-mounted playback device can be controlled to be synchronized with the external device based on the target time difference, so that the sound heard by the user can be as consistent as possible.
[0096] In the embodiments of the present application, the audio collection module 120 can include a clock counter, which can record the first collection time. In other embodiments, a sampling point can also be set, for example, an audio sampling point is set at the microphone, the analog-to-digital conversion module, and the digital filter module, and the time point passing through the sampling point is locked by a digital hardware circuit to obtain the first collection time.
[0097] Correspondingly, a corresponding relationship between the specified position of the second buffer in the charging box 100 and the playback time of the head-mounted playback device can be preset in the charging box 100, so as to determine the data output time in the specified position of the second buffer as the second playback time. In other embodiments, the second playback time can be fed back by the head-mounted playback device.
[0098] In the embodiments of the present application, the generated clock adjustment instruction can be sent to the head-mounted playback device, and after receiving the clock adjustment instruction, the head-mounted playback device can adjust its playback clock, so that the second playback time coincides with the first collection time after the target time difference, to realize synchronization of the playback time.
[0099] Before the above-mentioned synchronous playback control is performed, the first buffer needs to buffer data exceeding a preset data amount, so that when the synchronous playback control is performed, the playback time and the playback rate can be adjusted based on the buffered data to realize synchronization. Therefore, before the synchronous playback control is performed, the data of the target audio file can be obtained from the sound source device at a non-synchronous playback rate.
[0100] By controlling synchronous playback, the sound played by the head-mounted playback device can be consistent with the sound played by the external device, thereby providing a good listening experience for the user. However, if the external sound is small, after the user wears the head-mounted playback device, the user can not be able to hear the sound played by the external device due to the physical barrier of the head-mounted playback device. In such a case, whether to play synchronously has different effects.
[0101] Therefore, in some embodiments of the present application, the processor 110 can also be configured to determine that the reliability of the environmental audio signal is greater than a preset threshold before determining the target time difference between the collection of the environmental audio signal by the audio collection module 120 and the acquisition of the target audio file from the sound source device.
[0102] In the embodiments of the present application, the reliability of the environmental audio signal is represented by the signal strength, power or amplitude of the audio signal played by the external device in the environmental audio signal. When the reliability is less than or equal to the preset threshold, it indicates that the sound played by the external device is small, and the user can not be able to hear it, which will not affect the user's listening. Therefore, synchronous control can not be needed, so as to reduce the resource consumption caused by synchronous control and reduce power consumption.
[0103] Based on the same inventive concept, the embodiments of the present application also provide an audio playback method. Please refer to Figure 3 , Figure 3 The flowchart of an audio playback method provided by an embodiment of the present application. The audio playback method can be applied to the charging box provided by the foregoing embodiments. The audio playback method comprises: S310, collecting an environmental audio signal by an audio collection module.
[0104] S320, obtaining an identification result of the environmental audio signal obtained by an audio recognition module.
[0105] S330, obtaining a target audio file from a sound source device based on the identification result and the wireless module.
[0106] S340, sending the target audio file to the head-mounted playback device for playing through the wireless module.
[0107] In an embodiment, the audio playing method further comprises: determining target audio file information and a target position based on the environmental audio signal; the target position representing a position of the environmental audio signal in the target audio file; and sending the target audio file to the head-mounted playback device for playing through the wireless module, comprising: continuously obtaining, through the target audio file information and the target position, an audio data stream of the target audio file after the target position by the wireless module, so that the head-mounted playback device plays the target audio file starting from the target position.
[0108] In an embodiment, after obtaining the target audio file from the sound source device based on the recognition result and the wireless module, a target time difference between the environmental audio signal collected by the audio collection module and the audio data stream corresponding to the target audio file obtained from the sound source device is determined; and the playing of the target audio file by the head-mounted playback device is controlled to be synchronized with the external device based on the target time difference.
[0109] In an embodiment, the charging box comprises a first buffer and a second buffer. Before sending the target audio file to the head-mounted playback device for playing through the wireless module, the target audio file can also be temporarily stored through the first buffer. Determining the target time difference between the environmental audio signal collected by the audio collection module and the target audio file obtained from the sound source device comprises: when the target audio file is output to the head-mounted playback device, the target audio file is output to the second buffer; and the target time difference is determined by performing audio correlation processing on the environmental audio signal and the target audio file in the second buffer; the audio correlation processing is used to determine the time difference between the same features of the input data.
[0110] In an embodiment, the size of the second buffer is equal to the sum of the first buffer size and the second buffer size, the first buffer size is the audio buffer size of the head-mounted playback device, and the second buffer size is the buffer size corresponding to the preset maximum time delay of the sending data of the charging box to the head-mounted playback device.
[0111] In an embodiment, controlling the playing of the target audio file by the head-mounted playback device to be synchronized with the external device based on the target time difference comprises: adjusting the acquisition rate of the audio data stream of the target audio file continuously obtained by the wireless module from the sound source device based on the target time difference and a preset nominal frequency, so that the acquisition rate of the target audio file obtained by the wireless module matches the playing rate of the target audio file by the external device.
[0112] In an embodiment, the playback of the target audio file by the head-mounted playback device is controlled to be synchronized with the external device based on the target time difference, including: generating a clock adjustment instruction based on the target time difference; wherein the clock adjustment instruction is used to instruct the head-mounted playback device to adjust a playback clock, so that a second playback time of the playback clock matches a sum of a first collection time and the target time difference, the first collection time being a collection time of the audio collection module on the environmental audio signal, and the second playback time being a playback time of the head-mounted playback device on the target audio file.
[0113] In an embodiment, the method further includes: before determining the target time difference between the collection of the environmental audio signal by the audio collection module and the acquisition of the target audio file from the sound source device, determining that the reliability of the environmental audio signal is greater than a preset threshold.
[0114] In an embodiment, the audio recognition module is configured with an audio recognition model, and the acquisition of the recognition result of the environmental audio signal by the audio recognition module includes: identifying the environmental audio signal through the audio recognition model to obtain the recognition result.
[0115] In an embodiment, the sound source device includes an audio recognition model, and the acquisition of the recognition result of the environmental audio signal by the audio recognition module includes: sending the environmental audio signal to the sound source device through the wireless module, and acquiring the recognition result of the environmental audio signal by the audio recognition model from the sound source device.
[0116] In an embodiment, before the collection of the environmental audio signal by the audio collection module, the method further includes: receiving a recognition instruction; in response to the recognition instruction, controlling the audio collection module to collect the environmental audio signal to enable the recognition of the environmental audio signal.
[0117] Based on the same inventive concept, the embodiments of the present application also provide an audio playback system, which includes the charging box and the head-mounted playback device provided by any of the preceding embodiments.
[0118] The charging box is used to charge the head-mounted playback device, and the head-mounted playback device can be a headset, smart glasses, etc. The headset can be a truly wireless headset, including but not limited to an in-ear headset, an over-ear headset, a semi-in-ear headset, etc.
[0119] The functions of the charging box and the head-mounted playback device can refer to the preceding embodiments, which will not be described here.
[0120] In the embodiments of the present application, it should be understood that the disclosed method and device can also be implemented in other manners. The embodiments described above are merely exemplary. The each of the functional modules in the embodiments of the present application can be integrated into one independent part, or each module can exist alone, or two or more modules can be integrated into one independent part.
[0121] The above described embodiments can be freely combined without conflict, and the combined embodiments are within the protection scope of the present application.
[0122] The above describes only the specific embodiments of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all the changes or replacements are within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0123] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes the elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
Claims
1. A charging case, characterized in that, The application discloses a charging box for charging a head-mounted playback device; the charging box comprises a processor, an audio acquisition module, a wireless module and an audio recognition module; The processor is connected with the audio acquisition module, the wireless module and the audio recognition module respectively; the audio recognition module is connected with the audio acquisition module and the wireless module respectively; The wireless module is used for being communicatively connected with the head-mounted playback device and a sound source device; The audio acquisition module is used for acquiring an environmental audio signal; the environmental audio signal comprises an audio signal of audio played by an external device; The audio recognition module is used for obtaining an identification result of the environmental audio signal; the identification result is used for representing a target audio file comprising the audio played by the external device; The processor is further used for obtaining the target audio file from the sound source device through the identification result and the wireless module; The processor is further used for sending the target audio file to the head-mounted playback device for playing through the wireless module.
2. The charging case of claim 1, wherein, The audio recognition module is further used for determining target audio file information and a target position based on the environmental audio signal; the target position represents a position of the environmental audio signal in the target audio file; The processor is used for continuously obtaining an audio data stream after the target position in the target audio file based on the target audio file information and the target position, so that the head-mounted playback device plays the target audio file starting from the target position.
3. The charging case of claim 1, wherein, The processor is further used for: determining a target time difference between the environmental audio signal acquired by the audio acquisition module and an audio data stream corresponding to the target audio file obtained from the sound source device; and controlling the playing of the target audio file by the head-mounted playback device to be synchronized with the external device based on the target time difference.
4. The charging case of claim 3, wherein, The charging box comprises a first buffer and a second buffer; the first buffer is connected with the second buffer and the wireless module; the second buffer is greater than or equal to the size of an audio buffer of the head-mounted playback device, so that the data buffered by the second buffer is more than or equal to the data buffered by the head-mounted playback device; The first buffer is used for temporarily storing the target audio file before the target audio file is sent to the head-mounted playback device; The first buffer is further used for outputting the target audio file to the second buffer when the target audio file is output to the head-mounted playback device; The processor is used for performing audio correlation processing on the environmental audio signal and the target audio file in the second buffer to determine the target time difference; the audio correlation processing is used for determining a time difference between the same features of input data.
5. The charging case of claim 4, wherein, The size of the second buffer is equal to the sum of a first buffer size and a second buffer size; the first buffer size is the size of the audio buffer of the head-mounted playback device, and the second buffer size is a buffer size corresponding to a preset maximum time delay of data sent by the charging box to the head-mounted playback device.
6. The charging case of claim 3, wherein, The processor is used for: adjust the acquisition rate of the wireless module for acquiring the audio data stream of the target audio file from the sound source device based on the target time difference and a preset nominal frequency, so that the acquisition rate of the wireless module for acquiring the target audio file matches the playing rate of the target audio file by the external device.
7. The charging case of claim 3, wherein, The processor is configured to: generate a clock adjustment instruction based on the target time difference, wherein the clock adjustment instruction is used to instruct the head-mounted playback device to adjust a playback clock, so that a second playback time point of the playback clock matches a sum of a first acquisition time point and the target time difference, the first acquisition time point being an acquisition time point of the audio acquisition module for acquiring the environmental audio signal, and the second playback time point being a playback time point of the head-mounted playback device for playing the target audio file.
8. The charging case of claim 3, wherein, The processor is further configured to: determine that the reliability of the environmental audio signal is greater than a preset threshold before determining the target time difference between the acquisition of the environmental audio signal by the audio acquisition module and the acquisition of the target audio file from the sound source device.
9. The charging case of any one of claims 1-8, wherein, The audio recognition module is configured with an audio recognition model, and the audio recognition module is configured to recognize the environmental audio signal through the audio recognition model to obtain the recognition result.
10. The charging case of any one of claims 1-8, wherein, The sound source device comprises an audio recognition model; The audio recognition module is configured to send the environmental audio signal to the sound source device through the wireless module, and acquire the recognition result of the environmental audio signal by the audio recognition model from the sound source device.
11. The charging case of any one of claims 1-8, wherein, The processor is further configured to: receive an identification instruction; in response to the identification instruction, control the audio acquisition module to acquire the environmental audio signal to enable identification of the environmental audio signal.
12. An audio playback system, characterized by The application relates to a charging box and a head-mounted playback device. The charging box is configured to charge the head-mounted playback device. A processor applied to a charging box, the charging box being configured to charge a head-mounted playback device, the charging box comprising a processor, an audio acquisition module, a wireless module and an audio recognition module.
13. An audio playing method, characterized in that, The audio playback method comprises: acquiring an environmental audio signal through the audio acquisition module, the environmental audio signal comprising an audio signal of audio played by an external device; acquiring a recognition result of the environmental audio signal by the audio recognition module, the recognition result being used to represent a target audio file comprising the audio signal played by the external device; acquiring the target audio file from a sound source device based on the recognition result and the wireless module; sending the target audio file to the head-mounted playback device for playing through the wireless module. The application relates to a charging box and a head-mounted playback device. The charging box is configured to charge the head-mounted playback device. A processor applied to a charging box, the charging box being configured to charge a head-mounted playback device, the charging box comprising a processor, an audio acquisition module, a wireless module and an audio recognition module. The audio playback method comprises: acquiring an environmental audio signal through the audio acquisition module, the environmental audio signal comprising an audio signal of audio played by an external device; acquiring a recognition result of the environmental audio signal by the audio recognition module, the recognition result being used to represent a target audio file comprising the audio signal played by the external device; acquiring the target audio file from a sound source device based on the recognition result and the wireless module; sending the target audio file to the head-mounted playback device for playing through the wireless module.