Audio parameter adjustment method and system, electronic device and earphone
Patent Information
- Application Number
- CN202610963544.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本申请实施例提供一种音频参数调整方法、系统、电子设备及耳机,以解决相关技术调音策略和降噪曲线与真实环境往往是失配的,手动调节的技能门槛要求相对较高,人工调整难度较大,导致用户体验不佳的技术问题
本申请的有益效果:本申请实施例提出的一种音频参数调整方法、系统、电子设备及耳机,通过获取环境感知数据并生成音频参数调整提示词,基于该音频参数调整提示词确定建议音频参数策略,根据该建议音频参数策略对音频参数进行更新,以该方法应用于真无线耳机为例,能够基于真无线耳机的实际使用场景的情况来确定相应的建议音频参数策略,与真实环境更为匹配,不需要用户手动调节EQ、ANC,降低了调节的技能门槛和难度,提升了用户体验。
Smart Images

Figure CN122802832A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio adjustment technology, and in particular to an audio parameter adjustment method, system, electronic device, and headphones. Background Technology
[0002] True wireless earbuds, also known as TWS (True Wireless Stereo) earbuds, are stereo earbuds that require no physical connection between the left and right earbuds and pair with devices such as mobile phones via Bluetooth technology. Their core technologies include Bluetooth transmission, main control chips, and various audio codecs, and they generally feature active noise cancellation, in-ear detection, and touch controls.
[0003] In related technologies, the "adaptive EQ (Equalizer)" of true wireless earbuds often only supports parameter adjustments for scenarios such as movement or stillness. "Adaptive ANC (Active Noise Cancellation)" only switches between coarse-grained modes like "commuting / indoor / outdoor." However, due to the complex environmental noise in earbud usage scenarios, the tuning strategies and noise cancellation curves in these technologies are often mismatched with real-world environments. Although users can manually adjust EQ and ANC, the skill required for manual adjustment is relatively high, making it difficult and resulting in a poor user experience. Summary of the Invention
[0004] This application provides an audio parameter adjustment method, system, electronic device, and headphones to address the technical problem that related technical tuning strategies and noise reduction curves are often mismatched with the real environment, the skill threshold for manual adjustment is relatively high, and the difficulty of manual adjustment is large, resulting in a poor user experience.
[0005] In a first aspect, embodiments of this application provide an audio parameter adjustment method, the method comprising: acquiring environmental perception data, generating audio parameter adjustment prompts based on the environmental perception data, and determining a suggested audio parameter strategy based on the audio parameter adjustment prompts; Update the audio parameters based on the suggested audio parameter strategy.
[0006] In conjunction with the first aspect, in one embodiment of this application, generating a suggested audio parameter strategy based on the audio parameter adjustment prompt word includes: sending the audio parameter adjustment prompt word to a smart terminal device, determining the suggested audio parameter strategy based on the audio parameter adjustment prompt word and providing feedback through the smart terminal device, wherein the audio parameter adjustment prompt word is generated through a wearable audio playback device; updating the audio parameters based on the suggested audio parameter strategy includes: receiving the suggested audio parameter strategy and updating the audio parameters of the wearable audio playback device based on the suggested audio parameter strategy.
[0007] In conjunction with the first aspect, in one embodiment of this application, updating audio parameters based on the suggested audio parameter strategy includes any one of the following: when the suggested audio parameter strategy includes an audio parameter identifier and a future parameter corresponding to at least one future time period, upon reaching the start time of the future time period, controlling the replacement of the audio parameter corresponding to the audio parameter identifier with the future parameter of the future time period; when the suggested audio parameter strategy includes an audio parameter identifier and a suggested parameter corresponding to the audio parameter identifier, controlling the replacement of the audio parameter corresponding to the audio parameter identifier with the suggested parameter; when the suggested audio parameter strategy includes an audio parameter identifier and a suggested parameter corresponding to the audio parameter identifier, obtaining the audio parameter corresponding to the audio parameter identifier, and if the audio parameter is different from the suggested parameter, controlling the replacement of the audio parameter corresponding to the audio parameter identifier with the suggested parameter.
[0008] In conjunction with the first aspect, in one embodiment of this application, the method for determining the environmental perception data includes: collecting environmental image data, identifying image content in the environmental image data, and determining environmental image content description text; collecting environmental sound data, identifying sound content in the environmental sound data, and determining environmental sound content description text; and determining the environmental perception data based on at least one of the environmental image content description text and the environmental sound content description text.
[0009] In conjunction with the first aspect, in one embodiment of this application, identifying image content in the environmental image data and determining environmental image content description text includes: if the environmental image data includes a text label image, identifying the text content of the text label image; if the environmental image data includes an object image, identifying the object structure of the object image and generating object description content; if the environmental image data includes multiple vehicle images, determining vehicle dynamics and generating vehicle dynamics description content; if the environmental image data includes multiple pedestrian images, determining pedestrian dynamics and generating pedestrian dynamics description content; generating image content based on at least one of the text content, object description content, vehicle dynamics description content, and pedestrian dynamics description content; and generating the environmental image content description text based on the image content.
[0010] In conjunction with the first aspect, in one embodiment of this application, updating audio parameters based on the suggested audio parameter strategy includes: updating equalizer parameters and / or active noise reduction parameters based on the suggested audio parameter strategy.
[0011] Secondly, this application provides an audio parameter adjustment method applied to a smart terminal device. The method includes: receiving an audio parameter adjustment prompt word sent by a wearable audio playback device; determining a suggested audio parameter strategy based on the audio parameter adjustment prompt word; and feeding back the suggested audio parameter strategy to the wearable audio playback device so that the wearable audio playback device can update the audio parameters based on the suggested audio parameter strategy.
[0012] In conjunction with the second aspect, in one embodiment of this application, determining a suggested audio parameter strategy based on the audio parameter adjustment prompt word includes: matching the audio parameter adjustment prompt word with a preset target scene; if the match is successful, determining the preset parameter strategy corresponding to the successfully matched preset target scene as the suggested audio parameter strategy; if the match fails, collecting current environmental noise data and obtaining the audio parameters of the wearable audio playback device; generating a noise feature vector based on the current environmental noise data and determining a parameter feature vector based on the audio parameters; generating a feature vector to be predicted based on the noise feature vector and the parameter feature vector; inputting the feature vector to be predicted into a pre-trained lightweight neural network to obtain an output result, the output result including at least one future parameter for a future time period; and generating the suggested audio parameter strategy based on the audio parameter identifier of the audio parameters and the output result.
[0013] Thirdly, embodiments of this application also provide an earphone, including a microphone, a speaker, a communication module, an audio processing module, an image acquisition device, a control module, and a processor. The image acquisition device is used to acquire environmental image data; the microphone is used to acquire environmental sound data; the processor is used to identify image content in the environmental image data and determine environmental image content description text, and / or identify sound content in the environmental sound data and determine environmental sound content description text, determine environmental perception data based on at least one of the environmental image content description text and environmental sound content description text, and then generate audio parameter adjustment prompts based on the environmental perception data; the communication module is used to send the audio parameter adjustment prompts to a smart terminal device, determine a suggested audio parameter strategy based on the audio parameter adjustment prompts and provide feedback, and receive the suggested audio parameter strategy; the control module is used to update the audio parameters of the audio processing module of the wearable audio playback device based on the suggested audio parameter strategy; and the speaker is used to play the processed audio transmitted by the smart terminal device after processing by the audio processing module.
[0014] Fourthly, embodiments of this application also provide an audio parameter adjustment system, the system comprising a wearable audio playback device and a smart terminal device, wherein: the wearable audio playback device is used to acquire environmental perception data, generate audio parameter adjustment prompts based on the environmental perception data, and send the audio parameter adjustment prompts to the smart terminal device; the smart terminal device is used to generate suggested audio parameter strategies based on the audio parameter adjustment prompts and provide feedback; the wearable audio playback device is further used to receive the suggested audio parameter strategies and update the audio parameters of the wearable audio playback device based on the suggested audio parameter strategies.
[0015] Fifthly, embodiments of this application also provide an electronic device, the electronic device including a first memory, a first processor, and an audio parameter adjustment program stored in the first memory and executable on the first processor, wherein when the audio parameter adjustment program is executed by the first processor, it implements the steps of the audio parameter adjustment method described in any of the embodiments of the second aspect above.
[0016] In a sixth aspect, embodiments of this application also provide an earphone, the earphone including a second memory, a second processor, and an audio parameter adjustment program stored in the second memory and executable on the second processor, wherein when the audio parameter adjustment program is executed by the second processor, it implements the steps of the audio parameter adjustment method as described in any of the embodiments of the first aspect above.
[0017] In a seventh aspect, this application provides a computer-readable storage medium storing instructions that, when executed by a processor, can implement any possible implementation method as described in the first or second aspect. Eighthly, this application provides a computer program product that may contain computer instructions that, when executed on a processor, can implement any possible implementation method as described in the first or second aspect. The beneficial effects of this application are as follows: The audio parameter adjustment method, system, electronic device, and headphones proposed in this application acquire environmental perception data and generate audio parameter adjustment prompts. Based on these prompts, a suggested audio parameter strategy is determined, and the audio parameters are updated according to the suggested audio parameter strategy. Taking true wireless headphones as an example, this method can determine the corresponding suggested audio parameter strategy based on the actual usage scenario of the true wireless headphones, which is more in line with the real environment. It eliminates the need for users to manually adjust EQ and ANC, reduces the skill threshold and difficulty of adjustment, and improves the user experience. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0019] In the attached diagram: Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application; Figure 2 A flowchart illustrating an audio parameter adjustment method provided in one embodiment of this application; Figure 3 A flowchart illustrating an audio parameter adjustment prompt generation method provided in one embodiment of this application; Figure 4 A flowchart illustrating a proposed audio parameter strategy generation method provided in an embodiment of this application; Figure 5 A flowchart illustrating a proposed audio parameter strategy generation method provided in an embodiment of this application; Figure 6 A schematic flowchart illustrating a method for adjusting audio parameters according to an embodiment of this application; Figure 7 A flowchart illustrating an audio parameter adjustment method provided in one embodiment of this application; Figure 8 This is a schematic diagram of an audio parameter adjustment system provided in an embodiment of this application. Detailed Implementation
[0020] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0021] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. The drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the shape, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.
[0023] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. For example... Figure 1 As shown, the headphones establish a communication connection with one or more devices such as mobile phones, tablets, laptops, and smartwatches, and can then play audio from these devices. The headphones can also send audio adjustment prompts to preset objects such as mobile phones, tablets, laptops, and smartwatches, so that the preset objects can execute the audio adjustment prompts to obtain suggested audio parameters. These suggested audio parameters are then fed back to the headphones, which update the processing parameters of the audio processing module based on these suggested audio parameters.
[0024] In some embodiments, when the earphones establish a wireless connection and establish a communication channel with a preset object such as a mobile phone, tablet, laptop, or smartwatch for the first time, authentication can be performed through channel authentication (such as pairing). If authentication is successful, the earphones send a message carrying authentication data to the preset object. The method by which the earphones establish a communication connection with the preset object such as a mobile phone, tablet, laptop, or smartwatch is not limited.
[0025] As an example, authentication can be achieved by establishing a pathway during connection. On this pathway, the smart terminal sends authentication data, which the earphone receives, verifies, and if authentication is successful, sends authentication success data back to the smart terminal device. Data interaction can then occur along this pathway.
[0026] As an example, the headphones can also control the image acquisition device to acquire environmental images (videos or pictures), and then obtain image content by performing text recognition, object structure recognition, vehicle dynamic recognition, pedestrian dynamic recognition, etc. on the environmental images. The image content is then sent to the headphones, which in turn generate audio adjustment prompts based on the image content.
[0027] As another example, the headphones can also actively capture and recognize audio to obtain audio content, and then generate audio adjustment prompts based on the audio content.
[0028] As another example, headphones can also generate ambient content based on both audio and image content, and then generate audio adjustment prompts based on the ambient content.
[0029] In some embodiments, the image acquisition device can be integrated into the headphones. Since mobile phones and other devices may be placed in bags or other spaces unsuitable for image data acquisition during actual use, image data may be distorted if images are acquired through mobile phones or other devices. However, by integrating a camera into the headphones, it can be further ensured that real environmental images can be acquired when the headphones are in use.
[0030] It should be noted that the above scenario is merely an example of an application scenario provided by the embodiments of this application. The method can also be applied to other scenarios according to the user's needs. The embodiments of this application do not limit the actual form of various devices, components, etc. included in this scenario. In the specific application of the solution, it can be set according to actual needs. It should be noted that the collection and processing of data such as environmental perception data in this application must strictly comply with the requirements of relevant national laws and regulations in actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0031] The inventors discovered that the existing TWS earphone device EQ and ANC technologies in related technologies have the following shortcomings: Crude scene recognition: Some headphones' "adaptive EQ" can only recognize simple scenes, such as movement or stillness, and "adaptive ANC" only switches in a coarse-grained manner according to "commuting / indoor / outdoor". Neither can distinguish specific types of environmental noise, such as low-frequency rumble in the subway, high-frequency tire noise on the bus, and high-frequency reverberation of human voices in a coffee shop, resulting in a mismatch between the tuning strategy and noise reduction curve and the real environment.
[0032] Dynamic scene transition delay: When suddenly moving from a quiet environment to a noisy street, adjustments such as EQ and ANC lag (approximately 1-3 seconds).
[0033] Processing latency: Insufficient computing power in low-end chips and complex EQ and ANC algorithms lead to audio-visual desynchronization, which is more obvious in game and video scenarios. High tuning threshold: Most users do not know how to manually adjust the EQ, and the logic of automatic dynamic EQ is simple, resulting in insufficient trust.
[0034] To address the aforementioned issues, this application provides an audio parameter adjustment method that utilizes an AI (Artificial Intelligence) model to dynamically adjust the audio EQ / ANC of TWS earphones. This method involves acquiring environmental images and / or sounds at the earphone end, processing the environmental image and sound data locally to generate audio parameter adjustment prompts, and then having a smart terminal device connected to the earphone decide on the recommended audio parameter strategy to guide the earphone in adjusting its audio parameters. This eliminates the need for manual adjustment of the audio EQ / ANC by the user, simplifying the process and avoiding the audio-visual desynchronization issues caused by insufficient computing power in low-end chips that cannot effectively run complex EQ and ANC algorithms or provide timely audio parameter suggestions. It also avoids the lag in EQ and ANC adjustments during sudden scene changes, supports finer-grained scene segmentation, and enables timely generation and adjustment of parameters based on actual environmental conditions, improving the matching degree between the tuning strategy and noise reduction curve and the real environment.
[0035] Please see Figure 2 , Figure 2 A flowchart illustrating an audio parameter adjustment method provided in an embodiment of this application is shown below. Figure 2 As shown, the method includes the following steps: Step S210: Obtain environmental perception data and generate audio parameters based on the environmental perception data to adjust the prompt words.
[0036] In some embodiments, the method for determining environmental perception data includes: collecting environmental image data, identifying image content in the environmental image data, and determining environmental image content description text; collecting environmental sound data, identifying sound content in the environmental sound data, and determining environmental sound content description text; and determining environmental perception data based on at least one of the environmental image content description text and the environmental sound content description text.
[0037] The method for identifying sound content from environmental sound data can employ speech recognition technology known to those skilled in the art, which will not be elaborated upon here. For example, if the voice "Please sit tight and hold on, next stop XX" is recognized, then this sound content can be included as part of the environmental perception data, facilitating the subsequent determination of the current scene where the headphones are located.
[0038] As an example, this method can be applied to wearable audio playback devices, including but not limited to one or more devices such as true wireless earbuds, over-ear headphones, ear-hook headphones, bone conduction headphones, neckband headphones, over-ear speakers, and smart audio glasses.
[0039] In some implementations, identifying the sound content of environmental sound data can determine the sound type, such as dog barks, train horns, keyboard clicks, conversations, etc., as well as the specific content of the conversation. Then, environmental sound content description text is organized based on this sound content. The sound content can be filled using a preset environmental sound content description text template, or the sound content can be directly used as the environmental sound content description text. Other methods known to those skilled in the art can also be used to convert the sound content into environmental sound content description text.
[0040] As an example, environmental perception data can be templated according to different perception sources. For instance, one template can be used for sound-related data, another for image-related data, and a separate template can be set for data from both sound and image sources. If the environmental perception data comes from two or more sources, such as sound and images, preliminary verification can be performed using multi-source data to identify any inconsistencies for subsequent analysis.
[0041] As an example, a wearable audio playback device may include at least one of an image acquisition device and a sound acquisition device. If the image acquisition device is integrated into the wearable audio playback device, its placement should ideally be such that it can capture images of the surrounding environment. The sound acquisition device may be a device such as a microphone from true wireless earbuds.
[0042] As an example, the determination of the environmental image content description text can be achieved by attaching a video chip to the wearable audio playback device, or by setting a video processing module in the main chip of the wearable audio playback device. The detection is only performed when the headphones are worn and the function is turned on, and the video module detection is not performed continuously, but at intervals, which can save power and computing power consumption.
[0043] As an example, taking true wireless earbuds as the wearable audio playback device, when the earbud case is opened or the earbuds are removed from the earbud case, the earbuds establish a connection and a pathway with the smart terminal device. The smart terminal device sends a pathway authentication message to the earbuds. If the authentication is successful, the earbuds send authentication data back to the smart terminal device.
[0044] As an example, continuing to use true wireless earbuds as the wearable audio playback device, after the true wireless earbuds complete pairing and communication connections with the smart terminal device, the detection module of the true wireless earbuds can be triggered to acquire environmental perception data, triggering the collection of at least one of environmental image data and environmental sound data.
[0045] As another example, true wireless earbuds can also be triggered by voice commands to activate the detection module and acquire environmental perception data, including at least one of environmental image data and environmental sound data. The voice command triggering method can be implemented in a manner known to those skilled in the art, such as "Hey, please adjust parameter XX for me."
[0046] As another example, true wireless earbuds can also be triggered by pressure commands to activate their detection module and acquire environmental awareness data, including at least one of environmental image data and environmental sound data. For instance, a pressure sensor on the earbuds can detect pressure input. If the pressure value exceeds a preset threshold, it indicates that the current audio output of the earbuds may be unclear to the user, such as when the user presses the earbuds. In this case, a series of steps, including acquiring environmental awareness data, are triggered to further adjust the audio parameters and improve audio clarity.
[0047] As another example, a preset interval could be set to trigger the true wireless earbuds to activate the detection module and acquire environmental perception data, including at least one of environmental image data and environmental sound data. For instance, audio parameters could be adjusted every five minutes.
[0048] As an example, when the wearable audio playback device is being worn and the function of dynamically adjusting audio parameters is enabled, step S210 is triggered.
[0049] Other known triggering actions may be used by those skilled in the art; the above are merely examples.
[0050] In some embodiments, identifying image content in environmental image data and determining environmental image content description text includes: if the environmental image data includes a text label image, identifying the text content of the text label image; if the environmental image data includes an object image, identifying the object structure of the object image and generating object description content; if the environmental image data includes multiple vehicle images, determining vehicle dynamics and generating vehicle dynamics description content; if the environmental image data includes multiple pedestrian images, determining pedestrian dynamics and generating pedestrian dynamics description content; generating image content based on at least one of the text content, object description content, vehicle dynamics description content, and pedestrian dynamics description content; and generating environmental image content description text based on the image content.
[0051] The identification of text labels, objects, vehicles, pedestrians, etc. in environmental image data can be achieved in a manner known to those skilled in the art, such as by using a lightweight detection model, which will not be elaborated here.
[0052] Similar to environmental sound content description text, environmental image content description text can also be pre-set with corresponding templates or content filling rules. Alternatively, image content can be directly used as environmental image content description text.
[0053] Please see Figure 3 , Figure 3 A flowchart illustrating an audio parameter adjustment prompt generation method provided in an embodiment of this application is shown below. Figure 3 As shown, taking a wearable audio playback device as headphones and audio parameters as EQ and ANC parameters as an example, if the headphones are worn and the dynamic adjustment EQ / ANC function is enabled, the video detection module is triggered. It reads the collected environmental image data, including identifiers, object structures, vehicle and pedestrian dynamics, and enables microphone monitoring. It then performs microphone recognition on the collected environmental sound data to determine the sound content. Based on the recognition results, it generates corresponding environmental prompts (audio parameter adjustment prompts), and the headphones generate prompts accordingly. For example, adjusting EQ / ANC is needed to enter a subway / coffee shop, or there might be an explosion.
[0054] In some embodiments, the audio parameter adjustment prompt can be based on environmental perception data and the processing module identifier of the audio processing module that needs to be adjusted. The processing module identifier can be one, two, or more, and can be set by those skilled in the art as needed. As an example, the processing module identifier in the audio parameter adjustment prompt can be an EQ module identifier (equalizer identifier) and an ANC module identifier (active noise cancellation identifier).
[0055] In other embodiments, the audio parameter adjustment prompt can be based on environmental perception data and audio parameter identifiers of the audio parameters to be adjusted. These audio parameter identifiers can be one, two, or more, and can be specifically set by those skilled in the art as needed. For example, the audio parameter identifiers in the audio parameter adjustment prompt can be EQ identifiers (equalizer identifiers) and ANC identifiers (active noise cancellation identifiers).
[0056] Step S220: Adjust the prompt words based on the audio parameters to determine the suggested audio parameter strategy.
[0057] In some embodiments, generating a suggested audio parameter strategy based on audio parameter adjustment prompts includes: sending audio parameter adjustment prompts to a smart terminal device; determining and feeding back the suggested audio parameter strategy based on the audio parameter adjustment prompts via the smart terminal device; and generating the audio parameter adjustment prompts via a wearable audio playback device. In this scenario, updating the audio parameters based on the suggested audio parameter strategy includes: receiving the suggested audio parameter strategy; and updating the audio parameters of the wearable audio playback device based on the suggested audio parameter strategy.
[0058] As an example, the suggested audio parameter strategy can be determined locally on the headphones. In this case, the computing chip deployed locally on the headphones or other wearable audio playback devices has high performance, which can meet the computing power requirements of this method. However, if the computing power requirement of the headphones is high, insufficient local computing power may lead to slow generation of the suggested audio parameter strategy, resulting in problems such as audio-visual desynchronization and adjustment lag. To address this issue, audio parameter adjustment prompts can be sent to the smart terminal device, which can then handle the determination of the suggested audio parameter strategy. This leverages the computing power of the smart terminal device to quickly determine the suggested audio parameter strategy, improving response efficiency. The specific methods for determining the suggested audio parameter strategy locally on the headphones and using a smart terminal device are similar, the difference being the executing entity. The following description uses the example of determining the suggested audio parameter strategy on a smart terminal device to illustrate the specific method for determining the suggested audio parameter strategy.
[0059] In some embodiments, the smart terminal device determines and feeds back a suggested audio parameter strategy based on an audio parameter adjustment prompt word, including: the smart terminal device receiving an audio parameter adjustment prompt word sent by a wearable audio playback device; determining a suggested audio parameter strategy based on the audio parameter adjustment prompt word; and feeding back the suggested audio parameter strategy to the wearable audio playback device so that the wearable audio playback device can update the audio parameters based on the suggested audio parameter strategy.
[0060] Following the above embodiments, the method for generating suggested audio parameter strategies based on audio parameter adjustment prompts includes: matching the audio parameter adjustment prompts with a preset target scene; if the match is successful, determining the preset parameter strategy corresponding to the successfully matched preset target scene as the suggested audio parameter strategy; if the match fails, collecting current environmental noise data and obtaining audio parameters from the wearable audio playback device; generating a noise feature vector based on the current environmental noise data and determining a parameter feature vector based on the audio parameters; generating a feature vector to be predicted based on the noise feature vector and the parameter feature vector; inputting the feature vector to be predicted into a pre-trained lightweight neural network to obtain an output result, the output result including at least one future parameter for a future time period; and generating a suggested audio parameter strategy based on the audio parameter identifier of the audio parameters and the output result.
[0061] The current environmental noise data can be time-series data over a relatively long period of time or data at a specific point in time, which can be set by those skilled in the art as needed.
[0062] It is understandable that after receiving an audio parameter adjustment prompt, the smart terminal device will perform scene recognition based on the prompt. If it identifies the current environment as a pre-defined special scenario, such as an explosion scenario (sight precedes sound, such as fireworks), then it is considered to have a corresponding preset target scenario. The suggested audio parameter strategy can be directly determined using the preset parameter strategy corresponding to that target scenario. For example, a value might be changed initially, then restored to its original value after a few seconds. In this case, the suggested audio parameter strategy could be to replace the current audio parameter with a preset parameter value, and after a few seconds, switch back to the initial audio parameter. A corresponding preset descriptive word can be set for each target scenario. By comparing the similarity between the preset descriptive word and the audio parameter adjustment prompt, it can be determined whether a corresponding target scenario exists and, specifically, which target scenario it is. Another scenario is that the current environment might not be a pre-defined target scenario. In this case, it is necessary to predict the audio parameters based on the actual environmental conditions. As an example, the smart terminal device can collect current environmental noise data and obtain the corresponding noise feature vector. The extraction method for this noise feature vector can be implemented using methods known to those skilled in the art, and will not be elaborated here. Then, the audio parameters currently being used by the wearable audio playback device are obtained, resulting in a parameter feature vector. The specific method for extracting the parameter feature vector can be implemented using methods known to those skilled in the art, and will not be elaborated upon. The noise feature vector and the parameter feature vector are then concatenated and input into a pre-trained lightweight neural network to obtain future parameters for future time periods. These future parameters are then sent to the headphones. Once the start time of the future time period is reached, the control module on the headphones replaces the audio parameters with the future parameters, thus significantly improving the problem of parameter adjustment lag.
[0063] As an example, the lightweight neural network can employ existing network structures such as LSTM (Long Short-Term Memory) neural networks. When training this lightweight neural network, sample data can be used for training until the model converges. The sample data includes sequences of sample data across multiple sample time periods, and these sequences contain sample noise features and sample parameter features. As an example, the sample data can be collected by adjusting audio parameters after a professional wears a wearable audio playback device and plays substitute sample noise audio. Alternatively, sample data can be collected through other methods known to those skilled in the art.
[0064] In some embodiments, after obtaining the output, if the future parameter of the first future time period closest to the current time is the same as the current audio parameter of the wearable audio playback device, then the future time period and future parameter can be deleted, and a suggested audio parameter strategy can be generated based on the remaining output and the processing module identifier. In this way, the wearable audio playback device does not need to perform a secondary judgment on the suggested audio parameter strategy; it only needs to execute the suggested audio parameter strategy, thus avoiding unnecessary parameter adjustments.
[0065] Please see Figure 4 , Figure 4 A flowchart illustrating a proposed audio parameter strategy generation method provided in an embodiment of this application is shown below. Figure 4 As shown, due to the pre-established communication connection between the wearable audio playback device and the smart terminal device, upon receiving the audio parameter adjustment prompt, the smart terminal device can trigger AI deep computing to analyze the content of the prompt (analyzed content). Based on this prompt (the prompt word), it performs deep thinking and activates corresponding algorithm logic. For example, after matching a preset target scene, it activates dynamic EQ / ANC changes for the preset scene (corresponding preset parameter strategy). If no preset target scene is matched, the smart terminal device activates microphone environmental noise detection and starts a neural network algorithm to calculate EQ / ANC parameters. Based on the comparison between the output parameter result and the input parameter (the current EQ / ANC parameters of the headphones), it obtains the EQ / ANC parameter strategy result. Then, the AI result (one of the two results mentioned above, audio parameter strategy is recommended) is pushed to the headphones through a private protocol.
[0066] As an example, the preset target scene can be a scene where the sound changes in a short time. For example, the background sound of fireworks may disappear after a period of time (e.g., 1 minute). At this time, a temporary parameter value can be set for the preset target scene, and after a period of time (e.g., also 1 minute), it can be restored to the previous audio parameter.
[0067] Please see Figure 5 , Figure 5 A flowchart illustrating a proposed audio parameter strategy generation method provided in an embodiment of this application is shown below. Figure 5As shown, if no preset target scene is matched, during the input acquisition phase, the microphone of the smart terminal device collects ambient noise, then performs preprocessing and feature extraction to obtain a noise feature vector. Meanwhile, the audio parameters (EQ / ANC parameters) currently being executed by the audio processing module of the wearable audio playback device are also sent to the smart terminal device, thus obtaining EQ feature vectors and / or ANC feature vectors. During the neural network processing phase, taking the presence of one of the two feature vectors as an example, the noise feature vector is concatenated with either the EQ or ANC feature vector, and then input into a lightweight neural network to obtain the output result. This yields the corresponding EQ / ANC parameters. If the audio processing module includes an active noise cancellation module and an equalizer, the noise feature vector is concatenated with the EQ feature vector and then input into a lightweight neural network trained for the EQ parameters to obtain the corresponding output result. Similarly, the noise feature vector is concatenated with the ANC feature vector and then input into a lightweight neural network trained for the ANC parameters to obtain the corresponding output result. The processing module identifier in the suggested audio parameter strategy distinguishes which audio processing module is being adjusted. As an example, the audio parameters currently executed by the audio processing module of a wearable audio playback device can be obtained through the interaction between the headphones and the smart terminal device after the smart terminal device and headphones are first connected.
[0068] As an example, Figure 5 The EQ / ANC parameters collected during the input acquisition phase can be obtained by acquiring audio parameters from devices such as headphones, or by the smart terminal processing the environmental noise according to the scene (e.g., through the EQ / ANC algorithm in the relevant technologies), or by querying the external network through the description of the environmental noise to obtain the EQ / ANC parameters given by other users, professionals or existing large language models in the scene description.
[0069] In some embodiments, one or more scene models can be trained to determine the current environmental scene based on audio parameter adjustment prompts, matching the current environmental scene with a preset target scene, and then using the preset parameter strategy of the corresponding environmental scene as the suggested audio parameter strategy. When suggesting audio parameter strategies, for fixed-change scenes and short-term change scenes, such as explosion scenes (sight before sound), a value is changed initially and restored to the original value after a few seconds. For unknown scenes, a neural network is needed for prediction. As an example, the scene model can use a large language model from existing technologies. By constructing descriptive prompts for different preset scenes and training the large language model, the large language model can be equipped with the ability to recognize the corresponding preset scene based on audio parameter adjustment prompts. If recognition fails, the strategy for scenarios where matching between the audio parameter adjustment prompts and the preset target scene fails (a pre-defined abnormal scene execution strategy) is used for subsequent processing.
[0070] As an example, the semantic similarity between the current input scene description and the descriptions of multiple preset target scenes can be calculated based on the content of the prompt word. If the similarity is higher than a first similarity threshold (e.g., greater than 95%), then the preset parameter strategy of the corresponding target scene is used as the output result. If the similarity is higher than a second similarity threshold (e.g., greater than 75%) but lower than the first similarity threshold, it is not an indication of output failure. Instead, the reasoning ability of a large model can be used to adjust the scene for at least one preset target scene corresponding to the highest similarity match, and adaptively adjust the preset parameter strategy corresponding to that target scene to derive the reasoning parameter strategy for the scene corresponding to the current prompt word as the output result. Furthermore, if the similarity is lower than or equal to the second similarity threshold, then a lightweight neural network is used to determine the corresponding strategy.
[0071] Step S230: Update the audio parameters based on the suggested audio parameter strategy.
[0072] As mentioned in the above embodiments, if the audio parameter adjustment prompt is generated by the wearable audio playback device, then updating the audio parameters based on the suggested audio parameter strategy further includes: receiving the suggested audio parameter strategy and updating the audio parameters of the wearable audio playback device based on the suggested audio parameter strategy.
[0073] In some embodiments, audio parameters are updated based on a proposed audio parameter strategy, including any one of the following: When the suggested audio parameter strategy includes an audio parameter identifier and a future parameter for at least one future time period corresponding to the audio parameter identifier, when the start time of the future time period is reached, the control replaces the audio parameter corresponding to the audio parameter identifier with the future parameter of the future time period. When the suggested audio parameter strategy includes an audio parameter identifier and a suggested parameter corresponding to the audio parameter identifier, control the replacement of the audio parameter corresponding to the audio parameter identifier with the suggested parameter. If the suggested audio parameter strategy includes an audio parameter identifier and the suggested parameters corresponding to the audio parameter identifier, the audio parameter corresponding to the audio parameter identifier is obtained. If the audio parameter is different from the suggested parameter, the audio parameter corresponding to the audio parameter identifier is replaced with the suggested parameter.
[0074] For example, taking the audio processing module corresponding to the audio parameters as an EQ module (such as an equalizer), if a corresponding preset target scene is matched, the corresponding suggested audio parameter strategy is obtained. This suggested strategy might be to immediately change the EQ parameter from Y to X, and then, after 10 seconds, change it back to Y (or change it to the default parameter). In this case, the processing module identifier for the suggested audio parameter strategy is the EQ module identifier, future time period 1 is from the current time to 10 seconds after the current time, with the corresponding future parameter being X, and future time period 2 is from the current time to 11 seconds after the current time, with the corresponding future parameter being Y (the initial audio parameter, or the default audio parameter).
[0075] For example, if the audio processing module corresponding to the audio parameters is an ANC module (such as an active noise cancellation module), the suggested audio parameter strategy might be to immediately change the ANC parameter from N to Q. In this case, the processing module identifier for the suggested audio parameter strategy is the ANC module identifier, and the suggested parameter is Q.
[0076] During the above process, the headphones do not need to verify the parameters; they directly execute the suggested audio parameter strategy. However, to ensure more reasonable parameter adjustments, the suggested or future parameters can be compared with the initial audio parameters before executing them. If they differ, the audio parameters are then replaced with the suggested or future parameters. In other words, one approach is to first determine if there has been a change, and then update accordingly; the other is to update based on the parameters in the suggested audio parameter strategy regardless of whether there has been a change.
[0077] In some embodiments, the audio processing module includes at least one of an active noise cancellation module and an equalizer.
[0078] In some embodiments, updating audio parameters based on a proposed audio parameter strategy includes updating equalizer parameters (EQ parameters) and / or active noise cancellation parameters (ANC parameters) based on the proposed audio parameter strategy.
[0079] In some embodiments, the wearable audio playback device's dynamically adjusted active noise cancellation or dynamically adjusted dynamic equalizer function is enabled. Frequency processing is performed on the received audio data using audio parameters to process different frequency components of the target sound source's audio signal. Examples include EQ and noise reduction. EQ, short for Equalizer, is used to adjust the timbre by gaining or attenuating one or more frequency bands of the audio signal. Noise reduction can involve filtering, such as removing certain frequency bands of the audio signal to reduce noise.
[0080] Please see Figure 6 , Figure 6A schematic flowchart illustrating a method for adjusting audio parameters according to an embodiment of this application is shown below. Figure 6 As shown, taking headphones as an example of a wearable audio playback device, firstly, the smart terminal device establishes a connection and a communication channel with the headphones. The smart terminal device sends a channel authentication request to the headphones. If the headphones successfully authenticate, they send relevant messages along with authentication data. The headphones then activate their detection module, using either their integrated image acquisition device or other external image acquisition devices connected to the headphones' communication network to collect environmental image data (such as...). Figure 6 , Figure 3 (Step "1" shown in the diagram) can also collect ambient sound data through the headset's microphone if needed, thereby generating environmental perception data and obtaining audio parameter adjustment prompts (the prompts in the diagram). The headset sends the audio parameter adjustment prompts to the smart terminal device (such as...) via a private protocol. Figure 6 , Figure 3 (As shown in step "2"), the smart terminal device (i.e., the smart device) initiates AI deep computing to obtain suggested audio parameter strategies (such as... Figure 6 , Figure 4 (As shown in step "3"), the smart terminal device pushes the AI results (suggested audio parameter strategies) to the headphones (e.g., via a private protocol) Figure 6 , Figure 4 As shown in step "4" (the headphones determine whether to update the EQ / ANC parameters based on the result (by comparing them with the existing audio parameters). If the results differ from the existing audio parameters, then the EQ / ANC parameters need to be updated. The headphones adjust in real time and update the EQ / ANC parameters accordingly. The EQ / ANC parameter update process can be performed at least once within a connection cycle until the connection is lost.
[0081] As a concrete example, taking headphones as a wearable audio playback device, when a user wearing headphones walks to a subway station, the headphone's video module captures and recognizes images, generates an audio parameter adjustment prompt message: "Approaching the subway, EQ / ANC adjustment needed," and sends it to the smart terminal device. After deep consideration of the prompt message, the smart terminal device selects to activate the microphone to detect ambient noise and obtain the current ambient noise data, and activates a neural network algorithm to pre-calculate the EQ / ANC parameters, obtain a suggested audio parameter strategy, and pushes the AI result (suggested audio parameter strategy) to the headphones through a private protocol. The headphones receive the result and update the latest EQ / ANC accordingly.
[0082] As another concrete example, taking headphones as a wearable audio playback device, the video module (image acquisition device) of the headphones, when worn, collects environmental image data and identifies an explosion scene (such as an image of fireworks). It generates a prompt that an explosion is imminent and requires EQ / ANC adjustment, which is sent to the smart terminal device. The smart terminal device deeply considers the prompt and activates the EQ / ANC (preset parameter strategy) corresponding to the preset target scene. The preset parameter strategy involves first adjusting the audio parameters to EQ_A and / or ANC_A values, then reverting to the original EQ and / or ANC values after 3 seconds. The AI result (preset parameter strategy) is pushed to the headphones via a private protocol. Upon receiving this, the headphones update their EQ / ANC according to the strategy: first adjusting the EQ parameters to EQ_A, then restoring the original EQ parameters (or restoring to the default EQ parameters) after 3 seconds; then adjusting the ANC parameters to ANC_A, and restoring the original ANC parameters (or restoring to the default ANC parameters) after 3 seconds. The explosion scene could be fireworks, a building exploding, or other scenarios where the explosion is seen first and then the sound is heard. Considering the extremely short-lived changes in the explosion scene, a value is adjusted first, and then restored to the default value after a few seconds.
[0083] The audio parameter adjustment method provided in the above embodiments acquires environmental perception data (which can be directly collected by a wearable audio playback device, or collected and transmitted to the wearable audio playback device by other devices, or a combination of the above two methods to obtain environmental perception data) and generates audio parameter adjustment prompts. Based on the audio parameter adjustment prompts, a suggested audio parameter strategy is determined, and the audio parameters are updated according to the suggested audio parameter strategy. Taking the application of this method to true wireless earbuds as an example, it can determine the corresponding suggested audio parameter strategy based on the actual usage scenario of the true wireless earbuds, which is more in line with the real environment. It eliminates the need for users to manually adjust EQ and ANC, reduces the skill threshold and difficulty of adjustment, and improves the user experience.
[0084] As an example, an environmental awareness data is collected by a wearable audio playback device, and an audio parameter adjustment prompt is generated. This prompt is then sent to a smart terminal device, which generates a suggested audio parameter strategy based on the prompt and provides feedback. Subsequently, the wearable audio playback device receives the suggested audio parameter strategy and updates the audio parameters of the audio processing module accordingly. Taking this method as an example when applied to true wireless earbuds, it can determine the corresponding suggested audio parameter strategy based on the actual usage scenario of the earbuds, making it more compatible with the real environment. This reduces the skill threshold and difficulty of adjustment, improves the user experience, and avoids the problem of audio-visual desynchronization caused by insufficient computing power of the earbud chip forcing operation.
[0085] The method provided in the above embodiments uses AI model results to dynamically adjust EQ / ANC, which has a wider range of scene recognition. It uses intelligent terminal devices with higher computing power to determine the suggested audio parameter strategy, avoiding the problem of audio-visual desynchronization caused by the headphone chip's insufficient computing power forcing operation. The solution provided in this application embodiment determines the suggested audio parameter strategy faster, making parameter switching faster and latency lower. The above method does not require manual adjustment by the user, resulting in a better user experience. It provides more refined and accurate scene recognition and also supports predictive suggested audio parameter strategies, making the switching speed faster and more timely in terms of experience.
[0086] Please see Figure 7 , Figure 7 A flowchart illustrating an audio parameter adjustment method provided in an embodiment of this application is shown below. Figure 7 As shown, taking the application of this method to a smart terminal device as an example, the smart terminal device includes, but is not limited to, one or more devices such as mobile phones, tablets, laptops, and smartwatches. The method includes the following steps: Step S710: Receive audio parameter adjustment prompts sent by the wearable audio playback device; Step S720: Determine the suggested audio parameter strategy by adjusting the prompt words based on the audio parameters; In step S730, the suggested audio parameter strategy is fed back to the wearable audio playback device so that the wearable audio playback device can update the audio parameters based on the suggested audio parameter strategy.
[0087] In some embodiments, generating a suggested audio parameter strategy based on adjusting prompts based on audio parameters includes: matching the prompts with a preset target scene based on the adjusted audio parameters; if the match is successful, determining the preset parameter strategy corresponding to the successfully matched preset target scene as the suggested audio parameter strategy; if the match fails, collecting current environmental noise data and obtaining audio parameters from a wearable audio playback device; generating a noise feature vector based on the current environmental noise data and determining a parameter feature vector based on the audio parameters; generating a feature vector to be predicted based on the noise feature vector and the parameter feature vector; inputting the feature vector to be predicted into a pre-trained lightweight neural network to obtain an output result, the output result including at least one future parameter for a future time period; and generating a suggested audio parameter strategy based on the audio parameter identifier of the audio parameters and the output result.
[0088] The audio parameter adjustment method provided in the above embodiments generates a suggested audio parameter strategy based on the audio parameter adjustment prompts sent by the wearable audio playback device through a smart terminal device and feeds it back to the wearable audio playback device. After receiving the suggested audio parameter strategy, the wearable audio playback device updates the audio parameters according to the suggested audio parameter strategy, which can improve the efficiency and accuracy of audio parameter adjustment of the wearable audio playback device. Taking the method applied to true wireless earbuds as an example, it can determine the corresponding suggested audio parameter strategy based on the actual usage scenario of the true wireless earbuds, which is more in line with the real environment. It does not require users to manually adjust EQ and ANC, reducing the skill threshold and difficulty of adjustment and improving the user experience. The determination of suggested audio parameter strategy is executed by the smart terminal device, which does not place high requirements on the chip at the earbud end, avoiding the problem of audio-visual desynchronization caused by the earbud chip's insufficient computing power forcing operation.
[0089] about Figure 7 The specific limitations of the provided audio parameter adjustment methods can be found in the limitations of the audio parameter adjustment methods mentioned above, and will not be repeated here.
[0090] In one embodiment, an audio parameter adjustment system is provided, which is used to perform the audio parameter adjustment method provided in any of the above embodiments. Please refer to [link to previous document]. Figure 8 , Figure 8 A schematic diagram of the structure of an audio parameter adjustment system provided in an embodiment of this application is shown below. Figure 8 As shown, the audio parameter adjustment system 800 includes a wearable audio playback device 801 and a smart terminal device 802. The wearable audio playback device 801 is used to acquire environmental perception data, generate audio parameter adjustment prompts based on the environmental perception data, and send the audio parameter adjustment prompts to the smart terminal device 802. The smart terminal device 802 is used to determine a suggested audio parameter strategy based on the audio parameter adjustment prompts and provide feedback. The wearable audio playback device 801 is also used to receive the suggested audio parameter strategy and update its audio parameters based on the suggested audio parameter strategy.
[0091] For specific limitations regarding the audio parameter adjustment system, please refer to the limitations on the audio parameter adjustment method described above, which will not be repeated here. Each module in the aforementioned audio parameter adjustment system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware within or independently of the processor in the electronic device, or stored in software within the memory of the electronic device, so that the processor can call and execute the corresponding operations of each module.
[0092] In this embodiment, the audio parameter adjustment system is essentially set up with multiple modules to execute the audio parameter adjustment method in any of the above embodiments. The specific functions and technical effects can be referred to in the above embodiments, and will not be repeated here.
[0093] In one embodiment, an earphone is provided, comprising a microphone, a speaker, a communication module, an audio processing module, an image acquisition device, a control module, and a processor. The image acquisition device is used to acquire environmental image data; the microphone is used to acquire environmental sound data; the processor is used to identify image content in the environmental image data and determine environmental image content description text, and / or identify sound content in the environmental sound data and determine environmental sound content description text, determining environmental perception data based on at least one of the environmental image content description text and environmental sound content description text, and then generating audio parameter adjustment prompts based on the environmental perception data; the communication module is used to send the audio parameter adjustment prompts to a smart terminal device, and the smart terminal device determines and feeds back a suggested audio parameter strategy based on the audio parameter adjustment prompts, and receives the suggested audio parameter strategy; the control module is used to update the audio parameters of a wearable audio playback device based on the suggested audio parameter strategy; and the speaker is used to play processed audio transmitted from the smart terminal device after processing by the audio processing module.
[0094] As an example, the earphones could be TWS earphones.
[0095] For specific limitations regarding headphones, please refer to the section above. Figures 2 to 6 The limitations of the provided audio parameter adjustment methods will not be elaborated further here. The various modules in the aforementioned headphones can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.
[0096] In this embodiment, the earphone is essentially equipped with multiple modules to perform the above-mentioned tasks. Figures 2 to 6 The audio parameter adjustment method in any of the provided embodiments can be referred to the above embodiments for specific functions and technical effects, and will not be repeated here.
[0097] This application also provides an electronic device, including a first processor, a first memory, and a communication bus; the communication bus connects the first processor and the first memory; the first processor executes a computer program stored in the first memory to achieve the above-described functionality. Figure 7 The method described in any of the provided embodiments.
[0098] This application also provides an electronic device, which includes a first memory, a first processor, and an audio parameter adjustment program stored in the first memory and executable on the first processor. When the audio parameter adjustment program is executed by the first processor, it implements... Figure 7 The steps of the audio parameter adjustment method provided in any of the provided embodiments.
[0099] This application embodiment also provides an earphone, which includes a second memory, a second processor, and an audio parameter adjustment program stored in the second memory and executable on the second processor. When the audio parameter adjustment program is executed by the second processor, it implements the following... Figures 2 to 6 The steps for adjusting the provided audio parameters.
[0100] This application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being used to cause a computer to perform the method provided in any of the above embodiments.
[0101] This application also provides a non-volatile readable storage medium storing one or more modules (programs) that, when applied to a device, enable the device to execute the instructions included in the steps provided in this application.
[0102] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0103] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0104] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0105] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0107] It should be understood that the terms "first," "second," etc., used in this application are used to distinguish similar objects and do not necessarily indicate a specific order or sequence. The technical features to which these terms are used can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.
[0108] It should be understood that, in the various embodiments of this application, unless the context clearly indicates otherwise, "one or more" or "at least one" means one or more (including two).
[0109] It should be understood that although the flowcharts provided in the embodiments of this application indicate the various steps with arrows, the order indicated by the arrows does not necessarily limit the implementation order of these steps. Those skilled in the art can perform these steps in other orders according to different implementation scenarios and requirements.
[0110] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A method for adjusting audio parameters, characterized in that, The method includes: Acquire environmental perception data, and generate audio parameters to adjust prompt words based on the environmental perception data; Based on the audio parameters, adjust the prompt words to determine the suggested audio parameter strategy; Update the audio parameters based on the suggested audio parameter strategy.
2. The audio parameter adjustment method as described in claim 1, characterized in that, The method for generating a suggested audio parameter strategy based on the audio parameter adjustment prompts includes: sending the audio parameter adjustment prompts to a smart terminal device; determining the suggested audio parameter strategy based on the audio parameter adjustment prompts and providing feedback through the smart terminal device; wherein the audio parameter adjustment prompts are generated by a wearable audio playback device. The step of updating audio parameters based on the suggested audio parameter strategy includes: receiving the suggested audio parameter strategy and updating the audio parameters of the wearable audio playback device based on the suggested audio parameter strategy.
3. The audio parameter adjustment method as described in claim 1, characterized in that, The audio parameters are updated based on the proposed audio parameter strategy, including any one of the following: When the suggested audio parameter strategy includes an audio parameter identifier and a future parameter for at least one future time period corresponding to the audio parameter identifier, when the start time of the future time period is reached, the audio parameter corresponding to the audio parameter identifier is replaced with the future parameter of the future time period. When the suggested audio parameter strategy includes an audio parameter identifier and a suggested parameter corresponding to the audio parameter identifier, control the replacement of the audio parameter corresponding to the audio parameter identifier with the suggested parameter; When the suggested audio parameter strategy includes an audio parameter identifier and a suggested parameter corresponding to the audio parameter identifier, the audio parameter corresponding to the audio parameter identifier is obtained. If the audio parameter is different from the suggested parameter, the audio parameter corresponding to the audio parameter identifier is replaced with the suggested parameter.
4. The audio parameter adjustment method as described in claim 1, characterized in that, The methods for determining the environmental perception data include: Collect environmental image data, identify the image content in the environmental image data, and determine the descriptive text of the environmental image content; Collect ambient sound data, identify the sound content in the ambient sound data, and determine the descriptive text of the ambient sound content; The environmental perception data is determined based on at least one of the environmental image content description text and the environmental sound content description text.
5. The audio parameter adjustment method as described in claim 3, characterized in that, Identify the image content in the environmental image data and determine the environmental image content description text, including: If the environmental image data includes text-labeled images, identify the text content of the text-labeled images; If the environmental image data includes object images, identify the object structure in the object images and generate object description content; If the environmental image data includes multiple vehicle images, determine the vehicle dynamics and generate a description of the vehicle dynamics. If the environmental image data includes multiple pedestrian images, determine the pedestrian dynamics and generate a description of the pedestrian dynamics; Image content is generated based on at least one of the text content, object description content, vehicle dynamic description content, and pedestrian dynamic description content; Generate environmental image content description text based on the image content.
6. The audio parameter adjustment method as described in claim 1, characterized in that, Update the audio parameters based on the proposed audio parameter strategy, including: Update the equalizer parameters and / or active noise reduction parameters based on the proposed audio parameter strategy.
7. A method for adjusting audio parameters, characterized in that, Applied to smart terminal devices, the method includes: Receive audio parameter adjustment prompts sent by wearable audio playback devices; Based on the aforementioned audio parameters, adjust the prompt words to determine a suggested audio parameter strategy; The suggested audio parameter strategy is fed back to the wearable audio playback device so that the wearable audio playback device can update the audio parameters based on the suggested audio parameter strategy.
8. The audio parameter adjustment method as described in claim 7, characterized in that, Based on the aforementioned audio parameters, a strategy for determining suggested audio parameters by adjusting prompt words includes: Adjust the prompt words according to the audio parameters to match the preset target scene; If a match is successful, the preset parameter strategy corresponding to the preset target scenario that is successfully matched will be determined as the suggested audio parameter strategy; If the matching fails, the system collects current ambient noise data and obtains the audio parameters of the wearable audio playback device; generates a noise feature vector based on the current ambient noise data and determines a parameter feature vector based on the audio parameters; generates a feature vector to be predicted based on the noise feature vector and the parameter feature vector; inputs the feature vector to be predicted into a pre-trained lightweight neural network to obtain an output result, the output result including at least one future parameter for a future time period; and generates the suggested audio parameter strategy based on the audio parameter identifier of the audio parameters and the output result.
9. An audio parameter adjustment system, characterized in that, The audio parameter adjustment system includes a wearable audio playback device and a smart terminal device, wherein: A wearable audio playback device is used to acquire environmental perception data, generate audio parameter adjustment prompts based on the environmental perception data, and send the audio parameter adjustment prompts to a smart terminal device. The intelligent terminal device is used to determine a suggested audio parameter strategy based on the audio parameters and adjust the prompt words accordingly, and then provide feedback. The wearable audio playback device is also configured to receive the suggested audio parameter strategy and update the audio parameters of the wearable audio playback device based on the suggested audio parameter strategy.
10. An electronic device, characterized in that, The electronic device includes a first memory, a first processor, and an audio parameter adjustment program stored in the first memory and executable on the first processor. When the audio parameter adjustment program is executed by the first processor, it implements the steps of the audio parameter adjustment method as described in any one of claims 7 to 8.
11. An earphone, characterized in that, The headphones include a second memory, a second processor, and an audio parameter adjustment program stored in the second memory and executable on the second processor. When the audio parameter adjustment program is executed by the second processor, it implements the steps of the audio parameter adjustment method as described in any one of claims 1 to 6.