Non-contact audio processing method and system, wearable device and storage medium
By preprocessing and neural network identification of audio signals and environmental data of contactless audio devices and adjusting audio signals according to the environment type, the problem that existing devices cannot provide high-quality audio is solved, achieving a better auditory experience and meeting high-quality audio needs.
Patent Information
- Application Number
- CN202510351627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
Existing contactless audio devices cannot effectively identify the external environment and adjust the audio signal according to the environment, resulting in the inability to provide a high-quality audio experience and cannot meet the users' growing high-quality audio needs.
By preprocessing the first audio signal and environmental data, a neural network model is input to identify the environment type in which the user is located, processing the second audio signal according to the environment type, determining the target adjustment parameters and adjusting to obtain the target audio signal.
Through the identification of the external environment and the adjustment of the audio signal, high-quality audio signals are provided, which improves the user's auditory experience and meets the users' growing demand for high-quality audio.
Smart Images

Figure CN120224080A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of audio processing, and more specifically, relates to a non-contact audio processing method and system, a wearable device, and a storage medium. Background Art
[0002] Existing non-contact audio devices, although getting rid of the drawbacks brought by direct contact with the ear canal, have shortcomings in audio processing methods. Non-contact audio devices often cannot recognize the external environment and are difficult to adjust the audio signal according to the external environment to transmit high-quality audio signals to users; this reduces the user's auditory experience and cannot meet the growing demand for high-quality audio from users. Summary of the Invention The purpose of the present disclosure is to provide a non-contact audio processing method and system, a wearable device, and a storage medium, which improve the user's auditory experience and meet the growing demand for high-quality audio from users.
[0003] In the first aspect of the embodiments of the present disclosure, a non-contact audio processing method is provided, including: Preprocessing the first audio signal and environmental data to obtain first data; Inputting the first data into a neural network model to obtain the environmental type where the user is located; Processing the second audio signal according to the environmental type to obtain a target audio signal.
[0004] Optionally, the processing the second audio signal according to the environmental type to obtain a target audio signal includes: Determining the target adjustment parameter of the second audio signal based on the environmental type; Determining the adjustment step size of the target adjustment parameter based on the frequency domain characteristics of the environmental type; Adjusting the target adjustment parameter of the second audio signal based on the adjustment step size of the target adjustment parameter to obtain a target audio signal.
[0005] Optionally, the determining the adjustment step size of the target adjustment parameter based on the frequency domain characteristics of the environmental type includes: Selecting a corresponding target audio processing strategy from multiple audio processing strategies based on the frequency domain characteristics of the environmental type; The determination process of the target audio output strategy includes: In response to the frequency domain characteristics of the environmental type satisfying the low-frequency characteristics, adjusting the target adjustment parameter of the second audio signal based on the first low-frequency step size; In response to the frequency domain characteristics of the environmental type satisfying the medium-frequency characteristics, adjusting the target adjustment parameter of the second audio signal based on the first medium-frequency step size; In response to the frequency domain feature of the environment type satisfying the high-frequency feature, the target adjustment parameter of the second audio signal is adjusted based on the first high-frequency step size.
[0006] Optionally, the non-contact audio processing method further includes: Performing a fast Fourier transform on the first audio signal to obtain an audio frequency domain feature; Dividing the audio frequency domain feature into multiple sub-bands; the sub-bands include a low-frequency sub-band, a middle-frequency sub-band, and a high-frequency sub-band; Assigning weights to each sub-band based on the proportion of each sub-band in the audio frequency domain feature; Performing a weighted calculation on the weights assigned to each sub-band to obtain a frequency domain feature value; Determining the frequency domain feature of the environment type based on the frequency domain feature value.
[0007] Optionally, the determining the frequency domain feature of the environment type based on the frequency domain feature value includes: In response to the frequency domain feature value being less than or equal to the first threshold, determining that the frequency domain feature of the environment type is a low-frequency feature; In response to the frequency domain feature value being greater than the first threshold and the frequency domain feature value being less than the second threshold, determining that the frequency domain feature of the environment type is a middle-frequency feature; In response to the frequency domain feature value being greater than the second threshold, determining that the frequency domain feature of the environment type is a high-frequency feature.
[0008] Optionally, the target adjustment parameter includes at least one of an equalizer gain, a dynamic range compression ratio, and a noise suppression threshold. When the frequency domain feature of the environment type is a low-frequency feature, the priority weight of the equalizer gain is the reciprocal of the low-frequency sub-band energy; When the frequency domain feature of the environment type is a middle-frequency feature, the priority weight of the noise suppression threshold is proportional to the high-frequency sub-band energy; When the frequency domain feature of the environment type is a high-frequency feature, the priority weight of the dynamic range compression ratio is determined by the variance of the middle-frequency sub-band energy.
[0009] Optionally, the determining the target adjustment parameter of the second audio signal based on the environment type includes: Obtaining user preference data, where the user preference data includes equalization parameters or noise reduction intensities previously selected by the user; Obtaining the target adjustment parameter based on a weighted calculation of the environment type and the user preference data.
[0010] In a second aspect of the embodiments of the present disclosure, a non-contact audio processing system is provided, including: The first computing module is configured to preprocess the first audio signal and environmental data to obtain first data; The second computing module is configured to input the first data into a neural network model to obtain the environmental type where the user is located; The third computing module is configured to process the second audio signal according to the environmental type to obtain a target audio signal.
[0011] In a third aspect of the embodiments of the present disclosure, a wearable device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above non-contact audio processing method are implemented.
[0012] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above non-contact audio processing method are implemented.
[0013] The beneficial effects of the non-contact audio processing method, system, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows: By identifying the external environment, the audio signal is adjusted, and high-quality audio signals are transmitted to the user; the auditory experience of the user is improved, and the growing demand for high-quality audio of the user is met. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a flowchart of a non-contact audio processing method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of a non-contact audio processing system provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0016] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0017] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described by way of specific embodiments Figures 1-3 for illustration.
[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a non-contact audio processing method provided for an embodiment of the present disclosure. The method includes: S101: Preprocess the first audio signal and environmental data to obtain first data.
[0019] In this embodiment, the first audio signal is obtained by a microphone collecting the sounds around the user in real time. The content of the first audio signal contains information on environmental characteristics; for example, the sound of vehicle traffic on the road, the noisy voices in the mall, the sound of keyboard tapping in the office, etc. The environmental data around the user is collected in real time by multiple sensors; for example, the temperature around the user is collected by a temperature sensor, the noise data of the environment around the user is collected using a sound level meter, and the ambient sound pressure around the user is measured using a sound pressure meter.
[0020] In this embodiment, preprocessing the first audio signal and environmental data to obtain first data includes: Performing noise reduction, filtering, normalization, etc. on the first audio signal; performing calibration and normalization on the environmental data; and fusing the preprocessed first audio signal and environmental data to obtain first data.
[0021] In this embodiment, the manner of fusing the preprocessed first audio signal and environmental data includes: assigning different weights according to the importance of the data, and performing weighted fusion to obtain first data; for example, determining through experience or previous experiments that the weight of the audio feature vector is w1 and the weight of the environmental data feature vector is w2 (w1 + w2 = 1), multiplying the corresponding elements of the two feature vectors by their respective weights and then adding them to obtain the fused feature vector.
[0022] Performing principal component analysis (PCA) on the preprocessed audio features and environmental data features respectively, transforming them into new principal component spaces, obtaining the dimensionality-reduced feature representations; and then splicing or combining these dimensionality-reduced features in other ways.
[0023] S102: Input the first data into the neural network model to obtain the environmental type where the user is located.
[0024] In this embodiment, the obtained first data is input into a neural network model trained with a large amount of audio data and environmental data under different environments; the neural network model calculates the environmental type where the user is located through the forward propagation algorithm. Among them, the neural network model has learned the characteristic patterns of audio signals and environmental data under different environmental types; the neural network model can use a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN). The environmental types include indoor quiet environment, outdoor noisy environment, transportation hub environment, industrial production environment, etc.
[0025] S103: Process the second audio signal according to the environmental type to obtain the target audio signal.
[0026] In this embodiment, according to the identified environmental type, for the second audio signal, appropriate target adjustment parameters of the second audio signal are selected for processing.
[0027] In this embodiment, the target adjustment parameters of the second audio signal are determined based on the environmental type; the adjustment step size of the target adjustment parameters is determined based on the frequency domain characteristics of the environmental type; the target adjustment parameters of the second audio signal are adjusted based on the adjustment step size of the target adjustment parameters to obtain the target audio signal.
[0028] Among them, the method for determining the target adjustment parameters of the second audio signal based on the environmental type includes: First, establish a mapping relationship database in advance; the database details the range of basic audio adjustment parameters corresponding to different environmental types. When the environmental type is obtained, the initial audio adjustment parameter set corresponding to it is retrieved from the above mapping relationship database; the target adjustment parameters of the second audio signal are determined; Second, use feature extraction algorithms such as Short-Time Fourier Transform (STFT) or Mel-frequency Cepstral Coefficients (MFCC) to obtain the real-time frequency domain and time domain characteristics of the audio signal; according to these real-time characteristics, analyze the proportion of components such as speech, music, and noise in the audio signal.
[0029] III. Train a machine learning model. The machine learning model takes environmental type data, historical audio processing data, and user feedback data as inputs and outputs target adjustment parameters. During operation, the current environmental type data and relevant audio data processed previously are input into the machine learning model. The machine learning model performs calculations and predictions through internal algorithms (such as decision trees, neural networks, etc.) and outputs target adjustment parameters suitable for the current situation. The target adjustment parameters include at least one of the equalizer gain, dynamic range compression ratio, and noise suppression threshold.
[0030] In this embodiment, processing the second audio signal according to the environmental type to obtain the target audio signal includes: Determining the target adjustment parameter of the second audio signal based on the environmental type; Determining the adjustment step size of the target adjustment parameter based on the frequency domain characteristics of the environmental type; Adjusting the target adjustment parameter of the second audio signal based on the adjustment step size of the target adjustment parameter to obtain the target audio signal.
[0031] Specifically, to determine the adjustment step size of the target adjustment parameter, when the target adjustment parameter is the volume gain, gradually increase or decrease the volume gain according to the determined adjustment step size of the target adjustment parameter until the audio signal most suitable for the current environment is obtained; when the target adjustment parameter is the dynamic range compression ratio, gradually increase or decrease the compression ratio according to the adjustment step size. The finely adjusted audio signal of the present invention is the target audio signal, meeting the user's demand for high-quality audio in different environments.
[0032] In this embodiment, determining the adjustment step size of the target adjustment parameter based on the frequency domain characteristics of the environmental type includes: Performing a fast Fourier transform on the first audio signal to obtain audio frequency domain characteristics; Dividing the audio frequency domain characteristics into multiple sub-bands; the sub-bands include a low-frequency sub-band, a middle-frequency sub-band, and a high-frequency sub-band; Assigning weights to each sub-band based on the proportion of each sub-band in the audio frequency domain characteristics; Performing a weighted calculation on the weights assigned to each sub-band to obtain a frequency domain feature value; Specifically, is the m energy of the m sub-band, n ={1, 2, 3... },
[0033] The energy entropy H is calculated as
[0034] Weight The calculation formula is
[0035] Frequency domain eigenvalue F The calculation formula is
[0036] Among them, is the weight of the m th sub - frequency band, p m is the proportion of the m th sub - frequency band in the audio frequency domain characteristics; is the adjustment factor. Perform a fast Fourier transform on the audio signal to convert it from the time domain to the frequency domain to obtain the frequency domain characteristics of the audio. Divide the obtained audio frequency domain characteristics into multiple sub - frequency bands, namely the low - frequency sub - band, the middle - frequency sub - band, and the high - frequency sub - band. Calculate the energy of each sub - frequency band and use the proportion of its total energy as the weight of the sub - frequency band. Select a suitable eigenvalue for each sub - frequency band, such as the average frequency of the sub - frequency band. Multiply the weight of each sub - frequency band by the corresponding eigenvalue and then sum to obtain the frequency domain eigenvalue.
[0037] In this embodiment, determine the frequency domain characteristics of the environment type based on the frequency domain eigenvalue; Select the corresponding target audio processing strategy from multiple audio processing strategies based on the frequency domain characteristics of the environment type; The determination process of the target audio output strategy includes: In response to the frequency domain characteristics of the environment type satisfying the low - frequency characteristics, adjust the target adjustment parameter of the second audio signal based on the first low - frequency step; In response to the frequency domain characteristics of the environment type satisfying the middle - frequency characteristics, adjust the target adjustment parameter of the second audio signal based on the first middle - frequency step; In response to the frequency domain characteristics of the environment type satisfying the high - frequency characteristics, adjust the target adjustment parameter of the second audio signal based on the first high - frequency step.
[0038] The target adjustment parameter includes at least one of the equalizer gain, the dynamic range compression ratio, and the noise suppression threshold, When the frequency domain characteristics of the environment type are low - frequency characteristics, the priority weight of the equalizer gain is the reciprocal of the low - frequency energy of the sub - band; When the frequency domain characteristics of the environment type are middle - frequency characteristics, the priority weight of the noise suppression threshold is proportional to the energy of the high - frequency sub - band; When the frequency domain characteristics of the environment type are high - frequency characteristics, the priority weight of the dynamic range compression ratio is determined by the variance of the energy of the middle - frequency sub - band.
[0039] Wherein the first low-frequency step size, the first intermediate-frequency step size, and the first high-frequency step size are respectively different preset values, and the first low-frequency step size is greater than the first intermediate-frequency step size, and the first intermediate-frequency step size is greater than the first high-frequency step size.
[0040] Specifically, set the first threshold to 1 and the second threshold to 2. In response to the frequency-domain eigenvalue being less than or equal to the first threshold, determine that the frequency-domain feature of the environmental type is a low-frequency feature; in response to the frequency-domain eigenvalue being greater than the first threshold and less than the second threshold, determine that the frequency-domain feature of the environmental type is an intermediate-frequency feature; in response to the frequency-domain eigenvalue being greater than the second threshold, determine that the frequency-domain feature of the environmental type is a high-frequency feature.
[0041] When it is determined that the frequency-domain feature of the environmental type satisfies the low-frequency feature, adjust the target adjustment parameter of the second audio signal based on the first low-frequency step size. Using the adaptive filter technology, adjust the parameters of the filter according to the first low-frequency step size. For example, for a low-pass filter, adjust its cut-off frequency according to the first low-frequency step size. If the first low-frequency step size is 10 Hz and the current cut-off frequency is 150 Hz, the adjusted cut-off frequency becomes 160 Hz, thereby enhancing or weakening the audio signal in the low-frequency part. In some industrial environments, the operation of large mechanical equipment will generate continuous and strong low-frequency noise; in enclosed spaces such as underground parking lots, the engine noise of vehicles, the friction sound between tires and the ground, etc. are also mainly concentrated in the low-frequency region. Using the first low-frequency step size to adjust the target adjustment parameter can more accurately process the low-frequency characteristics. If the step size is set too large, when adjusting the audio signal, it may overly change the characteristics of the low-frequency signal, resulting in serious audio distortion and loss of important audio details. For example, when adjusting the gain of an equalizer, an overly large step size may cause the low-frequency gain to increase or decrease excessively, making the sound unnatural, and the original normal low-frequency sound becomes dull or overly sharp.
[0042] When the frequency-domain characteristics of the environmental type satisfy the intermediate-frequency characteristics, the target adjustment parameter of the second audio signal is adjusted based on the first intermediate-frequency step. At this time, the priority weight of the noise suppression threshold is proportional to the energy of the high-frequency sub-band. When using the dynamic equalizer technology to automatically adjust the intermediate-frequency sub-band gain according to the first intermediate-frequency step, the noise suppression threshold will be preferentially adjusted according to the weight. For example, in the music playback scenario, when it is detected that the instrument sound in the intermediate-frequency part is prominent, the intermediate-frequency gain is appropriately reduced according to the first intermediate-frequency step, and at the same time, the noise suppression threshold is increased according to the weight to balance the overall audio effect; for example, in the voice call scenario, the intermediate-frequency voice band is enhanced according to the first intermediate-frequency step, and at the same time, the noise suppression threshold is adjusted according to this weight to improve the clarity of the voice. The intermediate-frequency signal includes rich environmental information and key sound elements. In a bustling market, as well as the daily conversations of people in the office, the slight operation sounds of printers and computer fans, the energy of these sounds is mostly concentrated in the intermediate-frequency range. If the target adjustment parameter is adjusted with an inappropriate step, these important audio information will be seriously damaged. For example, when adjusting the equalizer gain, an inappropriate step will cause the gain of the intermediate-frequency part to be abnormal, resulting in the originally clear human voice becoming blurred and the key information being difficult to be accurately received. Therefore, when the environment presents intermediate-frequency characteristics, using a special first intermediate-frequency step to adjust the audio can better adapt to the auditory perception of the human ear, ensuring that the sound heard by people is natural, clear, and in line with normal auditory habits.
[0043] When the frequency-domain characteristics of the environmental type satisfy the high-frequency characteristics, the target adjustment parameter of the second audio signal is adjusted based on the first high-frequency step. The priority weight of the dynamic range compression ratio is determined by the variance of the energy of the intermediate-frequency sub-band. First, calculate the variance σ² of the energy of the intermediate-frequency sub-band, and use this as the priority weight of the dynamic range compression ratio. When using the noise shaping technology to transfer the noise energy from the high-frequency region sensitive to the human ear to the relatively insensitive high-frequency region according to the first high-frequency step, the dynamic range compression ratio is preferentially adjusted according to the weight of σ², while maintaining the high-frequency details of the audio signal and improving the audibility of the audio.
[0044] At the construction site, the sharp and ear-piercing sounds emitted by tools such as cutters and drills, as well as the slight current sounds of electronic components in some environments with dense electronic devices, all belong to high-frequency sounds. If an inappropriate step size is used to adjust the target adjustment parameter, the information carried by these high-frequency signals will be severely damaged. Taking the adjustment of the noise suppression threshold as an example, if the step size is set improperly, it may over-suppress useful high-frequency sounds, resulting in a dry and lifeless audio; or it may not effectively suppress high-frequency noise, causing a large amount of interference noise to be mixed into the audio, affecting the normal sound reception. Therefore, when the environment exhibits high-frequency characteristics, using a dedicated first high-frequency step size to adjust the audio can more accurately match the human ear's perception of high-frequency sounds. Appropriate adjustment can make the sound clearer and more layered, avoiding the dullness of the sound or the loss of details caused by improper adjustment, thus meeting people's pursuit of sound quality.
[0045] In this embodiment, determining the target adjustment parameter of the second audio signal based on the environment type includes: Obtain user preference data, where the user preference data includes the equalization parameters or noise reduction intensity previously selected by the user; calculate the target adjustment parameter based on the weighted calculation of the environment type and the user preference data.
[0046] Specifically, the user preference data is stored in the local database or cloud server of the device. When it is necessary to obtain the user preference data, the user preference data includes, but is not limited to, the equalization parameters, noise reduction intensity, tone preference, volume adjustment habits, etc. previously selected by the user. The equalization parameters specifically cover the detailed adjustment records of the sound gain in different frequency bands (such as low frequency, middle frequency, and high frequency).
[0047] Calculating the target adjustment parameter based on the weighted calculation of the environment type and the user preference data includes: Assign weights to the environmental factors and user preferences respectively; when the environmental noise is large, the weight of the environmental factors is relatively increased; the weights of the adjustment parameters obtained according to the environment type and the user preferences are y1 and y2 respectively; calculate the target adjustment parameter according to the assigned weights and weighted calculation; For example, for the equalizer gain, assuming that the basic gain corresponding to the environmental factors is [G1, G2, G3] (corresponding to the low, middle, and high frequency bands respectively), and the gain corresponding to the user preference is [P1, P2, P3], then the target adjustment parameter is [y1G1 + y2P1, y1G2 + y2P1, y1G1 + y2P1].
[0048] The key adjustment parameter for the low-pass filter is the cut-off frequency fc. Assume that the cut-off frequency corresponding to the environmental type is fc1, and the cut-off frequency preferred by the user is fc2. Combining the weights y1 and y2, calculate the target cut-off frequency fc = y1fc1 + y2fc2 of the low-pass filter. At the same time, the quality factor Q of the low-pass filter also affects its filtering characteristics; assume that the basic quality factor corresponding to the environmental type is Q1 and the quality factor corresponding to the user preference is Q2, and calculate the target quality factor Q = y1Q1 + y2Q2 to optimize the processing effect of the low-pass filter on the audio signal.
[0049] For the dynamic range compression ratio, assume that the basic value of the dynamic range compression ratio corresponding to the environmental type is R1, the basic value of the dynamic range compression ratio corresponding to the user preference is R2, and the priority weight determined by the variance σ² is y3 (0 ≤ y3 ≤ 1), then the target dynamic range compression ratio R = y3 * R1 + (1 - y3) * R2.
[0050] It can be concluded from the above that the beneficial effects of the non-contact audio processing method, system, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows: By identifying the external environment, the audio signal is adjusted to transmit high-quality audio signals to the user; the user's auditory experience is improved and the growing demand for high-quality audio of the user is met. Also, according to the user's preference, the target adjustment parameters are adjusted in real time. If the user feedbacks that the high frequency is too harsh, the system automatically reduces the high frequency gain and recalculates the adjustment parameters until the user is satisfied.
[0051] Corresponding to the non-contact audio processing method in the above embodiments, Figure 2 is a structural block diagram of a non-contact audio processing system provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 and the non-contact audio processing system 20 includes: a first calculation module 21, a second calculation module 22, and a third calculation module 23; Among them, the first calculation module 21 is used to preprocess the first audio signal and environmental data to obtain first data; The second calculation module 22 is used to input the first data into the neural network model to obtain the environmental type where the user is located; The third calculation module 23 is used to process the second audio signal according to the environmental type to obtain the target audio signal.
[0052] The present invention adjusts the audio signal by identifying the external environment and transmits high-quality audio signals to the user; it improves the user's auditory experience and meets the growing demand for high-quality audio of the user.
[0053] See Figure 3 and Figure 3Schematic block diagram of a wearable device provided by an embodiment of the present disclosure. As Figure 3 shown, the wearable device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above non-contact audio processing system embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.
[0054] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0055] The input device 302 may include a touchpad, a fingerprint sensor (for collecting the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0056] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0057] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of the non-contact audio processing method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the wearable device described in the embodiments of the present disclosure, which will not be elaborated here.
[0058] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiment are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0059] The computer-readable storage medium can be the internal storage unit of the wearable device in any of the foregoing embodiments, such as the hard disk or memory of the wearable device. The computer-readable storage medium can also be an external storage device of the wearable device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the wearable device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the wearable device. The computer-readable storage medium is used to store the computer program and other programs and data required by the wearable device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0060] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0061] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the wearable device and unit described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0062] In several embodiments provided in the present application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical or other forms of connection.
[0063] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0064] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0065] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present disclosure, and these modifications or substitutions should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A non-contact audio processing method, characterized in that: include: Preprocessing the first audio signal and the environmental data to obtain first data; Inputting the first data into the neural network model to obtain the type of environment in which the user is located; The second audio signal is processed according to the environment type to obtain a target audio signal.
2. The non-contact audio processing method according to claim 1, characterized in that: The processing of the second audio signal according to the environment type to obtain the target audio signal includes: determining a target adjustment parameter of the second audio signal based on the environment type; Determining an adjustment step size of a target adjustment parameter based on a frequency domain characteristic of the environment type; The target adjustment parameter of the second audio signal is adjusted based on the adjustment step size of the target adjustment parameter to obtain a target audio signal.
3. The non-contact audio processing method according to claim 2, characterized in that: The step of determining the adjustment step of the target adjustment parameter based on the frequency domain characteristics of the environment type includes: Selecting a corresponding target audio processing strategy from a plurality of audio processing strategies based on the frequency domain characteristics of the environment type; The process of determining the target audio output strategy includes: In response to the frequency domain characteristics of the environment type satisfying the low frequency characteristics, adjusting the target adjustment parameter of the second audio signal based on the first low frequency step length; In response to the frequency domain feature of the environment type satisfying the intermediate frequency feature, adjusting the target adjustment parameter of the second audio signal based on the first intermediate frequency step size; In response to the frequency domain characteristics of the environment type satisfying the high frequency characteristics, a target adjustment parameter of the second audio signal is adjusted based on a first high frequency step size.
4. The non-contact audio processing method according to claim 3, characterized in that: Also includes: Performing a fast Fourier transform on the first audio signal to obtain an audio frequency domain feature; Dividing the audio frequency domain features into a plurality of sub-frequency bands; the sub-frequency bands include a low-frequency sub-frequency band, a mid-frequency sub-frequency band and a high-frequency sub-frequency band; assigning a weight to each sub-frequency band based on the proportion of each sub-frequency band in the audio frequency domain feature; Perform weighted calculation on the weights assigned to each sub-band to obtain the frequency domain eigenvalue; The frequency domain characteristics of the environment type are determined based on the frequency domain characteristic values.
5. The non-contact audio processing method according to claim 4, characterized in that: Determining the frequency domain feature of the environment type based on the frequency domain feature value includes: In response to the frequency domain feature value being less than or equal to a first threshold, determining that the frequency domain feature of the environment type is a low-frequency feature; In response to the frequency domain feature value being greater than a first threshold value and the frequency domain feature value being less than a second threshold value, determining that the frequency domain feature of the environment type is a medium frequency feature; In response to the frequency domain feature value being greater than a second threshold, it is determined that the frequency domain feature of the environment type is a high frequency feature.
6. The non-contact audio processing method according to claim 5, characterized in that: The target adjustment parameter includes at least one of an equalizer gain, a dynamic range compression ratio, and a noise suppression threshold, When the frequency domain feature of the environment type is a low-frequency feature, the priority weight of the equalizer gain is the inverse of the low-frequency energy of the sub-band; When the frequency domain feature of the environment type is a medium frequency feature, the priority weight of the noise suppression threshold is proportional to the energy of the high frequency sub-band; When the frequency domain feature of the environment type is a high frequency feature, the priority weight of the dynamic range compression ratio is determined by the variance of the energy of the intermediate frequency sub-band.
7. The non-contact audio processing method according to claim 2, characterized in that: The determining the target adjustment parameter of the second audio signal based on the environment type comprises: Acquiring user preference data, wherein the user preference data includes equalization parameters or noise reduction strengths selected by the user in the past; The target adjustment parameter is obtained based on a weighted calculation of the environment type and the user preference data.
8. A non-contact audio processing system, characterized in that: include: A first calculation module, used for preprocessing the first audio signal and the environmental data to obtain first data; A second calculation module is used to input the first data into the neural network model to obtain the type of environment in which the user is located; The third calculation module is used to process the second audio signal according to the environment type to obtain a target audio signal.
9. A wearable device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.