Sound processing method and device, equipment and storage medium

By obtaining environmental recordings and determining the audio and video location, users are reminded to pay attention to the occurrence of sound, which solves the problem of users ignoring the attention of sound at home while focusing, and improves safety.

CN120238802APending Publication Date: 2025-07-01GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311873225.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When users wear headphones or concentrate on their work, they may miss the attention sounds that occur at home, such as the laughter and cry of children, the barking of pets or the call of their families, resulting in some accidents.

Method used

By obtaining environmental recordings and when detecting the sound of attention, it determines its audio-visual position in the sound field space, and outputs the recording to the corresponding device for playback to remind the user to pay attention to the occurrence of the sound.

Benefits of technology

Effectively remind users to pay attention to sounds in the environment and reduce accidents caused by ignoring sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238802A_ABST
    Figure CN120238802A_ABST
Patent Text Reader

Abstract

The invention provides a sound processing method and device, equipment and a storage medium. The method comprises the following steps: a first device obtains a first environment record; the first device determines a sound image position of the first environment recording in a sound field space under the condition that the first environment recording comprises a concerned sound and the first device is in communication connection with at least one second device; wherein the at least one second device supports playing of audio signals of at least two sound channels; and the first device outputs the first environment record to the at least one second device for playing according to the sound image position of the first environment record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to electronic technologies, including but not limited to sound processing methods, devices, equipment, and storage media. Background Art

[0002] When a user wears headphones or is concentrating on work, due to physical noise reduction of the headphones, algorithmic noise reduction, or the user's focus on something, the user may miss the crying or laughing of a child, the barking of a pet, or the call of a family member. Because the user fails to respond or react in time, some accidents may occur. Summary of the Invention

[0003] The sound processing method, device, equipment, and storage medium provided by this application can remind the user of the attention sounds occurring in the environment, thereby helping to avoid some accidents.

[0004] In a first aspect, an embodiment of this application provides a sound processing method. The method is applied to a first device and includes: obtaining a first environmental recording; determining the sound image position of the first environmental recording in the sound field space when the first environmental recording includes an attention sound and the first device is communicatively connected to at least one second device; where the at least one second device supports the playback of audio signals with at least two channels; and outputting the first environmental recording to the at least one second device for playback according to the sound image position of the first environmental recording.

[0005] In a second aspect, an embodiment of this application provides a sound processing method. The method is applied to a third device and includes: obtaining a first environmental recording; and sending the first environmental recording to a first device based on determining that the first environmental recording includes an attention sound.

[0006] In a third aspect, an embodiment of this application provides a sound processing device. The device is applied to a first device and includes: a first obtaining module configured to obtain a first environmental recording; a first determining module configured to determine the sound image position of the first environmental recording in the sound field space when the first environmental recording includes an attention sound and the first device is communicatively connected to at least one second device; where the at least one second device supports the playback of audio signals with at least two channels; and an output module configured to output the first environmental recording to the at least one second device for playback according to the sound image position of the first environmental recording.

[0007] In a fourth aspect, an embodiment of this application provides a sound processing device. The device is applied to a third device and includes: a second obtaining module configured to obtain a first environmental recording; and a first sending module configured to send the first environmental recording to a first device based on determining that the first environmental recording includes an attention sound.

[0008] In a fifth aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in the first aspect, or when the processor executes the program, it implements the method described in the second aspect.

[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method provided in the embodiment of the present application.

[0010] In an embodiment of the present application, a first device acquires a first ambient recording; when the first ambient recording includes a sound of concern, the first device outputs the first ambient recording at a corresponding sound image position, so as to achieve the purpose of reminding the user that a sound of concern has occurred in the environment, and thus is beneficial to avoiding the occurrence of some accidents.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings herein are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] The flowcharts shown in the drawings are only exemplary illustrations, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0014] Figure 1 It is a schematic diagram of a network architecture that may be applicable to an embodiment of the present application;

[0015] Figure 2 It is a schematic diagram of the implementation process of the sound processing method provided in an embodiment of the present application Figure 1 ;

[0016] Figure 3 It is a schematic diagram of the implementation process of the sound processing method provided in an embodiment of the present application Figure 2 ;

[0017] Figure 4 It is a schematic diagram of the implementation process of the sound processing method provided in an embodiment of the present application Figure 3 ;

[0018] Figure 5 It is a schematic implementation process of the sound processing method provided by the embodiments of the present application. Figure 4 ;

[0019] Figure 6 It is a simple floor plan of a family provided by the embodiments of the present application;

[0020] Figure 7 It is a schematic diagram of the attention sound recognition algorithm process provided by the embodiments of the present application;

[0021] Figure 8 It is a schematic structure of the sound processing device provided by the embodiments of the present application Figure 1 ;

[0022] Figure 9 It is a schematic structure of the sound processing device provided by the embodiments of the present application Figure 2 ;

[0023] Figure 10 It is a schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Specific Embodiments

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail in conjunction with the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not used to limit the scope of the present application.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0026] In the following descriptions, references to "some embodiments", "this embodiment", "embodiments of the present application", and examples, etc., describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0027] The descriptions such as "first, second, third", etc. that appear in the embodiments of the present application are only for the purpose of indicating and distinguishing the described objects, without any order, nor do they represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.

[0028] The network architecture and service scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. As is known to those of ordinary skill in the art, with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0029] Figure 1 Fig. shows a network architecture to which the embodiments of the present application may be applicable. As Figure 1 shown, the network architecture provided in this embodiment includes: a first device 101 and multiple third devices 102 - 10N distributed in different rooms or areas. Any third device has the ability to record sounds occurring in the space environment where it is located, and any third device can send the recorded environmental audio to the first device 101. Among them, the first device 101 can be a handheld device (such as a mobile phone, walkie - talkie), a wearable device (such as a smart bracelet, smart watch), etc. The third devices 102 - 10N can be the same or different. The third device can be a story machine, an early education machine, a monitoring device, a speaker device, a handheld device different from the first device, a wearable device different from the first device, etc. The sound processing method provided by the embodiments of the present application can be applied to smart home scenarios, security monitoring scenarios, etc., in any scenario where some accidents may occur because users do not respond or process specific sounds in the environment in a timely manner.

[0030] The embodiments of the present application provide a sound processing method. Figure 2 Schematic diagram of the implementation process of the sound processing method provided by the embodiments of the present application Figure 1 As Figure 2 shown, the method may include the following steps 201 to step 203:

[0031] Step 201, the first device acquires a first environmental audio recording.

[0032] Step 202, when the first environmental audio recording includes a sound of interest and the first device is communicatively connected to at least one second device, the first device determines the sound image position of the first environmental audio recording in the sound field space; wherein, the at least one second device supports the playback of audio signals with at least two channels.

[0033] Step 203, the first device outputs the first environmental audio recording to the at least one second device for playback according to the sound image position of the first environmental audio recording.

[0034] In an embodiment of the present application, a first device acquires a first ambient recording; when the first ambient recording includes a sound of interest, the first device outputs the first ambient recording at a corresponding sound image position, so as to achieve the purpose of reminding the user that a sound of interest has occurred in the environment, and thus is beneficial to avoiding the occurrence of some accidents.

[0035] The following further optional implementation manners of the above steps and related terms, etc. will be described respectively.

[0036] In step 201, the first device acquires a first ambient recording.

[0037] In some embodiments, the first device acquiring the first ambient recording includes: the first device recording the sound occurring in the environment. In other embodiments, the first device acquiring the first ambient recording includes: the first device receiving the first ambient recording sent by a third device.

[0038] In an embodiment of the present application, the first ambient recording may be the sound recorded by the first device or the sound recorded by the third device, and the present application does not limit this.

[0039] For an embodiment where the first ambient recording is sent by a third device to the first device, the first ambient recording may be the sound of interest identified and filtered out by the third device, and the third device sends it to the first device only when it determines that the recorded first ambient recording includes the sound of interest. The first ambient recording may also be any sound recorded by the third device, and the third device does not identify whether the first ambient recording includes the sound of interest.

[0040] In an embodiment of the present application, identifying whether the first ambient recording includes the sound of interest may be completed by the first device or by the third device. In the solution where the third device identifies whether the first ambient recording includes the sound of interest, the first ambient recording received by the first device is recorded by the third device, and the third device identifies whether it includes the sound of interest. The third device sends the first ambient recording to the first device only when it determines that the recorded first ambient recording includes the sound of interest. Based on this, the first device may identify the first ambient recording again to determine whether the first ambient recording includes the sound of interest; if so, the first device determines the sound image position of the first ambient recording in the sound field space and executes step 203; otherwise, the first device does not process the first ambient recording. Of course, in other embodiments, for the solution where the third device identifies whether the first ambient recording includes the sound of interest, after receiving the first ambient recording sent by the third device, the first device may also no longer identify whether the first ambient recording includes the sound of interest, but directly determine the sound image position of the first ambient recording in the sound field space and execute step 203.

[0041] In some other embodiments, the task of identifying whether the first environmental recording includes the sound of interest may be performed by the first device. Instead of identifying whether the first environmental recording includes the sound of interest, the third device sends the recorded audio data to the first device after recording the sound in the space environment where it is located, and the first device identifies the received audio data to determine whether it includes the sound of interest.

[0042] In some embodiments, the method further includes: the first device, in response to receiving the first environmental recording, identifies the sound type of the first environmental recording; and determines whether the first environmental recording is the sound of interest according to the sound type of the first environmental recording; if so, performs step 202; otherwise, does not process the first environmental recording.

[0043] Further, in some embodiments, the first device may identify the sound type of the first environmental recording in the following manner: use a trained first Artificial Intelligence (AI) model to identify the first environmental recording to obtain the sound type of the first environmental recording.

[0044] It should be noted that the recognition result of the first environmental recording output by the trained first AI model may be the sound type of the sound of interest or a specific value, and this specific value represents the sound type that does not include the sound of interest. That is, based on the output result of the first AI model, it can be determined whether the first environmental recording is the sound of interest.

[0045] The sound data (i.e., sample audio files) used to train or update the first AI model may be one or more types of sample audio files of the sound of interest provided by the user or developer, or one or more types of sample audio files of the sound of interest selected by the user from several fixed types of sounds provided by the device.

[0046] In a possible implementation manner, the first device may first perform data cleaning on the first environmental recording, such as file format alignment, data augmentation, and / or spectrogram conversion, and then extract features from the first environmental recording after data cleaning through algorithms such as Fbank or MFCC. Based on the trained first AI model, the features after extraction are identified to obtain the sound type of the first environmental recording; in this way, it is beneficial to improve the recognition accuracy of the first environmental recording, thereby reducing the probability of misoutput of the first environmental recording. Of course, in another possible implementation manner, the first device may also not perform data cleaning on the first environmental recording, but directly extract features from the first environmental recording through algorithms such as Fbank or MFCC.

[0047] In the embodiments of the present application, the trained first AI model may be a model imported at the time of the first device's factory shipment, or an AI model updated later via an update service (RomUpdateService, RUS) system, or may also be obtained by retraining the trained second AI model in the first device based on at least one sample audio file of at least one sound type input by the user. The trained first AI model is trained based on one or more audio files focusing on sounds. The first AI model may be a neural network model, a decision tree, or other models with machine learning capabilities.

[0048] In some embodiments, before using the trained first AI model to identify the first environmental recording, the method further includes: the first device receiving at least one sample audio file of at least one sound type input based on a third UI interface; and the first device training the trained second AI model based on the at least one sample audio file of the at least one sound type to obtain the trained first AI model. That is, the first device supports the user to customize the ability to identify and focus on sounds.

[0049] In the embodiments of the present application, the trained second AI model may be a model imported at the time of the first device's factory shipment, or an AI model updated later via an update service (RomUpdateService, RUS) system, or may also be a model obtained by training multiple times in the first device based on sample audio files input by the user.

[0050] In some embodiments, after the first device trains the trained second AI model based on at least one sample audio file of at least one sound type input based on a third UI interface to obtain the trained first AI model, the first device may further send the model parameters of the trained first AI model to the third device; wherein, the trained first AI model is used by the third device to identify the sound type of the recorded environmental recording; thus, enabling the model parameters of the AI model used by the first device to be synchronized with the third device, which is beneficial to improving the accuracy of the third device in identifying the focused sound. In this way, for the solution where the third device only sends the focused sound to the first device after identifying it, the improvement in the accuracy of identifying the focused sound is beneficial to reducing the number of times the third device sends non-focused sounds to the first device, and is also beneficial to reducing the number of times the first device outputs non-focused sounds.

[0051] In step 202, when the first environmental recording includes a focused sound and the first device is communicatively connected to at least one second device, the first device determines the sound image position of the first environmental recording in the sound field space.

[0052] In the embodiments of the present application, the first ambient recording may be audio data / audio signals for one or more unit times. For example, the first ambient recording may be audio data for one second or multiple seconds. For the embodiment where the first ambient recording is audio data for one unit time, the fact that the first ambient recording includes the sound of interest means that the first ambient recording is the audio data of the sound of interest. For the embodiment where the first ambient recording is audio data for multiple unit times, the fact that the first ambient recording includes the sound of interest may mean that the first ambient recording only includes the audio data of the sound of interest (the first device or the third device may extract the audio data of the sound of interest from the recorded original ambient recording as the first ambient recording), and the fact that the first ambient recording includes the sound of interest may also mean that the first ambient recording includes the audio data of the sound of interest and the audio data of non-sound-of-interest. The present application does not limit this.

[0053] In some embodiments, the at least one second device supports the playback of audio signals with at least two channels. For example, the second device is a Bluetooth headset, a wired headset, a Bluetooth speaker / speaker, or any other device that supports holographic audio or spatial audio. The so-called holographic audio or spatial audio means that after the audio signals of different sound sources are superimposed and played, they have different sound image positions in the sound field space in terms of listening perception.

[0054] For the determination of the sound image position of the first ambient recording, in some embodiments, the first device may obtain the sound image position of the first ambient recording in the sound field space by any one of the following methods (1)-(3):

[0055] (1) The first device determines the sound image position of the first ambient recording according to the sound type of the first ambient recording; wherein, different sound types correspond to different sound image positions, and the sound type is the sound type of the sound of interest.

[0056] Further, in some embodiments, the first device may determine the sound image position of the first ambient recording in the following way: determine the sound image position of the first ambient recording from the mapping table according to the sound type of the first ambient recording; wherein, the mapping table records the type identifiers of at least one sound type and the corresponding sound image positions.

[0057] In some embodiments, the method further includes: the first device receives configuration information, the configuration information includes specifying the sound image positions of at least one sound type; and the first device configures or updates the sound image positions of the corresponding sound types in the mapping table according to the configuration information.

[0058] In the embodiments of the present application, the mapping table may be a predefined default value, that is, configured in advance by developers at the time of factory shipment, or may be configured later based on the RUS system. Of course, the first device also supports updating the mapping table. In the embodiments of the present application, the method for the first device to receive configuration information is not limited. It may receive configuration information input by the user based on the UI interface of the first device, or receive configuration information remotely sent by the developer based on the communication module of the first device. In this way, the user or the developer can customize the sound image position of any sound type through the configuration information.

[0059] (2) The sound image position of the first environmental recording is equal to a preconfigured fixed sound image position; wherein, the sound image positions of the first environmental recordings of different sound types are the same.

[0060] In this embodiment, the preconfigured fixed sound image position may be pre-set when the first device leaves the factory, or may be a sound image position updated later based on the RUS system, or may also be configured based on the first UI interface.

[0061] (3) The first device determines the sound image position of the first environmental recording according to the positional relationship between the third device and the first device. In this way, when the first device outputs the first environmental recording based on the sound image position of the first environmental recording, the user can distinguish the position where the first environmental recording / attention sound occurs in terms of the sense of hearing.

[0062] For example, the user of the first device is working in the living room. At this time, an attention sound (such as the beeping sound of a pot lid) occurs in the kitchen diagonally in front of the user's right. Then the first device can set the sound image position of the first environmental recording transmitted from the third device in the kitchen to the diagonally front right of the user's head, so that when the first environmental recording is output based on this sound image position, the user can hear that the sound occurs in the front right.

[0063] It can be understood that the above embodiments describe that in the case where the first device is connected to the second device, the sound image position of the first environmental recording can be determined first, and then the first environmental recording is output based on this sound image position. For the scenario where the first device is not connected to the second device, the first device can achieve the purpose of reminding the user to pay attention to the sound by playing it externally.

[0064] That is, in some other embodiments, the method further includes: when the first environmental recording includes an attention sound and the first device is not communicatively connected to at least one second device, outputting a reminder message and / or the first environmental recording; wherein, the at least one second device supports playing of audio signals with at least two channels. For example, the second device is a Bluetooth headset, a wired headset, a Bluetooth speaker / speaker, or any other device that supports holographic audio or spatial audio

[0065] Exemplarily, in some embodiments, when the first ambient recording includes the sound of interest, outputting the first ambient recording includes: playing the first ambient recording through a speaker or a specified fourth device; wherein the speaker is disposed in the first device.

[0066] In the embodiments of the present application, there is no limitation on the content of the reminder information, that is, the reminder method of the first device. In short, it is sufficient to be able to remind of the occurrence of the sound of interest.

[0067] Exemplarily, in some embodiments, the first device outputs reminder information, including any one of the following methods (4)-(6):

[0068] (4) The first device plays a specific ringtone through a speaker or a specified fourth device;

[0069] (5) The first device plays a pre-configured reminder voice through the speaker or the specified fourth device; for example, the reminder voice includes the location where the sound of interest occurs.

[0070] (6) The first device plays the first ambient recording after sound optimization processing through the speaker or the specified fourth device.

[0071] Exemplarily, in some embodiments, the sound optimization processing may include sound amplification processing via a sound upward compressor, human voice highlighting processing, sound reduction processing via a sound downward compressor, sound effect processing, and / or noise reduction processing, etc.

[0072] In the embodiments of the present application, the specified fourth device may be one device or multiple devices. For example, the specified device may be the mobile phone of the family member of the user of the first device, or the user's tablet computer, personal computer, etc.

[0073] In step 203, the first device outputs the first ambient recording to the at least one second device for playing according to the sound image position of the first ambient recording.

[0074] In some embodiments, the first device may implement step 203 as follows: performing rendering processing on the first ambient recording according to the sound image position of the first ambient recording to obtain a second ambient recording; wherein the second ambient recording includes at least two-channel audio signals; and outputting the second ambient recording to the at least one second device for playing.

[0075] Further, in some embodiments, before the first device outputs the second environmental recording, the method further includes: the first device determines the sound image positions of the first audio signals created by at least one application in the sound field space respectively; wherein, the sound image positions of different first audio signals are different, and the sound image positions of the first audio signals are different from those of the first environmental recording; and the first device performs rendering processing on the first audio signals according to the sound image positions of the first audio signals to obtain second audio signals;

[0076] The outputting the second environmental recording includes: mixing the second environmental recording and the second audio signals corresponding to the first audio signals created by the at least one application into an audio signal with at least two channels and then outputting the audio signal to the at least one second device for playing.

[0077] It can be understood that the first environmental recording is essentially an audio signal. Therefore, whether the first environmental recording is rendered according to its sound image position or the first audio signals are rendered according to their sound image positions, the processing methods are the same. In a possible implementation manner, the first device can input the audio signal to be rendered and the corresponding sound image positions into a spatial audio rendering algorithm, so that the spatial audio rendering algorithm performs spatial processing on the corresponding audio signals according to the sound image positions and outputs the audio signals with the sound image positions. For example, the spatial audio rendering algorithm selects the corresponding head-related transfer function (HRTF) according to the input sound image positions, and renders the corresponding audio signals through the HRTF function. Among them, the parameters of the HRTF functions corresponding to different sound image positions are different.

[0078] For the solution that the first device mixes the second environmental recording and the second audio signals corresponding to the first audio signals created by the at least one application into an audio signal with at least two channels and then outputs the audio signal to the at least one second device for playing, in a possible implementation manner, the first device can superimpose (such as linear superposition or non-linear superposition) the left-channel data of the second environmental recording and the left-channel data of the second audio signals corresponding to the first audio signals created by the at least one application to obtain left-channel data, and superimpose (such as linear superposition or non-linear superposition) the right-channel data of the second environmental recording and the right-channel data of the second audio signals corresponding to the first audio signals created by the at least one application to obtain right-channel data.

[0079] It should be noted that in the embodiments of the present application, there are no restrictions on the conditions under which the first device outputs the reminder information or the first environmental recording, or outputs the first environmental recording according to the sound image position of the first environmental recording.

[0080] In a possible implementation, when the first environmental recording includes the sound of interest and the first device is communicatively connected to at least one second device, the first device defaults to output the first environmental recording according to the sound image position of the first environmental recording; when the first environmental recording includes the sound of interest and the first device is not communicatively connected to at least one second device, the first device defaults to output a reminder message and / or the first environmental recording.

[0081] In another possible implementation, the first device can also support the user to independently and freely specify the way to remind the user of the sound of interest that occurs. Optionally, in some embodiments, the method further includes: the first device displays a second UI interface; for example, when the first environmental recording includes the sound of interest, the first device displays a second UI interface.

[0082] Wherein, the second UI interface is used to specify the reminder method of the first environmental recording; the reminder method includes any one of the following methods (7)-(9):

[0083] (7) Output the first environmental recording according to the sound image position of the first environmental recording;

[0084] (8) Output a reminder message;

[0085] (9) Output the first environmental recording.

[0086] Exemplarily, in some embodiments, the second UI interface may include a first option, a second option, and a third option; wherein, the first option is used to specify to remind the first environmental recording in the manner of (7) above; the second option is used to specify to remind the first environmental recording in the manner of (8) above; the third option is used to specify to remind the first environmental recording in the manner of (9) above. When the first environmental recording includes the sound of interest and the first device is communicatively connected to at least one second device, the first device can also combine the selected reminder method in the second UI interface to determine which reminder method to use to remind the first environmental recording. For example, when the first environmental recording includes the sound of interest, the first device is communicatively connected to at least one second device, and the first option is selected, then the first environmental recording is reminded in the manner of (7) above. Another example is that when the first environmental recording includes the sound of interest, the first device is communicatively connected to at least one second device, and the second option is selected, then the first environmental recording is reminded in the manner of (8) above. Another example is that when the first environmental recording includes the sound of interest, the first device is communicatively connected to at least one second device, and the third option is selected, then the first environmental recording is reminded in the manner of (9) above.

[0087] In the case where the first environmental recording includes a voice of concern, the first device is not communicatively connected to at least one second device, and the first option is selected, a reminder of the first environmental recording is made in the manner of (8) or (9) above. In the case where the first environmental recording includes a voice of concern, the first device is not communicatively connected to at least one second device, and the second option is selected, a reminder of the first environmental recording is made in the manner of (8) above. In the case where the first environmental recording includes a voice of concern, the first device is not communicatively connected to at least one second device, and the third option is selected, a reminder of the first environmental recording is made in the manner of (9) above.

[0088] Another embodiment of the present application provides a sound processing method. Figure 3 It is a schematic implementation flow of the sound processing method provided by the embodiment of the present application. Figure 2 , as Figure 3 shown, the method includes the following steps 301 to step 302:

[0089] Step 301, the third device obtains the first environmental recording;

[0090] Step 302, the third device sends the first environmental recording to the first device based on determining that the first environmental recording includes a voice of concern.

[0091] In the embodiment of the present application, the type of the third device is not limited, and it can be various types of devices with environmental sound recording functions and communication functions with other devices. For example, the third device can be a monitoring device, a television, a projector, an oil fume machine, a story machine, a Bluetooth speaker, an early education machine, a mobile phone, or a projector, etc.

[0092] In some embodiments, the method further includes: the third device identifies the sound type of the first environmental recording; and the third device determines whether the first environmental recording is the voice of concern according to the sound type of the first environmental recording.

[0093] Further, in some embodiments, the third device can identify the sound type of the first environmental recording in the following way: use the trained first AI model to identify the first environmental recording to obtain the sound type of the first environmental recording.

[0094] It should be noted that the recognition result of the first environmental recording output by the trained first AI model may be the sound type of the voice of concern or a specific value, and the specific value represents the sound type that does not include the voice of concern. That is, based on the output result of the first AI model, it can be determined whether the first environmental recording is the voice of concern.

[0095] The voice data (i.e., sample audio files) used to train or update the first AI model can be one or more types of sample audio files of the concerned voices provided by the user or developer, or one or more types of sample audio files of the concerned voices selected by the user from several fixed types of voices provided by the device.

[0096] In some embodiments, the method further includes: before identifying the first environmental recording by using the trained first AI model, receiving the model parameters of the trained first AI model sent by the first device.

[0097] How the first device obtains the trained first AI model has been described above, so it will not be elaborated here.

[0098] Another embodiment of the present application provides a voice processing method. Figure 4 It is a schematic implementation process of the voice processing method provided by the embodiment of the present application. Figure 3 As Figure 4 shown, the method includes the following steps 401 to 409:

[0099] Step 401, the third device acquires a first environmental recording.

[0100] Step 402, the third device determines whether the first environmental recording includes a concerned voice; if so, execute Step 403; otherwise, return to execute Step 401.

[0101] Step 403, the third device sends the first environmental recording to the first device.

[0102] In the embodiment of the present application, the first environmental recording can be audio data / audio signals of one or more unit times. For example, the first environmental recording can be audio data of one second or multiple seconds. For the embodiment where the first environmental recording is audio data of one unit time, that the first environmental recording includes a concerned voice means that the first environmental recording is audio data of the concerned voice. For the embodiment where the first environmental recording is audio data of multiple unit times, that the first environmental recording includes a concerned voice can mean that the first environmental recording only includes audio data of the concerned voice (the third device can intercept the audio data of the concerned voice in the recorded original environmental recording as the first environmental recording), and that the first environmental recording includes a concerned voice can also mean that the first environmental recording includes audio data of the concerned voice and non-concerned voices. The present application does not limit this.

[0103] Step 404, the first device receives the first environmental recording sent by the third device; wherein, the first environmental recording includes a concerned voice.

[0104] Step 405, the first device determines the sound image position of the first environmental recording in the sound field space;

[0105] Step 406, the first device performs rendering processing on the first environmental recording according to the sound image position of the first environmental recording to obtain a second environmental recording;

[0106] Step 407, the first device determines the sound image positions of the first audio signals created by at least one application in the sound field space respectively; wherein, the sound image positions of different first audio signals are different, and the sound image positions of the first audio signals are different from those of the first environmental recording;

[0107] Step 408, the first device performs rendering processing on the first audio signal according to the sound image position of the first audio signal to obtain a second audio signal.

[0108] It should be noted that in the embodiments of the present application, the execution order of steps 405 and 407 is not limited. The first device can execute steps 405 and 407 in parallel, or execute step 405 first and then step 407, or execute step 407 first and then step 405. Nor is the execution order of steps 406 and 408 limited. The first device can execute steps 406 and 408 in parallel, or execute step 406 first and then step 408, or execute step 408 first and then step 406. Nor is the execution order of steps 406 and 407 limited.

[0109] Step 409, the first device mixes and outputs the second environmental recording and the second audio signals corresponding to the first audio signals created by at least one application.

[0110] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0111] The sound processing method provided by the embodiments of the present application is applied to distributed multi-mobile terminals. Utilizing the listening capabilities of multi-terminal devices (i.e., an example of the third device) in a house, the sound in the environment is monitored at all times, and the user's mobile phone (the main device, also an example of the first device) is notified of the identified specified sound. When the user turns on holographic audio on the Bluetooth headset (i.e., an example of the second device), the main device renders a fixed position (i.e., the sound image position) in the sound field space for real-time playback of the sound monitored by the distributed multi-terminal devices and performs certain rendering processing.

[0112] For example, a camera in the kitchen (i.e., an example of the third device) monitors the boiling sound of the kettle and the beeping sound of the kettle lid. After identifying the sound of the kettle boiling in the environment through a sound recognition algorithm, it sends the environmental recording to the main device. If the main device has the children's voice following function enabled, it will play the environmental recording uploaded by the distributed terminal device. If the user turns on the holography and connects the headphones at the same time, the user can hear the rendered real-time environmental recording at a fixed position.

[0113] Figure 5 It is a schematic implementation process of the sound processing method provided by the embodiments of the present application Figure 4 , such as Figure 5 shown, the method includes the following steps 501 to step 507:

[0114] Step 501, the distributed terminal monitors the environmental sound; wherein, the distributed terminal is an example of the third device;

[0115] Step 502, the distributed terminal identifies whether the monitored environmental sound (i.e., the first environmental recording) is the concerned sound; if so, execute step 503; otherwise, return to execute step 502;

[0116] Step 503, the distributed terminal uploads the first environmental recording to the main device; wherein, the main device is an example of the first device;

[0117] Step 504, the main device receives the first environmental recording and determines whether it is in the external playback state; if so, execute step 505; otherwise, execute step 506;

[0118] Step 505, the main device gives an external reminder (or plays the first environmental recording);

[0119] Step 506, the main device determines the sound image position of the first environmental recording;

[0120] Step 507, the main device renders and plays the first environmental recording based on the sound image position of the first environmental recording.

[0121] (1) Distributed smart home monitoring

[0122] Figure 6 It is a simplified floor plan of a family provided by the embodiments of the present application, such as Figure 6As shown in the figure, the distributed terminal devices can be the story machine 601 and early education machine 602 in the baby room, the camera 603 in the kitchen, the Xiaobu Bluetooth speaker 604 in the living room, as well as the user's mobile phone 605 and the family member's mobile phone 606. Any mobile phone can be used as the master device. For example, the mobile phone 605 is used as the master device, and there is only one master device. The master device is connected to the earphone 607 (i.e., an example of the second device). The distributed terminal has the capabilities of recording and voice recognition algorithm, and can recognize the concerned sounds occurring in the scene (such as the crying and laughing of children, the barking of pets, and the calls of family members, etc.). The voice recognition algorithm is described in step (3).

[0123] (2) Holographic monitoring position setting

[0124] The holographic audio algorithm sets a fixed spatial position / sound image position for the concerned sound (this position can be specified by the user). When the distributed terminal device transmits the environmental recording, if the master device is not connected to the earphone, it will play the recording externally; if it is already connected to the earphone and the holographic switch is turned on, the environmental recording will be rendered and played at the specified sound image position. If other external devices are specified, the environmental recording will also be synchronously forwarded to other devices.

[0125] (3) Concerned sound recognition algorithm

[0126] There are two types of concerned sounds. One is the default mode, where the system provides several predefined sounds (such as the crying and laughing of children, the barking of pets, and the calls of family members, etc., which can be customized according to product needs), and there is already a trained neural network model (i.e., an example of the first AI model). After the user selects, there is no need to wait and it can be used directly. The other is the custom mode, which means that the user can record a piece of audio that they want to pay attention to. It can be recorded using the distributed terminal or the master device. Then, on the master device, a tuned model for the user's concerned sound is obtained through training, and then sent to other devices. The user can record multiple recordings, but they need to wait for the time to train the model.

[0127] The concerned sound recognition algorithm process is as Figure 7 shown, including the following steps 701 to step 704:

[0128] Step 701, obtain the recording input by the user;

[0129] Step 702, extract the features of the recording;

[0130] Step 703, input the extracted features into the network model;

[0131] Step 704, output the sound type of the recording based on the network model.

[0132] In a possible implementation, after the user uploads the concerned voice file, feature extraction will be performed first. First is data cleaning, such as file format alignment, data augmentation, and spectrogram conversion, and then features are extracted through algorithms such as Fbank or MFCC. The extracted features are fine-tuned by a pre-trained model to obtain an adjusted model.

[0133] The default mode will have pre-trained models for several specific voices built-in. The built-in model of the custom mode cannot be used directly, but it is also pre-trained and can be used after secondary training after the user uploads a recording.

[0134] Based on this technical solution, when the concerned voice occurs near the distributed terminal in a certain environment, the mobile phone system can give a reminder on the mobile phone or perform real-time recording rendering at a specific spatial position of the earphone.

[0135] In the embodiments of this application, the training of the neural network model can be performed in the cloud or on the user's mobile phone. The model for identifying environmental recordings as concerned voices can be placed on the distributed terminal, on any device with relatively strong computing power in the home, or on the main device.

[0136] In the embodiments of this application, the distributed terminal can communicate with the main device in ways such as WiFi, local area network, and Bluetooth.

[0137] In the embodiments of this application, the sound image position of the environmental recording can be fixed, the position of the distributed terminal relative to the main device, the relative position of the sound source in the environmental recording to the distributed terminal, or different concerned voices can be placed at different sound image positions.

[0138] In the embodiments of this application, after receiving the environmental recording, the main device can perform secondary voice recognition confirmation or not. The reminder for the user can be a specific ringtone, real-time recording playback, voice broadcast, and recording playback after certain voice processing (such as highlighting the human voice, noise reduction processing, etc.).

[0139] In the embodiments of this application, in the custom mode, the user can upload one file for one voice or multiple files, and can create multiple voices they are concerned about. The model will be trained once after each upload.

[0140] It should be noted that although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution; or, steps in different embodiments may be combined into a new technical solution.

[0141] Based on the foregoing embodiments, an embodiment of the present application provides a sound processing device. The device includes each module included and each unit included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be an AI acceleration engine (such as an NPU, etc.), a GPU, a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0142] Figure 8 Structural schematic of the sound processing device provided by the embodiment of the present application Figure 1 , the sound processing device 80 is applied to a first device, such as Figure 8 As shown, the sound processing device 80 includes:

[0143] A first acquisition module 801, configured to acquire a first environmental recording;

[0144] A first determination module 802, configured to determine the sound image position of the first environmental recording in the sound field space when the first environmental recording includes a sound of interest and the first device is communicatively connected to at least one second device; wherein, the at least one second device supports the playback of audio signals with at least two channels;

[0145] An output module 803, configured to output the first environmental recording to the at least one second device for playback according to the sound image position of the first environmental recording.

[0146] In some embodiments, the outputting the first environmental recording to the at least one second device for playback according to the sound image position of the first environmental recording includes: rendering the first environmental recording according to the sound image position of the first environmental recording to obtain a second environmental recording; wherein, the second environmental recording includes audio signals with at least two channels; and outputting the second environmental recording to the at least one second device for playback.

[0147] In some embodiments, the first determination module 803 is further configured to: before outputting the second environmental recording, determine the sound image positions of the first audio signals created by at least one application in the sound field space respectively; wherein, the sound image positions of different first audio signals are different, and the sound image positions of the first audio signals are different from those of the first environmental recording; and perform rendering processing on the first audio signals according to the sound image positions of the first audio signals to obtain second audio signals; the output module 803 is configured to mix the second environmental recording and the second audio signals corresponding to the first audio signals created by at least one application into an audio signal with at least two channels and then output the audio signal to the at least one second device for playing.

[0148] In some embodiments, determining the sound image position of the first environmental recording in the sound field space includes: determining the sound image position of the first environmental recording according to the sound type of the first environmental recording; wherein, the sound image positions corresponding to different sound types are different, and the sound type is the sound type of the concerned sound; or, the sound image position of the first environmental recording is equal to a pre-configured fixed sound image position; or, determining the sound image position of the first environmental recording according to the positional relationship between the third device and the first device.

[0149] Further, in some embodiments, determining the sound image position of the first environmental recording according to the sound type of the first environmental recording includes: determining the sound image position of the first environmental recording from a mapping table according to the sound type of the first environmental recording; wherein, the mapping table records the type identifiers of at least one sound type and the corresponding sound image positions.

[0150] In some embodiments, the sound processing device 80 further includes a configuration module; wherein, the first acquisition module 801 is configured to receive configuration information, and the configuration information includes specifying the sound image positions of at least one sound type; the configuration module is configured to configure or update the sound image positions of the corresponding sound types in the mapping table according to the configuration information.

[0151] In some embodiments, the pre-configured fixed sound image position is configured based on a first UI interface.

[0152] In some embodiments, the output module 803 is further configured to output a reminder message and / or the first environmental recording when the first environmental recording includes a concerned sound and the first device is not communicatively connected to at least one second device.

[0153] Further, in some embodiments, the output reminder information includes: playing a specific ringtone through a speaker or a specified fourth device; wherein the speaker is disposed in the first device; or playing a pre-configured reminder voice through the speaker or the specified fourth device; or playing the recorded sound of the first environment after sound optimization processing through the speaker or the specified fourth device.

[0154] Further, in some embodiments, outputting the recorded sound of the first environment includes: playing the recorded sound of the first environment through a speaker or a specified fourth device; wherein the speaker is disposed in the first device.

[0155] Further, in some other embodiments, the output reminder information and / or the recorded sound of the first environment includes: based on determining that the first device is not communicatively connected to at least one second device, when the recorded sound of the first environment includes a sound of concern, outputting the reminder information and / or the recorded sound of the first environment; wherein the at least one second device supports playing of at least two-channel audio signals.

[0156] In some embodiments, the sound processing device 80 further includes a display module configured to: when the recorded sound of the first environment includes a sound of concern, display a second UI interface; wherein the second UI interface is used to specify a reminder method for the recorded sound of the first environment; the reminder method includes any one of the following methods: outputting the recorded sound of the first environment according to the sound image position of the recorded sound of the first environment; outputting the reminder information; outputting the recorded sound of the first environment.

[0157] In some embodiments, the recorded sound of the first environment is recorded by the third device, and the third device identifies whether it includes the sound of concern.

[0158] In some embodiments, the first determination module is further configured to: in response to receiving the recorded sound of the first environment, identify the sound type of the recorded sound of the first environment; and determine whether the recorded sound of the first environment is the sound of concern according to the sound type of the recorded sound of the first environment.

[0159] In some embodiments, the identifying the sound type of the recorded sound of the first environment includes: using a trained first AI model to identify the recorded sound of the first environment to obtain the sound type of the recorded sound of the first environment.

[0160] In some embodiments, the sound processing device 80 further includes a training module; wherein, the first acquisition module 801 is further configured to receive at least one sample audio file of at least one sound type input based on a third UI interface before identifying the first environmental recording by using the trained first AI model; the training module is configured to train the trained second AI model based on at least one sample audio file of at least one sound type to obtain the trained first AI model.

[0161] In some embodiments, the sound processing device 80 further includes a second sending module, configured to send the model parameters of the trained first AI model to the third device; wherein, the trained first AI model is used by the third device to identify the sound type of the recorded environmental recording.

[0162] In some embodiments, the first device is a mobile phone, and the third device is a different device from the first device.

[0163] Another embodiment of the present application provides a sound processing device. Figure 9 For the structural schematic of the sound processing device provided by the embodiment of the present application Figure 2 , this sound processing device 90 is applied to a third device. As Figure 9 shown, this sound processing device 90 includes:

[0164] A second acquisition module 901, configured to acquire a first environmental recording;

[0165] A first sending module 902, configured to send the first environmental recording to a first device based on determining that the first environmental recording includes a sound of concern.

[0166] In some embodiments, the sound processing device 90 further includes an identification module, configured to identify the sound type of the first environmental recording; and determine whether the first environmental recording is the sound of concern according to the sound type of the first environmental recording.

[0167] In some embodiments, the identifying the sound type of the first environmental recording includes: identifying the first environmental recording by using the trained first AI model to obtain the sound type of the first environmental recording.

[0168] In some embodiments, the second acquisition module 901 is further configured to receive the model parameters of the trained first AI model sent by the first device before identifying the first environmental recording by using the trained first AI model.

[0169] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0170] It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, each functional unit in the various embodiments of the present application may be integrated in a processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware, or may be implemented in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.

[0171] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0172] The embodiments of the present application provide an electronic device, Figure 10 which is a schematic structural diagram of the electronic device provided by the embodiments of the present application. As Figure 10 shown, the electronic device 100 includes a memory 1001 and a processor 1002. The memory 1001 stores a computer program that can run on the processor 1002. When the processor 1002 executes the program, it implements the steps in the sound processing method implemented by the first device provided in the above embodiments, or when the processor 1002 executes the program, it implements the steps in the sound processing method implemented by the third device provided in the above embodiments. That is, the electronic device may be the first device or the third device.

[0173] It should be noted that the memory 1001 is configured to store instructions and applications executable by the processor 1002, and can also cache data to be processed or already processed by each module in the processor 1002 and the electronic device 100 (such as, image data, audio data, voice communication data, and video communication data), and can be implemented by a flash memory (FLASH) or a random access memory (RAM).

[0174] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.

[0175] An embodiment of the present application provides a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute the steps in the method provided in the above method embodiment.

[0176] It should be pointed out here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0177] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" or "in some embodiments" that appear throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the order numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0178] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone. These three situations.

[0179] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0180] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0181] The modules described above as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules; they can be located in one place or distributed to multiple network units; some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0182] In addition, in each embodiment of this application, the various functional modules can all be integrated in a processing unit, or each module can be separately used as a unit, or two or more modules can be integrated in a unit; the above-mentioned integrated modules can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0183] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.

[0184] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as removable storage devices, ROMs, magnetic disks, or optical discs that can store program codes.

[0185] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0186] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0187] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0188] The above is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A sound processing method, characterized in that, The method is applied to a first device, and the method includes: Obtain a first ambient recording; When the first ambient recording includes a sound of interest and the first device is communicatively connected to at least one second device, determine the sound image position of the first ambient recording in the sound field space; wherein, the at least one second device supports the playback of audio signals with at least two channels; Output the first ambient recording to the at least one second device for playback according to the sound image position of the first ambient recording.

2. The method according to claim 1, characterized in that, The outputting the first ambient recording to the at least one second device for playback according to the sound image position of the first ambient recording includes: Perform rendering processing on the first ambient recording according to the sound image position of the first ambient recording to obtain a second ambient recording; wherein, the second ambient recording includes audio signals with at least two channels; Output the second ambient recording to the at least one second device for playback.

3. The method according to claim 2, characterized in that, Before outputting the second ambient recording, the method further includes: Determine the sound image positions of first audio signals created by at least one application in the sound field space respectively; wherein, the sound image positions of different first audio signals are different, and the sound image positions of the first audio signals are different from those of the first ambient recording; Perform rendering processing on the first audio signals according to the sound image positions of the first audio signals to obtain second audio signals; The outputting the second ambient recording includes: Mix the second ambient recording and the second audio signals corresponding to the first audio signals created by the at least one application into audio signals with at least two channels and then output them to the at least one second device for playback.

4. The method according to claim 1, characterized in that, The obtaining the first ambient recording includes: Receive the first ambient recording sent by a third device.

5. The method according to any one of claims 1-4, characterized in that, The determining the sound image position of the first ambient recording in the sound field space includes: Determine the sound image position of the first ambient recording according to the sound type of the first ambient recording; wherein, different sound types correspond to different sound image positions, and the sound type is the sound type of the sound of interest; or, The sound image position of the first ambient recording is equal to a pre-configured fixed sound image position; or, Determine the sound image position of the first ambient recording according to the positional relationship between the third device and the first device.

6. The method according to claim 5, wherein The determining the sound image position of the first ambient recording according to the sound type of the first ambient recording includes: Determine the sound image position of the first ambient recording from a mapping table according to the sound type of the first ambient recording; wherein, the mapping table records the type identifiers of at least one sound type and the corresponding sound image positions.

7. The method according to claim 6, wherein The method further includes: Receive configuration information, where the configuration information includes specifying the sound image positions of at least one sound type; Configure or update the sound image positions of the corresponding sound types in the mapping table according to the configuration information.

8. The method according to claim 5, wherein The pre-configured fixed sound image position is configured based on a first UI interface.

9. The method according to claim 1 or 4, characterized in that, The method further includes: When the first ambient recording includes a sound of interest and the first device is not communicatively connected to at least one second device, output a reminder message and / or the first ambient recording.

10. The method according to claim 9, characterized in that, The output reminder information includes: Playing a specific ringtone through a speaker or a specified fourth device; wherein, the speaker is provided in the first device; or, Playing a pre-configured reminder voice through the speaker or the specified fourth device; or, Playing the recorded audio of the first environment after sound optimization processing through the speaker or the specified fourth device.

11. The method according to claim 9, wherein Outputting the recorded audio of the first environment includes: Playing the recorded audio of the first environment through a speaker or a specified fourth device; wherein, the speaker is provided in the first device.

12. The method according to any one of claims 9-11, characterized in that, The method further includes: Displaying a second UI interface; wherein, the second UI interface is used to specify the reminder method of the recorded audio of the first environment; the reminder method includes any one of the following methods: Outputting the recorded audio of the first environment according to the sound image position of the recorded audio of the first environment; Outputting the reminder information; Outputting the recorded audio of the first environment.

13. The method according to any one of claims 4-12, wherein The recorded audio of the first environment is recorded by the third device, and the third device identifies whether it includes the concerned sound.

14. The method according to claim 4 or 13, characterized in that The method further includes: Responding to the received recorded audio of the first environment, identifying the sound type of the recorded audio of the first environment; Determining whether the recorded audio of the first environment is the concerned sound according to the sound type of the recorded audio of the first environment.

15. The method according to claim 14, wherein The identifying the sound type of the recorded audio of the first environment includes: Identifying the recorded audio of the first environment by using a trained first AI model to obtain the sound type of the recorded audio of the first environment.

16. The method according to claim 15, characterized in that, Before identifying the recorded audio of the first environment by using a trained first AI model, the method further includes: Receiving at least one sample audio file of at least one sound type input based on a third UI interface; Training a trained second AI model based on the at least one sample audio file of the at least one sound type to obtain the trained first AI model.

17. The method according to claim 16, characterized in that, The method further includes: Sending the model parameters of the trained first AI model to the third device; wherein, the trained first AI model is used by the third device to identify the sound type of the recorded environment audio.

18. The method according to any one of claims 4-17, characterized in that, The first device is a mobile phone, and the third device and the first device are two different devices.

19. A method for sound processing, characterized in that, The method is applied to the third device, and the method includes: Obtaining the recorded audio of the first environment; Based on determining that the recorded audio of the first environment includes the concerned sound, sending the recorded audio of the first environment to the first device.

20. The method according to claim 19, wherein The method further includes: Identifying the sound type of the recorded audio of the first environment; Determining whether the recorded audio of the first environment is the concerned sound according to the sound type of the recorded audio of the first environment.

21. The method according to claim 20, characterized in that, The identifying the sound type of the recorded audio of the first environment includes: Identifying the recorded audio of the first environment by using a trained first AI model to obtain the sound type of the recorded audio of the first environment.

22. The method according to claim 21, wherein Before identifying the recorded audio of the first environment by using a trained first AI model, the method further includes: Receiving the model parameters of the trained first AI model sent by the first device.

23. A sound processing device, characterized in that, The device is applied to a first device, and the device includes: A first acquisition module configured to acquire a first environmental recording; A first determination module configured to determine the sound image position of the first environmental recording in the sound field space when the first environmental recording includes a sound of interest and the first device is communicatively connected to at least one second device; wherein the at least one second device supports the playback of audio signals with at least two channels; An output module configured to output the first environmental recording to the at least one second device for playback according to the sound image position of the first environmental recording.

24. A sound processing device, characterized in that, The device is applied to a third device, and the device includes: A second acquisition module configured to acquire a first environmental recording; A first transmission module configured to transmit the first environmental recording to the first device based on determining that the first environmental recording includes a sound of interest.

25. An electronic device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 18, or when the processor executes the program, it implements the method according to any one of claims 19 to 22.

26. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 18, or when the computer program is executed by a processor, it implements the method according to any one of claims 19 to 22.