Audio processing method and device, electronic equipment and computer readable storage medium
By setting speakers and screen sounding units on non-adjacent side borders of the smart terminal, combined with sound-scene reconstruction and channel distribution technology, the problems of low sound image clarity and poor picture matching caused by traditional speakers are solved, and clearer audio playback and better user experience are achieved.
Patent Information
- Application Number
- CN202410093839.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-08-01
AI Technical Summary
Traditional speakers are set on the side frame of the smart terminal, resulting in low clarity in the user's viewing direction, which is difficult to match the screen effect, affecting the user's viewing experience.
Speakers are set on two non-adjacent side frames of the electronic device, and combined with the screen sounding unit, through sound and scene reconstruction, channel allocation and target sound signal extraction, the orientation enhancement of the target sound signal is achieved, and the clarity of the sound image in the screen direction and matching the picture effect is improved.
It improves the spatial effect and picture consistency of audio playback, and improves the user's viewing experience and immersion.
Smart Images

Figure CN120416735A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to an audio processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, for application scenarios such as playing videos, music, making calls, navigation, and games on intelligent terminals, the traditional speakers on the intelligent terminals are still used to play sounds. Since the traditional speakers are usually arranged on the side frames of the intelligent terminals, the orientation of the speakers results in a lower clarity of the sound image in the visual direction where the user is viewing, making it difficult to match the picture effect and affecting the user's viewing experience. Summary of the Invention
[0003] This application provides an audio processing method, apparatus, electronic device, and computer-readable storage medium, which can solve the problem that the clarity of the sound image in the screen direction is relatively low, it is difficult to match the picture effect, and thus affect the user's viewing experience.
[0004] To achieve the above object, this application adopts the following technical solutions:
[0005] In a first aspect, an audio processing method is provided, which is applied to an electronic device. The electronic device is provided with a screen, a screen sound generating unit, a first speaker, and a second speaker. The first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device. The method may include:
[0006] The electronic device determines the sound scene space parameters corresponding to the target sound signal in the audio information based on the picture information and audio information of the video to be played; performs channel allocation on the audio information to obtain a three-channel signal, where the three-channel signal includes a first left-channel signal, a first right-channel signal, and a first center-channel signal; extracts the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal; controls the first speaker to play the second left-channel signal, the second speaker to play the second right-channel signal, and the screen sound generating unit to play the second center-channel signal according to the sound scene space parameters; where the second left-channel signal is the background sound signal after removing the target sound signal from the first left-channel signal, the second right-channel signal is the background sound signal after removing the target sound signal from the first right-channel signal, and the second center-channel signal is the sound signal after adding the target sound signal; the audio information includes the target sound signal and the background sound signal.
[0007] In the above manner, the electronic device determines the soundscape space parameters of the target sound signal based on the picture information and audio information of the video to be played, performs channel allocation on the audio information and extracts the target sound signal, and then based on the soundscape space parameters, controls the speaker and the screen sound generating unit to play their respective corresponding channel signals, plays the extracted target sound signal through the screen, realizes the directional enhancement of the target sound signal in the screen direction, and makes the clarity of the sound image of the target sound signal higher in the screen direction; in combination with the soundscape space parameters, controls the playback of each channel signal, makes the audio playback effect more matched with the picture effect, and improves the user viewing effect.
[0008] In a possible implementation manner of the first aspect, based on the picture information and audio information of the video to be played, determining the soundscape space parameters corresponding to the target sound signal in the audio information includes:
[0009] Based on the picture information and audio information, determining the sound pressure level corresponding to the audio information, the soundscape category where the target object that emits the target sound signal in the picture information is located, and the sound source information corresponding to the target sound signal.
[0010] Through the above manner, further determining the sound pressure level corresponding to the audio information of the video to be played, as well as the soundscape category and sound source information where the target object is located, can reconstruct the corresponding soundscape space parameters for different scenes, make the determined soundscape space parameters more adapted to the scene of the video to be played, provide a reliable data reference for later playback, make the sound corresponding to the later video playback more matched with the picture effect, and improve the user viewing experience.
[0011] In a possible implementation manner of the first aspect, based on the picture information and the audio information, determining the soundscape category where the target object that emits the target sound signal in the picture information is located includes:
[0012] Based on the picture information and audio information, matching the corresponding target object for the target sound signal; based on the size and display range of the target object in the picture information, determining the soundscape category where the target object is located; the soundscape category includes one of long shot, panoramic shot, medium shot, close shot, and extreme close-up.
[0013] Through the above manner, based on the division and confirmation of the soundscape category corresponding to the target object in the picture information, the subsequent soundscape reconstruction of the target sound signal based on the soundscape category can determine more adapted soundscape space parameters, be applicable to various video scenes, realize the real-time adjustment of the soundscape space parameters corresponding to the soundscape category with the change of the video scene, and improve the consistency and coordination of audio and picture during subsequent video playback.
[0014] In a possible implementation of the first aspect, the sound source information includes azimuth information and spatial information, and the soundscape space parameter includes the optimal signal-to-noise ratio of the target sound signal and the background sound signal; based on a preset soundscape space database, matching the soundscape space parameter corresponding to the sound pressure level, soundscape category, and sound source information includes:
[0015] According to the sound pressure level and the soundscape category, determine the signal-to-noise ratio range of the target sound signal and the background sound signal in the soundscape space database; according to the azimuth information and the spatial information, determine the optimal signal-to-noise ratio of the target sound signal and the background sound signal within the signal-to-noise ratio range.
[0016] Through the above method, based on the azimuth information of the sound source and the spatial information of the sound source, the optimal signal-to-noise ratio of the target sound signal is determined, making the playback effect of the target sound signal clearer and more realistic, and the playback soundscape corresponding to the target sound signal more matching the video scene.
[0017] In a possible implementation of the first aspect, according to the soundscape space parameter, controlling the screen sound generating unit to play the second center channel signal includes:
[0018] Based on the optimal signal-to-noise ratio, determine the gain value corresponding to the target sound signal; according to the gain value, adjust the sound pressure level of the target sound signal, and control the screen sound generating unit to play the second center channel signal at the adjusted sound pressure level.
[0019] Through the above method, based on the optimal signal-to-noise ratio, determine the gain value corresponding to the target sound signal, and by adjusting the sound pressure level of the target sound signal, realize the enhancement of the target sound signal in the video scene, improve the playback effect of the target sound signal in the corresponding video scene, and make the target sound signal played in the screen direction clearer.
[0020] In a possible implementation of the first aspect, performing channel allocation on the audio information to obtain a three-channel signal includes:
[0021] When the audio information is a stereo signal, use the first weight parameter group for channel allocation to obtain a three-channel signal; or, when the audio information is a multi-channel signal, use the second weight parameter group for channel allocation to obtain a three-channel signal; wherein, the first weight parameter group performs virtual sound signal processing on the stereo signal, and the second weight parameter is used for linear combination transformation processing of the multi-channel signal.
[0022] Through the above method, based on the first weight parameter group and the second weight parameter group, perform virtual sound image, frequency domain, and linear combination transformation processing on the audio information to obtain the processed three-channel signal, improving the spatial effect during the playback of the three-channel signal and making the playback of the audio more coordinated with the picture.
[0023] In a possible implementation of the first aspect, extracting the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal includes:
[0024] Based on the correlation or phase difference between the first left-channel signal and the first right-channel signal, extract the target sound signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal.
[0025] Through the above method, the target sound signal can be well separated from the background sound signal by the phase difference, improving the extraction purity of the target sound signal; the detailed features of the target sound signal can be better extracted by the correlation, making the effect of the extracted target sound signal more natural, overcoming reverberation interference to a certain extent; at the same time, the computational complexity is reduced compared with other extraction methods.
[0026] In a possible implementation of the first aspect, after extracting the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal, the method further includes:
[0027] Perform noise reduction processing on the second center-channel signal to obtain a noise-reduced target sound signal; perform sound quality equalization processing on the noise-reduced target sound signal to obtain the target sound signal to be played by the screen sound generating unit.
[0028] Through the above method, noise reduction processing is performed on the extracted target sound signal, improving the extraction accuracy of the target sound signal, reducing the interference of the background sound signal, and making the target sound signal purer and clearer when played.
[0029] In a second aspect, an audio processing device is provided, which is applied to an electronic device. The electronic device is provided with a screen, a screen sound generating unit, a first speaker, and a second speaker. The first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device; the device includes:
[0030] A soundscape reconstruction unit, configured to determine the soundscape space parameters corresponding to the target sound signal in the audio information based on the picture information and audio information of the video to be played;
[0031] A channel allocation unit, configured to perform channel allocation on the audio information to obtain a three-channel signal, where the three-channel signal includes a first left-channel signal, a first right-channel signal, and a first center-channel signal;
[0032] A signal extraction unit, configured to extract the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal;
[0033] A sound output unit, configured to control a first speaker to play a second left-channel signal, a second speaker to play a second right-channel signal, and a screen sound generating unit to play a second center-channel signal according to soundscape space parameters; wherein, the second left-channel signal is a background sound signal obtained by removing a target sound signal from a first left-channel signal, the second right-channel signal is a background sound signal obtained by removing the target sound signal from a first right-channel signal, and the second center-channel signal is a sound signal after adding the target sound signal; the audio information includes the target sound signal and the background sound signal.
[0034] In a third aspect, an electronic device is provided, which includes: one or more processors, and a memory; the memory is coupled to the one or more processors, and the memory is configured to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method provided in the first aspect or any possible implementation manner of the first aspect as described above.
[0035] In a fourth aspect, a chip system is provided, which is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are configured to call computer instructions to cause the electronic device to execute the method provided in the first aspect or any possible implementation manner of the first aspect as described above.
[0036] In a fifth aspect, a computer-readable storage medium is provided, the computer-readable storage medium includes instructions, and when the instructions run on an electronic device, the electronic device is caused to execute the method provided in the first aspect or any possible implementation manner of the first aspect as described above.
[0037] In a sixth aspect, a computer program product is provided, and when the computer program product runs on an electronic device, the electronic device is caused to execute the method provided in the first aspect or any possible implementation manner of the first aspect as described above.
[0038] It can be understood that the beneficial effects of the second to sixth aspects as described above can refer to the relevant descriptions in the first aspect, and will not be elaborated herein. Description of the Drawings
[0039] Figure 1 A schematic diagram of an application scenario of the audio processing method provided by an embodiment of the present application;
[0040] Figure 2 A schematic diagram of the overall process of the audio processing method provided by an embodiment of the present application;
[0041] Figure 3 A schematic diagram of the process of soundscape reconstruction provided by an embodiment of the present application;
[0042] Figure 4It is a schematic flowchart of channel processing provided by an embodiment of the present application;
[0043] Figure 5 It is a schematic diagram of video playback provided by an embodiment of the present application;
[0044] Figure 6 It is a schematic diagram of video playback provided by an embodiment of the present application;
[0045] Figure 7 It is a schematic flowchart of an audio processing method provided by an embodiment of the present application;
[0046] Figure 8 It is a structural block diagram of an audio processing device provided by an embodiment of the present application;
[0047] Figure 9 It is a hardware system architecture diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0048] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0049] In order to illustrate the technical solutions described in the present application, the following will be described through specific embodiments.
[0050] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0051] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0052] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0053] As used in this specification and the appended claims, the term "if" can be construed contextually as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be construed contextually to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".
[0054] Currently, traditional speakers are usually arranged on the side frames of smart terminals. For example, the top speaker and the bottom speaker of a mobile terminal (mobile phone). In scenarios such as playing videos, music, making calls, and navigation, the orientation of the speakers causes the sound image to be blurry and muddy in the direction where the user is viewing, with low clarity, affecting the user's viewing experience.
[0055] In view of the above problems, the embodiments of the present application provide an audio processing method. On the basis of not affecting the user's listening experience, soundscape reconstruction, channel allocation, and target sound signal extraction are introduced to achieve directional enhancement of the target sound signal, solve the defect that the sound image is blurry and muddy in the user's viewing direction, and improve the user's viewing experience and immersion.
[0056] For ease of understanding, the following specifically explains the audio processing method provided by the embodiments of the present application in conjunction with the accompanying drawings.
[0057] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of the application scenario of the audio processing method provided by the embodiments of the present application. As Figure 1 shown, the audio processing method provided by the embodiments of the present application can be applied to electronic devices such as mobile phones and tablets; taking a mobile phone as an example, as Figure 1 shown in the side view of the mobile phone in Figure (a) of, the mobile phone may include a bottom speaker, a top speaker, and a screen sound generating unit arranged below the screen, and the target sound signal is played through the screen by an exciter. Among them, the bottom speaker and the top speaker are arranged on two non-adjacent side frames of the electronic device, such as the bottom side and the top side.
[0058] As Figure 1 shown in Figure (b) of, for the top speaker, the bottom speaker, and the screen sound generating unit, the embodiments of the present application detect the picture information of the video to be played, determine the corresponding soundscape space parameters of the target object in the picture information, and achieve soundscape reconstruction of the video to be played; perform channel allocation on the audio information of the video to be played to obtain a three-channel sound signal, and achieve Figure 1Three-channel external speakers that match the electronic device shown in the figure; by extracting the target sound signal from the left and right channel signals in the three channels, adding the extracted target sound signal to the center channel, and finally playing the sound signals of the left and right channels by the bottom speaker and the top speaker respectively based on the posture of the electronic device, and playing the target sound signal of the center channel by the screen sound generating unit, the enhancement of the target sound signal in the visual direction (or the direction of the screen orientation) when the user watches the video and the enhancement of the sound field of the audio information are realized, the clarity of the target sound signal of the video to be played is improved, and the user experience of watching the video is enhanced.
[0059] In some embodiments, one or more exciters are further provided under the display screen of the electronic device to drive the front display screen and structure, using the display screen as a vibrating body to generate sound waves by vibration. Among them, the exciter can be a piezoelectric ceramic unit exciter or a micro-vibration unit exciter. A multi-layer piezoelectric ceramic sheet is attached to a metal sheet (vibration membrane), and an alternating voltage is applied to the vibration membrane, and it continuously bends up and down with the change of the voltage to drive the load structure to vibrate and generate sound. The micro-vibration unit exciter uses the interaction between the electric field and the magnetic field to generate a force field to drive the load structure to vibrate and generate sound.
[0060] It should be noted that Figure 1 Only the playback state of the mobile phone in the landscape screen is exemplarily described, that is, the sound signals of the left channel and the right channel respectively played by the bottom speaker and the top speaker when the top is facing right; on the contrary, in the landscape screen state when the top is facing left, the sound signal of the left channel can be played by the top speaker, and the sound signal of the right channel can be played by the bottom speaker; the embodiments of the present application are only exemplarily described and do not limit the channel signals played by the bottom speaker and the top speaker.
[0061] In addition, the soundscape space parameters determined through soundscape reconstruction are space parameters reconstructed based on the relationship between the sound source corresponding to the target sound signal and the space. The screen sound generating unit can also set multiple exciters at different positions. When playing the target sound signal of the center channel through the screen sound generating unit, it can also control one or more exciters to vibrate and generate sound based on the sound source position corresponding to the target sound signal and the space information, or control the exciter closest to the sound source position to vibrate and generate sound based on different sound source positions.
[0062] Through the above audio processing method, the audio information of the video to be played is channel - allocated to obtain a three - channel sound signal adapted to the electronic device. While enhancing the spatial effect of audio playback, it can be applied to the playback of different types of audio data, such as stereo, five - channel, seven - channel and other multi - channel audio. By extracting the target sound signal, the clarity of the target sound signal is improved. The target sound signal of the center channel is played through the screen sound - generating unit to enhance the target sound signal and expand the sound field in the direction of the screen orientation. Based on the soundscape space parameters determined by soundscape reconstruction, the audio playback method is controlled to make the audio playback effect more matched with the picture, so that the user's viewing effect is more immersive and the viewing experience is better.
[0063] The following will explain the specific implementation process of audio processing with reference to the accompanying drawings.
[0064] As Figure 2 shown, the overall flow diagram of the audio processing method provided by the embodiment of the present application. The electronic device detects the picture information and audio information of the video to be played, combines the sound source information and spatial information of the target object in the picture, performs soundscape reconstruction, and obtains the soundscape space parameters corresponding to the target. The electronic device performs channel allocation on the audio information to generate a three - channel signal, namely, a left - channel signal, a right - channel signal, and a center - channel signal. Combining the soundscape space parameters, it controls the playback of the three - channel signal corresponding to the top speaker, the bottom speaker, and the screen sound - generating unit. At the same time, before playback, the target sound signals in the left - channel signal and the right - channel signal are respectively extracted, and the extracted target sound signals are played as the center - channel signal. Before playing the center - channel signal through the screen sound - generating unit, the electronic device also performs noise reduction and sound quality equalization processing on the target sound signal in the center - channel signal, making the target sound signal clearer and more natural. Through the broadcast of the screen sound - generating unit, the sound field of the audio playback is also more three - dimensional.
[0065] Based on the above overall implementation process, the specific implementation of each step will be further introduced. As Figure 3 shown, the flow diagram of soundscape reconstruction provided by the embodiment of the present application. In the process of soundscape reconstruction, based on the picture information and audio information of the video to be played, the soundscape space parameters corresponding to the target sound signal in the audio information are determined. Specifically, it may include the following steps:
[0066] S31, Detect the picture information and audio information.
[0067] S32, Match and locate the target sound signal with the target object in the picture information.
[0068] In some embodiments, by detecting the audio information and video information of the video to be played, the voiceprint in the audio information is recognized, and the objects in the video information are recognized through computer vision calculation; based on the voiceprint category in the audio information and the corresponding recognized objects, the sound and the object are located and matched to determine the target sound signal corresponding to the target object. For example, if the sound signals recognized in the audio information include human voices, bird sounds, and water sounds, and the objects recognized in the video information include people, birds, and water, then the sound sources of each sound signal (i.e., the azimuth information of the object that emits the sound in the video) can be determined based on different categories of sound signals and the corresponding objects.
[0069] Exemplarily, when there is a human voice dialogue, the human voice is used as the target sound signal, and the bird sound and the water sound can both be used as background sound signals. When the video is a non-human voice dialogue, the foreground sound signal in the video is used as the target sound signal, such as the music sound in a short video of pure music, or a video without human voices that only shows some sounds of the production process (such as knocking sounds, friction sounds, applause sounds, explosion sounds, etc. in the video).
[0070] Exemplarily, when multiple human object are detected in the video information, the timestamp information of the change in the facial features (such as lip shape features) of the task object and the timestamp information of the human voice signal in the audio information can be matched to determine the human object speaking at different times and the human voices emitted.
[0071] Exemplarily, when the target object is a human object, the corresponding human voice signal can include language, laughter, sobbing, coughing, snoring during sleep, etc. Different types of human voices can correspond to different voiceprint features. Correspondingly, the scenarios where human voices are emitted include scenarios such as conversations, asides, monologues, and inner monologues. The background sound signal can be all sounds other than the human voice signal, such as wind sounds, rain sounds, bird chirping sounds, flowing water sounds, the noise of the city, etc.
[0072] Exemplarily, in the process of matching and locating the sound signal with the object in the video, various sound signals in the audio information can be separated and extracted based on the frequency and amplitude corresponding to the sound signal, and then based on the objects detected in the video, the frequency and amplitude characteristics of each sound signal are used to perform location matching with each object. Or the matching and location of the sound signal and the object in the video are realized based on a neural network model. For example, the neural network model can include a sound network, a visual network, and a location network. The sound features and image features are extracted through the sound network and the visual network respectively, and the extracted sound features and image features are input into the location network for sound source location based on the attention mechanism.
[0073] It should be noted that the above method for sound source localization is only an example. Sound source localization can also be performed based on other sound-picture alignment methods. For example, sound source separation can be carried out using the pan coefficient to estimate the approximate position of each sound source, so as to achieve the matching localization of the sound source and the picture.
[0074] S33. Determine the soundscape category (distant view, panoramic view, medium shot, close-up, extreme close-up) corresponding to the target object in the picture.
[0075] In some embodiments, based on the picture information and audio information, the target sound signal is matched with the corresponding target object. Through the above detection and recognition of the picture information and audio information, as well as the matching localization of the object in the picture information and the sound signal in the audio information, the target object matching the target sound signal can be determined; based on the size and display range of the target object in the picture information, the soundscape category where the target object is located is determined; where the soundscape category may include one of distant view, panoramic view, medium shot, close-up, and extreme close-up.
[0076] Exemplarily, taking the target object as a human object and the target sound signal as a human voice signal as an example, through the detection and recognition of the picture information, the display size and display range of the human object in the picture can be determined, and based on the size and display range presented by the target object in the picture, the soundscape category corresponding to the target object can be further determined.
[0077] Among them, the soundscape category can be the size and range presented by the photographed subject and the picture image in the film and television screen framework. As shown in Table 1, the scene types corresponding to each soundscape category are determined, and the reference features corresponding to each scene type are determined based on the scene type. Based on the matching of the recognized size and display range of the target object with the reference features corresponding to each scene type, the corresponding soundscape category is determined.
[0078] For example, taking the human object as the target object, when the human object appears as a full body and is very small in the picture, only accounting for one-tenth or less of the picture, it is determined that the soundscape category of the current target object is a distant view; when the human object appears as a full body and accounts for one-sixth to one-eighth of the picture area, or the entire picture of the scene appears in the picture, such as the entire picture of a meeting room, it is determined that the soundscape category of the current target object is a panoramic view; similarly, for medium shots, close-ups, and extreme close-ups, based on the preset reference features corresponding to each soundscape category, the soundscape category corresponding to the target in each frame of the video to be played is determined through picture detection and matching.
[0079] Exemplarily, when there are multiple human object in the picture, the soundscape category corresponding to each speaking human object can be determined based on the matching and positioning of the voice signal and the human object. For example, in the same frame of the picture, the soundscape categories corresponding to different human objects may also be different. Thus, for different human objects, the corresponding soundscape category is matched, and further processing of the audio information is performed on the soundscape category corresponding to their voice.
[0080] Exemplarily, when it is recognized that there is no human object in the picture, the soundscape category corresponding to it is determined according to the size and range presented by the target object matching the foreground sound signal in the picture. For example, for the explosion sound in a game scene, the soundscape category corresponding to the target object (such as the explosion special effect in the picture) is determined according to the size of the corresponding target object and the range presented in the picture.
[0081] It should be noted that the adaptation of each scene type in Table 1 is only for exemplary illustration, and corresponding settings can also be made according to the video type and the actual application scenario. For example, in a navigation scene, a video call scene, etc., corresponding reference features can be set to identify the soundscape category of the target object.
[0082] Table 1
[0083]
[0084] S34. Match the soundscape space parameters based on the soundscape database, the audio information, and the soundscape category.
[0085] S35. Determine the optimal signal-to-noise ratio of the target sound signal and the background sound signal.
[0086] S36. Generate the parameter file (log file) to be called.
[0087] In some embodiments, the soundscape database is set based on the different spatial effects, acoustic effects, and environmental atmospheres that can be created based on the relationship between the sound sources and the space in the video. For example, the soundscape database may include the sound pressure level (SPL), the optimal signal-to-noise ratio (SNR), the distance parameter of the sound source in the soundscape category, and the space parameter (such as the room impulse information).
[0088] In some embodiments, the electronic device determines the sound pressure level corresponding to the audio information and the sound source information corresponding to the target sound signal based on the picture information and the audio information of the video to be played; wherein, the sound source information includes the azimuth information (direction) and the space information (such as indoor or outdoor). Based on the preset soundscape space database, the soundscape space parameters corresponding to the sound pressure level, the soundscape category, and the sound source information are matched.
[0089] Among them, the sound pressure level is the magnitude of the sound pressure, used to measure the sound intensity. After sampling and quantifying the audio information, the effective sound pressure value is calculated and then converted into the sound pressure level; that is, the effective sound pressure is obtained by squaring the time-domain sample values of the audio information, averaging them, and then taking the square root. Based on the reference sound pressure, the sound pressure level of the target sound signal or the sound pressure value of the background sound signal is obtained.
[0090] Exemplarily, according to the soundscape category and sound source information corresponding to the target object, the distance parameter and space parameter in the soundscape database are matched; for example, the soundscape category is panoramic, the azimuth information in the sound source information is 30° in the front left, and the space information in the sound source information is indoor, then the distance parameter in the soundscape database is matched as distence = 5m, and the space parameter is the specific room impulse sequence.
[0091] Exemplarily, based on the audio information, the sound pressure level of the target sound signal (sound source signal) of the target object and the sound pressure level of the background sound signal are calculated. According to the sound pressure levels of the two and the soundscape category, the signal-to-noise ratio range is matched in the soundscape database; for example, the calculated sound pressure level of the target sound signal is SPL = 75dB, the sound pressure level of the background sound signal is SPL = 65dB, and the soundscape category is panoramic, then the signal-to-noise ratio range matched based on the soundscape database is SNR = 10dB - 15dB. According to the distance parameter matched by the azimuth information, the room impulse sequence matched by the space information, and the signal-to-noise ratio range, the best signal-to-noise ratio of the target sound signal and the background sound signal is matched in the soundscape database; for example, based on the distance parameter of distence = 5m, the indoor room impulse sequence, and the signal-to-noise ratio range SNR = 10dB - 15dB, the best signal-to-noise ratio is matched as 15dB.
[0092] In addition, the soundscape database can also include soundscape space parameters such as reverberation parameters, decorrelation, and frequency components. Based on the soundscape category, sound source information, and sound pressure level, the appropriate reverberation parameters, decorrelation, and frequency components can be matched in the soundscape database, and then the best signal-to-noise ratio of the target sound signal and the background sound signal is determined, and a parameter file to be called (such as a log file) is generated for subsequent control of video playback to control the playback mode of the audio information.
[0093] Among them, the reverberation parameter can also include the reverberation time, that is, the time when the sound source stops sounding after reaching the steady-state sound field; different scenarios or environments can correspond to different reverberation times; for example, a long reverberation time can increase the richness of the sound quality when playing a video, but too long will affect the clarity; a short reverberation time is beneficial to clarity, but too short will make the sound appear thin. Therefore, more suitable reverberation parameters can be matched based on different scenarios to improve the sound effect during video playback. By analyzing the frequency components, the gain, loss, reflection, etc. of the spatial information (indoor room) on sounds of different frequencies can be determined, and further the reverberation time can be optimized.
[0094] In the above manner, when the video is played externally, the soundscape theory is introduced. By reconstructing the film and television soundscape, a soundscape database for controlling the audio playback method during video external playback is set up, and the soundscape space parameters are matched based on the video's picture information and audio information, so that the subsequent enhanced effect of the target sound signal is more matched with the video's picture effect, improving the user's immersion in watching the movie.
[0095] Based on the above results of the reconstruction of the soundscape space, the process of channel allocation for the audio information will be continued to be introduced below.
[0096] As Figure 4 shown, the schematic diagram of channel allocation provided by the embodiment of the present application. In the embodiment of the present application, by performing channel allocation on the audio information, a three-channel signal is obtained, where the three-channel signal includes a first left-channel signal, a first right-channel signal, and a first center-channel signal.
[0097] Exemplarily, as Figure 4 shown, the audio in the video to be played can be a stereo audio signal and a multi-channel audio signal; among them, the multi-channel audio signal can include a 5.1-channel audio signal or a 7.1-channel audio signal. For example, the 5.1-channel audio signal includes signals in five channel directions: left (L), center (C), right (R), left rear (LS), and right rear (RS), and the 7.1-channel audio signal includes signals in seven channel directions: left (L), center (C), right (R), front left (FL), front right (FR), rear left (BL), and rear right (BR); the stereo audio signal includes signals in the left channel (L) and the right channel (R), and different position sound signals are transmitted through the left and right two channels to make the sound present a stereo effect.
[0098] In some embodiments, when the audio information is a stereo signal, the first weight parameter group is used for channel allocation to obtain a three-channel signal; or, when the audio information is a multi-channel signal, the second weight parameter group is used for channel allocation to obtain a three-channel signal; among them, the first weight parameter group performs virtual sound signal processing on the stereo signal, and the second weight parameter is used for linear combination transformation processing on the multi-channel signal.
[0099] Exemplarily, when performing channel allocation, a three-channel upmixing algorithm may be used for stereo audio signals, and a three-channel downmixing algorithm may be used for multi-channel audio signals.
[0100] Exemplarily, the upmix three-channel algorithm is a three-channel algorithm for stereo to surround sound conversion, which uses short-time Fourier transform to perform frequency domain processing on the stereo audio signal, analyzes the frequency content of the two channel signals in different time segments, and obtains the changes in the frequency composition of the signals at different moments; separates the channel signals in different directions based on the audio-visual system; and then uses virtual sound signal processing on the separated signals to obtain three-channel signals.
[0101] The audiovisual system can be reconstructed from a three-dimensional sound field based on a matrix decoder and head-related transfer function. The spatial position of the channel signals is determined based on the phase difference, intensity difference, and delay information between the channels of the multi-channel audio signal, achieving channel signal separation. Virtual sound signal processing involves analyzing the spatial properties of stereo signals (such as phase difference and intensity difference) and converting these properties into sounds that can be emitted by multiple virtual channel signals. The perceptual characteristics of the human auditory system are used to simulate the actual sound field distribution to achieve a multi-channel surround sound effect. The actual signal processing process can include, but is not limited to, filtering, delay, reverberation, and gain control.
[0102] For example, the process of processing the virtual sound signal can be implemented based on the following formula:
[0103]
[0104] Among them, X, Y, Z are coefficients, that is, the first weight parameter group (such as Figure 4 4); L and R represent the signals of the two channels of the stereo audio signal; L', C', and R' represent the left, center, and right channels of the resulting three-channel signal. The specific values of X, Y, and Z are not limited, as long as they satisfy the arrangement relationship of the parameter group matrix.
[0105] Exemplarily, the downmix three-channel algorithm is the process of converting a multi-channel audio signal into a three-channel audio signal. During the mixing process, the information of multiple channels is reasonably merged into three channels based on the spatial hearing characteristics of the human ear and the sound field layout of the original multi-channel audio signal, and the original sound positioning and stereoscopic sense are maintained as much as possible.
[0106] Exemplarily, when performing channel allocation processing on a multi-channel audio signal, the multi-channel audio signal is subjected to frequency-domain processing using the short-time Fourier transform, the frequency content of the signals of multiple channels in different time segments is analyzed to obtain the change in the frequency composition of the signal at different moments; the channel signals in different directions are separated based on the sound image system; and then the separated signals are subjected to linear combination transform processing to obtain a three-channel signal.
[0107] Among them, the linear combination transform processing is to perform weighted summation (or combined with phase adjustment) on the signals of each channel of the multi-channel audio signal to simulate the spatial sound field effect of a three-channel system. The process of performing linear combination transform processing on the multi-channel audio signal can be implemented based on the following formula:
[0108] For a 5.1-channel audio signal:
[0109]
[0110] Among them, L, C, R, LS, and RS are the signals of each channel of the 5.1-channel audio signal; L′, C′, and R′ are the left-channel signal, center-channel signal, and right-channel signal of the obtained three-channel signal respectively; (0.7, 0, 0, 0.35, -0.35) is the weight parameter group 1 as shown in Figure 4 as shown in, (0, 1, 0, 0, 0) is the weight parameter group 2 as shown in Figure 4 as shown in, (0, 0, 0.7, -0.35, 0.35) is the weight parameter group 3 as shown in Figure 4 as shown in, and the three groups of weight parameter groups constitute the second weight parameter group.
[0111] For a 7.1-channel audio signal:
[0112]
[0113] Among them, L, C, R, FL, FR, BL, and BR are the signals of each channel of the 7.1-channel audio signal; L′, C′, and R′ are the left-channel signal, center-channel signal, and right-channel signal of the obtained three-channel signal respectively; is the second weight parameter group.
[0114] As shown in Figure 4As shown, taking a 5.1-channel audio signal as an example, each separated channel signal is used as a dimension in the signal space. By setting different weight coefficients, the channel signals in different directions are weighted and summed. For example, in channel allocation, the signal of the original center channel is not mixed with the left and right channels, so the weight coefficients corresponding to other channels except the center channel are all zero. For other channels (left channel, right channel, left rear channel, right rear channel), the corresponding weight parameters can be set according to the auditory direction positioning, and the left and right channel signals in the three-channel signal are obtained by weighted summation.
[0115] Among them, as Figure 4 shown, in order to maintain the naturalness and directionality of the sound of the center channel signal, the separated center channel signal is processed in the frequency domain, such as compensating the frequency response or adjusting the phase, etc., so that the center channel signal obtained after channel allocation can more accurately reproduce the spatial position of the original center channel signal.
[0116] It should be noted that the weight parameters in formula (2) and formula (3) are only exemplary, and the weight parameters corresponding to each channel can also be flexibly adjusted based on the actual application scenario. As Figure 4 shown, in the process of performing channel allocation processing, it also includes combining the soundscape space parameters of soundscape reconstruction to perform channel signal allocation processing to obtain a three-channel signal.
[0117] Through the above method, in the process of channel allocation, the virtual sound signal processing is performed on the stereo audio signal and the linear combination transformation processing is performed on the multi-channel audio signal respectively through the optimized weight parameter group, and through the frequency domain processing of the short-time Fourier transform, the spatial effect after channel allocation is improved.
[0118] In some embodiments, as Figure 4 shown, after the above channel allocation, in order to enhance the playback effect of the target sound signal, the electronic device extracts the target sound signal from the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal and a second center channel signal.
[0119] Among them, the first left channel signal, the first right channel signal and the first center channel signal are the channel signals obtained after channel allocation. The second left channel signal is the background sound signal after the target sound signal is removed from the first left channel signal, the second right channel signal is the background sound signal after the target sound signal is removed from the first right channel signal, and the second center channel signal is the sound signal after the target sound signal is added; the audio information includes the target sound signal and the background sound signal.
[0120] In some embodiments, the electronic device may extract a target sound signal based on the correlation or phase difference between the first left-channel signal and the first right-channel signal, and obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal.
[0121] Exemplarily, for stereo audio information or multi-channel audio information, target sound signals (such as human voice signals) may exist in both the left-channel signal and the right-channel signal in the three-channel signals obtained after the above distribution, and there may also be differences in the relative positions and intensities of the target sound signals in the left channel and the right channel; by calculating the cross-correlation function of the two channels, the alignment degree of the target sound signal components on the time axis can be confirmed. For example, if the sound signals in a certain frequency band in the two channels are highly correlated, it can be determined that the sound signals in this frequency band are part of the target sound signal.
[0122] Correspondingly, due to the different positions of the microphones during the recording of the original audio, there will still be a certain phase difference between the left-channel signal and the right-channel signal after subsequent channel distribution. By detecting the phase difference relationship between the two channels, the target sound signal is extracted.
[0123] It should be noted that the process of extracting the target sound signal can also be implemented based on methods such as blind source separation, independent component analysis, non-negative matrix factorization, and AI large models, which are not specifically limited here.
[0124] Through the above method, the target sound signal can be well separated from the background sound signal by the phase difference, improving the extraction purity of the target sound signal; the detailed features of the target sound signal can be better extracted through the correlation, making the effect of the extracted target sound signal more natural, overcoming reverberation interference to a certain extent; at the same time, the computational complexity is reduced compared with other extraction methods.
[0125] In some embodiments, as Figure 4 shown, in order to make the extraction accuracy of the target sound signal after channel distribution higher and the background sound signal in the center channel after extraction removed more cleanly, after extracting the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal, the method further includes: performing noise reduction processing on the second center-channel signal to obtain a noise-reduced target sound signal; performing sound quality equalization processing on the noise-reduced target sound signal to obtain the target sound signal to be played by the screen sound-emitting unit.
[0126] Exemplarily, the implementation of noise reduction for the center-channel signal after extracting the target sound signal can be achieved through methods such as active noise control, passive noise suppression, digital signal processing noise reduction, and filter noise reduction.
[0127] In the above manner, the loudness of the target sound signal can be increased, making the target sound signal more natural and clear; at the same time, the expansion of the sound field is also more three-dimensional.
[0128] In some embodiments, as Figure 4 shown, after channel allocation and extraction of the target sound signal, the electronic device controls the first speaker to play the second left channel signal, the second speaker to play the second right channel signal, and the screen sound generating unit to play the second center channel signal according to the soundscape space parameters.
[0129] Exemplarily, based on the best signal-to-noise ratio in the soundscape space parameters, a gain value corresponding to the target sound signal is determined; according to the gain value, the sound pressure level of the target sound signal is adjusted, and the screen sound generating unit is controlled to play the second center channel signal at the adjusted sound pressure level.
[0130] Exemplarily, by combining the best signal-to-noise ratio of the target sound signal and the background sound in the soundscape reconstruction, the enhancement degree of the target sound signal is finally determined.
[0131] It should be noted that the target sound signal in the above embodiments can be a human voice signal corresponding to the person in the picture, or a foreground sound signal of other target objects.
[0132] As Figure 5 and Figure 6 shown, when the target object is a human object and the target sound signal is a human voice, the human voice is played through the screen sound generating unit based on the gain value, and the corresponding background sound is played through the left channel and the right channel respectively.
[0133] Exemplarily, as Figure 5 shown, the electronic device identifies that the soundscape category corresponding to the target object is a close-up through picture detection, and based on the soundscape category, the sound pressure level of the audio information, and the sound source information, matches the soundscape space parameters in the soundscape database, and controls the screen sound generating unit to play the human voice based on the soundscape space parameters, and the left channel and the right channel play the background sound.
[0134] Exemplarily, as Figure 5 shown, the electronic device identifies that the soundscape category corresponding to the target object is a panorama through picture detection, and based on the soundscape category, the sound pressure level of the audio information, and the sound source information, matches the soundscape space parameters in the soundscape database, and controls the screen sound generating unit to play the human voice based on the soundscape space parameters, and the left channel and the right channel play the background sound.
[0135] As Figure 5 shown, when the top speaker of the electronic device faces right and the bottom speaker faces left, the left channel background sound can be played by the bottom speaker, and the right channel background sound can be played by the top speaker; as Figure 6As shown, when the top speaker faces left and the bottom speaker faces right, the background sound of the left channel can be played by the top speaker, and the background sound of the right channel can be played by the bottom speaker; that is, the playback mode of the left and right channels is determined according to the state of the electronic device.
[0136] In the above manner, in scenarios such as movies or games, when there is human voice in the video, the directional enhancement of the human voice can be achieved through the screen sound generating unit, and the consistency between the human voice and the picture can be maintained. When the target object of the target sound signal to be emitted does not appear in the picture, that is, when the target object corresponding to the target sound signal is not located in the picture, after channel processing, the first left channel signal and the first right channel signal can be directly played by the top speaker and the bottom speaker respectively, and the first center channel signal after channel processing can be played by the screen sound generating unit. The sound field on the vertical plane can be increased through the screen sound, which can compensate for the lack of the stereo signal in the vertical direction, thereby further expanding the sound field and increasing the sound field sweet spot distance.
[0137] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the audio processing method provided by the embodiment of the present application; the execution subject of this method can be an electronic device, such as a mobile phone, a tablet, etc. The electronic device is provided with a screen, a screen sound generating unit, a first speaker and a second speaker, and the first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device; based on the same implementation principle as above, it will not be elaborated here; as Figure 7 shown, this method may include the following steps:
[0138] S71, based on the picture information and audio information of the video to be played, determine the soundscape space parameters corresponding to the target sound signal in the audio information.
[0139] S72, perform channel allocation on the audio information to obtain a three-channel signal, and the three-channel signal includes a first left channel signal, a first right channel signal, and a first center channel signal.
[0140] S73, extract the target sound signal from the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal, and a second center channel signal.
[0141] S74, according to the soundscape space parameters, control the first speaker to play the second left channel signal, the second speaker to play the second right channel signal, and the screen sound generating unit to play the second center channel signal.
[0142] Wherein, the second left channel signal is the background sound signal after the target sound signal is removed from the first left channel signal, the second right channel signal is the background sound signal after the target sound signal is removed from the first right channel signal, and the second center channel signal is the sound signal after the target sound signal is added; the audio information includes the target sound signal and the background sound signal.
[0143] Through the embodiments of the present application, the electronic device constructs a soundscape space database, and based on the picture information and audio information, matches the soundscape space parameters in the soundscape space database for the external speaker system, so as to achieve directional enhancement of the target sound signal (such as a human voice signal), making the playback effect of the played target sound signal (such as a human voice signal) more matched with the picture information and improving the user's viewing experience; through the processing of channel allocation, combined with the parameter matrix of the linear combination transformation and the added frequency domain processing, the channel allocation process is optimized, and the audio space effect after channel allocation is improved; through the extraction and noise reduction of the target sound signal in the signals of the left and right channels after channel allocation, the extraction of the target sound signal is more accurate, the background sound signal is removed more cleanly, the loudness of the target sound signal (such as a human voice signal) is increased, and the sound is more natural and fresh, realizing the expansion of the sound field and making the played sound more stereoscopic.
[0144] Corresponding to the audio processing method provided in the above embodiments, Figure 8 The structural schematic diagram of the audio processing device provided by the embodiments of the present application is shown. The device is applied to an electronic device, and the electronic device is provided with a screen, a screen sound generating unit, a first speaker and a second speaker. The first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device; for the convenience of description, only the parts related to the embodiments of the present application are shown.
[0145] Referring to Figure 8 , the audio processing device includes:
[0146] A soundscape reconstruction unit 81, configured to determine the soundscape space parameters corresponding to the target sound signal in the audio information based on the picture information and audio information of the video to be played;
[0147] A channel allocation unit 82, configured to perform channel allocation on the audio information to obtain a three-channel signal, and the three-channel signal includes a first left channel signal, a first right channel signal, and a first center channel signal;
[0148] A signal extraction unit 83, configured to extract the target sound signal in the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal, and a second center channel signal;
[0149] A sound output unit 84, configured to control a first speaker to play a second left-channel signal, a second speaker to play a second right-channel signal, and a screen sound generating unit to play a second center-channel signal according to soundscape space parameters;
[0150] Wherein, the second left-channel signal is a background sound signal obtained by removing a target sound signal from a first left-channel signal, the second right-channel signal is a background sound signal obtained by removing the target sound signal from a first right-channel signal, and the second center-channel signal is a sound signal after adding the target sound signal; the audio information includes the target sound signal and the background sound signal.
[0151] In a possible implementation manner, the soundscape reconstruction unit 81 is further configured to determine, based on the video information and the audio information, the sound pressure level corresponding to the audio information, the soundscape category to which a target object that emits the target sound signal in the video information belongs, and the sound source information corresponding to the target sound signal.
[0152] In a possible implementation manner, the soundscape reconstruction unit 81 is further configured to match a corresponding target object to the target sound signal based on the video information and the audio information; determine the soundscape category to which the target object belongs based on the size and display range of the target object in the video information; the soundscape category includes one of long shot, panorama, medium shot, close shot, and extreme close-up.
[0153] In a possible implementation manner, the sound source information includes azimuth information and spatial information, and the soundscape space parameter includes the optimal signal-to-noise ratio of the target sound signal and the background sound signal; the soundscape reconstruction unit 81 is further configured to determine, in a soundscape space database, the signal-to-noise ratio range of the target sound signal and the background sound signal according to the sound pressure level and the soundscape category; determine the optimal signal-to-noise ratio of the target sound signal and the background sound signal within the signal-to-noise ratio range according to the azimuth information and the spatial information.
[0154] In a possible implementation manner, the sound output unit 84 is further configured to determine a gain value corresponding to the target sound signal based on the optimal signal-to-noise ratio; adjust the sound pressure level of the target sound signal according to the gain value, and control the screen sound generating unit to play the second center-channel signal at the adjusted sound pressure level.
[0155] In a possible implementation manner, the channel allocation unit 82 is further configured to, when the audio information is a stereo signal, perform channel allocation using a first weight parameter group to obtain a three-channel signal; or, when the audio information is a multi-channel signal, perform channel allocation using a second weight parameter group to obtain a three-channel signal; wherein, the first weight parameter group performs virtual sound signal processing on the stereo signal, and the second weight parameter is used to perform linear combination transformation processing on the multi-channel signal.
[0156] In a possible implementation, the signal extraction unit 83 is configured to extract a target sound signal based on the correlation or phase difference between the first left-channel signal and the first right-channel signal, and obtain a second left-channel signal, a second right-channel signal, and a second center-channel signal.
[0157] In a possible implementation, the device further includes: a noise reduction and equalization unit, configured to perform noise reduction processing on the second center-channel signal to obtain a noise-reduced target sound signal; and perform sound quality equalization processing on the noise-reduced target sound signal to obtain a target sound signal to be played by the screen sound generation unit.
[0158] In the embodiments of the present application, by constructing a soundscape space database, based on the picture information and audio information, the soundscape space parameters for the external sound system in the soundscape space database are matched, so as to achieve the directional enhancement of the target sound signal (such as a human voice signal), make the playback effect of the played target sound signal (such as a human voice signal) more matched with the picture information, and improve the user viewing experience; through the processing of channel allocation, combined with the parameter matrix of the linear combination transformation and the added frequency-domain processing, the channel allocation process is optimized, and the audio space effect after channel allocation is improved; through the extraction and noise reduction of the target sound signal in the signals of the left and right channels after channel allocation, the extraction of the target sound signal is made more accurate, the background sound signal is removed more cleanly, the loudness of the target sound signal (such as a human voice signal) is increased, and the sound is more natural and fresh, so as to achieve the expansion of the sound field and make the played sound more stereo.
[0159] Figure 9 FIG. 100 shows a schematic structural diagram of an electronic device.
[0160] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, a subscriber identification module (SIM) card interface 195, and a screen sound generating unit 196, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0161] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0162] In some embodiments, a memory may also be provided in the processor 110 for storing instructions and data. The memory in the processor 110 is a cache memory. This memory may store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0163] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0164] In some embodiments, the I2S interface may be used for audio communication. The processor 110 may include multiple groups of I2S buses. The processor 110 may be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. The PCM interface may also be used for audio communication to sample, quantize, and encode analog signals. Both the I2S interface and the PCM interface can be used for audio communication.
[0165] In some embodiments, the GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. The GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0166] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0167] In some embodiments, the modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194. The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor.
[0168] In some embodiments, the display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. The electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0169] Exemplarily, the screen sound generating unit 196 includes at least one actuator. The screen sound generating unit 196 is used to drive the display screen 194 to vibrate and generate sound based on the screen sound generating technology through the actuator. As Figure 5 shown, the actuator receives the driving signal sent by the processor 110, vibrates in response to the driving signal, and the vibration signal generated by the actuator drives the display screen 194 to vibrate, thereby generating sound waves.
[0170] It should be noted that the embodiments of the present application do not specifically limit the number of actuators included in the screen sound generating unit 196, the type of the actuator, and the relative position between the actuator and the display screen 194, which can be adjusted according to actual usage requirements.
[0171] Exemplarily, taking the electronic device as a mobile phone as an example, the display screen 194 may include a flexible screen area, and the vibration component of the screen sound generating unit 196 is disposed on the flexible screen area. In some embodiments, the vibration component may be pasted on the back of the display surface of the flexible screen area, and the processor 110 is connected to the vibration component. The processor 110 is configured to output a driving signal to the vibration component to trigger the vibration of the vibration component. The vibration of the vibration component drives the vibration of the flexible screen area, and the sound is generated through the vibration of the flexible screen area.
[0172] In some embodiments, the vibration component may be a piezoelectric ceramic sheet. Piezoelectric ceramics is an information functional ceramic material that can convert mechanical energy and electrical energy into each other. It has the characteristic of changing its thickness according to the current after being energized and generating vibration, converting voltage into mechanical energy, and resonating with the frame of the screen sound generating unit (or the frame of the electronic device using the screen sound generating unit) in a micro-vibration manner to generate sound and achieve sound generation. It can be understood that the sound generation of the flexible screen area can be understood as achieving the sound generation effect based on the vibration of the flexible screen area.
[0173] In some embodiments, the flexible screen area includes at least two flexible screen sub-areas. The edges of the two flexible screen sub-areas are spliced with the non-flexible screen area, and a vibration component is respectively disposed on each flexible screen sub-area, so that the flexible screen sub-areas are independent of each other and will not interfere with each other during vibration sound generation, and the sound generation effect is good.
[0174] In some embodiments, the processor 110 processes the video to be played. Based on the picture information and audio information, the corresponding soundscape space parameters are matched in the soundscape space database, and through channel allocation and extraction of the target sound signal for the audio information, the background sound signal is played through the speaker 170A and the target sound signal is played through the screen sound generating unit 196 disposed on the display screen 194 to achieve the directional enhancement of the target sound signal.
[0175] In some embodiments, the video codec is used to compress or decompress digital video. The electronic device 100 may support one or more video codecs. The NPU is a neural-network (NN) computing processor. By drawing on the biological neural network structure, such as drawing on the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, voice recognition, audio extraction, etc.
[0176] In some embodiments, the digital signal processor may frame the audio information of the original video to obtain multiple frames of signals, then perform time-frequency conversion on each frame of signal to obtain frequency-domain signals, and then send the frequency-domain signals to the NPU. The NPU inputs the frequency-domain signals into the trained NN network model to separate and extract the target sound signals from the audio information.
[0177] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, and the application processor, etc. Such as music playback, recording, etc.
[0178] The audio module 170 is used to convert digital audio information into analog audio signals for output, and is also used to convert analog audio inputs into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0179] The speaker 170A, also known as the "loudspeaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or hands-free calls through the speaker 170A. The speaker 170A can also include a top speaker and a bottom speaker, which are respectively used to play the sound signals of the left and right channels.
[0180] In some embodiments, the speaker 170A can be disposed on the side frame of the electronic device. When the number of speakers 170A is multiple, the multiple speakers 170A can be disposed on one or more side frames.
[0181] Exemplarily, taking the electronic device as a mobile phone as an example, Figure 1 shows a schematic diagram of dual speakers disposed on two side frames (the top and bottom frames) of the mobile phone. As Figure 1 shown in the (b) figure of, there is a speaker disposed on the upper frame adjacent to the display screen 194. As Figure 1 shown in the (b) of, there is another speaker disposed on the lower frame adjacent to the display screen 194. Among them, the upper frame and the lower frame are two opposite frames of the mobile phone. It should be noted that in actual implementation, the position and number of speakers can be adjusted according to the actual design requirements of the product. For example, the dual speakers can be disposed on the left frame and the right frame, and only one speaker or multiple speakers can be disposed on one frame. For another example, the speaker can be disposed at the center position of the frame, and the speaker can also be disposed at the edge position of the frame. It should be understood that different layout orientations of the speakers will result in different sound fields generated by the speakers.
[0182] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a call or a voice message, the voice can be received by bringing the receiver 170B close to the human ear.
[0183] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0184] In some embodiments, when playing a video, the posture of the electronic device 100 is detected by the acceleration sensor 180E, and the processor 110 determines the channel signals respectively played by the top speaker and the bottom speaker based on the posture of the electronic device 100. For example, when the top of the electronic device faces the right hand side and the bottom faces the left hand side, the left channel sound signal is played through the bottom speaker, and the right channel sound signal is played through the top speaker; on the contrary, the right channel sound signal is played through the bottom speaker, and the left channel sound signal is played through the top speaker.
[0185] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0186] This embodiment also provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on the electronic device, the electronic device is enabled to execute the above related method steps to implement the method in the above embodiments.
[0187] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above related steps to implement the method in the above embodiments.
[0188] In addition, the embodiments of the present application also provide a device, which may specifically be a chip, a component or a module. The device may include a processor and a memory connected to each other; wherein, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory so that the chip executes the methods in the above method embodiments.
[0189] Among them, the electronic device (such as a mobile phone, etc.), computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0190] From the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0191] In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0192] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
Claims
1. An audio processing method, characterized in that, Applied to an electronic device, the electronic device is provided with a screen, a screen sound generating unit, a first speaker and a second speaker, and the first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device; the method includes: Based on the picture information and audio information of the video to be played, determine the soundscape space parameters corresponding to the target sound signal in the audio information; Perform channel allocation on the audio information to obtain a three-channel signal, and the three-channel signal includes a first left-channel signal, a first right-channel signal and a first center-channel signal; Extract the target sound signal from the first left-channel signal and the first right-channel signal to obtain a second left-channel signal, a second right-channel signal and a second center-channel signal; According to the soundscape space parameters, control the first speaker to play the second left-channel signal, the second speaker to play the second right-channel signal, and the screen sound generating unit to play the second center-channel signal; Wherein, the second left-channel signal is the background sound signal after removing the target sound signal from the first left-channel signal, the second right-channel signal is the background sound signal after removing the target sound signal from the first right-channel signal, and the second center-channel signal is the sound signal after adding the target sound signal; the audio information includes the target sound signal and the background sound signal.
2. The method according to claim 1, characterized in that, The determining the soundscape space parameters corresponding to the target sound signal in the audio information based on the picture information and audio information of the video to be played includes: Based on the picture information and the audio information, determine the sound pressure level corresponding to the audio information, the soundscape category where the target object that emits the target sound signal in the picture information is located, and the sound source information corresponding to the target sound signal; Based on a preset soundscape space database, match the soundscape space parameters corresponding to the sound pressure level, the soundscape category and the sound source information.
3. The method according to claim 2, wherein The determining the soundscape category where the target object that emits the target sound signal in the picture information is located based on the picture information and the audio information includes: Based on the picture information and the audio information, match the corresponding target object for the target sound signal; Based on the size and display range of the target object in the picture information, determine the soundscape category where the target object is located; the soundscape category includes one of long shot, panorama, medium shot, close shot, and extreme close-up.
4. The method according to claim 2, characterized in that The sound source information includes azimuth information and spatial information, and the soundscape space parameters include the optimal signal-to-noise ratio between the target sound signal and the background sound signal; the matching the soundscape space parameters corresponding to the sound pressure level, the soundscape category and the sound source information based on a preset soundscape space database includes: According to the sound pressure level and the soundscape category, determine the signal-to-noise ratio range between the target sound signal and the background sound signal in the soundscape space database; According to the azimuth information and the spatial information, determine the optimal signal-to-noise ratio between the target sound signal and the background sound signal within the signal-to-noise ratio range.
5. The method according to claim 4, characterized in that Controlling the screen sound generating unit to play the second center channel signal according to the soundscape space parameters includes: Determining a gain value corresponding to the target sound signal based on the optimal signal-to-noise ratio; Adjusting the sound pressure level of the target sound signal according to the gain value, and controlling the screen sound generating unit to play the second center channel signal at the adjusted sound pressure level.
6. The method according to claim 1, wherein Performing channel allocation on the audio information to obtain a three-channel signal, including: When the audio information is a stereo signal, performing channel allocation using a first weight parameter group to obtain the three-channel signal; or, When the audio information is a multi-channel signal, performing channel allocation using a second weight parameter group to obtain the three-channel signal; Wherein, the first weight parameter group performs virtual sound signal processing on the stereo signal, and the second weight parameter is used to perform linear combination transformation processing on the multi-channel signal.
7. The method according to claim 1, wherein Extracting the target sound signal from the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal, and a second center channel signal, including: Extracting the target sound signal based on the correlation or phase difference between the first left channel signal and the first right channel signal to obtain the second left channel signal, the second right channel signal, and the second center channel signal.
8. The method according to any one of claims 1 to 7, characterized in that, After extracting the target sound signal from the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal, and a second center channel signal, the method further includes: Performing noise reduction processing on the second center channel signal to obtain a noise-reduced target sound signal; Performing sound quality equalization processing on the noise-reduced target sound signal to obtain the target sound signal to be played by the screen sound generating unit.
9. An audio processing device, characterized in that, Applied to an electronic device, the electronic device is provided with a screen, a screen sound generating unit, a first speaker, and a second speaker, and the first speaker and the second speaker are arranged on two non-adjacent side frames of the electronic device; the device includes: A soundscape reconstruction unit for determining soundscape space parameters corresponding to the target sound signal in the audio information based on the picture information and audio information of the video to be played; A channel allocation unit for performing channel allocation on the audio information to obtain a three-channel signal, and the three-channel signal includes a first left channel signal, a first right channel signal, and a first center channel signal; A signal extraction unit for extracting the target sound signal from the first left channel signal and the first right channel signal to obtain a second left channel signal, a second right channel signal, and a second center channel signal; A sound output unit for controlling the first speaker to play the second left channel signal, the second speaker to play the second right channel signal, and the screen sound generating unit to play the second center channel signal according to the soundscape space parameters; Wherein, the second left channel signal is the background sound signal after the target sound signal is removed from the first left channel signal, the second right channel signal is the background sound signal after the target sound signal is removed from the first right channel signal, and the second center channel signal is the sound signal after the target sound signal is added; the audio information includes the target sound signal and the background sound signal.
10. An electronic device, characterized in that, The electronic device includes: one or more processors, and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
11. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions, when the instructions run on an electronic device, causing the electronic device to execute the method according to any one of claims 1 to 8.