Multimedia playing method and electronic device
By identifying the audio and sound field information of moving sound sources and combining multiple sound-generating devices for audio playback, the problem of mismatch between the audio playback effect and the real scene in existing technologies is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202311439008.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-10-31
AI Technical Summary
When using dual speakers for audio playback, existing electronic devices cannot effectively simulate the sound field of a real scene, resulting in a poor user experience.
By identifying the audio and sound field information of moving sound sources, and combining multiple target sound-generating devices to play audio, target audio information matching the real scene is generated, and the display screen is controlled to play the screen information synchronously.
It improves the matching degree between the audio playback effect and the real scene, enhancing the user's auditory experience.
Smart Images

Figure CN118466884B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of terminals, and in particular, to a multimedia playing method and an electronic device. BACKGROUND
[0002] With the massive popularization and application of electronic devices (such as mobile phones, tablet computers and the like), the electronic devices can support more and more applications and become more and more powerful. The electronic devices develop in the direction of diversification and individualization and become indispensable electronic appliances in users' lives.
[0003] External playing is the main playing mode of sound in the daily audio-visual entertainment scenarios of users of electronic devices, and in particular when using video application programs (APPs) and game APPs, the good or bad of the external playing experience directly affects the overall audio experience of the device. The demand for audio motion sound image is relatively large in the video APPs and the game APPs, and the accurate implementation of the motion sound image can bring the users a sound experience of being in the scene.
[0004] At present, in order to improve the external playing effect of audio playing, the electronic device is equipped with double loudspeakers, and in the process of the user using the electronic device for external playing, there is a certain sound field expansion effect. However, in the process of the external playing of the electronic device, the external playing effect of the audio played through the double loudspeakers is quite different from the experience in the real scene, resulting in poor user experience. SUMMARY
[0005] Embodiments of the present application provide a multimedia playing method and an electronic device, by combining the audio information of a motion sound source with sound field information, and playing the combined audio information by using a plurality of target sound emitting devices matched with the motion sound source, so that the audio playing effect of the motion sound source is more matched with the video picture.
[0006] To achieve the above object, embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a multimedia playing method is provided, applied to an electronic device including at least three sound emitting devices and a display screen, wherein one sound emitting device is located at the top of the electronic device, one sound emitting device is located at the bottom of the electronic device, and at least one sound emitting device is located between the top and the bottom of the electronic device, and the multimedia playing method includes:
[0008] The multimedia resource to be played is acquired, audio information and picture information are extracted, the moving sound source is identified from the audio information, the audio information of the moving sound source is separated from the audio information, the picture information is identified to determine sound field information and moving sound source information corresponding to the moving sound source, the sound field information is used to represent environmental parameters where the moving sound source is located, the moving sound source information includes sound source position information, a plurality of target sound emitting devices are determined from at least three sound emitting devices according to the sound source position information of the moving sound source, target audio information is generated according to the audio information and the sound field information of the moving sound source, the target sound emitting devices are controlled to play the target audio information, and the display screen is controlled to play the picture information of the multimedia resource.
[0009] It can be understood that when the electronic device plays the multimedia resource, the moving sound source is identified in real time, a plurality of target sound emitting devices for playing the moving sound source are determined, the plurality of target sound emitting devices are controlled to play the target audio information fused with the sound field information, so that the effect of the electronic device playing the audio is basically consistent with the real scene, and the user experience is improved.
[0010] In some embodiments, before the target audio information is generated according to the audio information and the sound field information of the moving sound source, the multimedia resource playing method can further include:
[0011] The motion angle range of the moving sound source is determined according to the marking information of the moving sound source, the marking information is used to represent the corresponding relationship among the sound source type, the sound source motion time and the sound source position of the moving sound source, the audio information of the moving sound source is rendered according to the motion angle range of the moving sound source to obtain rendered audio information, and the rendered audio information includes audio information played by different target sound emitting devices.
[0012] It can be understood that for the moving sound source with different motion angle ranges, the electronic device can use different rendering methods for rendering processing. For example, the electronic device can use a multi-directional amplitude planning (MDAP) method to render process the moving sound source moving in a space region of a front direction (-45°-45°). The electronic device can use a vector base amplitude planning (VBAP) method to render process the moving sound source moving in a space region of a right front direction (0°-90°) and a left front direction (-90°-0°).
[0013] In other embodiments, the audio information of the moving sound source is rendered according to the motion angle range of the moving sound source to obtain rendered audio information, including:
[0014] Based on the location information of the moving sound source and the location information of the target sound-emitting device, determine the gain coefficient of the audio information of the moving sound source played by each target sound-emitting device; adjust the audio information of the moving sound source according to the gain coefficient of the target sound-emitting device to determine the audio information played by the target sound-emitting device.
[0015] This can be understood as follows: in order to ensure that the position of the virtual sound image of the moving sound source is consistent with the position information of the sound source when multiple target sound-emitting devices play audio information from a moving sound source, the electronic device can determine the gain coefficient of each target sound-emitting device playing audio information from the moving sound source, and determine the audio information to be played by each target sound-emitting device based on the gain coefficient.
[0016] In other embodiments, the audio information of the moving sound source is rendered according to the range of its motion angle to obtain rendered audio information, and the method further includes:
[0017] Based on the location information of the moving sound source and the location information of the target sound-emitting device, the gain coefficient of the audio information of the moving sound source output by the target sound-emitting device is determined; the gain coefficients corresponding to the target sound-emitting device are added together and normalized to obtain the normalized gain coefficient; based on the normalized gain coefficient, the audio information of the moving sound source is adjusted to determine the audio information to be played by the target sound-emitting device.
[0018] In other embodiments, before determining the motion angle range of the moving sound source based on the marking information of the moving sound source, the method further includes:
[0019] Based on the sound source type, sound source movement time, and sound source location information, the labeling information of the moving sound source is generated.
[0020] In this embodiment, the electronic device can generate marker information for each moving sound source based on its sound source type, movement time, and location information, and then store the marker information for each moving sound source in a sound source marker file. In this way, when rendering the audio information of the moving sound sources, the electronic device can directly obtain the marker information from the sound source marker file to determine the movement trajectory of the moving sound sources based on the marker information.
[0021] In other embodiments, the motion sound source information also includes scene category and sound source spatial information. The image information is identified to determine the sound field information and motion sound source information corresponding to the motion sound source, including:
[0022] The system identifies the scene information to determine the scene category, location information, and spatial information of the moving sound source. The scene category, location information, and spatial information are then input into a matching database, which outputs the sound field information corresponding to the moving sound source. The matching database stores the correspondence between multiple moving sound source information and sound field information.
[0023] In other embodiments, the sound field information includes background sound pressure level, target signal-to-noise ratio, distance parameters, or room pulse sequence. Scene category, sound source location information, and sound source spatial information are input into a matching database. The matching database outputs sound field information corresponding to the moving sound source, including:
[0024] Based on the shot category and / or sound source location information, determine the distance parameter and room pulse sequence; based on at least one of the shot category, distance parameter, or room pulse sequence, determine the sound pressure level of the moving sound source and the background sound pressure level; based on the shot category and / or distance parameter, adjust the background sound pressure level to determine the target background sound pressure level; based on the sound pressure level of the moving sound source and the target background sound pressure level, determine the target signal-to-noise ratio.
[0025] The target signal-to-noise ratio refers to the optimal signal-to-noise ratio when playing motion sound sources.
[0026] In other embodiments, target audio information is generated based on the audio information and sound field information of the moving sound source, including:
[0027] Based on the target signal-to-noise ratio, the loudness values of the audio information of the moving sound source and the background audio information are adjusted;
[0028] The audio information of the moving sound source after loudness adjustment is mixed with the background audio information to obtain the target audio information.
[0029] This can be understood as the electronic device adjusting the loudness values of the moving sound source's audio information and the background audio information based on the target signal-to-noise ratio of the moving sound source. This makes the external playback effect more realistic when the target sound-emitting device in the electronic device plays the target audio information.
[0030] In other embodiments, the target sound-emitting device is a preset number. The target sound-emitting device is determined from at least three sound-emitting devices based on the sound source location information of the moving sound source, including: calculating the distance value between the moving sound source and each sound-emitting device based on the sound source location information of the moving sound source.
[0031] Based on the distance between the moving sound source and each sound-generating device, a preset number of sound-generating devices are selected from at least three sound-generating devices as target sound-generating devices.
[0032] In other embodiments, a target sound-generating device is determined from at least three sound-generating devices based on the sound source location information of the moving sound source, including:
[0033] Based on the location information of the moving sound source, calculate the distance between the moving sound source and each sound-generating device;
[0034] For each sound-generating device, if the distance between the moving sound source and the sound-generating device is less than a distance threshold, then the sound-generating device is determined to be the target sound-generating device.
[0035] This can be understood as follows: during the movement of a sound source, the electronic device can determine the target sound-emitting device for playing the audio information of the sound source in real time, and adjust the target sound-emitting device in real time according to the sound source location information, so that the effect of the electronic device playing the audio is more matched with the effect of the picture.
[0036] In other embodiments, before controlling the target sound-emitting device to play target audio information, the method further includes:
[0037] Enhance the target audio information with sound effects.
[0038] Therefore, electronic devices can use a combination of variable speed and constant pitch algorithm and resampling to process the rendered audio information to obtain enhanced audio information, thereby realizing the Doppler effect of audio.
[0039] In other embodiments, at least three sound-generating devices are loudspeakers, or at least three sound-generating devices are screen sound-generating devices, or at least three sound-generating devices are a combination of loudspeakers and screen sound-generating devices.
[0040] It can be understood that regardless of whether the sound-generating device in the electronic device is a speaker or a screen sound device, the audio effect of playing multimedia resources in this application can be achieved.
[0041] In a second aspect, this application provides an electronic device, comprising: at least three sound-emitting devices, wherein the sound-emitting devices are used to play audio information; a display screen, wherein the display screen is used to display image information; one or more processors; and a memory; wherein the memory stores one or more computer programs, the one or more computer programs including instructions, which, when executed by the electronic device, cause the electronic device to perform the multimedia playback method as described in any one of the first aspects above.
[0042] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the multimedia playback method as described in any one of the first aspects.
[0043] Fourthly, this application provides a computer program product, which includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the multimedia playback method as described in any one of the first aspects.
[0044] It is understood that the electronic device described in the second aspect, the computer storage medium described in the third aspect, and the computer program product described in the fourth aspect are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here. Attached Figure Description
[0045] Figure 1 A schematic diagram of a mobile phone with dual speakers provided in an embodiment of this application;
[0046] Figure 2 A front view of an electronic device having four sound-emitting devices, provided as an embodiment of this application;
[0047] Figure 3 A front view of another electronic device having four sound-emitting devices provided in an embodiment of this application;
[0048] Figure 4 A cross-sectional schematic diagram of an electronic device having four sound-emitting devices, provided for an embodiment of this application;
[0049] Figure 5 A cross-sectional schematic diagram of another electronic device having four sound-emitting devices provided for an embodiment of this application;
[0050] Figure 6 A circuit connection example diagram of a sound-generating device provided in an embodiment of this application;
[0051] Figure 7 A circuit connection example diagram of another sound-generating device provided in an embodiment of this application;
[0052] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0053] Figure 9 A software structure diagram of an electronic device provided in an embodiment of this application;
[0054] Figure 10 A schematic diagram of a soundscape theory provided for an embodiment of this application;
[0055] Figure 11 A flowchart illustrating a multimedia playback method provided in an embodiment of this application;
[0056] Figure 12An example diagram of region division provided in an embodiment of this application;
[0057] Figure 13 Example of angular range sensing of a moving sound source provided in the embodiments of this application Figure 1 ;
[0058] Figure 14 Example of angular range sensing of a moving sound source provided in the embodiments of this application Figure 2 ;
[0059] Figure 15 Example of angular range sensing of a moving sound source provided in the embodiments of this application Figure 3 ;
[0060] Figure 16 An example diagram illustrating the motion trajectory of a moving sound source provided in an embodiment of this application;
[0061] Figure 17 An example diagram illustrating virtual sound source determination provided in this application embodiment;
[0062] Figure 18 An example diagram illustrating an audio playback scenario provided in an embodiment of this application;
[0063] Figure 19 This is a flowchart illustrating another multimedia playback method provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0065] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this application, unless otherwise stated, "a plurality of" means two or more.
[0066] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0067] Currently, electronic devices are typically equipped with dual speakers in order to allow users to experience a stereo audio playback effect. Figure 1 This is a schematic diagram of a mobile phone with dual speakers provided as an embodiment of this application. Figure 1 As shown, the electronic device is equipped with two speakers, namely a top speaker 101 located at the top of the electronic device and a bottom speaker 102 located at the bottom of the electronic device.
[0068] In some scenarios, such as Figure 1 As shown, when playing movies or games on an electronic device in landscape mode, the left and right sides of the device correspond to the left and right audio speaker areas, respectively. However, there is no sound in the center of the screen, meaning the top and bottom speakers are missing. This results in the user experiencing sound only from the left and right edges of the screen, leading to poor focus and an inconsistent experience with real-world scenarios, thus resulting in a poor user experience.
[0069] It needs to be explained that, Figure 1 The positions of the top speaker 101 and bottom speaker 102 in the electronic device shown are merely illustrative. The top speaker 101 may be located in the middle of the top of the electronic device, or it may be located on the left or right side of the top of the electronic device. The bottom speaker 102 may be located in the middle of the bottom of the electronic device, or it may be located on the left or right side of the bottom of the electronic device. The specific positions of the top speaker 101 and bottom speaker 102 in the electronic device are not limited in this embodiment.
[0070] Correspondingly, when the electronic device is in portrait mode and uses the top speaker 101 and the bottom speaker 102 to play audio, although an upper sound field and a lower sound field can be generated, a left sound field and a right sound field are missing, that is, a stereo effect cannot be generated.
[0071] It is evident that when only the top speaker 101 and the bottom speaker 102 are provided in the electronic device, no matter whether the electronic device is in landscape or portrait mode, it is impossible to generate a sound field in the four directions of top, bottom, left and right. In other words, the electronic device cannot have a surround sound field listening experience, resulting in a poor user experience.
[0072] Therefore, this application provides a multimedia playback method applied to an electronic device. When playing multimedia resources, the electronic device can determine the audio information of a moving sound source based on the audio information of the multimedia resource to be played, and then determine multiple target sound-emitting devices from at least three sound-emitting devices based on the sound source location of the moving sound source. Then, the electronic device combines the audio information of the moving sound source with the sound field information to obtain the target audio information, and controls the multiple target sound-emitting devices to play the target audio information. Simultaneously, the electronic device controls the display screen to show the video information of the multimedia resource.
[0073] Therefore, during the movement of the sound source, multiple target sound-emitting devices for playing audio are adjusted in real time according to the sound source location information. In addition, sound field information is introduced when playing audio information, making the audio information played by the target sound-emitting devices more in line with the audio effect of the real scene.
[0074] This multimedia playback method can be applied to scenarios where electronic devices play audio externally. The electronic device contains at least three sound-generating devices, which can all be speakers, all be screen-based sound generators, or a combination of speakers and screen-based sound generators. For example, a screen-based sound generator can be installed below the screen of the electronic device. When sound is needed, the screen-based sound generator drives the screen and its structure in front, using the screen as a vibrator to generate sound waves, which are then transmitted to the listener's ear.
[0075] Compared to dual speakers, the electronic device in this application embodiment can be equipped with multiple sound-generating devices. During the movement of the sound source, the audio is played using a matched target sound-generating device, making the sound played by the electronic device more focused. Thus, this solution allows the sound perceived by the user to approximate the sound in the real scene, improving the user experience.
[0076] For example, the multimedia playback method provided in this application can be applied to electronic devices with audio playback functions, such as mobile phones, tablets, personal computers (PCs), personal digital assistants (PDAs), smartwatches, netbooks, wearable electronic devices, augmented reality (AR) devices, virtual reality (VR) devices, in-vehicle devices, smart cars, and smart speakers. This application does not impose any limitations on this. The electronic device can also be a foldable screen device (e.g., a foldable phone), which is also not limited here.
[0077] For example, Figure 2This is a front view of an electronic device with four sound-emitting devices, provided as an embodiment of this application. The front of the electronic device, i.e., the screen of the electronic device, faces the user. Figure 2 As shown, the electronic device includes a top speaker 201, a left screen sound-emitting device 202, a right screen sound-emitting device 203, and a bottom speaker 204.
[0078] Figure 3 A front view of another electronic device with four sound-emitting devices provided in an embodiment of this application. (See attached diagram.) Figure 3 As shown, the electronic device includes a top speaker 301, a left screen sound-emitting device 302, a right screen sound-emitting device 303, and a bottom speaker 304.
[0079] It should be noted that, Figure 2 and Figure 3 The positions of the left and right screen sound-emitting devices shown in the illustration are merely examples. The specific positions of the left and right screen sound-emitting devices in this embodiment are not limited; they can achieve sound generation simply by being located below the display screen. The aforementioned screen sound-emitting devices can be piezoelectric ceramic screen sound-emitting exciters, voice coil screen sound-emitting exciters, magnetically levitated screen sound-emitting exciters, or micro-vibration unit exciters, etc. The specific type of screen sound-emitting device is also not limited in this embodiment.
[0080] For piezoelectric ceramic screen sound exciters, multiple layers of piezoelectric ceramic sheets are attached to a metal sheet, called a diaphragm. Applying alternating voltages to the diaphragm causes it to bend up and down in response to the voltage changes, driving the load structure to vibrate and produce sound. Micro-vibration unit exciters, also called linear vibrators, operate on a principle similar to linear motors, utilizing the interaction of electric and magnetic fields to generate a force field. Piezoelectric ceramic unit exciters generally perform poorly with low-frequency signals, while micro-vibration unit exciters offer a more balanced and flat frequency response within the speech range, resulting in better sound quality.
[0081] Figure 4 A cross-sectional schematic diagram of an electronic device having four sound-emitting devices is provided for an embodiment of this application, as shown below. Figure 4 The cross-sectional views shown in (a) and (b) show that the left and right screen sound-generating devices are both located below the screen. The electronic device drives the screen to vibrate by applying driving signals to the left and right screen sound-generating devices, thereby pushing the air in front of and behind the screen to produce sound.
[0082] In the embodiments of this application, such as Figure 5As shown, when the two screen-emitting devices in an electronic device are stationary, the device does not produce sound. When both screen-emitting devices protrude simultaneously, they push the screen forward, thus moving the air in front of and behind the screen to produce sound. Similarly, when both screen-emitting devices bend simultaneously, they push the screen backward, thus moving the air behind the screen to produce sound. In other words, the electronic device produces sound by driving the screen to vibrate back and forth through its screen-emitting devices.
[0083] also, Figure 2 to Figure 5 This is a schematic diagram of an electronic device with four sound-emitting devices according to an embodiment of this application. The number of sound-emitting devices in the electronic device is not limited in this embodiment; the electronic device may also have three sound-emitting devices, or more. For example, the electronic device may include one top speaker, one bottom speaker, and one screen sound-emitting device. Alternatively, the electronic device may include two top speakers, two bottom speakers, and two screen sound-emitting devices. Yet another example is that the electronic device may include one top speaker, one bottom speaker, and three screen sound-emitting devices.
[0084] In this embodiment of the application, an electronic device is provided with four sound-generating devices as an example. (Refer to...) Figure 6 and Figure 7 The circuit connection method of each sound-generating device is illustrated. These four sound-generating devices include a top speaker, a bottom speaker, a left-side screen sound-generating device, and a right-side screen sound-generating device.
[0085] As one possible implementation, the left and right screen sound-emitting devices in the electronic device are connected in parallel, and both screen sound-emitting devices share the same power amplifier (PA). For example... Figure 6 As shown, the electronic device has three power amplifiers (PAs): PA1, PA2, and PA3. The audio signal processed by PA1 is output to the top speaker, the audio signal processed by PA2 is output to the bottom speaker, and the audio signal processed by PA3 can be output to either the left or right screen speaker. In other words, the electronic device simultaneously controls three speaker devices through the three PAs. Specifically, the electronic device achieves three-channel sound by simultaneously emitting sound from the top speaker, bottom speaker, and left screen speaker. Alternatively, the electronic device can achieve three-channel sound by simultaneously controlling the top speaker, bottom speaker, and right screen speaker.
[0086] It should be noted that, Figure 6The audio signal processed by PA1 and output to the top speaker, the audio signal processed by PA2 and output to the bottom speaker, and the audio signal processed by PA3 and output to the left / right screen sound-emitting device shown are merely exemplary implementations and are not limited in this embodiment. For example, the audio signal processed by PA1 can also be output to the bottom speaker or the left / right screen sound-emitting device, the audio signal processed by PA2 can also be output to the top speaker or the left / right screen sound-emitting device, the audio signal processed by PA3 can also be output to the top speaker or the bottom speaker, and so on.
[0087] As another possible implementation, the left and right screen sound-emitting devices in the electronic device are connected in series, with each screen sound-emitting device having its own independent driver. For example... Figure 7 As shown, the electronic device has four power amplifiers (PAs): PA1, PA2, PA3, and PA4. The audio signal processed by PA1 is output to the top speaker, the audio signal processed by PA2 is output to the bottom speaker, the audio signal processed by PA3 is output to the left screen speaker, and the audio signal processed by PA4 is output to the right screen speaker. In other words, the electronic device can simultaneously control four speaker devices to produce sound, supporting four-channel external playback. Specifically, the electronic device uses the top speaker, bottom speaker, left screen speaker, and right screen speaker to simultaneously produce sound, achieving four-channel audio output.
[0088] It should be noted that, Figure 7 The audio signal processed by PA1 to PA4 shown in the figure is output to four sound-generating devices. This is only an exemplary implementation scheme and is not limited in this embodiment.
[0089] like Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0090] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, screen sound device 171, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc.
[0091] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0092] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0093] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0094] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0095] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0096] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0097] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0098] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0099] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0100] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.
[0101] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0102] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0103] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0104] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.
[0105] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance).
[0106] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0107] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0108] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.
[0109] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0110] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0111] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. Wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. GNSS can include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0112] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0113] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0114] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0115] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0116] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0117] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0118] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0119] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0120] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0121] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0122] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0123] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0124] The screen sound-generating device 171 is used to generate sound by driving the screen to vibrate through vibration.
[0125] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0126] In this embodiment of the application, in order to improve the audio playback effect of the electronic device, the electronic device 100 may be provided with multiple speakers. For example, the electronic device 100 may be provided with three speakers.
[0127] For example, in this embodiment of the application, during the process of the electronic device playing audio, the audio signal is played by the speaker 170A of the audio module 170, and at the same time, the screen sound device 171 drives the screen (i.e., the display screen 194) to emit screen sound to play the audio signal. The number of speakers 170A and screen sound devices 171 can be one or more. For example, the electronic device can be configured with two speakers 170A and two screen sound devices 171.
[0128] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0129] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0130] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and detach from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as making calls and data communication.
[0131] The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of an electronic device.
[0132] Figure 9 This is a software structure diagram of an electronic device provided in an embodiment of this application.
[0133] Understandably, a layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system may include an application layer (referred to as the application layer), an application framework layer (referred to as the framework layer), system libraries, and a kernel layer.
[0134] The application layer described above may include a series of application packages.
[0135] like Figure 9 As shown, the application package may include system applications. System applications refer to applications installed on the electronic device before it leaves the factory. For example, system applications may include programs such as desktop, camera, gallery, calendar, music, notes, and weather.
[0136] Application packages can also include third-party applications, which are applications that users download and install from app stores (or app markets). Examples include map applications, food delivery applications, reading applications (such as e-books), social networking applications, and travel applications.
[0137] The application framework layer described above provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0138] likeFigure 9 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0139] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0140] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, and more.
[0141] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0142] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).
[0143] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0144] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating the phone, and flashing indicator lights.
[0145] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0146] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0147] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0148] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0149] The Surface Manager is used to manage the display subsystem and provides the blending of two-dimensional and three-dimensional layers for multiple applications.
[0150] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0151] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0152] A 2D graphics engine is a drawing engine for 2D drawing.
[0153] In this embodiment of the application, the system library may further include: a resource acquisition module, a resource processing module, a sound source identification module, a sound source separation module, an information acquisition module, and an audio processing module.
[0154] The resource acquisition module can be used to acquire multimedia resources to be played. For example, the resource acquisition module can acquire game resources to be played.
[0155] The resource processing module can be used to process multimedia resources to obtain audio and video information.
[0156] The sound source recognition module can be used to identify sound sources in audio information, so as to identify moving sound sources from the audio information.
[0157] The sound source separation module can be used to separate the audio information corresponding to the moving sound source from the audio information.
[0158] The information acquisition module can process video information to obtain moving sound source information. For example, based on video information, the information acquisition module can determine the motion angle, location information, spatial information, and scene category of the moving sound source, etc. The information acquisition module is also used to determine the sound field information corresponding to the moving sound source. This sound field information may include background sound pressure level, target signal-to-noise ratio, distance setting parameters, or room pulse sequence, etc.
[0159] The audio processing module can determine the target sound-emitting device for playing the moving sound source based on the distance between the sound source's location information and each sound-emitting device. The audio processing module can also render the audio information of the moving sound source, and then combine the rendered audio information with the sound field information to obtain the target audio information.
[0160] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera driver, audio driver, and sensor driver.
[0161] The audio driver is used to drive the sound-generating device to play audio information.
[0162] The display driver is used to drive the display screen to show the visual information corresponding to the audio information.
[0163] To ensure clarity and brevity in the description of the following embodiments, the applicable scenarios of this solution will be briefly introduced first.
[0164] like Figure 10 As shown, soundscape theory combines landscape architecture, soundscape studies, psychoacoustics, and architectural and environmental acoustics to study the interrelationships between people, the environment, and sound, which differs from traditional noise control. Soundscape emphasizes perception rather than just physical quantities, considers positive and harmonious sounds rather than just noise, and views the sound environment as a resource.
[0165] The main factors determining shot size are: scene space, distance, and lens focal length. Shot size refers to the difference in the size of the subject as seen through the camera's viewfinder due to varying distances between the camera and the subject. Generally, shot sizes can be categorized into five types, from farthest to closest: long shot, full shot, medium shot, close-up, and extreme close-up. Table 1 below introduces each shot size and its corresponding applicable scene types.
[0166] Table 1
[0167]
[0168] The technical solutions involved in the following embodiments can all be implemented in electronic devices with the above-described hardware structure and software architecture.
[0169] In this embodiment, after acquiring the multimedia resources to be played, the electronic device processes the multimedia resources to obtain audio and video information. The electronic device performs sound source identification on the audio information, and after identifying moving sound sources, separates the audio information corresponding to the moving sound sources. The electronic device identifies the video information to determine the sound field information and moving sound source information corresponding to the moving sound sources (e.g., scene category, sound source type, sound source location information, the correspondence between sound source movement time and movement position, etc.). Then, based on the sound source location information of the moving sound sources, the electronic device determines multiple target sound-emitting devices. Further, the electronic device combines the rendered audio information with the sound field information to obtain target audio information, and then controls the target sound-emitting devices to play the target audio information and controls the display screen to play the video information.
[0170] The following section uses a mobile phone as an example to describe the multimedia playback method provided in the embodiments of this application in detail with reference to the accompanying drawings. Figure 11 This is a flowchart illustrating a multimedia playback method provided in an embodiment of this application, as shown below. Figure 11 As shown, the multimedia playback method may include steps S1100 to S1180.
[0171] Step 1100: The mobile phone obtains the multimedia resources to be played.
[0172] The multimedia resources refer to the resources to be played on the mobile phone. Multimedia resources can be game resources or video resources. For example, multimedia resources can be movie resources, game resources, etc. This application embodiment does not specifically limit the specific type of multimedia resources.
[0173] Furthermore, the multimedia resources can be resources stored locally on the mobile phone, or resources downloaded in real time, etc. The source of the multimedia resources to be played is not limited in this application embodiment.
[0174] In this embodiment of the application, the mobile phone can obtain multimedia resources to be played. These multimedia resources include audio information and video information.
[0175] Step 1110: The mobile phone extracts the multimedia resources to be played to obtain audio and video information.
[0176] In this embodiment, the mobile phone can separate the audio and video frames of the multimedia resource to be played to obtain audio information and video information. The audio information includes the audio stream signal to be played. The video information includes each frame corresponding to the multimedia resource to be played.
[0177] In the embodiments of this application, the audio stream signal included in the multimedia resources can be two-channel audio, 5.1-channel audio, or 7.1-channel audio, etc. Two-channel audio refers to audio including a left channel and a right channel. 5.1-channel audio refers to audio including a front left channel, a front right channel, a center channel, a rear left surround channel, a rear right surround channel, and a subwoofer channel. 7.1-channel audio refers to audio including a left front channel, a right front channel, a left surround channel, a right surround channel, a front center channel, a left rear channel, a right rear channel, and a subwoofer channel.
[0178] Step 1120: The mobile phone performs sound source identification on the audio information and identifies the motion sound source.
[0179] Among them, a moving sound source refers to a sound source that is in motion. For example, a moving sound source can be a human voice, a bird's voice, an explosion, a car's voice, a train's voice, or the sound of other moving objects, etc.
[0180] In this embodiment, the mobile phone can perform sound source identification on the audio information to identify various types of sound sources contained in the audio information. For example, the sound source identification module can identify the types of sound sources from the audio information, including: human voices, bird calls, gunshots, airplane sounds, car sounds, horseback riding sounds, footsteps, rain sounds, and thunder sounds, etc. Then, the mobile phone identifies motion sound sources from the various types of sound sources contained in the audio information.
[0181] In some embodiments, the mobile phone may pre-store a set of motion sound sources, which includes various types of motion sound sources. After identifying various types of sound sources from the audio information, the mobile phone matches the identified sound sources with the motion sound sources in the motion sound source set. If the mobile phone determines that the motion sound source set includes the sound sources identified from the audio information, then the mobile phone determines that the sound sources identified from the audio information are motion sound sources. If the mobile phone determines that the motion sound source set does not include the sound sources identified from the audio information, then the mobile phone determines that the sound sources identified from the audio information are not motion sound sources.
[0182] In other embodiments, the mobile phone can input audio information into a trained sound source recognition model, which outputs the various types of motion sound sources it has identified.
[0183] The aforementioned sound source recognition model is trained using sound clips corresponding to different types of moving sound sources. After training, the sound source recognition model has the ability to identify different types of moving sound sources. For example, human voice clips can be pre-collected, and then the collected human voice clips can be used to train a neural network to obtain a sound source recognition model that can be used to identify human voice sound sources. Similarly, bird call clips can be pre-collected, and then the collected bird call clips can be used to train a neural network to obtain a sound source recognition model that can be used to identify bird call sound sources. Likewise, car sound clips can be pre-collected, and then the collected car sound clips can be used to train a neural network to obtain a sound source recognition model that can be used to identify car sound sources.
[0184] It should be noted that after identifying various types of moving sound sources, the sound source recognition model described above can input these types of moving sound sources into a sound source classifier, which then separates the different types of moving sound sources. Alternatively, the sound source recognition model can directly output the separated types of moving sound sources.
[0185] After identifying different types of sound sources from audio information, the mobile phone can also use synchronous matching image recognition technology to detect different types of sound sources and determine moving sound sources from among them. For example, the mobile phone can use deep learning algorithms such as histogram of oriented gradient (HOG) features describing texture features, kernel support vector machine (SVM) learning algorithms to solve linear and nonlinear problems, and convolutional neural networks (CNN) to identify moving sound sources from different types of sound sources.
[0186] Furthermore, the number of motion sound sources that the mobile phone can identify from the audio information can be one or more; there is no limitation here. For example, the mobile phone can identify only human voices from the audio information, or it can also identify human voices, bird sounds, and car sounds from the audio information, and so on.
[0187] Step 1130: The mobile phone extracts the audio information corresponding to the motion sound source from the audio information.
[0188] In this embodiment of the application, after the mobile phone identifies the motion sound source from the audio information, the mobile phone can separate the audio information corresponding to the identified motion sound source from other background sounds, so as to separate the audio information corresponding to the motion sound source from the audio information.
[0189] Optionally, when the mobile phone identifies multiple motion sound sources from the audio information, the mobile phone can separate the audio information corresponding to the multiple motion sound sources from the audio information to obtain the audio information corresponding to each of the multiple motion sound sources.
[0190] Step 1140: The mobile phone identifies the screen information and obtains motion sound source information.
[0191] The aforementioned motion sound source information may include scene category, sound source type, sound source spatial information (e.g., whether the motion sound source is indoors or outdoors), sound source location information, or sound field information, etc. Among these, sound source type refers to the type to which the sound source belongs. For example, the sound source type corresponding to a human voice is a motion sound source. The sound source type corresponding to a machine sound source in a workshop is a static sound source. Sound source location information may include the initial and final positions of the motion sound source.
[0192] A sound field refers to the environment in which a sound source is located. A sound field can be a quiet environment (e.g., a street at night or a quiet library) or a non-quiet environment (e.g., a noisy street, supermarket, zoo, shopping mall). Sound field information can include background sound pressure level, target signal-to-noise ratio, distance parameters, or room pulse sequences. Background sound pressure level refers to the volume of background noise in the environment where the moving sound source is located. Distance parameters refer to the distance between the moving sound source and the viewer. Room pulse sequences are used to indicate environmental information, such as whether the room is open or not, and whether the room is sound-absorbing or not.
[0193] In some embodiments, the sound source location information can be determined by the position of the sound-emitting object corresponding to the moving sound source in the frame. For example, the sound source location information can be the coordinate value of the sound-emitting object corresponding to the moving sound source in the video frame.
[0194] Optionally, the mobile phone can analyze each frame of the video frame based on video semantic analysis to obtain the coordinates of the sound-emitting object displayed on the phone screen. Video semantic analysis refers to the process of extracting semantic components from the video. The mobile phone determines whether a moving sound source exists in the video frame displayed on the phone screen; when a moving sound source is present, the mobile phone identifies the coordinates of the moving sound source on the phone screen.
[0195] In other embodiments, the mobile phone can determine the location information of a sound source by the area where the sound-emitting object corresponding to the moving sound source is located. For example, the display interface on the mobile phone can be divided into multiple areas. Figure 12 This is an example diagram of region division provided for an embodiment of this application. For example... Figure 12As shown in (a), the mobile phone display interface is divided into four areas. The mobile phone can determine the location information of the moving sound source based on the area where the sound-emitting object corresponding to the moving sound source is located. If the sound-emitting object corresponding to a moving sound source is located in multiple areas, any one of these multiple areas can be used to represent the sound source location information, or the area with the largest area occupied by the moving sound source can be used to represent the sound source location information.
[0196] like Figure 12 As shown in (b), assuming the moving sound source is a human voice, the current location of speaker 1201 occupies regions 1 and 3. Speaker 1201 occupies the largest area in region 1, so the mobile phone can use region 1 to represent the sound source location information of speaker 1201. Speaker 1202 is currently located within region 4, so the mobile phone can use region 4 to represent the sound source location information of speaker 1202.
[0197] In this embodiment of the application, the mobile phone can determine the sound field information based on the scene category, sound source location information and sound source spatial information in the identified motion sound source information.
[0198] Optionally, the mobile phone pre-stores a matching database corresponding to motion sound source information and sound field information. After determining the scene category, sound source location information, and sound source spatial information in the motion sound source information, the scene category, sound source location information, and sound source spatial information are input into the matching database, and the matching database outputs the corresponding sound field information. That is, the mobile phone inputs the motion sound source information into the matching database, and the matching database outputs the background sound pressure level, target signal-to-noise ratio, distance parameters, or room pulse sequence based on the pre-stored correspondence between motion sound source information and sound field information.
[0199] This can be understood as follows: the matching database pre-stores the correspondence between motion sound source information and sound field information. The mobile phone inputs the motion sound source information into the matching database, and the matching database can accurately query the sound field information corresponding to the motion sound source information and output the sound field information corresponding to the motion sound source information.
[0200] For example, suppose the matching database pre-stores the correspondence between moving sound source information and background sound pressure level, target signal-to-noise ratio, distance parameters, and room pulse sequences. The mobile phone identifies the screen information and determines that the moving sound source corresponds to a panoramic view, a sound source location of 30° left front, and an indoor location. After the mobile phone inputs the panoramic view, sound source location, and sound source location into the matching database, the matching database determines the distance parameters and room pulse sequences based on the pre-stored correspondence between moving sound source information and sound field information. For example, the matching database determines the distance parameter to be 5 meters based on the sound source location. The matching database can determine the sound pressure level of the moving sound source (assuming the moving sound source signal has a sound pressure level of 75 dB) and the background sound pressure level (e.g., the background sound pressure level is 65 dB) based on at least one of the panoramic view, distance parameters, or room pulse sequences. Then, the matching database adjusts the background sound pressure level based on the panoramic view and / or distance parameters to determine the target background sound pressure level. The matching database determines the target signal-to-noise ratio (SNR) based on the sound pressure level (SPL) of the moving sound source and the target background sound SPL. The sound field information output by the matching database includes the background sound SPL, the target SNR, distance parameters, and the room pulse sequence. The mobile phone can obtain the sound field information output by the matching database, i.e., the background sound SPL, the target SNR, distance parameters, and the room pulse sequence. The target SNR refers to the optimal SNR between the moving sound source and the background sound.
[0201] In this embodiment, the mobile phone identifies the moving sound source from the audio information and processes the image information to obtain the moving sound source information. Based on this information, it determines the corresponding tagging information for the moving sound source. Then, the mobile phone tags the moving sound source according to the tagging information. This tagging information can be stored in a sound source tagging file.
[0202] As one possible implementation, a mobile phone can generate tagging information for a moving sound source based on the sound source type, sound source movement time, and sound source location information, so as to tag the moving sound source to be tagged.
[0203] For example, a mobile phone can use 64-bit, 32-bit, or 16-bit binary encoding to mark moving sound sources. For instance, assuming the first two bits represent the sound source type, if the phone determines the moving sound source to be marked is a static sound source, the first two bits can be "00". If the phone determines the sound source to be marked is a moving sound source, the first two bits can be marked as "01". The phone can use the third to eighth bits to represent the sound source's movement time; for example, the third to fifth bits are used to mark the start time of the moving sound source's movement, and the sixth to eighth bits are used to mark the end time of the moving sound source's movement. The phone can use the ninth to twelfth bits to represent the sound source's location information; for example, the ninth and tenth bits are used to mark the start position of the moving sound source's movement, and the eleventh and twelfth bits are used to mark the end position of the moving sound source's movement.
[0204] It should be noted that, assuming the mobile phone determines to use 16-bit binary to mark the sound source, after the mobile phone determines the marking information corresponding to the first to twelfth bits, the thirteenth to sixteenth bits can be padded with 0s so that the marking information corresponding to the moving sound source is a 16-bit binary value.
[0205] Step 1150: The mobile phone identifies multiple target sound-emitting devices based on the motion sound source information.
[0206] In this embodiment, the mobile phone can determine the coordinates of the moving sound source during its movement based on the sound source location information in the moving sound source information. Furthermore, the mobile phone can determine multiple target sound-emitting devices corresponding to different movement positions of the moving sound source based on the coordinates of the moving sound source during its movement and the coordinates of the sound-emitting devices.
[0207] The aforementioned target sound-emitting device refers to a sound-emitting device used to play audio information from a moving sound source. The number of target sound-emitting devices can be preset. For example, when a mobile phone has 4 sound-emitting devices, the number of target sound-emitting devices can be set to 3. Alternatively, when a mobile phone has 5 sound-emitting devices, the number of target sound-emitting devices can be set to 3 or 4.
[0208] In some embodiments, the mobile phone can determine the location information of the moving sound sources based on the marking information of each moving sound source in the sound source marking file. The mobile phone can determine the Euclidean distance between the sound-emitting device and the moving sound source based on the coordinate values of the moving sound source during its movement and the coordinate values of the sound-emitting device. Then, based on the Euclidean distance between each sound-emitting device and the moving sound source, the mobile phone determines a preset number of target sound-emitting devices from a plurality of sound-emitting devices.
[0209] The Euclidean distance mentioned above refers to the actual distance between two points in m-dimensional space. For example, assuming the coordinates of a sound-emitting device in a mobile phone are (x1, y1) and the coordinates of a moving sound source are (x2, y2), the Euclidean distance between the sound-emitting device and the moving sound source is...
[0210] For example, as Figure 2 As shown, the mobile phone includes a top speaker 201, a left-side screen sound-emitting device 202, a right-side screen sound-emitting device 203, and a bottom speaker 204. Assuming the coordinates of the top speaker 201 are (x3, y3) and the coordinates of the moving sound source are (x4, y4), the mobile phone can calculate the Euclidean distance between the top speaker 201 and the moving sound source as follows:
[0211]
[0212] Similarly, using the aforementioned Euclidean distance calculation method, the mobile phone can calculate the Euclidean distance d2 between the left screen speaker 202 and the moving sound source, the Euclidean distance d3 between the right screen speaker 203 and the moving sound source, and the Euclidean distance d4 between the bottom speaker 204 and the moving sound source. Based on the Euclidean distances between the four speaker devices and the moving sound source, the mobile phone can select the three speaker devices with the smallest Euclidean distances as target speaker devices to play the audio information corresponding to the moving sound source.
[0213] In some embodiments, if the mobile phone determines that the Euclidean distance between the top speaker 201, the left screen sound device 202, the right screen sound device 203, and the bottom speaker 204 and the motion sound source is equal, the mobile phone can use all four sound devices as target sound devices.
[0214] It should be noted that when a mobile phone identifies multiple moving sound sources from audio information, the phone can calculate the Euclidean distance between each moving sound source and the sound-emitting device to determine the target sound-emitting device corresponding to each moving sound source.
[0215] In other embodiments, after the mobile phone determines the distance between the moving sound source and the sound-emitting device, for each sound-emitting device, if the distance between the moving sound source and the sound-emitting device is less than a distance threshold, then the sound-emitting device is determined to be the target sound-emitting device.
[0216] For example, suppose a mobile phone has four sound-emitting devices and a distance threshold of 10. The mobile phone determines that the distance values between the moving sound source and the four sound-emitting devices are 12, 8, 5 and 7 respectively. The mobile phone can use the sound-emitting devices with distance values of 8, 5 and 7 as target sound-emitting devices.
[0217] Psychological research on the perception of the angular range of a moving sound source reveals that users' perception of the angular range of a moving sound source varies depending on the spatial region, rotational speed, and signal source. For example, when the actual angle of motion of a moving sound source is 90°, different testers may perceive the angle of motion as equal to 90°, greater than 90°, or less than 90°. In other words, the angle perceived by the human ear as the motion of a sound source does not necessarily match the actual angle of motion.
[0218] Depend on Figure 13 It can be seen that when the moving sound source actually moves 90°, the percentage of the perceived angle of the moving sound source varies among the 15 test subjects after a preset number of tests (e.g., 280 times). For example... Figure 13 As shown, among the 15 test subjects, a small percentage perceived the motion of the sound source as a 90° angle. In most cases, the test subjects perceived the motion of the sound source as a 90° angle or greater. For example, when the sound source moved from directly in front of the test subject (0°) to the test subject's right (90°), meaning the actual angle of motion of the sound source was 90°, the test subject might perceive the motion as a 70° angle, i.e., a less than 90° angle.
[0219] Depend on Figure 14 As shown in (a), when the moving sound source moves 90° in different spatial regions, the perceived angle of the moving sound source by the tester also varies; that is, the perceived angle of the moving sound source by the tester may be equal to 90°, greater than 90°, or less than 90°. Figure 14 As shown in (b), when the moving sound source moves 90° in different spatial regions and the rotation speed of the moving sound source is different, the perceived angle of the moving sound source by the tester is also different. That is, the perceived angle of the moving sound source by the tester may be equal to 90°, greater than 90° or less than 90°.
[0220] In other words, by Figure 14 As shown in Figures (a) and (b), the perceived angle of a moving sound source differs depending on the spatial region. For example, when the sound source moves at a certain angle directly in front (-45°~45°), to the right front (0°~90°), and to the left front (-90°~0°), the perceived angle also differs. Furthermore, the perceived angle also differs when the sound source moves at different rotational speeds in different spatial regions. Figure 15As shown in (a) and (b), when a moving sound source moves 90° at different rotational speeds in different spatial regions, the perceived angle of the moving sound source by the tester in different regions also varies. That is, the perceived angle of the moving sound source by the tester may be equal to 90°, greater than 90°, or less than 90°.
[0221] This can be understood as follows: because moving sound sources rotate at different rates in different spatial regions, the angle of motion perceived by the user varies. Therefore, for moving sound sources in different spatial regions, different target sound-emitting devices can be used for virtual sound image localization, ensuring that the sound emitted by the target sound-emitting device corresponds to the position of the sound-emitting object displayed on the screen. This maintains spatial consistency between the sound and image in the video playback. Consequently, the position of the moving sound source displayed on the phone is essentially consistent with the position of the virtual sound image, creating a more realistic stereo environment and enhancing the user's audiovisual experience.
[0222] Virtual sound image, also known as virtual sound source or sensory sound source, or simply sound image, refers to the ability of a listener to perceive the spatial location of a sound source when sound is played aloud, thus forming a sound image. This sound image is essentially the visual representation of a sound field in the human brain. For example, a person with their eyes closed can immerse themselves in a sound field and imagine the state of the sound source through auditory perception, such as its direction, volume, and distance.
[0223] For example, such as Figure 16 As shown, assuming the human voice source moves from the initial position A to the target position D, during the movement of the human voice source, the mobile phone can select the three closest sound-emitting devices from the four sound-emitting devices to play audio based on the distance between the human voice source and each sound-emitting device, so as to realize the virtual sound image positioning of the human voice source.
[0224] For example, when a human voice source moves from initial position A to position B, after the mobile phone determines the distances of sound-emitting devices 1601, 1620, 1630, and 1640 from the human voice source, the mobile phone determines that the three sound-emitting devices closest to the human voice source are sound-emitting devices 1610, 1620, and 1630. In this case, the mobile phone can use sound-emitting devices 1610, 1620, and 1630 to play the audio information corresponding to the human voice source to achieve virtual sound image localization.
[0225] For example, when the human voice source moves to position C, the mobile phone determines that the three closest sound-emitting devices to the human voice source are sound-emitting device 1620, sound-emitting device 1630, and sound-emitting device 1640. In this case, the mobile phone can use sound-emitting devices 1620, 1630, and 1640 to play the audio information corresponding to the human voice source to achieve virtual sound image positioning.
[0226] Step 1160: After determining the motion angle of the motion sound source based on the motion sound source information, the mobile phone renders the audio information of the motion sound source to obtain the rendered audio information.
[0227] In this embodiment, the mobile phone can determine the motion angle and spatial region of the moving sound source based on the sound source marker information in the sound source marker file. Furthermore, the mobile phone uses different rendering methods to process the audio information of the moving sound source in different spatial regions to obtain the rendered audio information.
[0228] Algorithms for rendering audio information include MDAP and VBAP.
[0229] The VBAP method is a generalization of discrete-pair amplitude signal feeding in the horizontal plane or a special sagittal plane to three-dimensional space. This method achieves local audio translation through three adjacent loudspeakers without affecting the binaural time difference in low frequencies or the spectrum in high frequencies. Therefore, the VBAP method provides more accurate sound localization in three-dimensional space.
[0230] The MDAP method is based on VBAP, but uses more directions to accurately position the virtual sound image. It calculates the gain factors of multiple translation vectors in the required translation direction, adds the gain factors of each speaker, normalizes them using a normalization formula, and then plays them back.
[0231] In this embodiment, the mobile phone can use the MDAP method to render sound sources moving directly in front (-45° to 45°) in the spatial region. The mobile phone can also use the VBAP method to render sound sources moving to the right front (0° to 90°) and left front (-90° to 0°) in the spatial region.
[0232] This can be understood as follows: In order to ensure that the position of the virtual sound image of the moving sound source is consistent with the position information of the sound source when multiple target sound-emitting devices play audio information of the moving sound source, the mobile phone can determine the gain coefficient of each target sound-emitting device playing audio information of the moving sound source based on the position information of the moving sound source and the position information of the target sound-emitting device. Based on the gain coefficient of the target sound-emitting device, the audio information of the moving sound source is adjusted to determine the audio information played by the target sound-emitting device.
[0233] The following example demonstrates the rendering process using a mobile phone to render a moving sound source using the VBAP method.
[0234] Assuming a mobile phone has four sound-emitting devices, SP1, SP2, SP3, and SP4, the phone determines the three sound-emitting devices closest to the moving sound source as the target sound-emitting devices based on the distances between each of the four devices and the moving sound source. For example, the target sound-emitting devices are SP1, SP2, and SP3. (See [link to relevant documentation]). Figure 17 Assume the three sound-generating devices SP1, SP2, and SP3 are located in the phone as follows: Figure 17 As shown, user 1700 is watching a video on their mobile phone. The phone can determine the position of the virtual sound source VSP corresponding to the moving sound source based on the position information of the sound-emitting devices SP1, SP2, and SP3, thereby achieving virtual sound image localization.
[0235] For example, assuming the position of user 1700's head is the origin 0, the position of the virtual sound source VSP in a three-dimensional coordinate system with the vertical, horizontal, and z axes as the x-axis, y-axis, and z-axis directions, respectively, is represented by the vector P in the figure, which starts from the origin 0.
[0236] Since vector P is a three-dimensional vector, it can be represented by the linear sum of three-dimensional vectors extending from the origin 0 along the directions of sound-generating device SP1, sound-generating device SP2, and sound-generating device SP2 respectively. That is, vector P can be represented by the linear sum of vectors L1, L2, and L3. Vector P can be represented by the following formula 1.
[0237] P = g1L1 + g2L2 + g3L3 (Formula 1)
[0238] In Formula 1 above, P represents the vector of the virtual sound source in the three-dimensional coordinate system; L1, L2, and L3 represent the normal vector values of the sound-emitting devices SP1, SP2, and SP3, respectively; g1, g2, and g3 are the gain coefficients. The gain coefficients indicate the gain of the audio played from each target sound-emitting device. The gain coefficients can be calculated using Formula 2 below.
[0239]
[0240] The MDAP method, building upon the VBAP method, uses more directions to precisely locate the virtual sound image position of the moving sound source. The MDAP method calculates the gain factors of multiple translation vectors in the required translation directions, sums the gain factors of each sound-generating device, and then normalizes the result using a normalization formula before playback. The normalization formula is as follows:
[0241]
[0242] Where p is the loudness value of the virtual sound source point; in an anechoic chamber, p is set to 1, meaning the loudness of the virtual sound source points is equal; in a real room with a certain amount of reverberation, p is usually set to 2.
[0243] It should be noted that the specific implementation process of rendering audio information of moving sound sources using the MDAP or VBAP method on mobile phones can be found in existing implementation methods, and will not be described in detail here.
[0244] Optionally, the mobile phone can also enhance the rendered audio information of moving sound sources. For example, the phone can use a combination of variable speed and constant pitch algorithms and resampling to process the rendered audio information to obtain enhanced audio information, thereby achieving the Doppler effect in audio. The Doppler effect refers to the change in wavelength or frequency due to the relative motion between the observer and the sound source. When the sound source moves relative to a stationary observer, or when the observer moves relative to a stationary sound source, the pitch heard by the observer changes. If the sound source is stationary, the sound heard by the human ear will have the same pitch as the sound emitted by the source.
[0245] Among them, the variable-speed-invariant-pitch algorithm refers to maintaining the pitch and semantics while slowing down or speeding up the speech. In other words, it adjusts the duration of the audio information. There are three types of variable-speed-invariant-pitch algorithms: time-domain methods, frequency-domain methods, and parametric methods.
[0246] In this embodiment, when upsampling the rendered audio information of a moving sound source, the mobile phone can first resample the rendered audio information and then perform speed adjustment. When downsampling the rendered audio information, the mobile phone can first perform speed adjustment and then resample the speed-adjusted information. Thus, by combining resampling with speed adjustment of the rendered audio information, the Doppler effect of the audio information is enhanced.
[0247] It should be noted that the execution order of steps 1150 and 1160 in this embodiment is not limited. Step 1150 can be executed first and then step 1160 can be executed first and then step 1150 can be executed, or steps 1150 and 1160 can be executed simultaneously.
[0248] Step 1170: The mobile phone combines the rendered audio information with the sound field information to obtain the target audio information.
[0249] Step 1180: The mobile phone controls the target sound-emitting device to play the target audio information.
[0250] In this embodiment, the mobile phone can determine the target audio information based on the rendered audio information and sound field information. That is, the mobile phone can adjust the loudness values of the moving sound source's audio information and the background audio information based on the target signal-to-noise ratio. Then, it mixes the loudness-adjusted moving sound source's audio information with the background audio information to obtain the target audio information. Thus, by combining the rendered audio information with the sound field information and improving the spatial sound effects of the moving sound source, the mobile phone not only optimizes the ambient sound field but also enhances the user's immersion and realism.
[0251] As an example, such as Figure 18 As shown, assume the mobile phone has four sound-emitting devices: a top speaker 1810, a screen speaker 1820, a screen speaker 1830, and a bottom speaker 1840. A video played on the phone shows a car moving from position E to position F. During this movement, the car's sound source moves from its initial position E to its target position F. As the car's sound source moves, the phone can select the three closest sound-emitting devices as the target sound sources based on their distance from each device, thus achieving virtual sound image localization of the car's sound source.
[0252] For example, when the car sound source is at location E, after the mobile phone determines the distances of the top speaker 1810, screen sound device 1820, screen sound device 1830, and bottom speaker 1840 from the car sound source, the mobile phone determines that the three sound devices closest to the car sound source are the top speaker 1810, screen sound device 1820, and screen sound device 1830. In this case, the mobile phone can use the top speaker 1810, screen sound device 1820, and screen sound device 1830 to play the audio information of the car sound to achieve virtual sound image positioning.
[0253] For example, when the car sound source moves to position F, the mobile phone determines that the three closest sound-emitting devices to the car sound source are the screen sound emitter 1820, the screen sound emitter 1830, and the bottom speaker 1840. In this case, the mobile phone can use the screen sound emitter 1820, the screen sound emitter 1830, and the bottom speaker 1840 to play the audio information of the car sound to achieve virtual sound image positioning.
[0254] like Figure 18As shown, assuming a video playing on the phone shows a car driving while two boys are playing a ball, the phone can determine the target sound source for playing the sounds of the boys playing based on the distance between the two boys and the sound-emitting devices in the video. For example, when the two boys move to area G, the phone determines that the three sound sources closest to the boys are the top speaker 1810, the screen sound source 1820, and the screen sound source 1830. In this case, the phone can use the top speaker 1810, the screen sound source 1820, and the screen sound source 1830 to play the audio information of the two boys playing, thus achieving virtual sound image localization.
[0255] When playing audio information from car sounds and the sounds of two boys playing, the sound-generating device in a mobile phone can combine it with environmental sound field information (such as wind sounds). By combining the sound field information with the audio playback, a more realistic sound effect can be presented.
[0256] In this embodiment, after the mobile phone identifies the moving sound source and sound field information in the video frame, it determines the target sound-emitting device for playing audio based on the position of the moving sound source and each sound-emitting device. Then, the mobile phone renders the audio information of the moving sound source to obtain rendered audio information, and combines the rendered audio information with the sound field information to obtain the target audio information, which is then played through the target sound-emitting device. Compared to dual speakers, in this solution, the electronic device determines the target sound-emitting device for playing the audio of the moving sound source based on the position of the moving sound source and each sound-emitting device, and combines it with ambient sound playback. This allows the position of the sound perceived by the user to be approximately the same as the position of the moving sound source in the video frame, improving the accuracy of the moving sound source localization and enhancing the user experience.
[0257] In some embodiments, taking a movie as an example of the multimedia resource to be played, the following will describe the multimedia playback method provided in this application embodiment in conjunction with the interaction between various modules in the mobile phone's system library. For details, please refer to... Figure 19 .
[0258] like Figure 19 As shown, the multimedia playback method may further include the following steps:
[0259] Step 1900: The mobile phone's resource acquisition module acquires the movie resources to be played.
[0260] The movie resources to be played can be movie resources pre-cached on the mobile phone, movie resources downloaded in real time from the server, etc., and there are no restrictions here.
[0261] Step 1910: The resource acquisition module sends the movie resources to be played to the resource processing module.
[0262] In this embodiment, after the resource acquisition module acquires the movie resource to be played, it sends the movie resource to the resource processing module. Correspondingly, the resource processing module receives the movie resource.
[0263] Step 1920: The resource processing module processes the movie resources to obtain audio and video information.
[0264] Step 1930: The resource processing module sends audio information to the sound source recognition module and video information to the information acquisition module.
[0265] Correspondingly, the sound source recognition module receives the audio information sent by the resource processing module, and the information acquisition module receives the video information sent by the resource processing module.
[0266] Step 1940: The sound source recognition module performs sound source recognition on the audio information to identify moving sound sources.
[0267] In this embodiment of the application, the sound source recognition module can perform sound source recognition on the audio information corresponding to the movie resource in order to identify the moving sound source in the audio information to be played.
[0268] Similarly, the sound source recognition module identifies the sound source of the audio information. The number of identified moving sound sources can be one or more, depending on the actual number of identified moving sound sources.
[0269] Step 1950: The sound source separation module separates the audio information corresponding to the moving sound source from the audio information.
[0270] For example, after the sound source recognition module identifies the moving sound source as a moving car sound source from the audio information, the sound source separation module can separate the audio information corresponding to the moving car sound source from the audio information.
[0271] Step 1960: The information acquisition module identifies the image information and determines the motion sound source information.
[0272] The process by which the information acquisition module determines the motion sound source information corresponding to the motion sound source can be found in the implementation process of step 1140 above, and will not be elaborated here.
[0273] It should be noted that the motion source information may differ for different moving sound sources. For example, the sound source location and sound field information for rain sound and car sound are not the same.
[0274] Step 1970: The audio processing module determines multiple target sound-emitting devices for playing the moving sound source based on the sound source marking information of the moving sound source in the sound source marking file.
[0275] The process by which the audio processing module determines the target sound-emitting device can be found in step 1150 above, and will not be repeated here.
[0276] Step 1980: The audio processing module determines the motion angle of the moving sound source based on the sound source marker information of the moving sound source in the sound source marker file, and then renders the audio information of the moving sound source.
[0277] The process by which the audio processing module renders the audio information of the moving sound source can be found in the implementation process of step 1160 above, and will not be elaborated here.
[0278] It should be noted that the execution order of steps 1970 and 1980 in this embodiment is not limited. Step 1970 can be executed first and then step 1980, or step 1980 can be executed first and then step 1970, or steps 1970 and 1980 can be executed simultaneously.
[0279] In step 1990, the audio processing module combines the rendered audio information with the sound field information to obtain the target audio information.
[0280] Step 1991: The audio driver controls multiple target sound-emitting devices to play target audio information.
[0281] In this embodiment, the audio driver corresponding to the target sound-emitting device can control the target sound-emitting device to play target audio information. Simultaneously, the display driver controls the display screen to display the image information corresponding to the target audio information.
[0282] It's important to explain that each sound-generating device can correspond to one audio driver. For example, when a phone has a top speaker, a bottom speaker, and two screen-based sound-generating devices, the top speaker corresponds to one audio driver, the bottom speaker corresponds to one audio driver, and each of the two screen-based sound-generating devices has its own audio driver. Alternatively, one audio driver can correspond to multiple sound-generating devices. For example, when a phone has a top speaker, a bottom speaker, and two screen-based sound-generating devices, the top speaker corresponds to one audio driver, the bottom speaker corresponds to one audio driver, and the two screen-based sound-generating devices correspond to one audio driver.
[0283] It is understood that the aforementioned electronic devices, etc., include hardware structures and / or software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this invention.
[0284] This application embodiment can divide the above-mentioned electronic device into functional modules according to the method example described above. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0285] Embodiments of this application also provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the multimedia playback method in the above embodiments.
[0286] Embodiments of this application also provide a computer program product, which includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the aforementioned method steps to implement the multimedia playback method in the above embodiments.
[0287] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the apparatus to perform the multimedia playback method executed by the electronic device in the above method embodiments.
[0288] In this embodiment, the electronic device, computer-readable storage medium, computer program product or device are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0289] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0290] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0291] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0292] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A multimedia playback method, characterized in that, The method, applied to an electronic device comprising at least three sound-generating devices and a display screen, wherein one sound-generating device is located at the top of the electronic device, one sound-generating device is located at the bottom of the electronic device, and at least one sound-generating device is located between the top and bottom of the electronic device, comprises: Acquire the multimedia resources to be played and extract the audio and video information; After identifying the motion sound source from the audio information, the audio information of the motion sound source is separated from the audio information; The image information is identified to determine the motion sound source information of the moving sound source; the motion sound source information includes shot category, sound source location information, and sound field information; the sound field information includes background sound pressure level, target signal-to-noise ratio, distance parameter, or room pulse sequence; the sound field information is used to characterize the environmental parameters of the moving sound source; wherein, identifying the image information to determine the motion sound source information of the moving sound source includes: identifying the image of the image information to determine the shot category and / or the sound source location information of the environment where the moving sound source is located; determining the distance parameter and the room pulse sequence based on the shot category and / or the sound source location information; determining the sound pressure level of the moving sound source and the background sound pressure level based on at least one of the shot category, the distance parameter, or the room pulse sequence; adjusting the background sound pressure level based on the shot category and / or the distance parameter to determine the target background sound pressure level; and determining the target signal-to-noise ratio based on the sound pressure level of the moving sound source and the target background sound pressure level. Based on the sound source location information of the moving sound source, a plurality of target sound-generating devices are determined from the at least three sound-generating devices; Target audio information is generated based on the audio information of the moving sound source and the sound field information; The system controls the target sound-emitting device to play the target audio information and controls the display screen to play the video information of the multimedia resources.
2. The method according to claim 1, characterized in that, Before generating target audio information based on the audio information of the moving sound source and the sound field information, the method further includes: Based on the marking information of the moving sound source, the range of motion angles of the moving sound source is determined; the marking information is used to characterize the sound source type, the correspondence between the sound source motion time and the sound source position; Based on the range of motion angles of the moving sound source, the audio information of the moving sound source is rendered to obtain rendered audio information, which includes audio information to be played from different target sound-emitting devices.
3. The method according to claim 2, characterized in that, The step of rendering the audio information of the moving sound source based on its motion angle range to obtain rendered audio information includes: Based on the sound source location information of the moving sound source and the location information of the target sound-emitting device, determine the gain coefficient of each target sound-emitting device for playing the audio information of the moving sound source; Based on the gain coefficient corresponding to the target sound-emitting device, the audio information of the moving sound source is adjusted to determine the audio information to be played by the target sound-emitting device.
4. The method according to claim 2, characterized in that, The step of rendering the audio information of the moving sound source according to the range of its motion angle to obtain the rendered audio information further includes: Based on the sound source location information of the moving sound source and the location information of the target sound-emitting device, determine the gain coefficient of the audio information of the moving sound source output by the target sound-emitting device; The gain coefficients corresponding to the target sound-emitting devices are added together and then normalized to obtain the normalized gain coefficients. The audio information of the moving sound source is adjusted according to the normalized gain coefficient to determine the audio information to be played by the target sound-emitting device.
5. The method according to any one of claims 2-4, characterized in that, Before determining the motion angle range of the moving sound source based on the marking information of the moving sound source, the method further includes: Based on the sound source type, sound source movement time, and sound source location information corresponding to the moving sound source, the marking information of the moving sound source is generated.
6. The method according to any one of claims 1-4, characterized in that, The motion sound source information also includes sound source spatial information. The step of identifying the image information to determine the motion sound source information includes: The image information is identified to determine the scene category, sound source location information, and sound source spatial information of the environment in which the moving sound source is located; The scene category, sound source location information, and sound source spatial information are input into the matching database, and the sound field information corresponding to the moving sound source is output through the matching database. The matching database stores the correspondence between multiple moving sound source information and sound field information.
7. The method according to claim 1, characterized in that, The step of generating target audio information based on the audio information of the moving sound source and the sound field information includes: Based on the target signal-to-noise ratio, the loudness values of the audio information of the moving sound source and the background audio information are adjusted; The target audio information is obtained by mixing the audio information of the moving sound source after adjusting the loudness value with the background audio information.
8. The method according to any one of claims 1-4, characterized in that, The target sound-generating device is a preset number, and the step of determining the target sound-generating device from the at least three sound-generating devices based on the sound source location information of the moving sound source includes: Based on the sound source location information of the moving sound source, calculate the distance value between the moving sound source and each sound-emitting device; Based on the distance between the moving sound source and each sound-generating device, a preset number of sound-generating devices are determined from the at least three sound-generating devices as the target sound-generating devices.
9. The method according to any one of claims 1-4, characterized in that, The step of determining the target sound-generating device from the at least three sound-generating devices based on the sound source location information of the moving sound source includes: Based on the sound source location information of the moving sound source, calculate the distance value between the moving sound source and each sound-emitting device; For each sound-generating device, if the distance between the moving sound source and the sound-generating device is less than a distance threshold, then the sound-generating device is determined to be the target sound-generating device.
10. The method according to any one of claims 1-4, characterized in that, The at least three sound-generating devices are all loudspeakers, or the at least three sound-generating devices are all screen sound-generating devices, or the at least three sound-generating devices are a combination of the loudspeakers and the screen sound-generating devices.
11. An electronic device, characterized in that, include: At least three sound-generating devices, which are used to play audio information; The display screen is used to display screen information; One or more processors; Memory; The memory stores one or more computer programs, the one or more computer programs including instructions that, when executed by the electronic device, cause the electronic device to perform the multimedia playback method as described in any one of claims 1-10.
12. A computer-readable storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the multimedia playback method as described in any one of claims 1-10.
Citation Information
Patent Citations
Audio playing method and electronic equipment
CN116347320A
Device and method for enhancing sound quality of video
WO2022059869A1