Delivery method, electronic equipment and related device
The first electronic device sends audio and loudness information to the second device, adjusts the audio loudness of the second device, solves the problem of loudness inappropriate between devices and improves the user experience.
Patent Information
- Application Number
- CN202410217993.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-08-26
AI Technical Summary
When projecting screens or sounds between devices, the loudness of video or audio played by device 2 is too large or too small, resulting in poor user experience.
The first electronic device transmits audio information and loudness information to the second electronic device by receiving the delivery command, adjusts the audio loudness of the second electronic device so that the loudness difference between it and the first electronic device is less than or equal to the threshold value, ensuring adaptation to the user.
It realizes the adaptation of audio loudness between devices and improves user experience.
Smart Images

Figure CN120547397A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of terminal technology, and in particular to a delivery method, electronic equipment, and related devices. Background Art
[0002] As electronic devices become increasingly versatile, they can now cast screens or audio. For example, device 1 can cast screens or audio to device 2, and device 2 can play the video or audio from device 1.
[0003] Currently, the loudness of video or audio played by device 1 is adapted to the user. However, when device 1 projects the screen or sound to device 2, the loudness of the video or audio played by device 2 is too loud or too soft, resulting in a poor user experience. Summary of the Invention
[0004] The embodiments of the present application provide a delivery method, electronic device, and related apparatus, which can ensure that the loudness of the video or audio played by the delivery device is adapted to the user in a delivery scenario, thereby improving the user experience.
[0005] In a first aspect, embodiments of the present application provide a delivery method. The execution entity of the delivery method may be a first electronic device, or a chip or processor in the first electronic device. The following description uses the first electronic device as an example. In this method, the first electronic device receives a first delivery instruction, which instructs the delivery of audio to a second electronic device. In response to the first delivery instruction, the first electronic device may send first audio information and first loudness information to the second electronic device.
[0006] The first audio information is used to indicate the first audio, and the first audio information may include the first audio and / or a download address of the first audio. The first audio information is used by the second electronic device to obtain the first audio so that the second electronic device can play the first audio.
[0007] The first loudness information is used to indicate a first loudness when the second electronic device plays the first audio. Specifically, the first loudness information is used to instruct the second electronic device to play the first audio at the first loudness based on information about the first audio. The difference between the first loudness and the second loudness is less than or equal to a first threshold, and the second loudness is the loudness when the first electronic device plays the first audio.
[0008] In some embodiments, the first loudness is a preset loudness adapted to the user. In some embodiments, the first loudness may be a loudness calculated by the first electronic device based on the loudness of the audio played by the second electronic device, and reference may be made to the relevant description in the following possible implementations.
[0009] In an embodiment of the present application, in a scenario where a first electronic device plays audio or video to a second electronic device, the first electronic device can instruct the second electronic device to play the first audio at a first loudness. In this way, the difference between the loudness of the audio played by the second electronic device and the loudness of the audio played by the first electronic device is less than or equal to a first threshold. It can be said that the loudness of the audio played by the second electronic device is consistent with the loudness of the audio played by the first electronic device, is adapted to the user, and is neither too loud nor too soft, thereby improving the user experience.
[0010] In one possible implementation, in response to the first delivery instruction, the first electronic device may first send information about a second audio file to the second electronic device. The information about the second audio file is used to instruct the second electronic device to play the second audio file. The information about the second audio file is used to indicate the second audio file, and the information about the second audio file may include the second audio file and / or a download address for the second audio file. The information about the second audio file is used by the second electronic device to obtain the second audio file so that the second electronic device can play the second audio file.
[0011] The first electronic device can record the second audio played by the second electronic device to obtain a third audio, where the loudness of the third audio is the third loudness. In some embodiments, for example, the first electronic device can be configured with a hardware module for obtaining the loudness of audio. While the first electronic device is recording the second audio, the first electronic device can use this hardware module to obtain the loudness of the third audio. In some embodiments, the first electronic device can also obtain the loudness of the third audio by matching the second audio with the third audio. For details, see the relevant description in the following possible implementations.
[0012] When the difference between the third loudness and the fourth loudness is greater than the first threshold, the first electronic device sends the first loudness information to the second electronic device. In other words, the first electronic device may first send the second audio information to the second electronic device and record the second audio played by the second electronic device to obtain the third audio. The first electronic device then compares the loudness of the third audio with the difference between the fourth loudness. When the difference between the third and fourth loudness is greater than the first threshold, the first electronic device may send the first loudness information to the second electronic device so that the second electronic device can play the first audio at the first loudness when playing the first audio.
[0013] In some embodiments, the fourth loudness is a preset loudness; or, the fourth loudness is related to the loudness of the audio played by the first electronic device in the past. Alternatively, in some embodiments, the fourth loudness is the loudness of the second audio played by the first electronic device.
[0014] In summary, the fourth loudness is a loudness adapted to the user. In the embodiment of the present application, the purpose of the first electronic device comparing the difference between the loudness of the third audio and the fourth loudness is to determine the difference between the loudness of the second audio played by the second electronic device and the fourth loudness. When the difference is greater than the first threshold, it indicates that the loudness of the audio played by the second electronic device is too loud or too soft and is not adapted to the user. In this way, the first electronic device can enable the first loudness of the first audio played by the second electronic device to be adapted to the user by sending the first loudness information to the second electronic device. It is understandable that when the difference is less than or equal to the first threshold, it indicates that the loudness of the audio played by the second electronic device is adapted to the user, and the second electronic device does not need to adjust the loudness of the audio.
[0015] In some embodiments, the first electronic device may initiate recording in response to the delivery instruction and stop recording after a first preset duration, where the first preset duration is greater than the duration of the second audio. Alternatively, the first electronic device may initiate recording upon successful connection with the second electronic device and stop recording after the first preset duration. It should be understood that this connection between the first electronic device and the second electronic device can transmit the first audio.
[0016] In this implementation, the start time for the first electronic device to record the second audio played by the second electronic device can be flexibly set, and the second audio played by the second electronic device can be recorded, which can ensure that the first electronic device determines the difference between the loudness of the audio played by the first electronic device and the loudness of the audio played by the second electronic device.
[0017] The above embodiment describes how a first electronic device can send first loudness information to a second electronic device, so that the second electronic device can play the first audio at a first loudness that is adapted to the user. The following describes possible implementations of the first loudness information:
[0018] Method 1: When the difference between the third loudness and the fourth loudness is greater than the first threshold, the first electronic device may adjust the amplitude of the first audio. The first loudness information includes the first audio after the amplitude is adjusted.
[0019] It should be understood that the amplitude of audio is related to the energy of the audio, and the energy of audio is related to the loudness of the audio. When the third loudness of the second audio played by the second electronic device is larger or smaller, the first electronic device can adjust the amplitude of the first audio so that the loudness is moderate when the first audio is played using this amplitude. The first electronic device can send the first audio with the adjusted amplitude to the second electronic device. In this way, when the second electronic device plays the first audio, because the amplitude of the first audio is adjusted, the loudness of the first audio played by the second electronic device will not be too loud or too soft compared to the third loudness, and can be adapted to the user.
[0020] Method 2: When the difference between the third loudness and the fourth loudness is greater than the first threshold, the first electronic device may adjust the volume of the first electronic device. The first loudness information includes the volume of the first electronic device.
[0021] The first electronic device adjusts the volume of the first electronic device. In the projection scenario, when the first electronic device adjusts its own volume, the operation can also be synchronized to the second electronic device. In other words, when the first electronic device adjusts its own volume, the first electronic device can send first loudness information to the second electronic device, and the first loudness information includes the volume of the first electronic device.
[0022] In other words, when the first electronic device adjusts the volume of the first electronic device, the second electronic device will also adjust the volume of the second electronic device based on the operation of adjusting the volume of the first electronic device. In this way, by adjusting the volume of the first electronic device, the second electronic device can adjust the volume of the second electronic device to an appropriate volume. Because the volume and loudness of the second electronic device are mapped to each other, the second electronic device can adjust the volume of the second electronic device to an appropriate volume, so that the loudness of the audio played by the second electronic device can be adapted to the user.
[0023] Mode three: the first loudness information is specifically used to instruct the second electronic device to adjust the audio loudness to the first loudness.
[0024] In some embodiments, the first loudness information includes the first loudness. Accordingly, upon receiving the first loudness information, the second electronic device may adjust the loudness of the audio to the first loudness, so that when the second electronic device plays the first audio, the loudness of the first audio is the first loudness.
[0025] For example, the second electronic device may pre-store a mapping relationship between the volume and loudness of the second electronic device. The second electronic device may adjust the volume of the second electronic device to the volume mapped to the first loudness based on the first loudness and the mapping relationship between the volume and loudness. In this way, the second electronic device may play the first audio at the volume mapped to the first loudness, and the loudness of the first audio is the first loudness.
[0026] The embodiments of the present application provide multiple methods for adjusting the loudness of the first audio played by the second electronic device, which are easy to implement and enable the loudness of the first audio played by the second electronic device to adapt to the user.
[0027] In one possible implementation, while the second electronic device is playing a first audio signal, the first electronic device may record the first audio signal played by the second electronic device to obtain a fourth audio signal, where the loudness of the fourth audio signal is the fifth loudness. For example, the first electronic device may record the first audio signal played by the second electronic device while the first electronic device is moving; alternatively, the first electronic device may periodically record the first audio signal played by the second electronic device.
[0028] When the difference between the fifth loudness and the second loudness is greater than a first threshold, the first electronic device may send second loudness information to the second electronic device, where the second loudness information is used to instruct the second electronic device to play the first audio at the first loudness. In this way, the second electronic device may play the first audio at the first loudness in response to the second loudness information. Thus, the difference between the loudness of the first audio played by the second electronic device (the first loudness) and the second loudness is less than or equal to the first threshold.
[0029] In this implementation, after the first electronic device plays the first audio to the second electronic device, the first electronic device may continue to detect the difference between the loudness of the first audio played by the second electronic device and the second loudness, so that the first electronic device can send second loudness information to the second electronic device based on the loudness difference, and promptly adjust the loudness of the first audio played by the second electronic device to the first loudness.
[0030] In one possible implementation, the first electronic device can adjust the loudness of the audio played by the second electronic device when switching the audio being played, and enable the loudness of the audio played by the second electronic device to be consistent with the loudness of the audio played by the first electronic device. In this scenario, the first electronic device can receive an instruction to switch audio. The instruction to switch audio is used to indicate the switching of the audio being played. The instruction can be triggered by the user or actively triggered after an audio is played. This embodiment of the present application does not limit this. After the first electronic device receives the instruction to switch audio, it can trigger the second electronic device to play the fifth audio. The fifth audio played by the second electronic device can be the sixth loudness, and the difference between the sixth loudness and the seventh loudness is less than or equal to the first threshold.
[0031] In some embodiments, the seventh loudness can refer to the description of the fourth loudness. In some embodiments, the seventh loudness is the loudness of the first electronic device when playing the fifth audio.
[0032] In one possible implementation, after a first electronic device plays audio to a second electronic device, the user can trigger a switch in the audio playback device. For example, the user can trigger the first electronic device to play audio to a third electronic device. In this scenario, the first electronic device can receive a second playback instruction, which instructs the first electronic device to play audio. The second playback instruction can refer to the description of the first playback instruction.
[0033] In response to the second delivery instruction, the first electronic device can trigger the third electronic device to play the first audio. In response to the second delivery instruction, the first electronic device can send information about the first audio and third loudness information to the third electronic device. When the third electronic device switches to playing the first audio, the third electronic device can continue playing the first audio, following the playback progress of the second electronic device. In response, the information about the first audio is used to indicate the audio that follows the first audio played by the second electronic device.
[0034] The third loudness information is used to instruct the third electronic device to play the first audio at the eighth loudness based on the information about the first audio. The third loudness information can refer to the description related to the first loudness information. In this way, in response to the third loudness information, the third electronic device can play the first audio at the eighth loudness based on the information about the first audio.
[0035] The difference between the eighth loudness and the first loudness is less than or equal to the first threshold. In other words, when the first electronic device receives the second delivery instruction, it can adopt the delivery method provided in the embodiment of the present application. For example, the first electronic device records the first audio or prompt sound (such as the second audio) played by the third electronic device, obtains the loudness difference between the audio played by the second electronic device and the third electronic device, and then adjusts the loudness of the audio played by the third electronic device so that the loudness of the audio played by the third electronic device is consistent with the loudness of the audio played by the second electronic device.
[0036] The following describes a method in which the first electronic device obtains the third loudness of the third audio and the first electronic device plays the fourth loudness of the second audio:
[0037] In some embodiments, the first electronic device may match the second audio with the third audio. When the second audio and the third audio are successfully matched, the first electronic device obtains the loudness of the third audio and the loudness of the second audio when the first electronic device plays the second audio.
[0038] In one possible implementation, the first electronic device matching the second audio with the third audio includes: the first electronic device acquiring a first audio feature of the second audio and a second audio feature of the third audio. The first electronic device may match the second audio with the third audio based on the first audio feature and the second audio feature.
[0039] In some embodiments, taking the first audio feature as an example, the first audio feature may include, but is not limited to, audio content, audio amplitude over a period of time, and frequency. In some embodiments, for example, when the similarity between the first audio feature and the second audio feature is greater than or equal to a similarity threshold, the second audio and the third audio are successfully matched; and when the similarity between the first audio feature and the second audio feature is less than the similarity threshold, the second audio and the third audio are unmatched.
[0040] In some embodiments, taking the first audio feature as an example, the first audio feature may include but is not limited to: short-time Fourier transform feature, Mel-frequency cepstral coefficient, or filter bank feature.
[0041] In this embodiment, the first audio feature and the second audio feature are represented in the form of a spectrum, the first audio feature corresponds to a first spectrum, and the second audio feature corresponds to a second spectrum. The first electronic device can match the second audio with the third audio by referring to the following description:
[0042] The first electronic device may divide the first spectrum into N regions and the second spectrum into N regions in the same division manner, where N is an integer greater than 1. The first electronic device obtains a first anchor point for each region in the first spectrum and a second anchor point for each region in the second spectrum, where the anchor point is a point corresponding to maximum energy in the region.
[0043] The first electronic device matches the first anchor point and the second anchor point. When the number of matched anchor points is greater than a second threshold, the second audio and the third audio are determined to be successfully matched. In one possible implementation, the first spectrum and the second spectrum are both used to represent the relationship between the time, frequency, and energy of the audio. When the frequency of the first anchor point is the same as the frequency of the second anchor point, and the difference between the time of the first anchor point and the time of the second anchor point is within a second preset duration, the first anchor point and the second anchor point are determined to be successfully matched.
[0044] First, a method for a first electronic device to obtain the loudness of a third audio signal is introduced:
[0045] The successfully matched second anchor point is designated as the second target anchor point. The second electronic device may obtain the energy sum of the second anchor points within a preset range of the second target anchor point, and perform a weighted summation on the energy sum based on the frequency of the second target anchor point and the mapping relationship between frequency and weight to obtain the energy of the second target anchor point. The energy of the second target anchor point is used to represent the loudness of the third audio.
[0046] Second, a method for the second electronic device to obtain the loudness of the second audio played by the first electronic device is introduced:
[0047] The first anchor point that is successfully matched is the first target anchor point. The first electronic device can obtain the energy mean of the first anchor points within the preset range of the first target anchor point, and perform a weighted sum of the energy sums based on the frequency of the first target anchor point and the mapping relationship between the frequency and the weight to obtain the energy of the first target anchor point.
[0048] Because the first electronic device may not actually be playing the first audio, the first electronic device can obtain the energy of the first target anchor point. The energy of the first target anchor point is the energy represented in the first spectrum, not the actual loudness of the first electronic device playing the first audio. In this embodiment of the present application, the first electronic device can obtain the loudness of the first audio when the first electronic device is playing the first audio based on the energy of the first target anchor point.
[0049] Exemplarily, the first electronic device stores a first mapping relationship between the volume, the energy of the first target anchor point, and the loudness when the first electronic device plays audio. The first electronic device can obtain the loudness, such as the fourth loudness, when the first electronic device plays the second audio based on the energy of the first target anchor point, the volume of the first electronic device, and the first mapping relationship.
[0050] In one possible implementation, the second audio and the first audio are continuous audio in a segment; or, the second audio is: audio used to indicate that the first electronic device and the second electronic device are successfully connected, and the first audio is audio projected from the first electronic device to the second electronic device.
[0051] Among them, when the second audio and the first audio are continuous audio in a segment of audio, the first electronic device can also play the second audio before receiving the first delivery instruction. In this way, the first electronic device can respond to the first delivery instruction and send the second audio information to the second electronic device so that the second electronic device can play the second audio.
[0052] Among them, when the second audio is: audio used to indicate that the first electronic device and the second electronic device are successfully connected, and the first audio is audio cast by the first electronic device to the second electronic device, before the first electronic device receives the first cast instruction, the first electronic device can also play the first audio, so that the first electronic device responds to the first cast instruction. In this way, the first electronic device can send information about the second audio to the second electronic device in response to the first cast instruction so that the second electronic device can play the second audio. In this example, the second audio can be a prompt tone.
[0053] In this implementation, while the first electronic device is playing audio, the user can trigger the first electronic device to broadcast audio or video to the second electronic device. In other words, the broadcast method provided in this embodiment of the application can be applied to scenarios where the first electronic device is playing audio or not playing audio, and this embodiment of the application does not limit this.
[0054] In a second aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory is used to store code instructions, and the processor is used to run the code instructions to execute the method described in the first aspect or any possible implementation of the first aspect.
[0055] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run on a computer, the computer executes the method described in the first aspect or any possible implementation of the first aspect.
[0056] In a fourth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when the computer program is run on a computer, enables the computer to execute the method described in the first aspect or any possible implementation of the first aspect.
[0057] In a fifth aspect, the present application provides a chip or chip system, comprising at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected via a line, and the at least one processor is configured to execute a computer program or instruction to perform the method described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip may be an input / output interface, a pin, or a circuit.
[0058] In one possible implementation, the chip or chip system described above in this application further includes at least one memory, in which instructions are stored. The memory may be a storage unit within the chip, such as a register, a cache, etc., or a storage unit of the chip (e.g., a read-only memory, a random access memory, etc.).
[0059] It should be understood that the second to fifth aspects of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic diagram of a scenario applicable to the delivery method provided in an embodiment of the present application;
[0061] Figure 2 A schematic diagram of an interface for setting the volume for the current user;
[0062] Figure 3 A flowchart of an embodiment of the delivery method provided in the embodiments of the present application;
[0063] Figure 4 A schematic diagram of a scenario for delivering audio provided in an embodiment of the present application;
[0064] Figure 5A A flowchart of another embodiment of the delivery method provided in the embodiment of the present application;
[0065] Figure 5B A flowchart of another embodiment of the delivery method provided in the embodiment of the present application;
[0066] Figure 6 A schematic diagram of anchor point matching provided in an embodiment of the present application;
[0067] Figure 7 A flowchart of another embodiment of the delivery method provided in the embodiment of the present application;
[0068] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0069] For ease of understanding, the following first introduces the relevant terms and concepts involved in the embodiments of the present application:
[0070] 1. Screen Casting: Device 1 casts its screen to device 2, or device 1 casts its screen to device 2. Screen casting can be understood as playing a video from device 1 on device 2, where the video may include audio. For example, when device 1 plays video 1 and device 1 casts its screen to device 2, device 2 can play video 1. In some embodiments, device 2 can continue playing video 1 at the same pace as device 1, or it can start playing video 1 from the beginning, although this is not a limitation in the present embodiment.
[0071] 2. Sound casting: Device 1 casts sound to device 2, or it can be said that device 1 casts sound to device 2. Sound casting can be understood as: playing the audio from device 1 on device 2 instead of playing the video image. For example, when device 1 plays audio 1 and device 1 casts sound to device 2, device 2 can play audio 1. In some embodiments, device 2 can continue to play audio 1 according to the playback progress of audio 1 played by device 1, or device 2 can play audio 1 from the beginning, and the embodiments of the present application do not limit this. In the following embodiments, "device 2 continues to play audio 1 according to the playback progress of audio 1 played by device 1" is used as an example for explanation.
[0072] 3. Casting: Casting in the embodiments of this application includes screen casting and sound casting. That is, in the embodiments of this application, casting from device 1 to device 2 can be understood as device 1 casting audio to device 2, or as device 1 casting video to device 2. In both scenarios, device 2 can play the audio from device 1. This embodiment of this application focuses on the loudness of the audio played from device 1 by device 2, and does not restrict the screen casting or sound casting scenario.
[0073] 4. Volume: refers to the relative value of the volume on the electronic device. In some embodiments, the volume of the current electronic device can be normalized to 0-100%, for example, the volume of the electronic device is adjusted to 50%.
[0074] 5. Loudness: The actual volume of the sound heard by the user. In some embodiments, the loudness of the audio played by the electronic device can be measured by a sound pressure meter and calculated.
[0075] If the first electronic device and the second electronic device have the same loudness, the user hears the audio played by the two electronic devices at the same volume. However, due to hardware differences between different electronic devices, even if the volume of the two electronic devices is the same, the loudness of the audio played by the two electronic devices is different, and the user hears the audio played by the two electronic devices at different volumes.
[0076] For example, the hardware of the mobile phone and the speaker (such as speakers, etc.) are different. The maximum loudness supported by the mobile phone is 40 phons and the volume is 0-100%. The loudness of the mobile phone and the volume of the mobile phone are in a first mapping relationship. The maximum loudness supported by the speaker is 80 phons and the volume is 0-100%. The loudness of the speaker and the volume of the speaker are in a second mapping relationship. Because the maximum loudness supported by the speaker is greater than the maximum loudness supported by the mobile phone, when the volume of the mobile phone and the speaker is 50%, the loudness of the audio 1 played by the mobile phone is less than the loudness of the audio 1 played by the speaker. In this way, even if the volume of the two electronic devices is the same, but the actual loudness of the same audio is different, the volume of the audio played by the two electronic devices heard by the user is also different.
[0077] 6. The delivery method provided in the embodiments of the present application is applicable to scenarios in which a first electronic device delivers audio or video to a second electronic device, and scenarios in which, after the first electronic device delivers audio or video to the second electronic device, the audio or video is switched to a third electronic device for delivery.
[0078] For example, taking the first electronic device as a mobile phone and the second electronic device as a speaker, the delivery method provided in the embodiments of the present application is applicable to a scenario where the mobile phone delivers audio or video to the speaker. Alternatively, taking the first electronic device as a mobile phone and the second electronic device as a vehicle computer, the delivery method provided in the embodiments of the present application is applicable to a scenario where the mobile phone delivers audio or video to the vehicle computer.
[0079] For example, taking the first electronic device as a mobile phone, the second electronic device as a speaker, and the third electronic device as a smart screen as an example, the delivery method provided in the embodiment of the present application is applicable to the scenario where the mobile phone delivers audio or video to the speaker and then switches to the smart screen to deliver audio or video.
[0080] In the embodiment of the present application, the first electronic device may be an electronic device with a projection function, for example, the first electronic device may include but is not limited to: a mobile phone, a personal digital assistant (PDA), a handheld device with a wireless communication function, a computing device or a wearable device, etc., and the embodiment of the present application does not specifically limit the form of the first electronic device. Exemplarily, the second electronic device is an electronic device with an audio playback function, for example, the second electronic device may include but is not limited to: a mobile phone, a speaker, a car device, or a wearable device, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in a smart home, etc., and the embodiment of the present application does not specifically limit the form of the second electronic device.
[0081] In some embodiments, the first electronic device and the second electronic device may be referred to as user equipment (UE), a terminal, or the like. The first electronic device and the second electronic device may be interconnected via a communication network to enable audio and video delivery. The communication network may be, but is not limited to, a short-range communication network such as a Wi-Fi hotspot network, a Wi-Fi peer-to-peer (P2P) network, a Bluetooth network, a ZigBee network, or a near-field communication (NFC) network.
[0082] Figure 1A schematic diagram of a scenario in which the delivery method provided in the embodiment of the present application is applicable. Figure 1 , this scenario may include: a first electronic device 11 and a second electronic device 12. Figure 1 The following description is made by taking the first electronic device 11 as a mobile phone and the second electronic device 12 as a speaker as an example. In some embodiments, for example, the second electronic device 12 may also be a television, a car-mounted device, etc.
[0083] Take the example of playing audio from a mobile phone to a speaker, refer to Figure 1 In a, the mobile phone can play audio 1. The user operates the mobile phone to trigger the mobile phone to play audio 1 to the speaker. Figure 1 In b, after the audio is successfully delivered, the speaker can play audio 1.
[0084] Among them, the user operates the mobile phone to trigger the mobile phone to play audio 1 to the speaker. For example, a play control can be displayed on the user interface of the mobile phone. The user operates the play control to trigger the mobile phone to play audio 1 to the speaker. For example, the user can also use voice commands to instruct the mobile phone to play audio 1 to the speaker. The embodiment of the present application does not limit the way in which the user triggers the first electronic device 11 to play audio and video to the second electronic device 12. It should be understood that Figure 1 The example a in the figure takes the delivery control on the user interface of a mobile phone operated by the user as an example.
[0085] For example, refer to Figure 1 In a, when the mobile phone plays audio 1, the loudness is moderate (such as the loudness of audio 1 is loudness 1), which is adapted to the user and the user experience is high. When the mobile phone successfully delivers audio to the speaker, according to the current delivery method, the speaker will play audio 1 from the mobile phone at the default volume or the volume of the last audio playback. If the default volume or the volume of the last audio playback is too high or too low, the loudness of the speaker when playing audio 1 will also be too high or too low, resulting in a poor user experience. For example, refer to Figure 1 In b, when the speaker plays audio 1, the loudness of audio 1 is loudness 3, which is much louder than loudness 1 and is too loud.
[0086] In some embodiments, when a mobile phone plays Audio 1 to a speaker, it can send the phone's volume to the speaker, so that the speaker can adjust its volume to match that of the phone's speaker and play Audio 1. However, based on the above definition of loudness, due to differences in hardware between the mobile phone and the speaker, even if the volume of the mobile phone and the speaker is the same, the loudness of the same audio (Audio 1) played by the two may be different. This may cause the loudness of Audio 1 played by the speaker to be too loud or too soft after the playback, resulting in a poor user experience.
[0087] In some embodiments, a volume management function may be configured in the second electronic device, for example, the volume management function may be configured in an application (APP) of the second electronic device. In this example, the user can manually set the volume of the audio played by the second electronic device from other electronic devices in the APP, that is, when other electronic devices play audio to the second electronic device, the second electronic device can play the audio at the volume set by the user, and can adapt to the user. For example, the electronic device that plays audio to the second electronic device may include a first electronic device. When the first electronic device plays audio to the second electronic device, the second electronic device can play the audio at the volume set by the user, which can avoid the problem of the second electronic device playing the audio too loud or too soft after the play.
[0088] For example, refer to Figure 2 , the user interface of the APP may include an electronic device identifier 21 and a volume setting control 22. The electronic device identifier 21 is used to indicate: an electronic device that plays audio to a second electronic device. The user operates the volume setting control 22 to set the volume. For example, Figure 2 In , the user sets the volume of the audio played by speaker 1 when the mobile phone is connected to speaker 1 and audio is cast to speaker 1.
[0089] However, in this example, the user needs to set the volume of each electronic device that plays audio to the second electronic device, which is complicated. In addition, some electronic devices are not pre-configured with a volume management function, so the audio playback loudness during playback cannot be adjusted.
[0090] In response to the above problem, an embodiment of the present application provides a delivery method. In a scenario where a first electronic device delivers audio or video to a second electronic device, the loudness of the audio played by the second electronic device can be adjusted so that the difference between the loudness of the audio played by the second electronic device and the loudness when the first electronic device plays the audio is less than or equal to a first threshold. That is, the difference between the loudness of the audio played by the second electronic device and the loudness when the first electronic device plays the audio is small. In this way, in the delivery scenario, the loudness of the audio played by the second electronic device can be adapted to the user, thereby improving the user experience.
[0091] The following describes the delivery method provided by the embodiments of the present application in conjunction with specific embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. Figure 3 A flowchart of an embodiment of the delivery method provided in the embodiment of the present application.
[0092] Reference Figure 3 The delivery method provided in the embodiment of the present application may include:
[0093] S301: A first electronic device receives a first delivery instruction.
[0094] In some embodiments, the user can operate the first electronic device to trigger the first electronic device to receive the first delivery instruction. For example, the user can operate the interface of the first electronic device or use voice commands to trigger the first electronic device to receive the first delivery instruction, which is not described in detail in the embodiments of the present application. In some embodiments, the first electronic device can also receive the first delivery instruction in specific scenarios. For example, when the user starts the vehicle and the mobile phone is successfully connected to the on-board device in the vehicle, the mobile phone can receive the first delivery instruction, such as the mobile phone can deliver audio and video to the on-board device.
[0095] The embodiment of the present application does not limit the form in which the first electronic device receives the first delivery instruction.
[0096] The first delivery instruction is used to instruct the first electronic device to deliver a video to the second electronic device. In some embodiments, the first electronic device can deliver audio or video to the second electronic device, and the video can also include audio.
[0097] In some embodiments, the first electronic device may receive the first delivery instruction while the audio is playing. For example, the audio may be the first audio or the second audio in the following embodiments, and reference may be made to the description in the following embodiments. In some embodiments, the first electronic device may receive the first delivery instruction while the audio is not playing.
[0098] S302: The first electronic device establishes a connection with the second electronic device.
[0099] S302 may refer to the process of establishing a connection between the first electronic device and the second electronic device in the current delivery scenario, which will not be described in detail in this embodiment of the present application.
[0100] S303: The first electronic device sends audio information to the second electronic device.
[0101] The audio information may include: audio or a download address of the audio. The audio information is used to indicate the audio, which can be understood as the audio played by the first electronic device to the second electronic device.
[0102] In some embodiments, for example, if the first electronic device locally stores audio or can receive audio from another electronic device, the first electronic device can send the audio to the second electronic device. In some embodiments, for example, if the first electronic device downloads the audio according to the audio download address, the first electronic device can send the audio download address to the second electronic device.
[0103] Exemplarily, for example, when the audio is call audio, the first electronic device can receive audio from the other end of the call (i.e., the electronic device that is talking to the first electronic device). In the scenario where the first electronic device is projecting the screen to the second electronic device, after the first electronic device receives the audio from the other end of the call, the first electronic device can send the call audio to the second electronic device.
[0104] For example, when the audio is a call ringtone audio, the first electronic device can pre-store the call ringtone audio locally. When the first electronic device receives a call request from the other party, in the scenario where the first electronic device casts the screen to the second electronic device, the first electronic device can send the call ringtone audio to the second electronic device.
[0105] For example, when the audio is music, video, etc. stored locally on the first electronic device, in a scenario where the first electronic device projects the screen to the second electronic device, the first electronic device can send the audio to the second electronic device.
[0106] Exemplarily, for example, when the first electronic device plays audio and video online, the first electronic device can download the audio and video according to the download address of the audio and video. In the scenario where the first electronic device casts the screen to the second electronic device, the first electronic device can send the download address of the audio to the second electronic device, and the second electronic device can download the audio according to the download address of the audio to play the audio.
[0107] S304: The second electronic device plays audio according to the audio information. The loudness of the audio is a first loudness. The difference between the first loudness and the second loudness is less than or equal to a first threshold. The second loudness is the loudness of the audio when the first electronic device plays the audio.
[0108] In some embodiments, when the audio information includes audio, the second electronic device can receive the audio from the first electronic device, and the second electronic device can play the audio. In this example, in a call scenario, if the first electronic device is playing the call audio, and the user triggers the first electronic device to cast the call audio to the second electronic device, the first electronic device can then send the subsequent audio from the other end of the call to the second electronic device, so that the second electronic device can continue to play the call audio. Therefore, in the casting scenario, the user can hear a coherent audio playback.
[0109] In some embodiments, when the audio information includes the download address of the audio, the second electronic device can download the audio according to the download address and play the audio. In this example, for example, the first electronic device is playing audio, and the user triggers the first electronic device to deliver the call audio to the second electronic device, then the audio information may also include a playback progress, which is used to indicate the progress of the first electronic device playing the audio. For example, the progress can be expressed in playback time, which is not described in detail in the embodiments of the present application. In this example, the second electronic device can continue to play the audio according to the playback progress, so that in the delivery scenario, the user can hear continuous audio playback.
[0110] It should be understood that in the embodiment of the present application, when the second electronic device plays audio according to the audio information, the loudness of the audio may be a first loudness, and the difference between the first loudness and the second loudness is less than or equal to the first threshold. The second loudness is the loudness of the audio when the first electronic device plays the audio.
[0111] In some embodiments, the audio played by the second electronic device can be considered the first audio. In this embodiment, S304 can also be understood as: the second electronic device plays the first audio, the loudness of the first audio is the first loudness, the difference between the first loudness and the second loudness is less than or equal to the first threshold, and the second loudness is the loudness of the first electronic device when playing the first audio.
[0112] For example, in the delivery method provided in the embodiment of the present application, Figure 4 In the example, the first electronic device is a mobile phone and the second electronic device is a speaker. Figure 4 In a, the mobile phone plays audio 1, and the loudness of audio 1 is loudness 1. The user operates the mobile phone to trigger the mobile phone to play audio 1 to the speaker. For example, the user can operate the play control on the mobile phone interface to trigger the mobile phone to play audio 1 to the speaker. Figure 4 In b, after the audio delivery is successful, the speaker can play Audio 1. At this time, the loudness of Audio 1 played by the speaker is Loudness 2. The difference between Loudness 1 and Loudness 2 is less than or equal to the first threshold. That is, in the delivery scenario, the loudness of Audio 1 played by the speaker is consistent with the loudness of Audio 1 played by the mobile phone.
[0113] It should be understood that in the embodiment of the present application, the loudness of the audio 1 played by the speaker is consistent with the loudness of the audio 1 played by the mobile phone, which can be understood as: the difference between the loudness of the audio 1 played by the speaker and the loudness of the audio 1 played by the mobile phone is less than or equal to the first threshold.
[0114] In some embodiments, the first electronic device may also send the first loudness information to the second electronic device. For example, the first electronic device may simultaneously send the audio information and the first loudness information to the second electronic device, or the first electronic device may separately send the audio information and the first loudness information to the second electronic device. The audio information may be considered as the first audio information, i.e., the first electronic device sends the first audio information and the first loudness information to the second electronic device.
[0115] In this embodiment, the second electronic device may play the first audio at the first loudness in response to the first audio information and the first loudness information. Accordingly, in this embodiment, S303 may be replaced by: the first electronic device sends the first audio information and the first loudness information to the second electronic device.
[0116] The first loudness information is used to instruct the second electronic device to play the first audio at the first loudness based on the information about the first audio. Exemplarily, the first loudness information includes the first loudness or the volume corresponding to the first loudness. In some embodiments, the first loudness is a pre-set loudness adapted to the user. In some embodiments, the first loudness may be a loudness calculated by the first electronic device based on the loudness of the audio played by the second electronic device, as described in the following embodiments.
[0117] Accordingly, S304 may be replaced with: in response to the first loudness information, the second electronic device plays the audio at a first loudness according to the information of the first audio, where the difference between the first loudness and the second loudness is less than or equal to the first threshold, and the second loudness is the loudness of the first electronic device when playing the audio.
[0118] In an embodiment of the present application, in a scenario where a first electronic device is projected onto a second electronic device, the difference between the loudness of the audio played by the second electronic device and the loudness of the audio played by the first electronic device is less than or equal to a first threshold value. It can be said that the loudness of the audio played by the second electronic device is consistent with the loudness of the audio played by the first electronic device, is adapted to the user, is not too loud or too soft, and can improve the user experience.
[0119] The following is a specific example of the delivery method provided by the embodiment of the present application:
[0120] Example 1:
[0121] In some embodiments, when the first electronic device and the second electronic device are successfully connected, the second electronic device may play a prompt tone to indicate that the first electronic device and the second electronic device are successfully connected. For example, the prompt tone may be an audio signal that the first electronic device and the second electronic device are successfully connected, or a piano ring tone, etc., which is not limited in this embodiment of the present application.
[0122] In this example, when the first electronic device casts audio to the second electronic device, the second electronic device can first play the prompt sound and then play the audio from the first electronic device. In other words, in response to the screen casting instruction, the first electronic device can trigger the second electronic device to play the prompt sound and play the audio from the first electronic device. In this example, in response to the screen casting instruction, the first electronic device can send the second electronic device information of the second audio, and the second audio information is used to indicate the prompt sound. The second electronic device can play the prompt sound in response to the second audio information. In addition, the first electronic device can also send the first audio information and the first loudness information to the second electronic device, so that the second electronic device can play the first audio at the first loudness based on the first audio information.
[0123] The prompt tone can be regarded as the second audio, and the audio from the first electronic device can be regarded as the first audio, that is, the second electronic device can play the second audio as well as the first audio. The audio from the first electronic device can be understood as the audio played by the first electronic device to the second electronic device.
[0124] In one embodiment, the second audio (such as a prompt tone) may be N seconds long. The first electronic device may pre-store the duration (N seconds) of the second audio.
[0125] Exemplarily, in this example, when the first electronic device and the second electronic device are successfully connected, the first electronic device can turn on the microphone and record the second audio (such as a prompt tone) played by the second electronic device to obtain a third audio. Because when the first electronic device and the second electronic device are successfully connected, the first electronic device also needs to send the audio information of the second audio to the second electronic device, or the first electronic device also needs to send an instruction to play the second audio to the second electronic device, so there is a delay Δt1 between "the first electronic device and the second electronic device are successfully connected" and "the second electronic device plays the second audio". Therefore, the first electronic device records the second audio played by the second electronic device and obtains the third audio, and the duration of the third audio is greater than the duration of the second audio (N seconds). For example, the duration of the third audio is (Δt1+N) seconds.
[0126] Exemplarily, in this example, after the first electronic device and the second electronic device are successfully connected, the first electronic device can turn on the microphone when sending the audio information of the second audio to the second electronic device, or when the first electronic device sends the instruction to play the second audio to the second electronic device, and record the second audio (such as a prompt tone) played by the second electronic device to obtain the third audio. In this way, there is a delay Δt2 between "the first electronic device sends the audio information of the second audio to the second electronic device (or the first electronic device sends the instruction to play the second audio to the second electronic device)" and "the second electronic device plays the second audio". Therefore, the first electronic device records the second audio played by the second electronic device and obtains the third audio, and the duration of the third audio is greater than the duration of the second audio (N seconds). For example, the duration of the third audio is (Δt2+N) seconds.
[0127] In summary, it can be summarized as follows: the first electronic device can turn on the microphone in response to the first delivery instruction, record the second audio (such as a prompt tone) played by the second electronic device, and obtain the third audio. Among them, the first electronic device can turn on the microphone when it receives the first delivery instruction. Alternatively, the first electronic device can turn on the microphone when the connection with the second electronic device is successful. Alternatively, the first electronic device can turn on the microphone when sending audio information of the second audio information to the second electronic device (or the first electronic device sends an instruction to play the second audio to the second electronic device). The embodiment of the present application does not limit the moment when the first electronic device turns on the microphone.
[0128] In some embodiments, the first electronic device may pre-store a prompt tone, and the first electronic device may identify the starting point and ending point of the prompt tone based on audio recognition, and then record the prompt tone played by the second electronic device.
[0129] Example 2:
[0130] In some embodiments, when the audio delivered by the first electronic device to the second electronic device is call audio, call ringtone audio, or other audio, this type of audio can be called the first type of audio, and the first type of audio delivery requires high real-time performance. For example, in a call scenario, the first electronic device delivers audio to the second electronic device. If the second electronic device plays the audio with a delay, the call is interrupted, which affects the user experience. Therefore, in this example, in order to improve the real-time delivery of the first type of audio, when the first electronic device and the second electronic device establish a connection, the second electronic device may not play the prompt tone, but directly play the audio from the first electronic device, thereby reducing the waiting time for the user.
[0131] In some embodiments, when the audio delivered by the first electronic device to the second electronic device is music, audio in a video stream, etc., this type of audio can be called the second type of audio, and the second type of audio delivery does not require high real-time performance. For example, in a music delivery scenario, the first electronic device delivers audio to the second electronic device. If the audio playback of the second electronic device is delayed, the user can still hear the music after waiting for a while, which will not affect the user experience. Therefore, in this example, when the first electronic device and the second electronic device establish a connection, the second electronic device can play a prompt tone first, and then play the audio from the first electronic device.
[0132] In the scenario where the second electronic device plays a prompt tone, the following settings may be provided:
[0133] In some embodiments, the first electronic device may pre-store information about a prompt tone, which may include a prompt tone (i.e., the audio of the prompt tone) or a download address for the prompt tone. When the first electronic device and the second electronic device are successfully connected, the first electronic device may send the prompt tone information to the second electronic device so that the second electronic device can play the prompt tone based on the prompt tone information.
[0134] For example, when the prompt tone information is a prompt tone, the second electronic device can play the prompt tone after receiving the prompt tone. For example, when the prompt tone information is a download address of the prompt tone, the second electronic device can download the prompt tone according to the download address after receiving the download address of the prompt tone, and then play the prompt tone.
[0135] In some embodiments, the first electronic device may pre-store information about the prompt tone, and the second electronic device may also pre-store information about the prompt tone. When the first electronic device and the second electronic device are successfully connected, the first electronic device may send an instruction to the second electronic device to play the prompt tone. Accordingly, upon receiving the instruction from the first electronic device, the second electronic device may play the prompt tone based on the prompt tone information stored in the second electronic device.
[0136] It should be understood that in the embodiment of the present application, when the first electronic device plays the first type of audio to the second electronic device, after the first electronic device and the second electronic device are successfully connected, the second electronic device can directly play the audio from the first electronic device without playing the prompt tone. In other words, the second electronic device can play the first audio. In this example, the first N seconds of the first audio can be regarded as the second audio. In other words, in this example, the first audio and the second audio are both audio played by the first electronic device to the second electronic device.
[0137] For example, taking the first type of audio as call audio, the first electronic device plays the call audio to the second electronic device, and the second electronic device plays the call audio. The second electronic device can play the call audio first and split the call audio into the first N seconds of audio (such as audio 1) and the audio after N seconds (such as audio 2).
[0138] In some embodiments, audio 1 can be considered the second audio, and audio 2 can be considered the first audio. Both the first and second audio are call audio, and the second audio is continuous with the first audio. In other words, the first electronic device can first send the information of audio 1 to the second electronic device, and then send the information of audio 2, so that the second electronic device can play audio 1 and audio 2.
[0139] For example, the call audio played by the first electronic device to the second electronic device is "Hello, today is xxxxxx. Appointment is tomorrow yyyyy". The second electronic device can play the second audio "Hello, today is xxxxxx", followed by the first audio "Appointment is tomorrow yyyyy". Alternatively, the second electronic device can play the second audio "Hello, today is x", followed by the first audio "xxxxx. Appointment is tomorrow yyyyy".
[0140] In some embodiments, when the first electronic device plays a second type of audio to the second electronic device, after the first electronic device and the second electronic device are successfully connected, the second electronic device can play a prompt tone and the audio from the first electronic device. In other words, the second electronic device can play the second audio and the first audio.
[0141] Exemplarily, the second audio may be a prompt tone, and the second audio is used to indicate that the first electronic device and the second electronic device are successfully connected. The second audio may be: audio, such as music, cast by the first electronic device to the second electronic device. Exemplarily, the second audio may be part of the audio in the prompt tone (such as the first x seconds of audio in the prompt tone), and correspondingly, the first audio may be part of the audio in the prompt tone (such as the audio after x seconds in the prompt tone), and the audio cast by the first electronic device to the second electronic device. Exemplarily, the second audio may be: the prompt tone, and the first y seconds of audio in the audio cast by the first electronic device to the second electronic device, and the first audio may be: audio after y seconds of audio in the audio cast by the first electronic device to the second electronic device.
[0142] For example, taking the second type of audio as music, the first electronic device plays Music 1 to the second electronic device, and the second electronic device can play a prompt tone, such as a piano ringtone, and Music 1. In other words, the first electronic device can first send the prompt tone (such as the second audio) information to the second electronic device, and then send the Music 1 (the first audio) information, so that the second electronic device can play the prompt tone and Music 1.
[0143] In Example 2, when a first electronic device plays audio to a second electronic device, the audio types are different. The second electronic device may play a prompt tone or not before playing the audio. To detect the loudness of the audio played by the second electronic device, the first electronic device may record the audio played by the first electronic device.
[0144] In this example, because the second electronic device can play a prompt tone or not, in this example, the first electronic device can respond to the first delivery instruction to start recording, and stop recording after recording the first preset duration, thereby obtaining the third audio. In this example, regardless of whether the second electronic device plays a prompt tone, the first electronic device can respond to the first delivery instruction to turn on the microphone to record the audio played by the second electronic device. The timing of turning on the microphone of the first electronic device can refer to the description in Example 1.
[0145] In this embodiment of the present application, to ensure that the first electronic device can record the prompt sound played by the second electronic device, the first preset duration can be set to the duration of the prompt sound. For example, if the second audio is N seconds long, the first preset duration can be set to N seconds. In this example, due to the delay in delivery, the third audio recorded by the first electronic device may include part of the first audio played by the second electronic device.
[0146] Alternatively, taking into account the delivery delay issue (such as Δt1 and Δt2), the first preset duration can be set to "duration of the prompt tone + delivery delay", and the delivery delay can refer to the description in Example 1. For example, if the second audio is N seconds, the first preset duration can be (N+Δt) seconds. In this way, when the second electronic device plays the prompt tone, the first electronic device can record the prompt tone in its entirety. In this example, taking into account the delivery delay issue, the third audio recorded by the first electronic device can include the second audio played by the second electronic device.
[0147] Δt varies depending on when the first electronic device turns on its microphone. For example, when the first electronic device and the second electronic device are successfully connected, the first electronic device may turn on its microphone, and Δt may be Δt1. When the first electronic device sends audio information to the second electronic device, the microphone may be turned on, and Δt may be Δt2. When the first electronic device receives the first delivery instruction, the microphone may be turned on, and Δt may be Δt3.
[0148] It is understandable that, generally speaking, the delivery delay is less than 1 second. Therefore, in some embodiments, after the first electronic device turns on the microphone, it can record (N+1) seconds of audio, which can ensure that the second audio played by the first electronic device is fully recorded. In the embodiment of the present application, the purpose of ensuring the complete recording of the second audio played by the first electronic device is to facilitate the subsequent calculation of the loudness and loudness difference of the second audio and the third audio by the first electronic device. Please refer to the description in the following embodiment.
[0149] Referring to Examples 1 and 2 above, a first electronic device can record a second audio played by a second electronic device to obtain a third audio. In an embodiment of the present application, the first electronic device can obtain the loudness of the third audio, for example, the loudness of the third audio is the third loudness. In some embodiments, for example, a hardware module for obtaining the loudness of audio can be configured in the first electronic device. When the first electronic device is recording the second audio, the first electronic device can use the hardware module to obtain the loudness of the third audio.
[0150] In addition, the first electronic device may also obtain a difference between the third loudness and the fourth loudness.
[0151] In some embodiments, the fourth loudness is a preset loudness that is suitable for most users. In this embodiment of the present application, the purpose of the first electronic device obtaining the difference between the third and fourth loudnesses is to determine whether the difference between the third and fourth loudnesses is too large, so that the loudness of the second audio played by the second electronic device differs significantly from the fourth loudness. If the loudness of the second audio played by the second electronic device differs significantly from the fourth loudness, the loudness of the audio played by the second electronic device can be adjusted to suit the user.
[0152] In some embodiments, the fourth loudness is related to the loudness of audio played historically by the first electronic device. The loudness of audio played historically by the first electronic device may represent a loudness adapted to the user. Thus, the purpose of the first electronic device obtaining the difference between the third and fourth loudnesses can be described above.
[0153] Exemplarily, the fourth loudness is related to the loudness of the audio played historically by the first electronic device. It can be understood that, for example, the fourth loudness is the loudness with the largest number among the loudnesses of the audio played historically by the first electronic device, or, for example, the fourth loudness is the average or median value of the loudness of the audio played historically by the first electronic device.
[0154] In some embodiments, the fourth loudness is the loudness of the second audio played by the first electronic device. In some embodiments, the fourth loudness may be preconfigured in the first electronic device, or calculated by the first electronic device according to loudness calculation logic configured in the first electronic device. The loudness calculation logic used to calculate the fourth loudness may refer to the description in the following embodiments.
[0155] In an embodiment of the present application, the first electronic device can determine whether to adjust the loudness of the audio played by the second electronic device based on the difference between the third loudness and the fourth loudness. For example, when the difference between the third loudness and the fourth loudness is less than or equal to the first threshold value, it indicates that the difference between the loudness of the second electronic device playing the second audio and the loudness of the first electronic device playing the second audio is small. Similarly, when the first electronic device plays audio (such as the first audio) to the second electronic device, the difference between the loudness of the second electronic device playing the first audio and the loudness of the first electronic device playing the first audio is also small. For example, the loudness of the first audio played by the second electronic device is the first loudness, and the loudness of the first electronic device playing the first audio is the second loudness, and the difference between the first loudness and the second loudness is less than or equal to the first threshold value. Among them, the difference between the first loudness and the second loudness is less than or equal to the first threshold value, indicating that the difference between the loudness of the first audio played by the second electronic device and the loudness of the first audio played by the first electronic device is small.
[0156] In other words, when the difference between the third loudness and the fourth loudness is less than or equal to the first threshold, it indicates that when the first electronic device plays audio to the second electronic device, the difference between the loudness of the audio played by the second electronic device and the loudness of the audio played by the second electronic device is also small, that is, the loudness of the audio played by the second electronic device is moderate, neither too loud nor too soft. Therefore, in this example, there is no need to adjust the loudness of the audio played by the second electronic device.
[0157] In some embodiments, when the difference between the third loudness and the fourth loudness is greater than the first threshold, it indicates that the loudness of the second audio played by the second electronic device is significantly different from the loudness of the second audio played by the second electronic device. Similarly, when the first electronic device plays audio (such as the first audio) to the second electronic device, the loudness of the first audio played by the second electronic device is significantly different from the loudness of the first audio played by the first electronic device. For example, the difference between the first loudness and the second loudness is greater than the first threshold.
[0158] In other words, when the difference between the third loudness and the fourth loudness is greater than the first threshold value, it indicates that when the first electronic device plays audio to the second electronic device, the difference between the loudness of the audio played by the second electronic device and the loudness of the audio played by the second electronic device is large, that is, the loudness of the audio played by the second electronic device is too loud or too small. Therefore, in this example, it is necessary to adjust the loudness of the audio played by the second electronic device. The first electronic device can send the first loudness information to the second electronic device, so that the second electronic device can adjust the audio loudness based on the first loudness information. After the second electronic device adjusts the loudness of the audio, when the second electronic device plays the first audio, the difference between the loudness of the first audio and the loudness when the first electronic device plays the first audio can be less than or equal to the first threshold value. For example, the loudness of the first audio is the first loudness, and the difference between the first loudness and the second loudness (that is, the loudness when the first electronic device plays the first audio) is less than or equal to the first threshold value.
[0159] The manner in which the second electronic device adjusts the loudness of the audio played by the second electronic device may refer to the description in the following embodiments.
[0160] In some embodiments, combining the above examples 1 and 2, refer to Figure 5A The delivery method provided in the embodiment of the present application may include the following steps:
[0161] Step 1: The first electronic device plays audio to the second electronic device.
[0162] Step 2: The second electronic device plays audio, and the first electronic device records the audio played by the second electronic device (such as the second audio, or the audio recorded within the first preset time length).
[0163] Step 3: The first electronic device determines the loudness (eg, the third loudness) of the audio played by the second electronic device based on the recorded audio.
[0164] Step 4: The second electronic device adjusts the loudness of the audio played by the second electronic device.
[0165] Combined with Example 2, Figure 5B The following describes the audio delivery method provided by the embodiment of the present application, which distinguishes the types of audio delivered by the first electronic device to the second electronic device:
[0166] Step 1A: The first electronic device determines whether there is a prompt sound based on the audio type.
[0167] Step 2A: When there is a prompt sound, the second electronic device plays the prompt sound, and then plays the audio projected by the first electronic device to the second electronic device.
[0168] Step 3A: When there is no prompt tone, the second electronic device plays the audio projected by the first electronic device to the second electronic device.
[0169] After step 2A and step 3A, you can perform the following:
[0170] Step 4A: The first electronic device records (N+1) seconds of audio played by the second electronic device, where N seconds is the duration of the second audio and 1 second is the maximum duration of the delivery delay set in the embodiment of the present application.
[0171] Step 5A: The first electronic device obtains a loudness difference between the first electronic device and the second electronic device based on the recorded audio and the original audio of the audio segment.
[0172] The original audio of the audio segment can be regarded as unplayed audio. In this embodiment of the present application, the first electronic device can obtain the loudness of the recorded audio (such as the third loudness) and the loudness of the first electronic device when playing the audio segment (the fourth loudness). The first electronic device can use the difference between the third loudness and the fourth loudness as the loudness difference between the first electronic device and the second electronic device.
[0173] Step 6A: The second electronic device adjusts the loudness to be consistent with the loudness of the first electronic device.
[0174] In the embodiment of the present application, after the second electronic device adjusts the loudness of the audio, the loudness of the second electronic device when playing the first audio is the first loudness, and the difference between the first loudness and the second loudness when the first electronic device plays the first audio is less than or equal to the first threshold. That is, the difference between the loudness of the audio played by the second electronic device and the loudness of the audio played by the first electronic device is small, and the audio played by the second electronic device is moderate and not too loud or too soft, which can improve the user experience.
[0175] The following describes a process in which the first electronic device obtains the third loudness of the third audio, the fourth loudness when the first electronic device plays the second audio, and the second electronic device adjusts the loudness of the audio:
[0176] In some embodiments, when the second audio is a prompt sound, because the first electronic device can store information about the prompt sound, the first electronic device can obtain the second audio according to the information about the prompt sound.
[0177] In some embodiments, when the second audio is audio cast from the first electronic device to the second electronic device, the first electronic device can store the second audio locally, or the first electronic device can obtain the second audio according to the download address of the second audio.
[0178] In this way, the first electronic device can obtain the second audio. In some embodiments, the second audio can be regarded as the original audio S ori .
[0179] Based on the above examples 1 and 2, the first electronic device can record the second audio played by the second electronic device to obtain a third audio, which can be regarded as the recorded audio S rec .
[0180] It should be understood that in the embodiments of the present application, the purpose of loudness measurement is to enable a first electronic device to obtain the difference between the loudness of the first electronic device when playing audio and the loudness of the second electronic device when playing audio, which can also be understood as the loudness difference between the first electronic device and the second electronic device. In the embodiments of the present application, the first electronic device can obtain the loudness difference between the first electronic device and the second electronic device based on the loudness of the first electronic device when playing the second audio and the loudness of the second electronic device when playing the second audio.
[0181] In some embodiments, because the third audio recorded by the first electronic device includes not only the second audio but also audio recorded during the delivery delay, the first electronic device can first match the second and third audios. The purpose of this matching is to align the timing of the second and third audios to facilitate analysis of the loudness difference between the second and third audios.
[0182] In an embodiment of the present application, when the second audio and the third audio are successfully matched, the first electronic device can obtain the loudness of the third audio and the loudness of the second audio played by the first electronic device. When the second audio and the third audio are not matched, the first electronic device does not need to obtain the loudness of the third audio and the loudness of the second audio played by the first electronic device.
[0183] 1. The following first describes the process of the first electronic device matching the second audio and the third audio:
[0184] In some embodiments, the first electronic device may use audio recognition to obtain the first audio feature of the second audio and the second audio feature of the third audio, respectively. When the similarity between the first audio feature of the second audio and the second audio feature of the third audio is greater than or equal to a third threshold, the first electronic device determines that the second audio and the third audio are successfully matched. Furthermore, when the similarity between the second audio feature of the second audio and the second audio feature of the third audio is less than the third threshold, the first electronic device determines that the second audio and the third audio are unmatched.
[0185] Taking the second audio as an example, the first audio feature of the second audio may include but is not limited to: audio content, amplitude and frequency of the audio within a period of time, etc.
[0186] In some embodiments, taking the second audio as an example, the first audio feature of the second audio may include but is not limited to: short time Fourier transform (STFT) feature, mel-frequency cepstral coefficient (MFCC), or filter bank (Fbank) feature.
[0187] For example, the first electronic device may perform spectrum calculation on the second audio and the third audio, respectively, to obtain a first STFT feature of the second audio and a second STFT feature of the third audio. In this example, the first electronic device may match the second audio and the third audio based on the first STFT feature and the second STFT feature.
[0188] In some embodiments, after the first electronic device converts the second audio into a first STFT feature and converts the second audio into a second STFT feature using STFT, the first electronic device may further perform Mel filtering on the first STFT feature and the second STFT feature, respectively, to obtain a first Mel-frequency cepstral coefficient (MFCC) of the first STFT feature and a second MFCC of the second STFT feature. In this example, the first electronic device may match the second audio and the third audio based on the first MFCC and the second MFCC.
[0189] Alternatively, the first electronic device may further perform filter bank (Fbank) processing on the first STFT feature and the second STFT feature, respectively, to obtain a first Fbank feature of the first STFT feature and a second Fbank feature of the second STFT feature. In this example, the first electronic device may match the second audio and the third audio based on the first Fbank and the second Fbank.
[0190] In some embodiments, the first audio feature of the second audio (such as the first STFT feature, the first MFCC, or the first Fbank feature) and the second audio feature of the third audio (such as the second STFT feature, the second MFCC, or the second Fbank feature) can be represented in the form of spectra. For example, the first audio feature corresponds to the first spectrum, and the second audio feature corresponds to the second spectrum.
[0191] The following describes a process in which the first electronic device matches the second audio and the third audio based on the first spectrum and the second spectrum, taking the first spectrum corresponding to the first STFT feature and the second spectrum corresponding to the second STFT feature as examples:
[0192] After STFT processing, the audio frequency spectrum is obtained. The frequency spectrum represents the relationship between the time, frequency, and energy of the audio. Time is relative to the starting point of the audio. Accordingly, the first frequency spectrum represents the relationship between the time, frequency, and energy of the second audio frequency. The second frequency spectrum represents the relationship between the time, frequency, and energy of the third audio frequency.
[0193] Here, taking the first spectrum as an example, the process of the first electronic device processing the first spectrum is described:
[0194] Figure 6 The a in the figure shows the first spectrum, refer to Figure 6 In the a, the first spectrum is plotted on the horizontal axis as time and on the vertical axis as frequency. The grayscale in the first spectrum represents the energy at the corresponding time and frequency. Energy can be used to represent loudness; the greater the energy, the louder the audio played by the electronic device at that time.
[0195] In the embodiment of the present application, the first electronic device may divide the first spectrum into n regions and the second spectrum into n regions in the same division manner, where n is an integer greater than 1.
[0196] In some embodiments, the first electronic device may divide the first spectrum and the second spectrum using a uniform division method or an uneven division method. Uniform division means that each area after division is the same size. Uneven division means that areas of different sizes exist after division.
[0197] In some embodiments, each region after division may be square, rectangular, or triangular, etc., and the present application does not limit the shape of each region. The shape of each of the n regions in the first spectrum may be the same or different, and the shape of each of the n regions in the second spectrum may be the same or different.
[0198] After obtaining n regions of the first spectrum, the first electronic device can obtain a first anchor point for each region of the first spectrum. The first anchor point is the point corresponding to the maximum energy in the region. For example, the first electronic device can determine the point corresponding to the maximum energy in each region of the first spectrum. The point corresponding to the maximum energy in each region can be referred to as the first anchor point. In this way, the first electronic device can obtain n anchor points of the first spectrum.
[0199] In some embodiments, the first electronic device determines the point corresponding to the maximum energy in each region of the first spectrum and may record the time, frequency, and energy corresponding to the point, or the first electronic device may record the coordinates (t, f) of the point and the energy A(t, f), where t represents the time corresponding to the point, f represents the frequency corresponding to the point, and A represents the energy corresponding to the point.
[0200] Similarly, after obtaining n regions of the second spectrum, the first electronic device can obtain a second anchor point for each region of the second spectrum. The second anchor point is the point corresponding to the maximum energy in the region. For example, the first electronic device can determine the point corresponding to the maximum energy in each region of the second spectrum. The point corresponding to the maximum energy in each region can be referred to as the second anchor point. In this way, the first electronic device can obtain n anchor points of the second spectrum.
[0201] In some embodiments, the first electronic device determines the point corresponding to the maximum energy in each region of the second spectrum, and can record the time, frequency, and energy corresponding to the point, or the first electronic device can record the coordinates (t, f) of the point, and the energy A(t, f).
[0202] The first electronic device matches the first anchor point and the second anchor point. When the number of successfully matched anchor points is greater than or equal to a second threshold, the first electronic device determines that the second audio and the third audio are successfully matched.
[0203] Because the first spectrum and the second spectrum are divided in the same manner, there is a one-to-one correspondence between the n regions of the first spectrum and the n regions of the second spectrum. In some embodiments, the first electronic device matches the first anchor point and the second anchor point, which can be understood as matching the first anchor point and the second anchor point in each region corresponding to the first spectrum and the second spectrum.
[0204] For example, refer to Figure 6 In a, the area in the first row and first column of the first spectrum can be called region 1, and the area in the first row and first column of the second spectrum can be called region 1A. Region 1 and region 1A correspond to each other, and the first electronic device can match the first anchor point in region 1 and the second anchor point in region 1A. In this way, based on the region correspondence, the first electronic device can match n first anchor points with the corresponding n second anchor points.
[0205] The first electronic device may match the first anchor point with the second anchor point in the following manner:
[0206] In the embodiment of the present application, when the frequency of the first anchor point is the same as the frequency of the second anchor point, and the difference between the time of the first anchor point and the time of the second anchor point is a second preset duration, the first electronic device determines that the first anchor point and the second anchor point are successfully matched.
[0207] In some embodiments, the second preset duration may be 0. In this example, when the frequency of the first anchor point is the same as the frequency of the second anchor point, and the time of the first anchor point is equal to the time of the second anchor point, the first electronic device determines that the first anchor point and the second anchor point are successfully matched.
[0208] In some embodiments, since the start times of the second audio and the third audio cannot be guaranteed to be the same, a fixed time difference Δt may exist. Therefore, when the frequency of the first anchor point and the frequency of the second anchor point are the same, and the difference between the time of the first anchor point and the time of the second anchor point is a second preset duration (e.g., Δt), the first electronic device determines that the first anchor point and the second anchor point are successfully matched. For example, with reference to the description in the above embodiment, Δt can be Δt1 or Δt2.
[0209] In some embodiments, multiple second anchor points may have the same frequency as the same first anchor point. In this case, the first electronic device can obtain the time difference between each second anchor point and the first anchor point. Because there are multiple first anchor points and multiple second anchor points, the first electronic device can obtain the time difference between the same first anchor point and each second anchor point (the second anchor point has the same frequency as the first anchor point). The first electronic device can determine the number of identical differences and determine the first anchor point and second anchor point corresponding to the largest number of differences as a successful match.
[0210] For example, the first anchor point 1 and the second anchor point 2 are used as examples for description. There are three second anchor points with the same frequency as the first anchor point 1, and two second anchor points with the same frequency as the second anchor point 2. The first electronic device can obtain the time differences between the second anchor point and the first anchor point 1, for example, Δt1', Δt2', and Δt3', respectively. The first electronic device can obtain the time differences between the second anchor point and the first anchor point 2, for example, Δt3' and Δt4', respectively. Among them, Δt3' has the largest number of differences. The first electronic device can then determine that the first anchor point 1 and the second anchor point with a time difference of Δt3' from the first anchor point 1 are successfully matched, and the first electronic device can determine that the first anchor point 2 and the second anchor point with a time difference of Δt3' from the first anchor point 2 are successfully matched.
[0211] Reference Figure 6 In b, taking the first spectrum as an example, the first electronic device matches the anchor point of each area in the first spectrum with the anchor point of each area in the second spectrum. The successfully matched anchor points can be as follows Figure 6 The “×” in b indicates.
[0212] In some embodiments, when the number of successfully matched anchor points is greater than or equal to a second threshold, the first electronic device determines that the second audio and the third audio are successfully matched.
[0213] In some embodiments, the second threshold may be n×preset ratio, where n is the number of divided regions in the spectrum (or the total number of first anchor points). The preset ratio is pre-set.
[0214] As shown in the above example, the first spectrum and the second spectrum can be divided into n regions, and the points with the maximum energy in each region are matched. Because the first anchor point and the second anchor point are both the points with the maximum energy in the corresponding regions, the signal-to-noise ratio of the points with the maximum energy is high, and the background noise has less impact on the audio, which makes it easier for the first electronic device to accurately obtain the loudness of the third audio. The reason is that: in the embodiment of the present application, the first electronic device can use the energy of the anchor point with a high signal-to-noise ratio to obtain the loudness of the third audio. For the anchor point with a high noise ratio, the background noise has less impact on the audio. In this way, the loudness of the third audio obtained by the first electronic device based on the energy of the anchor point with a high signal-to-noise ratio is highly accurate and is not affected by background noise. Even in a noisy environment, the first electronic device can accurately obtain the loudness of the third audio.
[0215] In some embodiments, to further improve the signal-to-noise ratio of the anchor points, after obtaining n first anchor points and n second anchor points, the first electronic device can first remove first anchor points and second anchor points whose energies are less than an energy threshold. This ensures that the energies of the remaining first anchor points and second anchor points are all greater than the energy threshold, and the signal-to-noise ratio of the first anchor points and second anchor points is high. This can further improve the accuracy of the loudness of the third audio based on the matching results of the first and second anchor points.
[0216] In the embodiment of the present application, when the second audio and the third audio are successfully matched, the first electronic device can obtain the loudness of the third audio and the loudness of the second audio when the first electronic device plays it.
[0217] 2. The following describes a process in which the first electronic device obtains the third loudness of the third audio:
[0218] In some embodiments, the successfully matched second anchor point may be referred to as a second target anchor point.
[0219] In the embodiment of the present application, the first electronic device may perform weighted summation based on the frequency of the second target anchor point and the mapping relationship between frequency and weight to obtain the energy of the second target anchor point, and the energy of the second target anchor point is used to represent the third loudness.
[0220] Because the human ear responds differently to different frequencies, different weights can be set for different frequencies so that the user can hear audio at a moderate loudness. For example, the human ear is more sensitive to frequencies around 3000 Hz, so the weights for frequencies around 3000 Hz can be set higher, while other frequencies can be set lower. It is conceivable that more detailed weighting methods can be set for different frequencies, which will not be described in detail in the embodiments of the present application.
[0221] In this example, there may be multiple successfully matched second anchor points. The first electronic device may perform weighted summation on the energy of each second target anchor point based on the frequency of each second target anchor point and the mapping relationship between frequency and weight to obtain the energy of the second target anchor point. The energy of the second target anchor point is used to represent the loudness of the third audio, i.e., the third loudness.
[0222] In some embodiments, to more accurately calculate the loudness of the third audio, for each second target anchor point, the first electronic device may obtain the average energy of points within a preset size range centered on the second target anchor point. The first electronic device may use this average energy as the energy of the second target anchor point. Referring to this method, the first electronic device may obtain the energy of each second target anchor point.
[0223] In this embodiment, the preset size range can be square, rectangular, circular, etc., and the present application embodiment does not limit this. For example, the preset size range can be: the size range of a*b around the second target anchor point as the center. Correspondingly, the energy of the second target anchor point (i.e., the mean value within the preset size range) E rec (i) can be shown in the following formula 1:
[0224]
[0225] f i -b / 2≤f≤f i +b / 2
[0226] Among them, ti represents the time corresponding to the second target anchor point, fi represents the frequency corresponding to the second target anchor point, and A rec (t,f) represents the energy of each point within the preset size range.
[0227] In this example, the first electronic device can obtain the energy of each second target anchor point by referring to the same method. The first electronic device can perform a weighted summation of the energy of each second target anchor point based on the frequency of each second target anchor point and the mapping relationship between frequency and weight to obtain the energy of the second target anchor point.
[0228] For example, the energy E of the second target anchor point sec (representing the loudness of the third audio frequency) can be shown in the following formula 2:
[0229]
[0230] Where n represents the number of all second anchor points, and the number of successfully matched anchor points is greater than or equal to 1 and less than or equal to n. rec (i) represents the energy of each second target anchor point, Indicates the weight corresponding to the frequency of each second target anchor point.
[0231] In some embodiments, in order to facilitate calculation, the first electronic device may sec Take the logarithm and get L sec , L sec Used to indicate the third loudness of the third audio frequency.
[0232] 3. The following describes a process in which the first electronic device obtains the fourth loudness of the second audio played by the first electronic device:
[0233] In some embodiments, the first anchor point that is successfully matched may be referred to as a first target anchor point.
[0234] In an embodiment of the present application, for the first target anchor point, the first electronic device can perform weighted summation based on the frequency of the first target anchor point and the mapping relationship between frequency and weight to obtain the energy of the first target anchor point. Please refer to the relevant description of "the first electronic device obtains the energy of the second target anchor point" in "2".
[0235] For example, in the second spectrum, the energy M of each first target anchor point ori (i)=(t i ,f i ), the first electronic device can calculate the average energy of the points within a preset size range with the first target anchor point as the center, and use the average as the energy E of the first target anchor point ori (i) can be calculated by referring to the above formula 1.
[0236] Similar to obtaining the energy of the second target anchor point in "2", the first electronic device can perform weighted summation of the energy of each first target anchor point according to the frequency of each first target anchor point and the mapping relationship between frequency and weight to obtain the energy E' of the first target anchor point. main .
[0237] In some embodiments, in order to facilitate calculation, the first electronic device may calculate E' main Take the logarithm and get L′ main .
[0238] Unlike the energy of the first target anchor point, because the second electronic device plays the second audio, and the first electronic device records the second audio to obtain the third audio, the second electronic device is actually playing the second audio. Therefore, the energy of the second target anchor point obtained in "2" above can represent the loudness of the third audio, such as the third loudness. However, for the first electronic device, the first electronic device does not play the second audio. According to the method in "2" above, the first electronic device can obtain the energy of the first target anchor point. The energy of the first target anchor point does not represent the actual loudness of the first electronic device when playing the second audio.
[0239] In some embodiments, the volume of the first electronic device, the energy of the first target anchor point, and the loudness of the audio played by the first electronic device have a first mapping relationship, and the first mapping relationship can be expressed as the following formula 3:
[0240] L main =(α·L' main +β)·vol Formula 3
[0241] Vol represents the volume of the first electronic device, which can be the current volume of the first electronic device or the volume corresponding to a moderate loudness. α and β are known coefficients, L main Indicates the loudness of the second audio when the first electronic device plays it.
[0242] In some embodiments, the loudness L of the second audio played by the first electronic device is main , which can be called the fourth loudness.
[0243] In summary, after the first electronic device obtains the third loudness and the fourth loudness, the first electronic device can determine whether to adjust the loudness of the audio played by the second electronic device based on the difference between the third loudness and the fourth loudness. When the difference between the third loudness and the fourth loudness is less than or equal to the first threshold, it indicates that the difference between the loudness of the second audio played by the second electronic device and the loudness when the first electronic device plays the second audio is small, and the difference between the first loudness and the second loudness when the second electronic device plays the first audio is less than or equal to the first threshold, and the second electronic device does not need to adjust the loudness of the audio. When the difference between the third loudness and the fourth loudness is greater than the first threshold, it indicates that the difference between the loudness of the second audio played by the second electronic device and the loudness when the first electronic device plays the second audio is large, and the second electronic device needs to adjust the loudness of the audio.
[0244] The following describes how the second electronic device adjusts the loudness of the audio:
[0245] In the embodiment of the present application, because the first electronic device can obtain the difference between the third loudness and the fourth loudness, when the difference between the third loudness and the fourth loudness is greater than the first threshold, the first electronic device can send the first loudness information to the second electronic device. The first loudness information is used to instruct the second electronic device to play the first audio at the first loudness based on the information of the first audio.
[0246] It should be understood that when the second electronic device plays the second audio, the first electronic device can obtain the loudness difference between the first electronic device and the second electronic device based on the recorded third audio and the second audio it has obtained. The first electronic device can also adjust the loudness of the audio played by the second electronic device based on this loudness difference, so that the loudness of the first audio played by the second electronic device next is the first loudness. Wherein, the difference between the first loudness and the second loudness is less than or equal to the first threshold value, that is, the first electronic device can trigger the second electronic device to play the second audio at a moderate loudness, thereby improving the user experience.
[0247] The first loudness information is described in detail below:
[0248] Method 1:
[0249] In some embodiments, the first electronic device may adjust the amplitude of the first audio, so that the first electronic device may send first loudness information to the second electronic device, where the first loudness information includes the first audio after the amplitude is adjusted.
[0250] It should be understood that the amplitude of audio is related to the energy of the audio, and the energy of audio is related to the loudness of the audio. When the third loudness of the second audio played by the second electronic device is larger or smaller, the first electronic device can adjust the amplitude of the first audio so that the loudness is moderate when the first audio is played using this amplitude. The first electronic device can send the first audio with the adjusted amplitude to the second electronic device. In this way, when the second electronic device plays the first audio, because the amplitude of the first audio is adjusted, the loudness of the first audio played by the second electronic device will not be too loud or too soft compared to the third loudness, and can be adapted to the user.
[0251] Exemplarily, there is a mapping relationship between the amplitude of the audio and the loudness of the audio. The first electronic device can adjust the amplitude of the first audio to the amplitude corresponding to the first loudness based on the first loudness and the mapping relationship between the amplitude and the loudness. In this way, when the second electronic device plays the adjusted first audio, the loudness of the first audio can be the first loudness. In some embodiments, the first loudness can be equal to the fourth loudness, or the difference between the first loudness and the fourth loudness is less than a fourth threshold, and the fourth threshold is less than the first threshold.
[0252] The first loudness may be a loudness pre-set in the first electronic device and adapted to the user's hearing. Alternatively, the first loudness may be calculated by the first electronic device based on the fourth loudness and a fourth threshold.
[0253] Method 2:
[0254] In some embodiments, the first electronic device adjusts the volume of the first electronic device. In the casting scenario, when the first electronic device adjusts its own volume, the operation can also be synchronized to the second electronic device. In other words, when the first electronic device adjusts its own volume, the first electronic device can send first loudness information to the second electronic device, where the first loudness information includes the volume of the first electronic device.
[0255] In other words, when the first electronic device adjusts the volume of the first electronic device, the second electronic device will also adjust the volume of the second electronic device based on the volume adjustment operation of the first electronic device. In this way, by adjusting the volume of the first electronic device, the second electronic device can adjust the volume of the second electronic device to an appropriate volume. Because the volume and loudness of the second electronic device are mapped to each other, the second electronic device can adjust the volume of the second electronic device to an appropriate volume, so that the loudness of the audio played by the second electronic device can be adapted to the user.
[0256] Method 3:
[0257] In some embodiments, the first loudness information is specifically used to instruct the second electronic device to adjust the audio loudness to the first loudness.
[0258] In the embodiment of the present application, the first electronic device does not need to perform an operation of processing the first audio, but directly sends the first loudness information to the second electronic device, where the first loudness information is used to instruct the second electronic device to adjust the audio loudness to the first loudness.
[0259] Exemplarily, the first loudness information includes the first loudness. Accordingly, when the second electronic device receives the first loudness information, it can adjust the loudness of the audio to the first loudness, so that when the second electronic device plays the first audio, the loudness of the first audio is the first loudness. Exemplarily, the second electronic device can pre-store a mapping relationship between the volume and loudness of the second electronic device. The second electronic device can adjust the volume of the second electronic device to the volume mapped by the first loudness based on the first loudness and the mapping relationship between the volume and loudness. In this way, the second electronic device can play the first audio at the volume mapped by the first loudness, and the loudness of the first audio is the first loudness.
[0260] The above describes the detailed process of the delivery method provided in the embodiment of the present application. Figure 7 A flow chart of the delivery method provided in the embodiment of the present application. Figure 7 The delivery method provided in the embodiment of the present application may include:
[0261] Step 1B: The first electronic device processes the second audio (original audio) to obtain a first frequency spectrum.
[0262] Step 2B: The first electronic device divides the first spectrum into n regions.
[0263] Step 3B: The first electronic device obtains a first anchor point in each area, where the first anchor point is a point corresponding to maximum energy in the area.
[0264] Step 4B: The first electronic device processes the third audio (recorded audio) to obtain a second frequency spectrum.
[0265] Step 5B: The first electronic device divides the second spectrum into n regions.
[0266] Step 6B: The first electronic device obtains a second anchor point in each area, where the second anchor point is a point corresponding to maximum energy in the area.
[0267] Step 7B: The first electronic device matches the first anchor point and the second anchor point.
[0268] Step 8B: When the anchor point is successfully matched, the first electronic device obtains the third loudness of the third audio and the fourth loudness when the first electronic device plays the first audio.
[0269] Step 9B: When the difference between the third loudness and the fourth loudness is greater than the first threshold, the second electronic device adjusts the loudness and plays the first audio.
[0270] The loudness of the first audio is a first loudness, a difference between the first loudness and the second loudness is less than or equal to a first threshold, and the second loudness is a loudness of the first audio played by the first electronic device.
[0271] Step 10B: When the anchor point matching fails, the second electronic device plays the first audio at a default volume or the volume at which the audio was last played.
[0272] It should be understood that steps 1B to 10B may refer to the description in the above embodiment, and the specific implementation methods and technical effects thereof may refer to the description in the above embodiment, which will not be repeated here.
[0273] As described in the above embodiment, when "a first electronic device plays audio to a second electronic device", a method for adjusting the loudness of audio played by the second electronic device is introduced. In some embodiments, after the second electronic device plays the first audio, the delivery method in the embodiment of the present application can continue to be used.
[0274] In some embodiments, while a second electronic device is playing a first audio signal, the first electronic device may record the first audio signal played by the second electronic device to obtain a fourth audio signal, where the loudness of the fourth audio signal is a fifth loudness. When the difference between the fifth loudness and the second loudness is greater than a first threshold, the first electronic device may send second loudness information to the second electronic device, where the second loudness information is used to instruct the second electronic device to play the first audio signal at the first loudness. In this way, the second electronic device may play the first audio signal at the first loudness in response to the second loudness information. Thus, the difference between the loudness of the first audio signal played by the second electronic device (the first loudness) and the second loudness is less than or equal to the first threshold.
[0275] In other words, in this embodiment, after the first electronic device plays the first audio to the second electronic device, the first electronic device can continue to detect the difference between the loudness of the first audio played by the second electronic device and the second loudness, so that the first electronic device can send second loudness information to the second electronic device based on the loudness difference, and promptly adjust the loudness of the first audio played by the second electronic device to the first loudness.
[0276] In this embodiment of the present application, the first electronic device may record (N+Δt) seconds of the first audio, or the first electronic device may record N seconds of the first audio, or the first electronic device may record M seconds of the first audio to obtain the fourth audio. M may be different from N. In this embodiment of the present application, after obtaining the fourth audio, the first electronic device may obtain the fifth loudness of the fourth audio. The specific obtaining process can refer to the description of "the first electronic device obtains the third loudness of the third audio" in the above embodiment.
[0277] When the difference between the fifth loudness and the second loudness is greater than the first threshold, the first electronic device may send the second loudness information to the second electronic device, so that the difference between the loudness of the first audio played by the second electronic device and the second loudness is less than or equal to the first threshold. For details of this process, refer to the description of "when the difference between the third loudness and the fourth loudness is greater than the first threshold, the first electronic device may send the first loudness information to the second electronic device" in the above embodiment.
[0278] Scenario 1:
[0279] In one possible scenario, for example, if a first electronic device records audio played by a second electronic device, the loudness of the recorded audio obtained by the first electronic device may vary as the distance between the first electronic device and the second electronic device varies during the movement of the first electronic device. Therefore, in this example, as the first electronic device moves, the first electronic device can record the first audio played by the second electronic device in real time to obtain a fourth audio. The first electronic device can also adjust the loudness of the first audio played by the second electronic device at any time based on the difference between the fifth loudness and the second loudness of the fourth audio, thereby adapting to the user's needs at any time and improving the user experience.
[0280] Scenario 2:
[0281] In one possible scenario, in order to adapt to the user, the first electronic device can periodically record the first audio played by the second electronic device to obtain the fourth audio, and the first electronic device can adjust the loudness of the first audio played by the second electronic device at any time according to the difference between the fifth loudness and the second loudness of the fourth audio, so as to adapt to the user at any time and improve the user experience.
[0282] Scenario 3:
[0283] In one possible scenario, in order to adapt to the user, the first electronic device can record the first audio played by the second electronic device when switching the audio to be played to obtain the fourth audio. The first electronic device can adjust the loudness of the first audio played by the second electronic device at any time based on the difference between the fifth loudness and the second loudness of the fourth audio, so that it can adapt to the user at any time and improve the user experience.
[0284] Among them, the audio that the first electronic device switches to play can be understood as: the first electronic device plays the next audio. For example, when the first electronic device plays music 1 to the second electronic device, when music 1 is finished, it will continue to play music 2. Then, when the first electronic device switches to playing music 2, the first electronic device can record music 2 played by the second electronic device to obtain a fourth audio. Then, the first electronic device can adjust the loudness of the first audio played by the second electronic device at any time based on the difference between the fifth loudness and the second loudness of the fourth audio, and can adapt to the user at any time to improve the user experience.
[0285] In the embodiment of the present application, the loudness of the audio played by the second electronic device can be adjusted not only when the audio is projected, but also after the loudness of the audio played by the second electronic device is adjusted, the loudness of the audio played by the second electronic device can be adjusted in real time or periodically.
[0286] In some embodiments, the first electronic device may adjust the loudness of the audio played by the second electronic device when switching the audio played, and enable the loudness of the audio played by the second electronic device to be consistent with the loudness of the audio played by the first electronic device.
[0287] In this scenario, the first electronic device can receive an instruction to switch audio. The instruction to switch audio is used to indicate the switching of the delivered audio. The instruction can be triggered by the user or actively triggered after an audio is played. This embodiment of the present application does not limit this. After the first electronic device receives the instruction to switch audio, it can trigger the second electronic device to play the fifth audio (for example, music 2). The fifth audio played by the second electronic device can be the sixth loudness, and the difference between the sixth loudness and the seventh loudness is less than or equal to the first threshold.
[0288] In some embodiments, the seventh loudness may refer to the description related to the fourth loudness.
[0289] In an embodiment of the present application, each time the audio to be delivered is switched, the first electronic device may use the delivery method of the present application to enable the loudness of the audio played by the second electronic device to be consistent with the loudness of the audio played by the first electronic device.
[0290] In some embodiments, after a first electronic device plays audio to a second electronic device, the user can trigger a switch in the audio playback device. For example, the user can trigger the first electronic device to play audio to a third electronic device. In this scenario, the first electronic device can receive a second playback instruction, which instructs the first electronic device to play audio. The second playback instruction can refer to the description of the first playback instruction.
[0291] In response to the second delivery instruction, the first electronic device can trigger the third electronic device to play the first audio. In response to the second delivery instruction, the first electronic device can send information about the first audio and third loudness information to the third electronic device. When the third electronic device switches to playing the first audio, the third electronic device can continue playing the first audio, following the playback progress of the second electronic device. In response, the information about the first audio is used to indicate the audio that follows the first audio played by the second electronic device.
[0292] The third loudness information is used to instruct the third electronic device to play the first audio at the eighth loudness based on the information about the first audio. The third loudness information can refer to the description related to the first loudness information. In this way, in response to the third loudness information, the third electronic device can play the first audio at the eighth loudness based on the information about the first audio.
[0293] The difference between the eighth loudness and the first loudness is less than or equal to the first threshold. In other words, when the first electronic device receives the second delivery instruction, it can use the method provided in the embodiment of the present application, for example, by recording the first audio or prompt sound (such as the second audio) played by the third electronic device, to obtain the loudness difference between the audio played by the second electronic device and the third electronic device, and then adjust the loudness of the audio played by the third electronic device so that the loudness of the audio played by the third electronic device is consistent with the loudness of the audio played by the second electronic device.
[0294] The loudness of the audio played by the second electronic device may be the loudness of the first audio played by the second electronic device, such as the first loudness. That is, the difference between the eighth loudness and the first loudness is less than or equal to the first threshold. In other words, the mobile phone can enable the loudness of the audio played by the speaker to be consistent with the loudness of the audio played by the mobile phone.
[0295] For example, the first electronic device is a mobile phone, the second electronic device is a speaker, and the third electronic device is a television. When the mobile phone plays music 1 to the speaker, the mobile phone can use the playback method provided in the embodiment of the present application to enable the speaker to play music 1 at a first loudness, and the difference between the first loudness and the second loudness (the loudness of the mobile phone playing music 1) is less than or equal to a first threshold. While the speaker is playing music 1, the user triggers the mobile phone to play music 1 to the television, wherein the television can continue playing music 1 by following the progress of the speaker playing music 1.
[0296] When the TV plays Music 1, the mobile phone can use the delivery method provided in the embodiment of the present application to enable the TV to play Music 1 at an eighth loudness, where the difference between the eighth loudness and the first loudness is less than or equal to the first threshold. In other words, the mobile phone can enable the loudness of the audio played on the TV to be consistent with the loudness of the audio played on the speaker.
[0297] It is conceivable that the first electronic device responds to the third delivery instruction, which instructs the fourth electronic device to deliver the audio. The first electronic device, in response to the third delivery instruction, can trigger the fourth electronic device to play the first audio. The first audio has a ninth loudness, and the difference between the ninth loudness and the eighth loudness is less than or equal to the first threshold.
[0298] In an embodiment of the present application, in a scenario where the delivery device is switched, the first electronic device can adopt the delivery method provided in an embodiment of the present application to enable the loudness of the audio played by the subsequent delivery device (such as the third electronic device) to be less than or equal to the first threshold value compared with the loudness of the audio played by the previous delivery device (such as the second electronic device).
[0299] It should be noted that the data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0300] In one embodiment, the present application also provides an electronic device, which may be the first electronic device or the second electronic device in the above embodiment. Figure 8 The electronic device may include: a processor 801 (e.g., a CPU) and a memory 802. The memory 802 may include a high-speed random-access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage. The memory 802 may store various instructions for performing various processing functions and implementing the method steps of the present application.
[0301] Optionally, the electronic device involved in this application may further include: a power supply 803, a communication bus 804, and a communication port 805. The communication port 805 is used to enable communication between the electronic device and other peripheral devices. In the embodiment of the present application, the memory 802 is used to store computer-executable program code, which includes instructions; when the processor 801 executes the instructions, the instructions cause the processor 801 of the electronic device to perform the actions in the above-mentioned method embodiment. The implementation principles and technical effects are similar and will not be repeated here.
[0302] Optionally, the electronic device involved in this application may further include: a display screen 806. The display screen 806 is used to display an interface of the electronic device.
[0303] It should be noted that the modules or components described in the above embodiments may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code through a processing element, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code, such as a controller. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0304] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0305] The term "plurality" in this article refers to two or more. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the previous and next associated objects are in an "or" relationship; in the formula, the character " / " indicates that the previous and next associated objects are in a "division" relationship. In addition, it should be understood that in the description of this application, words such as "first" and "second" are only used for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0306] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0307] It can be understood that in the embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
Claims
1. A delivery method, characterized in that: Applied to a first electronic device, the method includes: receiving a first delivery instruction, wherein the first delivery instruction is used to instruct delivery of audio to a second electronic device; First audio information and first loudness information are sent to the second electronic device, where the first loudness information is used to instruct the second electronic device to play the first audio at a first loudness based on the first audio information, where a difference between the first loudness and the second loudness is less than or equal to a first threshold, and the second loudness is the loudness of the first electronic device when playing the first audio.
2. The method according to claim 1, characterized in that Before sending the first audio information and the first loudness information to the second electronic device, the method further includes: Sending second audio information to the second electronic device, where the second audio information is used to instruct the second electronic device to play the second audio; Recording the second audio played by the second electronic device to obtain a third audio, where the loudness of the third audio is a third loudness; Sending the first loudness information to the second electronic device includes: When the difference between the third loudness and the fourth loudness is greater than the first threshold, the first loudness information is sent to the second electronic device.
3. The method according to claim 2, characterized in that The recording of the second audio played by the second electronic device includes: In response to the delivery instruction, starting recording; The recording stops when a first preset duration is recorded, and the first preset duration is greater than the duration of the second audio.
4. The method according to claim 2 or 3, characterized in that Before sending the first loudness information to the second electronic device, the method further includes: adjusting the amplitude of the first audio, wherein the first loudness information includes the first audio after the amplitude is adjusted; or The volume of the first electronic device is adjusted, where the first loudness information includes the volume of the first electronic device.
5. The method according to claim 2 or 3, characterized in that The first loudness information is specifically used to instruct the second electronic device to adjust the audio loudness to the first loudness.
6. The method according to any one of claims 2 to 5, characterized in that The fourth loudness is a preset loudness; or, The fourth loudness is related to the loudness of audio played in the past by the first electronic device.
7. The method according to any one of claims 2 to 5, characterized in that The fourth loudness is the loudness of the first electronic device when playing the second audio.
8. The method according to any one of claims 2 to 6, characterized in that The method further comprises: Recording the first audio played by the second electronic device to obtain a fourth audio, where the loudness of the fourth audio is a fifth loudness; When the difference between the fifth loudness and the second loudness is greater than the first threshold, second loudness information is sent to the second electronic device, where the second loudness information is used to instruct the second electronic device to play the first audio at the first loudness.
9. The method according to claim 8, characterized in that The recording of the first audio played by the second electronic device includes: During the movement of the first electronic device, recording the first audio played by the second electronic device; or The first audio played by the second electronic device is recorded periodically.
10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: Receive instructions for switching audio; The second electronic device is triggered to play a fifth audio, where the fifth audio has a sixth loudness, a difference between the sixth loudness and a seventh loudness is less than or equal to a first threshold, and the seventh loudness is the loudness of the first electronic device when playing the fifth audio.
11. The method according to any one of claims 1 to 10, characterized in that After sending the first audio information and the first loudness information to the second electronic device, the method further includes: receiving a second delivery instruction, wherein the second delivery instruction is used to instruct switching to a third electronic device to deliver the audio; Sending information about the first audio and third loudness information to the third electronic device, where the third loudness information is used to instruct the third electronic device to play the first audio at an eighth loudness based on the information about the first audio, where a difference between the eighth loudness and the first loudness is less than or equal to the first threshold.
12. The method according to any one of claims 2 to 9, characterized in that The method further comprises: matching the second audio and the third audio; When the matching is successful, the loudness of the third audio and the loudness of the second audio when the first electronic device plays are obtained.
13. The method according to claim 12, characterized in that The matching of the second audio and the third audio includes: Obtaining a first audio feature of the second audio and a second audio feature of the third audio; The second audio and the third audio are matched according to the first audio feature and the second audio feature.
14. The method according to claim 13, characterized in that The first audio feature and the second audio feature are represented in the form of spectrum, the first audio feature corresponds to a first spectrum, and the second audio feature corresponds to a second spectrum; The matching the second audio and the third audio according to the first audio feature and the second audio feature includes: Divide the first spectrum into N regions and the second spectrum into N regions in the same division manner, where N is an integer greater than 1; Obtaining a first anchor point for each region in the first spectrum and a second anchor point for each region in the second spectrum, where the anchor point is a point corresponding to maximum energy in the region; matching the first anchor point and the second anchor point; When the number of matched anchor points is greater than a second threshold, it is determined that the second audio and the third audio are matched successfully.
15. The method according to claim 14, characterized in that The first spectrum and the second spectrum are both used to represent the relationship between time, frequency, and energy of the audio, and matching the first anchor point and the second anchor point includes: When the frequency of the first anchor point is the same as the frequency of the second anchor point, and the difference between the time of the first anchor point and the time of the second anchor point is within a second preset duration, it is determined that the first anchor point and the second anchor point are successfully matched.
16. The method according to claim 14 or 15, characterized in that The successfully matched second anchor point is used as the second target anchor point. Obtaining the loudness of the third audio includes: Obtaining the energy sum of the second anchor point within the preset range of the second target anchor point; According to the frequency of the second target anchor point and the mapping relationship between the frequency and the weight, a weighted sum is performed on the energy sum to obtain the energy of the second target anchor point. The energy of the second target anchor point is used to represent the loudness of the third audio.
17. The method according to claim 14 or 15, characterized in that The first anchor point that is successfully matched is used as the first target anchor point; and obtaining the loudness of the first electronic device when playing the first audio includes: Obtaining the energy mean of the first anchor point within the preset range of the first target anchor point; performing a weighted summation on the energy sum according to the frequency of the first target anchor point and a mapping relationship between the frequency and the weight to obtain the energy of the first target anchor point; The loudness of the first audio when the first electronic device plays the first audio is obtained according to the energy of the first target anchor point.
18. The method according to any one of claims 2 to 9, characterized in that The second audio and the first audio are continuous audio in one audio segment; or, The second audio is audio used to indicate that the first electronic device and the second electronic device are successfully connected, and the first audio is audio projected from the first electronic device to the second electronic device.
19. The method according to any one of claims 2 to 9, characterized in that When the second audio and the first audio are continuous audio in a segment, before receiving the first delivery instruction, the method further includes: Playing the second audio; When the second audio is audio used to indicate that the first electronic device and the second electronic device are successfully connected, and the first audio is audio projected from the first electronic device to the second electronic device, before receiving the first projecting instruction, the method further includes: Play the first audio.
20. An electronic device, characterized in that: The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, and is configured to store computer program codes, where the computer program codes include computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 19.
21. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1-19.
22. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises computer instructions, and when the computer instructions are executed on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 19.
23. A computer program product, characterized in that The computer program product comprises a computer program code, and when the computer program code is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Data processing method and related equipment
CN114077412A
Sound effect adjusting method and electronic equipment
CN114666631A
Volume adjustment method, electronic equipment and computer readable storage medium
CN115562611A
Bluetooth master device, Bluetooth slave device and volume control method
CN115914704A
Method and system for adjusting volume, and electronic device
WO2022206825A1