Terminal device and voice broadcast volume adjustment method
By collecting mixed audio data with a detector, calculating the difference between the base volume and the reference volume, and combining the ambient noise volume to obtain the adjustment coefficient, the broadcast volume is automatically adjusted, which solves the problem of mismatch between the voice broadcast volume of the terminal device and improves the user's auditory experience.
Patent Information
- Application Number
- CN202410800091.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-06-19
AI Technical Summary
When the terminal device plays media assets and provides voice broadcasts, the volume settings are not matched with the volume of the external device and the media asset source, resulting in the voice broadcast being too loud or too soft, which affects the user's auditory experience.
By collecting mixed audio data through a detector, calculating the difference between the base volume and the reference volume, and combining it with the ambient noise volume to obtain an adjustment coefficient, the broadcast volume is automatically adjusted to meet the user's needs.
It enables automatic adjustment of voice broadcast volume based on multiple reference dimensions, improving the user's auditory experience and ensuring that the volume adapts to the external environment and user habits.
Smart Images

Figure CN118828103B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a terminal device and a method for adjusting the volume of voice broadcast. Background Art
[0002] Terminal devices refer to electronic devices with sound acquisition capabilities, such as smart TVs, mobile phones, smart speakers, computers, and robots. Taking smart TVs as an example, smart TVs are television products based on Internet application technology, equipped with open operating systems and chips, and featuring voice recognition modules, enabling two-way human-computer interaction to meet diverse and personalized user needs.
[0003] The terminal device has audio playback capabilities, enabling it to play media resources from different signal sources. For example, it can play audio data from locally stored media resources via speakers or other audio output devices, or it can play audio data from media resources provided by external devices. Some terminal devices also support voice broadcasting functions, such as wake-up and false wake-up responses, and TTS (Text to Speech) broadcasting, to assist users in voice interaction with the terminal device through voice broadcasting.
[0004] However, the volume of the audio data corresponding to the aforementioned media assets and voice broadcast functions is manually set by the user. When the volume set by the external device is low or the volume corresponding to the media asset source is low, the system volume set by the terminal device may be too loud; conversely, when the volume set by the external device is high or the volume corresponding to the media asset source is high, the system volume set by the terminal device may be too low. In such cases, if the user uses the voice broadcast function of the terminal device, the voice broadcast volume may also be too loud or too soft, affecting the user's auditory experience. Summary of the Invention
[0005] This application provides a terminal device and a method for adjusting the volume of voice broadcasts to solve the problem of voice broadcasts being too loud or too soft.
[0006] In a first aspect, some embodiments of this application provide a terminal device, including an audio output device, a detector, and a controller. The audio output device is configured to play output sound from internal audio data, including media asset audio data and broadcast audio data; the detector is configured to acquire mixed audio data, including audio data generated by acquiring the output sound and audio data generated by acquiring external ambient sound; the controller is configured to execute the following program steps:
[0007] Obtain the base volume set by the media asset audio data;
[0008] Calculate a first difference between the base volume and the reference volume, where the reference volume is the average volume of the mixed audio data over a target time period;
[0009] Detect ambient noise volume based on the internal audio data and the mixed audio data;
[0010] The adjustment coefficient is obtained based on the first difference value and the ambient noise volume.
[0011] The target volume is calculated using the adjustment factor, the reference volume, and the base volume.
[0012] Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
[0013] In some embodiments of this application, the controller is configured to: acquire the media asset audio data for a first preset time period to generate a first audio frame; extract a first signal value of the first audio frame; and calculate the basic volume based on the first signal value, wherein the basic volume is calculated based on the sum of the absolute values of the first signal values, or calculated based on the logarithm of the sum of the squares of the first signal values.
[0014] In some embodiments of this application, the controller is further configured to: acquire the mixed audio data for a second preset time period to generate a second audio frame; extract a second signal value of the second audio frame; calculate a sampling volume based on the second signal value, wherein the sampling volume is calculated based on the sum of the absolute values of the second signal values, or based on the logarithm of the sum of the squares of the second signal values; and calculate a reference volume based on the sampling volume.
[0015] In some embodiments of this application, the controller is configured to perform the calculation of the reference volume based on the sampled volume, and further configured to: calculate a first average value of the sampled volume over a third preset time period, the third preset time period being greater than a second preset time period; perform a weighted average on the first average value over a target time period to generate a second average value, the target time period being greater than the third preset time period; and associate the second average value with the third preset time period to generate the reference volume.
[0016] In some embodiments of this application, the controller performs the function of detecting ambient noise volume based on the internal audio data and the mixed audio data, and is configured to: acquire the mixed audio volume of the mixed audio data; calculate a second difference value between the base volume and the mixed audio volume; if the second difference value is less than a noise threshold, mark the ambient noise volume as a first noise state; if the second difference value is greater than or equal to the noise threshold, mark the ambient noise volume as a second noise state, wherein the ambient noise volume of the second noise state is higher than the ambient noise volume of the first noise state.
[0017] In some embodiments of this application, if the first difference value is greater than 0, the controller executes the process of obtaining an adjustment coefficient based on the first difference value and the ambient noise volume, configured to: obtain a first difference threshold, the first difference threshold being greater than 0; if the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is in the first noise state, then obtain a first adjustment coefficient; if the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is in the second noise state, then obtain a second adjustment coefficient; the first adjustment coefficient and the second adjustment coefficient are both less than 1, and the first adjustment coefficient is less than the second adjustment coefficient.
[0018] In some embodiments of this application, if the first difference value is less than 0, the controller executes an adjustment coefficient based on the first difference value and the ambient noise volume, configured to: obtain a second difference threshold, the second difference threshold being less than 0, and the absolute values of the second difference threshold and the first difference threshold being equal; if the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is in the first noise state, then obtain a third adjustment coefficient; if the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is in the second noise state, then obtain a fourth adjustment coefficient, the third adjustment coefficient and the fourth adjustment coefficient being greater than 1, and the third adjustment coefficient being less than the fourth adjustment coefficient.
[0019] In some embodiments of this application, the controller is configured to perform the calculation of a target volume using the adjustment coefficient, the reference volume, and the base volume, by: calculating a third average of the reference volume and the base volume; and generating the target volume based on the third average and the adjustment coefficient, wherein the target volume is the product of the third average and the target adjustment coefficient.
[0020] In some embodiments of this application, the controller is further configured to: detect the playback status of the broadcast audio data; if the playback status is a first playback status, execute the step of obtaining the basic volume set by the media asset audio data, wherein the first playback status is used to indicate that the broadcast audio data is not currently being played; if the playback status is a second playback status, when the playback status of the broadcast audio data changes to the first playback status, execute the step of obtaining the basic volume set by the media asset audio data, wherein the second playback status is used to indicate that the broadcast audio data is being played.
[0021] Secondly, some embodiments of this application also provide a method for adjusting the volume of voice broadcasts, which can be applied to the terminal device provided in the first aspect. The method includes the following steps:
[0022] Get the base volume settings for media asset audio data;
[0023] Calculate a first difference between the base volume and the reference volume, where the reference volume is the average volume of the mixed audio data over a preset time period;
[0024] Detect ambient noise volume based on internal audio data and the mixed audio data;
[0025] The adjustment coefficient is obtained based on the first difference value and the ambient noise volume.
[0026] The target volume is calculated using the adjustment factor, the reference volume, and the base volume.
[0027] Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
[0028] As can be seen from the above technical solutions, the terminal device and voice broadcast volume adjustment method provided in some embodiments of this application can obtain the base volume set by the media asset audio data and calculate a first difference value between the base volume and the reference volume. The reference volume is the average volume of the mixed audio data over a preset time period. Then, the ambient noise volume is detected based on the internal audio data and the mixed audio data, and an adjustment coefficient is obtained according to the first difference value and the ambient noise volume. A target volume is calculated using the adjustment coefficient, the reference volume, and the base volume, and the volume of the broadcast audio data is set to the target volume, so that the audio output device plays the broadcast audio data at the target volume. This method can automatically adjust the volume of voice broadcast in the terminal device by combining multiple reference dimensions of the terminal device, making the volume of voice broadcast more suitable for user needs, thereby improving the user's auditory experience. Attached Figure Description
[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0030] Figure 1 This application provides a system architecture diagram for a terminal device voice interaction scenario according to some embodiments;
[0031] Figure 2 This is a schematic diagram of the hardware configuration of a terminal device provided in some embodiments of this application;
[0032] Figure 3 This is a schematic diagram of the software configuration of a terminal device provided in some embodiments of this application;
[0033] Figure 4 A schematic diagram of a network architecture for voice interaction provided for some embodiments of this application;
[0034] Figure 5 This is a schematic diagram of the audio data playback process provided in some embodiments of this application;
[0035] Figure 6 A flowchart illustrating a voice broadcast volume adjustment method provided in some embodiments of this application;
[0036] Figure 7 A schematic flowchart illustrating the calculation of a reference volume provided for some embodiments of this application;
[0037] Figure 8 A schematic flowchart illustrating the detection of ambient noise volume provided in some embodiments of this application;
[0038] Figure 9 A schematic diagram illustrating the process of obtaining the adjustment coefficient provided in some embodiments of this application;
[0039] Figure 10 A flowchart illustrating the process of obtaining adjustment coefficients in conjunction with peripheral connection status, provided for some embodiments of this application;
[0040] Figure 11 This is a flowchart illustrating a method for adjusting volume based on the playback status of broadcast audio data, provided in some embodiments of this application. Detailed Implementation
[0041] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0042] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0043] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0044] Figure 1 An exemplary system architecture is shown that can be applied to the voice interaction scenarios of the terminal device in this application. For example... Figure 1 As shown, 10 is a server and 200 is a terminal device, including (smart TV 200a, mobile device 200b, smart speaker 200c).
[0045] In this application, server 10 and terminal device 200 communicate data through various communication methods. Terminal device 200 can be connected via local area network (LAN), wireless local area network (WLAN), and other networks. Server 10 can provide various content and interactive features to terminal device 200. For example, terminal device 200 and server 10 can send and receive information, and receive software updates.
[0046] In some embodiments, server 10 may be a server that provides various services, such as a backend server that supports audio data collected by terminal device 200. The backend server may analyze and process the received audio and other data, and feed back the processing results (e.g., endpoint information) to the terminal device. Server 10 may be a server cluster or multiple server clusters, and may include one or more types of servers.
[0047] In some embodiments, the terminal device 200 can be hardware or software. When the terminal device 200 is hardware, it can be various electronic devices with sound acquisition capabilities, including but not limited to smart speakers, smartphones, televisions, tablets, e-book readers, smartwatches, media players, computers, AI devices, robots, smart vehicles, etc. When the terminal devices 200, 201, and 202 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software programs or software modules (e.g., for providing sound acquisition services) or as a single software program or software module. No specific limitations are made here.
[0048] It should be noted that the voice broadcast volume adjustment method provided in this application embodiment can be executed by server 10, terminal device 20, or by both server 10 and terminal device 20. This application does not limit this.
[0049] Figure 2 A hardware configuration block diagram of a terminal device 200 according to an exemplary embodiment is shown. For example... Figure 2 The terminal device 200 shown includes at least one of the following: a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280. The controller includes a central processing unit, an audio processor, a graphics processor, RAM, ROM, and a first to an nth interface for input / output.
[0050] In some embodiments, the display 260 includes a display screen component for presenting an image, a driving component for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0051] In some embodiments, the display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.
[0052] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The terminal device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the terminal device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the terminal device 200 supports Bluetooth connection communication, it needs to have a communication device 220 with Bluetooth functionality.
[0053] In some embodiments, the communication device 220 enables the terminal device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize components such as data cables and interfaces to connect the terminal device 200 to external devices. Wireless connections utilize wireless signals or wireless networks to connect the terminal device 200 to external devices. The terminal device 200 can directly establish a connection with external devices or indirectly establish a connection through gateways, routers, or connection devices.
[0054] In some embodiments, the user input interface 280 can be used to receive instructions from user input.
[0055] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0056] In some embodiments, the sound acquisition device can be a microphone, also known as a "microphone" or "voice transducer," which can be used to receive the user's voice and convert the sound signal into an electrical signal. The terminal device 200 can be equipped with at least one microphone. In other embodiments, the terminal device 200 can be equipped with two microphones, which, in addition to acquiring sound signals, can also perform noise reduction. In still other embodiments, the terminal device 200 can also be equipped with three, four, or more microphones, enabling sound signal acquisition, noise reduction, sound source identification, and directional recording functions, etc.
[0057] In some embodiments, the microphone may be built into the terminal device 200, or the microphone may be connected to the terminal device 200 via wired or wireless means. Of course, the embodiments of this application do not limit the location of the microphone on the terminal device 200. Alternatively, the terminal device 200 may not include a microphone, that is, the microphone is not provided in the terminal device 200. The terminal device 200 may connect an external microphone (also called a microphone) via an interface (such as a USB interface 130). This external microphone can be fixed to the terminal device 200 by an external fastener (such as a camera bracket with a clip).
[0058] In some embodiments, the controller 250 controls the operation of the terminal device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the terminal device 200.
[0059] For example, controller 250 includes at least one of a central processing unit (CPU), an audio processor, a graphics processing unit (GPU), a RAM (Random Access Memory), a ROM (Read-Only Memory), a first to an nth interface for input / output, a communication bus, etc.
[0060] In order to perform user interaction, in some embodiments, the terminal device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources in the terminal device 200.
[0061] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for display devices.
[0062] In some examples, the terminal device's operating system is Android, such as... Figure 3 As shown, the terminal device 200 can be logically divided into an application layer (referred to as the "application layer") 21, a kernel layer 22, and a hardware layer 23.
[0063] In some embodiments, such as Figure 3 As shown, the hardware layer may include Figure 2 The controller 250, communication device 220, detector 230, etc., are shown. Application layer 21 includes one or more applications. These applications can be system applications or third-party applications. For example, application layer 21 includes a voice application, which can provide a voice interaction interface and services to enable the connection between the smart TV 200-1 and the server 10.
[0064] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21, and is used to manage and control hardware and software resources.
[0065] In some embodiments, kernel layer 22 includes a detector driver, which is used to send voice data collected by detector 230 to a voice application. For example, when the voice application in terminal device 200 is started and a communication connection is established between terminal device 200 and server 10, the detector driver is used to send the user-input voice data collected by detector 230 to the voice application. Then, the voice application sends query information containing this voice data to intent recognition module 202 in the server. Intent recognition module 202 is used to input the voice data sent by terminal device 200 into the intent recognition model.
[0066] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the terminal device 200 in this application embodiment. Depending on factors such as the function of the display device and the type of the operating system, the number of levels and the specific level type of the operating system may be expressed in other forms.
[0067] To clearly illustrate the embodiments of this application, the following description is provided in conjunction with... Figure 4 This application describes a speech recognition network architecture provided in its embodiments.
[0068] See Figure 4 , Figure 4 This is a schematic diagram of a voice interaction network architecture provided in an embodiment of this application. Figure 4 In this system, the terminal device receives input information and outputs the processing results of that information. The speech recognition module deploys a speech recognition service to convert audio into text; the semantic understanding module deploys a semantic understanding service to perform semantic parsing on the text; the business management module deploys a business instruction management service to provide business instructions; the language generation module deploys a language generation service (NLG) to convert instructions instructing the terminal device to execute into text language; and the speech synthesis module deploys a speech synthesis (TTS) service to process the text language corresponding to the instructions and send it to the speaker for playback. In one embodiment, Figure 4 The architecture shown can contain multiple entity service devices with different business services deployed, or one or more entity service devices can combine one or more functional services.
[0069] In some embodiments, the following describes the basis Figure 4 The process of processing information from input terminal devices in the architecture shown is described with an example, taking a query statement input via voice as an example:
[0070] [Speech Recognition]
[0071] After receiving a query statement input via voice, the terminal device can perform noise reduction and feature extraction on the audio of the query statement. The noise reduction process may include steps such as removing echoes and environmental noise.
[0072] [Semantic understanding]
[0073] Using acoustic and language models, natural language understanding is performed on the identified candidate text and associated contextual information. The text is parsed into structured, machine-readable information, including business domain, intent, slots, and other semantic information. An executable intent confidence score is obtained, and the semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.
[0074] [Business Management]
[0075] Based on the semantic parsing results of the query statement text, the semantic understanding module sends query instructions to the corresponding business management module to obtain the query results provided by the business service, as well as the actions required to "complete" the user's final request, and feeds back the device execution instructions corresponding to the query results.
[0076] [Language Generation]
[0077] Natural Language Generation (NLG) is configured to generate spoken text from information or instructions. Specifically, it can be categorized into casual conversation, task-oriented, knowledge-based question-answering, and recommendation-based systems. In casual conversation, NLG performs intent recognition and sentiment analysis based on context, then generates open-ended responses. In task-oriented conversations, learned strategies are used to generate responses, typically including clarifying needs, guiding the user, asking questions, confirming, and closing remarks. In knowledge-based question-answering conversations, the required knowledge (knowledge, entities, fragments, etc.) is generated based on question type identification and classification, information retrieval, or text matching. In recommendation-based conversation systems, user interests are matched, candidate recommendations are ranked, and then recommended content is generated for the user.
[0078] [Speech Synthesis]
[0079] The speech synthesis module is configured to present voice output to the user. It synthesizes voice output based on text provided by the digital assistant. For example, the generated dialogue response is in the form of a text string. The speech synthesis module then converts the text string into audible voice output.
[0080] It should be noted that, Figure 4 The architecture shown is merely an example and is not intended to limit the scope of protection of this application. Other architectures can also be used to achieve similar functions in the embodiments of this application. For example, all or part of the above process can be performed by the terminal device, which will not be elaborated here.
[0081] Based on the aforementioned terminal device 200, in some embodiments, the terminal device 200 can play media data provided by different signal sources, which can be audio data, video data, or audio-visual data. When the terminal device 200 is equipped with an audio output device 270, the audio data of the media data can be played through the audio output device; when the terminal device 200 is equipped with a display 260, video data can be played through the display 260.
[0082] In some embodiments, the media data can be data stored locally on the terminal device 200 or data provided by an external device connected to the terminal device 200. For example, the terminal device 200 can connect to an external device such as a set-top box or game console through the device interface 240, receive media data sent by the external device, and play the media data through the audio output device 270 and / or the display 260.
[0083] The following embodiments of this application describe the playback process of audio data, wherein the audio data can be an independent audio signal or an audio signal contained in audio and video data. For ease of distinction, the audio data corresponding to the media asset data in this application embodiment is represented as media asset audio data. The format of the media asset audio data can be various, such as WAV, MP3, WMA, AAC, Ogg Vorbis, RA, or APE, etc. This application does not impose any limitations on this.
[0084] like Figure 5 As shown, in some embodiments, the controller 250 of the terminal device 200 can be a motherboard 251. The motherboard 251 includes a System on Chip (SOC) and an Audio Power Amplifier (AMP). The audio output device 270 is a speaker SPK 271. The speaker SPK 271 is used to play audio data from the terminal device 200. The SOC is used to control the playback process of the audio data, such as acquiring audio data and setting the volume of the audio data. The amplifier AMP is used to amplify the loudness of the audio data according to the volume set by the SOC, so that the speaker SPK can play a sound of the corresponding loudness. The audio data processed by the amplifier AMP, which is the local audio stream to be played by the terminal device 200, is looped back to the SOC.
[0085] In some embodiments, the terminal device 200 can also collect mixed audio data via the detector 230. The sounds corresponding to the mixed audio data include the sounds played by the terminal device 200 and the external environmental sounds surrounding the terminal device 200. For example... Figure 5 As shown, detector 230 is a MIC (Microphone Array) array 231, which acquires and generates PDM (Pulse Density Modulation) audio data, i.e., mixed audio data. Then, the PDM audio data is sent to the SOC for processing.
[0086] In some embodiments, the terminal device 200 Figure 4 The application layer shown can also provide voice broadcast functionality through specific applications or services. For example, such as... Figure 4The voice application shown in the application layer can be configured with a voice broadcast function. Users can activate the voice broadcast function of the voice application through specific interactive actions, so that the terminal device 200 can play the audio data of the voice broadcast through the audio output device 270.
[0087] In some embodiments, a user can activate the voice broadcast function using a specific voice wake-up word, which triggers the terminal device 200 to play the corresponding broadcast audio data. For example, if the voice wake-up word is "HI, ABC", after the user outputs "HI, ABC" to the terminal device 200, the terminal device 200 can generate response broadcast audio data and play the broadcast audio data through the application or service of the voice broadcast function, thereby responding to the user's voice input.
[0088] In some embodiments, a user can activate the voice broadcast function by using a specific wake-up gesture or a control command input via a specific key, thereby triggering the terminal device 200 to play the corresponding broadcast audio data. For example, the terminal device 200 is equipped with a display 260 that supports touch functionality, allowing the user to trigger the voice broadcast function through a specific swipe gesture.
[0089] In some embodiments, users can input specific inquiry commands through voice interaction or other means. Terminal device 200 parses the intent of the inquiry command and responds to it via voice broadcasting. For example, a user can input the voice command "What time is it?" into terminal device 200. Terminal device 200 parses the voice command, generates broadcast audio data corresponding to the real-time time, and plays the broadcast audio data corresponding to the real-time time to respond to the user's voice command.
[0090] It is understood that the above-described voice playback function is merely an illustrative example and is not intended to limit the scope of the application. The voice broadcast in this embodiment may support more forms of interaction.
[0091] Accordingly, in some embodiments, the terminal device 200 can set the volume of the broadcast audio data. The volume of the broadcast audio data can be consistent with the volume of the media asset audio data, meaning the volumes of the broadcast audio data and the media asset audio data can be consistent with the system volume of the terminal device 200. Alternatively, the volume of the broadcast audio data can be a separately configured audio setting, meaning the terminal device 200 needs to set the volume of the broadcast audio data independently.
[0092] However, when the terminal device 200 simultaneously plays media asset audio data and broadcasts audio data, regardless of whether the volume of the broadcast audio data matches that of the media asset audio data, its volume setting is always manually configured by the user. Since the final loudness of the media asset audio data is also related to the volume of the media asset audio data source or the volume of the external device providing the media asset audio data, if the external device's volume setting is high or the media asset source's volume is high, the system volume setting of the terminal device 200 may be too low; conversely, if the external device's volume setting is low or the media asset source's volume is low, the system volume setting of the terminal device 200 may be too high. In this case, if the user uses the terminal device 200's voice broadcast function, the voice broadcast volume may also be too loud or too soft, resulting in a reduced auditory experience. For example, if the external environment of the terminal device 200 is noisy, a low voice broadcast volume may be difficult for the user to hear; conversely, if the external environment of the terminal device 200 is quiet, an excessively loud voice broadcast volume will also negatively impact the user experience.
[0093] To this end, some embodiments of this application provide a terminal device 200, which can automatically adjust the volume of the audio data being played based on multiple reference dimensions such as the sound of the external environment, the sound of the audio data being played locally, and user habits, so that the sound of the audio data being played is more in line with the user's needs, thereby improving the user's auditory experience.
[0094] like Figure 6 As shown, in some embodiments, the terminal device 200 may include an audio output device 270, a detector 230, and a controller 250. The audio output device 270 is configured to play output sound of internal audio data, including media asset audio data and broadcast audio data. The detector 260 is configured to collect mixed audio data, which includes audio data generated by collecting output sound and audio data generated by collecting external ambient sound. That is, the mixed audio data is generated based on the sound played by the audio output device 270 and the sound of the external environment. For example, the mixed audio data may be... Figure 5 The PDM audio data shown; such as Figure 6 As shown, the controller 250 is configured to perform the following program steps:
[0095] S6001: Obtain the base volume of the media asset audio data settings.
[0096] Terminal device 200 can obtain the base volume set in the media asset audio data. The base volume is the volume set by terminal device 200 for the media asset audio data, that is, the system volume of terminal device 200. For example, Figure 5The audio data playback process shown is such that, since the power amplifier AMP loops the audio data back to the SOC, the terminal device 200 can directly obtain the media asset audio data in the SOC through the application layer and calculate the basic volume corresponding to the media asset audio data.
[0097] Volume is used to represent the intensity of sound, and can also be expressed as loudness, intensity, or energy, which can be calculated from the signal amplitude within a sound frame. Therefore, in some embodiments, the terminal device 200 can acquire media audio data for a first preset time period to generate a first audio frame. Then, the first signal value of the first audio frame is extracted, and a base volume is calculated based on the first signal value. The base volume is calculated based on the sum of the absolute values of the first signal values, or based on the logarithm of the sum of the squares of the first signal values. The first preset time period can be customized, such as 1 minute or 30 seconds.
[0098] Regarding the calculation method of the sum of absolute values, in some embodiments, the volume is equal to the sum of the absolute values of the audio frame signal values, that is, the volume can be calculated based on the following formula:
[0099]
[0100] In the formula, volume is the volume, S i It is the i-th sample point in the audio frame, and n is the number of sample points in each audio frame. The base volume is the sum of the absolute values of the first audio frame.
[0101] In some embodiments, the volume is calculated by taking the logarithm of the sum, where the volume equals the sum of the squares of the signal values of the audio frames, then taking the logarithm base 10, and multiplying by 10. That is, the volume can be calculated based on the following formula:
[0102]
[0103] In the formula, volume is the volume, S i It is the i-th sampling point in the audio frame, and n is the number of sampling points in each audio frame. The base volume is calculated by taking the logarithm of the sum of the squared values of the first audio frame.
[0104] S6002: Calculate the first difference value between the base volume and the reference volume, wherein the reference volume is the average volume of the mixed audio data over a preset time period.
[0105] After obtaining the base volume corresponding to the media asset audio data, the terminal device 200 can determine how to set the volume of the broadcast audio data by comparing the base volume and the reference volume. The terminal device 200 can calculate the first difference between the base volume and the reference volume, such as calculating the difference between the base volume and the reference volume, or calculating the difference between the reference volume and the base volume. The reference volume is the average volume of the mixed volume data over a preset time period; that is, the reference volume is calculated based on the volume of the mixed audio data.
[0106] To better suit user habits, in some embodiments, the terminal device 200 can acquire mixed audio data over a second preset time period to generate a second audio frame. For example, as... Figure 5 In the audio data playback process shown, since the detector 230 sends the collected mixed audio data to the SOC, the terminal device 200 can directly obtain the mixed audio data of a preset time period in the SOC through the application layer to generate the second audio frame.
[0107] Similarly, after acquiring the second audio frame, the terminal device 200 can calculate the corresponding volume using the signal value of the second audio frame. That is, the terminal device 200 can extract the second signal value of the second audio frame and then calculate the sampled volume based on the second signal value. The sampled volume is calculated based on the sum of the absolute values of the second signal values, or by taking the logarithm of the sum of the average values of the second signal values. Then, a reference volume is calculated based on the sampled volume.
[0108] It is understood that the volume calculation method includes the sum of absolute values and the calculation method of taking the logarithm of the sum. For details, please refer to the basic volume calculation method provided in the above embodiments, which will not be repeated here.
[0109] In some embodiments, for calculating the reference volume, the terminal device 200 can calculate a first average value of the sampled volume within a third preset time period, and then perform a weighted average on the first average value within a target time period to generate a second average value. The third preset time period is longer than the second preset time period, and the target time period is longer than the third preset time period. For example, the second preset time period can be a specific duration of acquired mixed audio data, the third preset time period can be various time periods within a day, and the target time period can be a preset number of days. The second average value is then correlated with the third preset time period to generate the reference volume.
[0110] In other words, the terminal device 200 can calculate the average volume of mixed audio data at different times of the day, and then perform a weighted average of the calculated audio average over a preset number of days to obtain a reference volume for that time period based on user habits.
[0111] In some embodiments, the terminal device 200 can invoke the detector 230 to collect mixed audio data through a preset voice recording service. After the voice recording service is started, the terminal device 230 can copy and store the mixed audio data for a second preset time period, such as 1 minute, to generate a corresponding second audio frame. Then, the sampling volume corresponding to the mixed audio data is calculated using the second audio frame.
[0112] For example, the second preset time period is 1 minute, the third preset time period is each hour of the day, and the target time period is 30 days. Figure 7 As shown, after the terminal device 200 starts the voice recording service, it copies and stores the collected mixed audio data for 1 minute, and then calculates the sampling volume corresponding to the copied and stored mixed audio data. Then the mean of the audio during the time period T1-T2, i.e., the first average, is:
[0113]
[0114] In the formula, T1 is the start time of the time period to be calculated, T2 is the end time of the time period to be calculated, and the time span between T1 and T2 is 1 hour. t1 represents the start time of the calculation within the T1-T2 time period, t2 represents the end time of the calculation within the T1-T2 time period, and n represents the volume of the nth calculation within the T1-T2 time period.
[0115] The method for calculating n is as follows:
[0116]
[0117] In the formula, N represents the second preset time period. Using the formula for calculating the first average value, the terminal device 200 can calculate the average volume for different time periods within a day. The terminal device 200 can then perform a weighted average of the average volume for the same time period over 30 days to derive a reference volume based on user habits, as shown in the following formula:
[0118]
[0119] In the formula, X represents the target time period, and j represents the sampled value from 1 to X. Thus, for different time periods, the terminal device 200 can acquire different reference volumes and compare them with the real-time acquired base volume. The terminal device 200 can continuously calculate the reference volume value until the voice recording service stops running.
[0120] It should be noted that the values of the second preset time period, the third preset time period, and the target time period in the above example are merely one exemplary combination. In the embodiments of this application, the second preset time period, the third preset time period, and the target time period can be other values, provided that the second preset time period is less than the third preset time period and the target time period covers the entire third preset time period. This application does not impose any restrictions on this.
[0121] Accordingly, in some embodiments, after the terminal device 200 obtains the real-time base volume, it can obtain the reference volume corresponding to the current time period and calculate the difference between the base volume and the reference volume. For example, the reference volume is... The base volume obtained by terminal device 200 is Vy Terminal device 200 can calculate Vy- The difference Y is used as the first difference value. Based on the value of the first difference value, the terminal device 200 can determine whether the volume corresponding to the currently played media audio data differs too much from the user's usual volume.
[0122] S6003: Detect ambient noise volume based on internal audio data and the mixed audio data.
[0123] After calculating the first difference value, the terminal device 200 also detects the ambient noise volume of the external environment in which the terminal device 200 is located based on the internal audio data and the mixed audio data. The ambient noise volume is the volume corresponding to the external ambient sound. The terminal device 200 can determine whether the external environment is noisy by the height of the ambient noise volume, and adjust the volume of the broadcast audio data accordingly.
[0124] like Figure 8 As shown, in some embodiments, when detecting ambient noise volume, the terminal device 200 can acquire the mixed audio volume of the mixed audio data, and then calculate a second difference value between the base volume and the mixed audio volume. If the second difference value is less than a noise threshold, the ambient noise volume is marked as a first noise state; if the second difference value is greater than or equal to the noise threshold, the ambient noise volume is marked as a second noise state. The ambient noise volume in the second noise state is higher than that in the first noise state. The second noise state can be used to characterize a noisy external environment, while the first noise state can be used to characterize a quiet external environment. For example, the noise threshold can be 30~40dB.
[0125] For example, if the noise threshold is 40dB, the terminal device 200 can calculate the second difference value by mixing the audio data and the media asset audio data obtained by loopback. The second difference value is equal to the difference between the mixed audio volume and the base volume. Therefore, the second difference value is the volume corresponding to the external ambient sound. The terminal device 200 can determine whether the external ambient sound is noisy based on the value of the second difference value.
[0126] Since the audio data played by the terminal device 200 may be audio data sent by an external device, and the external device will have its corresponding volume settings, there will be multiple volume settings between the terminal device 200 and the external device. In order to adapt to the volume settings of the external device, the terminal device 200 is more likely to have a volume setting that is too high or too low. That is, when the terminal device 200 is connected to an external device, the volume is more likely to be too high or too low compared to the user's habits.
[0127] Therefore, in some embodiments, the terminal device 200 detects whether it is connected to an external device while detecting the ambient noise level. For example, the terminal device 200 can detect whether it is connected to an external device by listening to system broadcasts, checking the device connections of each device interface 240, or using MediaStore (media storage service). When the terminal device 200 is connected to an external device, it can mark the external device connection status for subsequent judgment and processing.
[0128] In some embodiments, the peripheral connection status can be implemented through specific flag bits and flag values. The flag values may include one or more types of characters, such as English characters, numeric characters, etc. The terminal device 200 can determine whether it is in a peripheral connection state, that is, whether it is connected to an external device, by reading the flag value of the flag bits.
[0129] S6004: Obtain the adjustment coefficient according to the first difference value and the ambient noise volume.
[0130] After calculating the first difference value and detecting the ambient noise volume, the terminal device 200 can obtain an adjustment coefficient based on the first difference value and the ambient noise volume. The adjustment coefficient is used to assist in setting the volume of the broadcast audio data. There can be multiple adjustment coefficients, and the terminal device 200 can determine the corresponding adjustment coefficient based on the first difference value and the ambient noise volume.
[0131] The following embodiments use the difference between the base volume and the reference volume as an example for illustration, but this is not a limitation. The first difference value can also be the difference between the reference volume and the base volume. The corresponding principle and logic are the same as those of the first difference value being the difference between the base volume and the reference volume, and will not be elaborated on in this application.
[0132] like Figure 9 As shown, in scenarios where the base volume is greater than the reference volume (i.e., the first difference value is greater than 0), in some embodiments, when the terminal device 200 obtains the adjustment coefficient, it obtains a first difference threshold. The first difference threshold, being greater than 0, is used to determine whether the difference between the current base volume and the user's habitual reference volume is too large. If the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is in a first noise state, the terminal device 200 obtains the first adjustment coefficient; if the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is in a second noise state, the terminal device 200 obtains a second adjustment coefficient. The first and second adjustment coefficients are both less than 1, with the first adjustment coefficient being less than the second adjustment coefficient.
[0133] In other words, if the base volume of the media audio data played by the terminal device 200 in real time is greater than a specific reference volume value, and the external environment is quiet, the terminal device 200 will obtain a first adjustment coefficient to slightly lower the volume of the audio data being played; if the external environment is noisy, the terminal device 200 will obtain a second adjustment coefficient to significantly lower the volume of the audio data being played.
[0134] It is worth noting that the small and large amplitudes described in the above embodiments are relative to the first and second adjustment coefficients.
[0135] In some embodiments, if the first difference value is greater than 0 but less than the first difference threshold, it indicates that the volume of the current media asset audio data is small compared with the user's habitual reference volume. In this case, there is no need to adjust the volume of the broadcast audio data, and the terminal device 200 does not perform any processing.
[0136] like Figure 9As shown, for scenarios where the base volume is less than the reference volume (i.e., the first difference value is less than 0), in some embodiments, when the terminal device 200 obtains the adjustment coefficient, it obtains a second difference threshold. The second difference threshold is less than 0, and the absolute values of the second difference threshold and the first difference threshold are equal. If the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is in the first noise state, the terminal device 200 obtains a third adjustment coefficient; and if the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is in the second noise state, it obtains a fourth adjustment coefficient. The third and fourth adjustment coefficients are both greater than 1, and the third adjustment coefficient is less than the fourth adjustment coefficient.
[0137] In other words, if the base volume of the media audio data played by the terminal device 200 in real time is lower than a specific reference volume value, and the external environment is quiet, the terminal device 200 will obtain a third adjustment coefficient to slightly increase the volume of the audio data being played; if the external environment is noisy, the terminal device 200 will obtain a fourth adjustment coefficient to significantly increase the volume of the audio data being played.
[0138] It is worth noting that the small and large amplitudes described in the above embodiments are relative to the third and fourth adjustment coefficients.
[0139] Similarly, in some embodiments, if the first difference value is greater than 0, but the first difference value is greater than the second difference threshold, it means that the volume of the current media asset audio data is small compared with the user's habitual reference volume. In this case, there is no need to adjust the volume of the broadcast audio data, and the terminal device 200 does not perform any processing.
[0140] For example, the first adjustment coefficient is 0.5, the second adjustment coefficient is 0.8, the third adjustment coefficient is 1.1, the fourth adjustment coefficient is 1.2, the first difference threshold is M, the second difference threshold is -M, and the first difference value between the base volume and the reference volume is Y. Then, when Y ≥ M > 0, if the external environment is quiet, the adjustment coefficient obtained by the terminal device 200 is 0.5; while if the external environment is noisy, the adjustment coefficient obtained by the terminal device 200 is 0.8. When Y ≤ -M < 0, if the external environment is quiet, the adjustment coefficient obtained by the terminal device 200 is 1.1; while if the external environment is noisy, the adjustment coefficient obtained by the terminal device 200 is 1.2.
[0141] To accommodate different audio output devices, in some embodiments, one of the second and third adjustment coefficients may have a value of 1. For example, the second adjustment coefficient may be 1, but in this case, the third adjustment coefficient must be greater than 1. Alternatively, the third adjustment coefficient may be 1, but the second adjustment coefficient must be less than 1.
[0142] In some embodiments, the terminal device 200 can also set different adjustment coefficients according to different peripheral connection states, namely, a peripheral-connected adjustment coefficient and a non-peripheral-connected adjustment coefficient. The absolute value of the peripheral-connected adjustment coefficient is greater than the absolute value of the non-peripheral-connected adjustment coefficient. That is, the first adjustment coefficient can include a first peripheral-connected adjustment coefficient and a first non-peripheral-connected adjustment coefficient; the second adjustment coefficient can include a second peripheral-connected adjustment coefficient and a second non-peripheral-connected adjustment coefficient; the third adjustment coefficient can include a third peripheral-connected adjustment coefficient and a third non-peripheral-connected adjustment coefficient; and the fourth adjustment coefficient can include a fourth peripheral-connected adjustment coefficient and a fourth non-peripheral-connected adjustment coefficient.
[0143] Correspondingly, the absolute value of the first adjustment coefficient with peripherals is greater than the absolute value of the first adjustment coefficient without peripherals, the absolute value of the second adjustment coefficient with peripherals is greater than the absolute value of the second adjustment coefficient without peripherals, the absolute value of the third adjustment coefficient with peripherals is greater than the absolute value of the third adjustment coefficient without peripherals, and the absolute value of the fourth adjustment coefficient with peripherals is greater than the absolute value of the fourth adjustment coefficient without peripherals.
[0144] For ease of explanation, some embodiments of this application combine peripheral connection status, ambient noise volume, and a first difference value. The first noise state represents a quiet environment, and the second noise state represents a noisy environment. The terminal device 200 is divided into 8 states as shown in the table below.
[0145] When Y≥M>0, the state of terminal device 200 is as follows:
[0146]
[0147] When Y≤-M<0, the state of terminal device 200 is as follows:
[0148]
[0149] For example, such as Figure 10 As shown, when the terminal device 200 obtains the adjustment coefficient, the adjustment coefficient corresponding to state A can be the first adjustment coefficient with peripheral device (0.6), the adjustment coefficient corresponding to state B can be the first adjustment coefficient without peripheral device (0.5), the adjustment coefficient corresponding to state C can be the second adjustment coefficient with peripheral device (0.9), the adjustment coefficient corresponding to state D can be the second adjustment coefficient without peripheral device (0.8), the adjustment coefficient corresponding to state E can be the third adjustment coefficient with peripheral device (1.2), the adjustment coefficient corresponding to state F can be the third adjustment coefficient without peripheral device (1.1), the adjustment coefficient corresponding to state G can be the fourth adjustment coefficient with peripheral device (1.4), and the adjustment coefficient corresponding to state H can be the fourth adjustment coefficient without peripheral device (1.3).
[0150] S6005: Calculate the target volume using the adjustment coefficient, the reference volume, and the base volume.
[0151] After obtaining the adjustment coefficient, the terminal device 200 calculates the target volume for broadcasting audio data by combining the adjustment coefficient, reference volume, and base volume. In this way, the calculated target volume takes into account both user habits and real-time volume, allowing the volume of the broadcast audio data to better suit the actual application scenario of the terminal device 200.
[0152] In some embodiments, when calculating the target volume, the terminal device 200 may calculate a third average of the reference volume and the base volume, and then generate the target volume based on the third average and an adjustment coefficient. The target volume is the product of the third average and the target adjustment coefficient.
[0153] For example, the target volume is The base volume is The reference volume corresponding to the basic volume time period is Terminal device 200 calculates the target volume using the following formula:
[0154]
[0155] In the formula, The adjustment coefficient obtained by terminal device 200.
[0156] S6006: Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
[0157] After calculating the target volume, the terminal device 200 sets the volume of the broadcast audio data to the target volume, enabling the audio output device 270 to play the voice broadcast function sound according to the target volume. In some embodiments, the terminal device 200 can set the volume of the broadcast audio data to the target volume through a SOC, and the power amplifier AMP amplifies the broadcast audio data according to the target volume, thereby enabling the audio output device to play the sound corresponding to the loudness of the target volume.
[0158] In some embodiments, the terminal device 200 may trigger the execution of steps S6001-S6006 according to a second preset time period, that is, the terminal device 200 automatically executes steps S6001-S6006 every second preset time period. For example, if the second preset time period is 1 minute, the terminal device 200 will trigger the execution of steps S6001-S6006 every minute.
[0159] In order not to interfere with the voice broadcast function that is currently in use, such as Figure 11As shown, in some embodiments, the terminal device 200 also detects the playback status of the broadcast audio data. If the playback status is a first playback status, the step of obtaining the base volume set by the media asset audio data is executed, i.e., step S6001; while if the playback status is a second playback status, the step of obtaining the base volume set by the media asset audio data is executed when the playback status of the broadcast audio data changes to the first playback status. The first playback status indicates that the broadcast audio data is not currently being played, and the second playback status indicates that the broadcast audio data is currently being played.
[0160] In other words, if the terminal device 200 triggers the execution of steps S6001-S6006, the terminal device 200 is performing voice broadcasting, that is, playing the broadcast audio data corresponding to the voice broadcasting function. The terminal device 200 then waits for the voice broadcasting to end before setting the newly calculated target volume.
[0161] In some embodiments, when the terminal device 200 detects the playback status of the audio data, it can use the AudioManager class to detect whether the terminal device 200 is currently playing audio data. Alternatively, the terminal device 200 can check running services or processes to detect whether it is currently playing audio data. For example, the terminal device 200 can use ActivityManagerService or PackageManager to query the list of currently running services or processes to see if the service or process for voice playback is included. Alternatively, the terminal device 200 can also listen to system broadcasts or logs to detect whether it is currently playing audio data. For example, the terminal device 200 can register a BroadcastReceiver to listen for notifications used to implement the voice playback function.
[0162] In some embodiments, the terminal device 200 is equipped with a display 260, and when the user interface corresponding to the voice broadcast is displayed on the display 260 during voice broadcast, the terminal device 200 can also determine whether the terminal device 200 is currently playing broadcast audio data by taking a screenshot and analyzing the user interface.
[0163] It should be noted that the methods for detecting whether the terminal device 200 is currently playing audio data in the above embodiments are merely illustrative examples. Other methods, such as third-party libraries or tools, can also be used to detect whether the terminal device 200 is currently playing audio data in this application embodiment. This application does not impose any restrictions on this.
[0164] Based on the aforementioned terminal device 200, some embodiments of this application also provide a method for adjusting the volume of voice broadcasting, which can be applied to the terminal device 200 provided in the above embodiments. For example... Figure 6As shown, the method includes the following steps:
[0165] S6001: Obtain the base volume of the media asset audio data settings;
[0166] S6002: Calculate the first difference value between the base volume and the reference volume, wherein the reference volume is the average volume of the mixed audio data over a preset time period;
[0167] S6003: Detect ambient noise volume based on internal audio data and the mixed audio data;
[0168] S6004: Obtain the adjustment coefficient according to the first difference value and the ambient noise volume;
[0169] S6005: Calculate the target volume using the adjustment coefficient, the reference volume, and the base volume;
[0170] S6006: Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
[0171] As can be seen from the above technical solutions, the terminal device and voice broadcast volume adjustment method provided in some embodiments of this application can obtain the base volume set by the media asset audio data and calculate a first difference value between the base volume and the reference volume. The reference volume is the average volume of the mixed audio data over a preset time period. Then, the ambient noise volume is detected based on the internal audio data and the mixed audio data, and an adjustment coefficient is obtained according to the first difference value and the ambient noise volume. A target volume is calculated using the adjustment coefficient, the reference volume, and the base volume, and the volume of the broadcast audio data is set to the target volume, so that the audio output device plays the broadcast audio data at the target volume. This method can automatically adjust the volume of voice broadcast in the terminal device 200 by combining multiple reference dimensions of the terminal device 200, making the volume of the voice broadcast more suitable for user needs, thereby improving the user's auditory experience.
[0172] The same or similar parts between the various embodiments in this specification can be referred to each other, and will not be repeated here.
[0173] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of the embodiments of the present invention.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
[0175] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.
Claims
1. A terminal device, characterized in that, include: An audio output device is configured to output sound that plays internal audio data, including media asset audio data and broadcast audio data. The detector is configured to acquire mixed audio data, which includes audio data generated by acquiring the output sound and audio data generated by acquiring external ambient sound. The controller is configured as follows: Obtain the base volume set by the media asset audio data; Calculate a first difference between the base volume and the reference volume, where the reference volume is the average volume of the mixed audio data over a target time period; Detect ambient noise volume based on the internal audio data and the mixed audio data; The adjustment coefficient is obtained based on the first difference value and the ambient noise volume. The target volume is calculated using the adjustment factor, the reference volume, and the base volume. Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
2. The terminal device according to claim 1, characterized in that, The controller executes the acquisition of the media asset audio data and sets the base volume, which is configured as follows: Acquire the media asset audio data for a first preset time period to generate a first audio frame; Extract the first signal value of the first audio frame; The base volume is calculated based on the first signal value. The base volume is calculated based on the sum of the absolute values of the first signal values, or based on the logarithm of the sum of the squares of the first signal values.
3. The terminal device according to claim 1, characterized in that, The controller is also configured to: Acquire the mixed audio data for a second preset time period to generate a second audio frame; Extract the second signal value of the second audio frame; The sampling volume is calculated based on the second signal value, which is calculated based on the sum of the absolute values of the second signal value, or based on the logarithm of the sum of the squares of the second signal value. The reference volume is calculated based on the sampled volume.
4. The terminal device according to claim 3, characterized in that, The controller performs the calculation of the reference volume based on the sampled volume and is also configured to: Calculate the first average value of the sampled volume within a third preset time period, wherein the third preset time period is greater than the second preset time period; A weighted average is applied to the first average value within the target time period to generate a second average value, wherein the target time period is longer than the third preset time period; The second average value is correlated with the third preset time period to generate the reference volume.
5. The terminal device according to claim 1, characterized in that, The controller is configured to detect ambient noise volume based on the internal audio data and the mixed audio data. Obtain the mixed audio volume of the mixed audio data; Calculate a second difference value between the base volume and the mixed audio volume; If the second difference value is less than the noise threshold, the ambient noise volume is marked as a first noise state; If the second difference value is greater than or equal to the noise threshold, the ambient noise volume is marked as a second noise state, where the ambient noise volume of the second noise state is higher than that of the first noise state.
6. The terminal device according to claim 5, characterized in that, If the first difference value is greater than 0, the controller executes an adjustment coefficient based on the first difference value and the ambient noise volume, configured as follows: Obtain a first difference threshold, wherein the first difference threshold is greater than 0; If the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is the first noise state, then the first adjustment coefficient is obtained; If the first difference value is greater than or equal to the first difference threshold, and the ambient noise volume is the second noise state, then a second adjustment coefficient is obtained; the first adjustment coefficient and the second adjustment coefficient are less than 1, and the first adjustment coefficient is less than the second adjustment coefficient.
7. The terminal device according to claim 6, characterized in that, If the first difference value is less than 0, the controller executes an adjustment coefficient based on the first difference value and the ambient noise volume, configured as follows: Obtain a second difference threshold, the second difference threshold being less than 0, and the absolute values of the second difference threshold and the first difference threshold being equal; If the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is the first noise state, then a third adjustment coefficient is obtained; If the first difference value is less than or equal to the second difference threshold, and the ambient noise volume is the second noise state, then a fourth adjustment coefficient is obtained, wherein the third adjustment coefficient and the fourth adjustment coefficient are greater than 1, and the third adjustment coefficient is less than the fourth adjustment coefficient.
8. The terminal device according to claim 1, characterized in that, The controller is configured to calculate the target volume using the adjustment factor, the reference volume, and the base volume. Calculate the third average of the reference volume and the base volume; The target volume is generated based on the third average value and the adjustment coefficient, wherein the target volume is the product of the third average value and the adjustment coefficient.
9. The terminal device according to claim 1, characterized in that, The controller is also configured to: Detect the playback status of the broadcast audio data; If the playback state is the first playback state, the step of obtaining the basic volume set by the media asset audio data is executed. The first playback state is used to indicate that the broadcast audio data is not currently being played. If the playback state is the second playback state, when the playback state of the broadcast audio data changes to the first playback state, the step of obtaining the basic volume set by the media asset audio data is executed. The second playback state is used to indicate that the broadcast audio data is being played.
10. A method for adjusting the volume of a voice broadcast, characterized in that, Applied to the terminal device according to any one of claims 1-9, the method includes: Get the base volume settings for media asset audio data; Calculate a first difference between the base volume and the reference volume, where the reference volume is the average volume of the mixed audio data over a preset time period; Detect ambient noise volume based on internal audio data and the mixed audio data; The adjustment coefficient is obtained based on the first difference value and the ambient noise volume. The target volume is calculated using the adjustment factor, the reference volume, and the base volume. Set the volume of the broadcast audio data to the target volume so that the audio output device plays the sound of the broadcast audio data at the target volume.
Citation Information
Patent Citations
Method and device for automatically adjusting equipment voice broadcast volume
CN111427536A
Volume adjustment method and device, equipment and storage medium
CN112151025A