Audio switching control method and device, equipment and storage medium

By acquiring and fusing audio signals, hardware status, and scene event information in the vehicle audio system, predicting audio interruption events and adaptively calculating transition parameters, the problem of plosive sounds and auditory gaps during audio switching under multiple concurrent audio sources is solved, resulting in smoother audio switching and a better user experience.

CN121744221APending Publication Date: 2026-03-27WUTONG AUTOLINK (BEIJING) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing in-vehicle audio systems lack a unified management and control mechanism in scenarios with multiple concurrent audio sources and real-time interruptions, resulting in popping sounds and auditory gaps during audio switching. Furthermore, existing solutions cannot dynamically adapt to user needs, affecting user experience.

Method used

By acquiring the current audio signal, the computing power status of the vehicle hardware, and scene event information, feature fusion is performed to predict audio interruption events and time windows. Based on this, transition parameters are adaptively calculated, and smooth transition processing is performed to achieve proactive prediction and optimization of audio switching.

Benefits of technology

It improves the smoothness of audio switching, reduces auditory gaps, enhances the user experience, and optimizes system resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744221A_ABST
    Figure CN121744221A_ABST
Patent Text Reader

Abstract

The invention provides an audio switching control method and device, equipment and a storage medium, and the method comprises the steps: obtaining a current audio signal, a vehicle-mounted hardware computing power state and scene event information, carrying out the feature fusion, obtaining a fusion perception feature, predicting an audio interruption event and a corresponding time window based on the fusion perception feature, and carrying out the prediction of the audio interruption event. Determining corresponding audio transition processing parameters according to the audio interruption event, the corresponding time window, the current audio signal and the computing power state of the vehicle-mounted hardware, and performing smooth transition processing on the interrupted audio and the current audio signal based on the audio transition processing parameters to complete audio switching; according to the method, the current audio signal, the computing power state of the vehicle-mounted hardware and the scene event information are acquired, feature fusion is performed to predict the audio interruption event and the time window, and the audio transition processing parameters are adaptively determined and the smoothing processing is executed based on the prediction, so that active prediction and optimization of audio switching are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cockpit audio control technology, and in particular to an audio switching control method, device, equipment and storage medium. Background Technology

[0002] As automobiles become increasingly intelligent and connected, in-vehicle infotainment systems have evolved from simple audio playback devices into the core of integrated intelligent cockpits. However, multiple audio sources, such as navigation prompts, multimedia music, telephone calls, voice-synthesized announcements, button feedback sounds, and emergency alarm sounds, often play concurrently within the vehicle, making the acoustic environment more complex. High-fidelity music playback may be interrupted by navigation commands, online audio content playback may be interrupted by emergency calls, or information broadcasts may be interrupted by switching driving modes. This multi-source concurrency and real-time interruption scenario places higher demands on the in-vehicle audio system's real-time scheduling and smooth switching capabilities.

[0003] Currently, there is a lack of unified management and control mechanisms for switching control solutions in audio interruption scenarios. These solutions are designed only for single audio source types or specific scenarios, failing to establish an integrated processing framework covering all audio sources and scenarios. Secondly, existing processing modes are mostly passive responses and post-event remedies, mainly initiating the correction process only after the abrupt change in audio switching has already occurred. Due to processing delays, it is difficult to completely eliminate the popping sounds and auditory gaps generated at the moment of switching, affecting the continuity of the experience. On the other hand, existing solutions mostly use fixed processing parameters, which cannot be appropriately optimized for different scenarios. As vehicle usage time increases and hardware performance changes, as well as the personalized operating habits formed by different users, existing static solutions cannot be adapted and optimized to actual needs, resulting in a deterioration in user experience. Summary of the Invention

[0004] The purpose of this application is to provide an audio switching control method, apparatus, device, and storage medium to solve the above-mentioned technical problems.

[0005] This application provides an audio switching control method, which includes: acquiring a current audio signal, vehicle hardware computing power status, and scene event information; performing feature fusion on the current audio signal, vehicle hardware computing power status, and scene event information to obtain fused perception features, and predicting an audio interruption event and a corresponding time window based on the fused perception features, wherein the audio interruption event refers to an audio signal switching event that interrupts the current audio signal with an interrupting audio associated with the scene event information; determining corresponding audio transition processing parameters based on the audio interruption event and the corresponding time window, as well as the current audio signal and the vehicle hardware computing power status; and performing smooth transition processing on the interrupting audio and the current audio signal based on the audio transition processing parameters to complete the audio switching, wherein the interrupting audio is determined based on the scene event information.

[0006] In one embodiment of this application, feature fusion is performed on the current audio signal, the vehicle hardware computing power status, and scene event information to obtain fused perception features. This includes: quantizing and aligning the real-time amplitude, sampling rate, and number of channels of the current audio signal with the computing power occupancy rate of the vehicle hardware computing power status, and then concatenating them to generate a first feature vector, wherein the vehicle hardware computing power status represents the computing power occupancy rate of the audio digital signal processor; converting the scene event information into event identifiers according to a preset event type encoding rule to obtain a second feature vector, wherein the scene event information is a scene related to audio source interruption determined by listening to system broadcasts, vehicle bus signals, or user interaction events; and aligning and concatenating the first feature vector and the second feature vector in the time dimension to generate fused perception features.

[0007] In one embodiment of this application, predicting audio interruption events and corresponding time windows based on the fused sensing features includes: inputting the fused sensing features into a hybrid prediction framework for processing, the hybrid prediction framework including a rule engine, a decision tree model, and a long short-term memory network model; based on the rule engine, matching the scene event information in the fused sensing features with preset audio source priority rules, judging high-confidence audio interruption events, and outputting rule judgment results; through the decision tree model, performing probabilistic inference based on the current audio signal features and historical operation data in the fused sensing features, and outputting the probability value of the audio interruption event; according to the long short-term memory network model, analyzing the corresponding time dependencies of the scene event information in the fused sensing features, and outputting the predicted probability of the audio interruption event and the corresponding time window; and performing weighted processing based on the rule judgment result, the probability value of the audio interruption event, the predicted probability of the audio interruption event, and the corresponding time window to obtain the audio interruption event and the corresponding time window.

[0008] In one embodiment of this application, the hybrid prediction framework further includes: a tiered response based on the predicted probability of an audio interruption event, wherein the tiered response includes: if the predicted probability of an audio interruption event is greater than or equal to a first threshold, then triggering the early execution of a smooth transition process based on a preset advance duration within the prediction time window; if the predicted probability of an audio interruption event is less than the first threshold but greater than or equal to a second threshold, then monitoring the amplitude of the current audio signal, and triggering the early execution of a smooth transition process based on a preset advance duration when the amplitude of the current audio signal is greater than or equal to a preset amplitude threshold; and if the predicted probability of an audio interruption event is less than the second threshold, then only monitoring the amplitude of the current audio signal.

[0009] In one embodiment of this application, determining the corresponding audio transition processing parameters based on the audio interruption event and the corresponding time window, as well as the current audio signal and the computing power status of the vehicle hardware, includes: determining a base duration for fade-in and fade-out processing based on the duration of the interrupted audio; determining an attenuation curve based on the frequency of the interrupted audio, the attenuation curve including a logarithmic attenuation curve and an exponential attenuation curve, the logarithmic attenuation curve being used for interrupted audio with a high proportion of low frequencies, and the exponential attenuation curve being used for interrupted audio with a high proportion of high frequencies; calculating the fade-out processing step size and the fade-in processing slope of the current audio signal based on the sampling rate of the current audio signal, the target amplitude change requirement, and the computing power status of the vehicle hardware, the target amplitude change requirement characterizing the expected amplitude change value of the current audio signal from the current amplitude to the target amplitude; and determining the base duration, attenuation curve, fade-out processing step size, and fade-in processing slope as audio transition processing parameters.

[0010] In one embodiment of this application, determining the base duration for fade-in / fade-out processing based on the duration of the interrupted audio includes: if the duration of the interrupted audio is less than a first duration threshold, then a first base duration is used; if the duration of the interrupted audio is greater than a second duration threshold, then a second base duration is used, wherein the first duration threshold is less than the second duration threshold; if the duration of the interrupted audio is less than or equal to a third duration threshold, then a third base duration is used, wherein the third duration threshold is less than the first duration threshold.

[0011] In one embodiment of this application, performing smooth transition processing on interrupted audio and current audio signal based on the audio transition processing parameters includes: performing linear fade-out processing on the current audio signal based on fade-out processing step size and attenuation curve; when the amplitude of the current audio signal attenuates to a preset auditory neglect threshold, performing linear fade-in processing on the interrupted audio according to fade-in processing slope and attenuation curve; when the interrupted audio reaches the target amplitude based on the target amplitude change requirement, completing the switch from the current audio signal to the interrupted audio, and releasing the audio processing resources associated with the current audio signal.

[0012] In one embodiment of this application, the audio switching control method further includes: upon completion of audio switching, acquiring the switched audio signal output during the audio switching process, and determining an objective smoothness score based on the number of amplitude abrupt changes between adjacent sampling points in the switched audio signal; acquiring an audio switching experience score fed back by the human-computer interaction interface, and determining the audio switching experience score as a subjective smoothness score; and optimizing the audio switching control based on the objective smoothness score and the subjective smoothness score, wherein the audio switching control optimization includes optimizing the prediction logic of audio interruption events and optimizing the calculation logic of audio transition processing parameters.

[0013] In one embodiment of this application, optimizing audio switching control based on the current objective smoothness score and the current subjective smoothness score includes: constructing a reinforcement learning model with the optimization objectives of maximizing the subjective smoothness score and minimizing the objective smoothness score, wherein the reinforcement learning model uses the prediction logic of the current audio interruption event and the calculation logic of the audio transition processing parameters as variables; normalizing the current subjective smoothness score and the current objective smoothness score respectively, and substituting them into a preset reward function to calculate the reward value for the current audio switching, wherein the reward function is configured to provide positive incentives for the subjective smoothness score that shows an upward trend and negative incentives for the objective smoothness score that shows an upward trend; and training the reinforcement learning model based on the reward value to optimize the prediction logic of the audio interruption event and the calculation logic of the audio transition processing parameters.

[0014] This application embodiment also provides an audio switching control device, which includes: an audio signal processing module, used to acquire the current audio signal, the vehicle hardware computing power status, and scene event information; perform feature fusion on the current audio signal, the vehicle hardware computing power status, and the scene event information to obtain fused perception features, and predict audio interruption events and corresponding time windows based on the fused perception features, wherein the audio interruption event refers to an audio signal switching event in which an interrupting audio associated with the scene event information interrupts the current audio signal; and an audio smoothing switching module, used to determine corresponding audio transition processing parameters based on the audio interruption event and the corresponding time window, as well as the current audio signal and the vehicle hardware computing power status; and perform smoothing transition processing on the interrupting audio and the current audio signal based on the audio transition processing parameters to complete the audio switching, wherein the interrupting audio is determined according to the scene event information.

[0015] This application also provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the audio switching control method as described in any of the above embodiments.

[0016] This application also provides a computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a computer's processor, cause the computer to perform the audio switching control method as described in any of the above embodiments.

[0017] The beneficial effects of this invention are as follows: This application provides an audio switching control method, apparatus, device, and storage medium. By acquiring the current audio signal, the vehicle's hardware computing power status, and scene event information, and performing feature fusion, a fused perception feature is obtained. Based on the fused perception feature, an audio interruption event and its corresponding time window are predicted. According to the audio interruption event, the corresponding time window, the current audio signal, and the vehicle's hardware computing power status, corresponding audio transition processing parameters are determined. Based on the audio transition processing parameters, a smooth transition processing is performed on the interrupted audio and the current audio signal to complete the audio switching. This application achieves proactive prediction and optimization of audio switching by acquiring the current audio signal, the vehicle's hardware computing power status, and scene event information, performing feature fusion to predict the audio interruption event and time window, and adaptively calculating transition parameters based on this to perform smooth processing.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application; Figure 2 This is a flowchart illustrating an exemplary embodiment of an audio switching control method according to this application; Figure 3 This is a flowchart illustrating a specific audio switching control method according to an exemplary embodiment of this application; Figure 4 This is a schematic diagram of an audio switching control device shown in an exemplary embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer system for an electronic device, as illustrated in an exemplary embodiment of this application. Detailed Implementation

[0020] The embodiments of this application will be described below with reference to the accompanying drawings and specific examples. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be understood that the preferred embodiments are only for illustrating this application and are not intended to limit the scope of protection of this application.

[0021] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the illustrations only show the components related to this application and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.

[0023] The term "and / or" used in this application describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship.

[0024] In in-vehicle infotainment systems, concurrent multi-source audio and real-time interruptions can cause popping sounds and auditory gaps during audio switching. However, existing solutions are primarily designed for single audio source types or specific scenarios, failing to establish an integrated processing framework covering all audio sources and scenarios. Furthermore, their processing modes are mostly passive responses and post-event remedies, initiating correction processes only after a sudden audio switching event occurs. Due to processing delays, it's difficult to completely eliminate the acoustic abrupt changes during switching. In addition, existing solutions use fixed processing parameters, lacking dynamic adaptation, which affects the smoothness of audio switching and the real-time response of the system. For example, if music is playing while the vehicle is in motion, and the navigation system triggers a turn prompt, the existing switching control mechanism forcibly interrupts the current audio signal the moment the navigation prompt intervenes, causing popping sounds during the switching process, interrupting the auditory experience, reducing the clarity of navigation instructions, and resulting in incomplete information reception. Based on this, this application proposes an audio switching control method. By acquiring the current audio signal, the computing power status of the vehicle hardware, and scene event information, feature fusion is performed to predict audio interruption events and time windows. Based on this, transition parameters are adaptively calculated and smoothing processing is performed to achieve proactive prediction and optimization of audio switching, improve audio switching smoothness, reduce auditory breaks, and thus enhance the user experience.

[0025] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application.

[0026] Reference Figure 1 As shown, the system architecture may include a vehicle 110 and a computer device 120. The computer device 120 acquires the current audio signal, the onboard hardware computing power status, and scene event information from the vehicle 110 to obtain fused perception features. Based on these features, it predicts audio interruption events and corresponding time windows. Then, based on the interruption events, time windows, current audio signal, and onboard hardware computing power status, it determines corresponding audio transition processing parameters. Based on these parameters, it performs smooth transition processing on the interrupted audio and the current audio signal to complete the audio switching. The vehicle 110 includes an in-vehicle infotainment system with audio-visual functions and at least a microphone to collect audio signals within the cabin. The computer device 120 refers to a computing power support terminal device that carries the program implementation environment for the audio switching control method, including but not limited to tablets, microcomputers, embedded computers, and cloud servers.

[0027] Figure 2 This is a flowchart illustrating an exemplary embodiment of the present application of an audio switching control method, which can... Figure 1It can be executed in the implementation environment described above, but it can also be implemented in other implementation environments. No specific limitations are imposed on the aforementioned implementation environments here. (See also...) Figure 2 As shown, the flowchart of this audio switching control method includes at least steps S210 to S240, which are described in detail below: In step S210, the current audio signal, the computing power status of the vehicle hardware, and scene event information are obtained.

[0028] The current audio signal refers to the audio stream being played or output in the in-vehicle environment, such as multimedia music, radio programs, or telephone calls, representing the user's current primary auditory focus; the aforementioned in-vehicle hardware computing power status refers to the computing resource usage of the in-vehicle infotainment system or related audio processing modules, such as CPU utilization and memory usage; the aforementioned scenario event information represents various events occurring in the in-vehicle environment that may lead to audio switching or interruption, including but not limited to navigation commands, incoming call notifications, vehicle warnings, user voice commands, or driving mode switching.

[0029] In one embodiment of this application, the current audio signal can be obtained by monitoring the digital signal stream of the vehicle audio output channel, and by sampling or reading data from the audio buffer, including but not limited to information such as the real-time amplitude, sampling rate, and number of channels of the current audio signal; the computing power status of the vehicle hardware is obtained by querying the vehicle system resource manager, including reading CPU load, memory usage, DSP chip occupancy, and HAL layer soft coding load; scene event information is obtained by monitoring event messages from the vehicle's internal communication bus, operating system broadcasts, or user human-machine interaction operations.

[0030] In step S220, the current audio signal, the computing power status of the vehicle hardware, and the scene event information are fused to obtain fused perception features, and the audio interruption event and the corresponding time window are predicted based on the fused perception features.

[0031] In one embodiment of this application, the aforementioned audio interruption event refers to an audio signal switching event that interrupts the current audio signal and is associated with scene event information.

[0032] In one embodiment of this application, the real-time amplitude, sampling rate, and number of channels of the current audio signal are quantized, aligned, and formatted with the computing power occupancy rate of the vehicle hardware computing power status, and then concatenated to generate a first feature vector. The vehicle hardware computing power status represents the computing power occupancy rate of the audio digital signal processor.

[0033] In one embodiment of this application, scene event information is converted into event identifiers according to a preset event type encoding rule to obtain a second feature vector. The scene event information is a scene related to the interruption of the sound source, which is determined by listening to system broadcasts, vehicle bus signals or user interaction events. The first feature vector and the second feature vector are aligned and spliced ​​in the time dimension to generate fused perception features.

[0034] In the embodiments of this application, audio signal features and hardware performance indicators from different sources with different data types and dimensions are converted into a unified numerical representation that can be processed subsequently, reducing data complexity and improving processing efficiency. As one possible implementation, the real-time amplitude can be normalized to the range of 0 to 1, and the sampling rate and number of channels can be directly included as integer values. The computing power utilization rate can also be normalized and arranged into a fixed-length array based on a predetermined order to obtain a first feature vector. In some other feasible environments, short-time Fourier transform can be performed on the real-time amplitude to extract spectral features, and the sampling rate and number of channels can be encoded, and the computing power utilization rate can be segmented and quantized, then combined into a first feature vector through vector concatenation.

[0035] The aforementioned scene event information is stored and represented as a text string, which needs to be converted into a structured event identifier. The aforementioned preset event type encoding rules define how to map each specific scene event to a unique numerical value or numerical vector, while the second feature vector represents the set of numerical identifiers after encoding conversion. Specifically, each event type can be converted into a binary vector to represent the existence of the event type; or a unique integer value can be assigned to each event type using label encoding.

[0036] In the embodiments of this application, the first feature vector and the second feature vector may be generated at different time points or have different sampling frequencies. To ensure that the subsequent prediction model accurately understands the information correlation, they must be aligned in the time dimension. In specific implementation, a fixed time window can be defined, and all relevant audio, hardware, and scene event data can be collected within this window and sorted and expressed according to timestamps.

[0037] In one embodiment of this application, fused perceptual features are input into a hybrid prediction framework for processing. This hybrid prediction framework includes a rule engine, a decision tree model, and a long short-term memory network model.

[0038] The aforementioned hybrid prediction framework is a software module integrating multiple prediction models and / or rules, deployed on the main control unit of the in-vehicle infotainment system, and can communicate with the audio processing unit via the in-vehicle network. The aforementioned rule engine is a component used to execute predefined business rules, matching input data with rules in a rule base; the aforementioned decision tree model is a supervised learning algorithm that uses a tree structure to represent decision rules, used for classification or regression prediction of new data; the aforementioned long short-term memory network model is a recurrent neural network that learns and remembers long-term dependencies, and can be used to process time-series data and capture temporal correlations between events.

[0039] Specifically, based on a rule engine, scene event information from the fused perception features is matched with preset audio source priority rules to determine high-confidence audio interruption events and output the rule determination result. Using a decision tree model, probabilistic inference is performed based on the current audio signal features and historical operation data from the fused perception features to output the probability value of the audio interruption event. Based on a long short-term memory network model, the time dependencies of scene event information in the fused perception features are analyzed to output the predicted probability of the audio interruption event and its corresponding time window. The rule determination result, the probability value of the audio interruption event, the predicted probability of the audio interruption event, and the corresponding time window are weighted and processed to obtain the audio interruption event and its corresponding time window.

[0040] This application's solution incorporates fused perceptual features into a hybrid prediction framework. A rule engine rapidly identifies high-confidence audio interruption events based on preset priority rules, providing deterministic judgments. A decision tree model performs probabilistic inference based on current audio signal features and historical operation data, capturing the impact of user behavior patterns and audio signal characteristics on the probability of interruption events. Furthermore, a long short-term memory network model further analyzes the temporal dependencies of scene event information, predicting the probability of audio interruption events and providing corresponding time windows. Finally, the different prediction results are weighted to obtain comprehensive and accurate predictions of audio interruption events and their time windows, overcoming the limitations of a single prediction mechanism and enabling the system to more accurately predict various complex audio interruption scenarios.

[0041] In one embodiment of this application, the hybrid prediction framework further includes a hierarchical response based on the predicted probability of the audio interruption event. The hierarchical response includes: If the predicted probability of an audio interruption event is greater than or equal to the first threshold, then the smooth transition processing will be triggered in advance within the predicted time window based on a preset advance duration. If the predicted probability of an audio interruption event is less than the first threshold and greater than or equal to the second threshold, then monitor the amplitude of the current audio signal, and trigger the early execution of smooth transition processing based on the preset advance duration when the amplitude of the current audio signal is greater than or equal to the preset amplitude threshold. If the predicted probability of an audio interruption event is less than the second threshold, then only the amplitude of the current audio signal is monitored.

[0042] Please refer to Figure 3 As shown, Figure 3 This is an exemplary embodiment of the present application illustrating a specific audio switching control method's switching decision flowchart. In one specific embodiment, the signal acquisition period can be selected as 10ms. During this period, audio features including the current audio signal, hardware computing power including the vehicle's hardware computing power status, and scene event information are acquired and fused to obtain fused perception features. Based on the fused perception features, a three-level prediction is performed within a hybrid prediction framework using a rule engine, a decision tree model, and an LSTM network (Long Short-Term Memory network model) to obtain the fused prediction result. Then, based on probability judgment, a graded response is performed, categorized as high probability, medium probability, and low probability. The first threshold can be selected as 80%, and the second threshold can be selected as 30%. When the prediction probability is greater than or equal to the first threshold, the smooth transition processing is triggered within a 50-150ms interval based on a preset advance time within the prediction time window. When the prediction probability is less than the first threshold but greater than or equal to the second threshold, the amplitude of the current audio signal is monitored, and the smooth transition processing is triggered within a 50-150ms interval based on a preset advance time when the amplitude of the current audio signal is greater than or equal to 20 dB of a preset amplitude threshold. If the amplitude of the current audio signal is less than 20 dB of the preset amplitude threshold, only the amplitude of the current audio signal is monitored. When the prediction probability is less than the second threshold, only the amplitude of the current audio signal is monitored.

[0043] By introducing a hierarchical response mechanism into the hybrid prediction framework, the system can intelligently adjust the triggering strategy for smooth transition processing based on the predicted probability of audio interruption events. When the prediction probability is high, the system proactively executes the smooth transition in advance to ensure timely and smooth switching. When the prediction probability is medium, the system uses the real-time amplitude of the current audio signal for auxiliary judgment to avoid unnecessary premature operation while maintaining a certain level of responsiveness. When the prediction probability is low, the system only monitors the amplitude to avoid misoperation and resource waste caused by low-confidence predictions. This hierarchical response strategy, closely integrated with the prediction results of the hybrid prediction framework, forms a more complete and intelligent audio switching control process, enabling the system to optimize resource utilization efficiency while ensuring user experience.

[0044] In step S230, the corresponding audio transition processing parameters are determined based on the audio interruption event and the corresponding time window, as well as the current audio signal and the computing power status of the vehicle hardware.

[0045] In one embodiment of this application, the base duration of fade-in / fade-out processing is determined based on the duration of the interrupted audio; if the duration of the interrupted audio is less than a first duration threshold, the first base duration is used; if the duration of the interrupted audio is greater than a second duration threshold, the second base duration is used; if the duration of the interrupted audio is less than or equal to a third duration threshold, the third base duration is used; the third duration threshold is less than the first duration threshold, and the first duration threshold is less than the second duration threshold.

[0046] The aforementioned baseline duration is the basic time parameter for fade-in / fade-out processing, used to set the time frame for the transition process. Specifically, it can be categorized into three types—short, medium, and long—based on the duration of the interrupted audio, corresponding to very short, moderate, and slightly long baseline durations, respectively. This ensures that short prompts can quickly intervene and exit, while longer audio information can achieve a smoother transition. This duration threshold and baseline duration can be pre-configured in the system or adjusted according to user preferences; no specific baseline duration parameter is limited here.

[0047] In one embodiment of this application, an attenuation curve is determined based on the frequency of the interrupted audio. The attenuation curve includes a logarithmic attenuation curve and an exponential attenuation curve. The logarithmic attenuation curve is used for interrupted audio with a high proportion of low frequencies, and the exponential attenuation curve is used for interrupted audio with a high proportion of high frequencies.

[0048] The aforementioned decay curves define the specific trajectory of the audio signal's amplitude change over time during fade-in and fade-out processes. Because the human ear perceives different amplitude changes in audio signals of different frequencies differently, using a logarithmic decay curve for interrupted audio with a higher proportion of low frequencies makes the amplitude change more audibly smoother; while using an exponential decay curve for interrupted audio with a higher proportion of high frequencies may make it appear or disappear more quickly and clearly. The system can perform spectral analysis on the interrupted audio to identify its main frequency components and then select the most suitable decay curve. The terms "high low-frequency proportion" or "high high-frequency proportion" refer to interrupted audio where the low-frequency signal portion accounts for more than 50% or equal to the frequency signal portion, and vice versa.

[0049] In one embodiment of this application, based on the sampling rate of the current audio signal, the target amplitude change requirement, and the computing power status of the vehicle hardware, the fade-out processing step size of the current audio signal and the fade-in processing slope of the interrupted audio are calculated. The target amplitude change requirement represents the expected amplitude change value of the current audio signal from the current amplitude to the target amplitude.

[0050] Since the sampling rate determines the number of samples that can be adjusted in amplitude per unit time, a higher sampling rate allows for smaller step sizes and finer transitions. The target amplitude change requirement clarifies the target volume to which the current audio signal needs to be attenuated, and the target volume to which the audio needs to be amplified after an interruption, reflecting the desired auditory effect. The onboard hardware computing power represents the actual hardware constraints; higher computing power allows for smaller step sizes, enabling smoother and faster transitions. Taking all three factors into account, the optimal step size and slope are output.

[0051] In one embodiment of this application, the reference duration, decay curve, fade-out processing step size, and fade-in processing slope are determined as audio transition processing parameters. Based on this, the determination of audio transition processing parameters in this application can be dynamically and finely adjusted according to the characteristics of the interrupted audio, the attributes of the current audio signal, and the actual computing power of the vehicle hardware. This improves the smoothness and naturalness of audio switching, avoids abruptness or discomfort caused by parameter mismatch, and ensures that important information such as navigation and telephone calls can be clearly and gently introduced without affecting the main audio experience in the vehicle environment, thus optimizing the user's auditory perception experience.

[0052] In step S240, a smooth transition process is performed on the interrupted audio and the current audio signal based on the audio transition processing parameters to complete the audio switching.

[0053] In one embodiment of this application, the interrupted audio is determined based on scene event information. In some feasible environments, all audio sources can be pre-traversed to collect attributes such as PCM format, sampling rate, and duration to establish an audio source attribute table; the computing power and supported encoding formats of different hardware platforms can be collected to construct a hardware-audio source adaptation matrix; the active or passive interruption scene can be classified and the trigger signal feature scene-signal-audio source association mapping table can be labeled.

[0054] The aforementioned hardware-audio source adaptation matrix is ​​a structured database table. Its rows represent the chip models of different hardware platforms, while the columns represent different audio source types, such as navigation, TTS (text-to-speech), and multimedia. Each cell defines the optimal configuration or capability range that can be used to process a specific audio source on a specific hardware platform. Since different car models or configurations have varying computing power, DSP capabilities, and supported audio encoding formats in their cockpit chips, constructing this hardware-audio source adaptation matrix allows for the determination of the highest supported algorithm complexity by querying the matrix after detecting the current chip model and whether the current audio signal is high-fidelity multimedia music, thus preventing system overload.

[0055] The aforementioned scenario-signal-audio source association mapping table is an event rule base providing decision support. It records the mapping relationships between three types of information: scenario (service type of interruption event, such as passive interruption, active interruption, etc.); signal (original trigger signal captured in vehicle bus, system broadcast, network protocol); and audio source (audio source indicating the audio source affected by the aforementioned signal). For example, if an SPP signal emitted by the Bluetooth protocol stack is detected, querying the mapping table can clearly identify it as an incoming call scenario, which will interrupt the currently playing multimedia audio source. Embodiments of this application can determine the audio interruption based on the aforementioned pre-built scenario-signal-audio source association mapping table.

[0056] In one embodiment of this application, the current audio signal is linearly faded out based on the fade-out processing step size and attenuation curve. When the amplitude of the current audio signal attenuates to a preset auditory ignoring threshold, the interrupted audio is linearly faded in according to the fade-in processing slope and attenuation curve. When the interrupted audio reaches the target amplitude based on the target amplitude change requirement, the switch from the current audio signal to the interrupted audio is completed, and the audio processing resources associated with the current audio signal are released.

[0057] Linear fade-out processing gradually reduces the amplitude of the current audio signal according to a predetermined fade-out step size and attenuation curve. This is calculated and processed using a coefficient that decreases linearly over time at each audio sample point. Linear fade-in processing, on the other hand, is calculated and processed using a coefficient that increases linearly over time at each audio sample point. The aforementioned preset auditory neglect threshold refers to a threshold that is difficult for the human ear to perceive or that the volume is negligible. It can be specifically set or adjusted according to the auditory characteristics of the human ear. The aforementioned target amplitude is the desired volume level that needs to be achieved before interrupting the audio.

[0058] The aforementioned release of audio processing resources related to the current audio signal includes reclaiming and releasing the computing resources, memory space, processor resources, etc. allocated to the current audio signal after the current audio signal has faded out and been completely replaced by the interrupted audio. This includes, but is not limited to, closing the relevant audio decoder, stopping the audio stream processing thread, and clearing the audio buffer.

[0059] This application's solution achieves smooth audio switching by precisely controlling the fade-out of the current audio signal and interrupting the fade-in process, while avoiding audio gaps or overlaps. Secondly, it promptly releases audio processing resources associated with the current audio signal, optimizing system resource utilization and enhancing computational support for subsequent audio processing tasks.

[0060] In one embodiment of this application, the method further includes, upon completion of audio switching, acquiring the switched audio signal output during the audio switching process, and determining an objective smoothness score based on the number of amplitude abrupt changes between adjacent sampling points in the switched audio signal; and acquiring an audio switching experience score fed back by the human-computer interaction interface, and determining the audio switching experience score as the subjective smoothness score.

[0061] In one embodiment of this application, a reinforcement learning model is constructed with the optimization objectives of maximizing the subjective smoothness score and minimizing the objective smoothness score. The model is then trained based on the current subjective and objective smoothness scores to iteratively optimize the prediction logic for audio interruption events and the calculation logic for audio transition processing parameters.

[0062] In one embodiment of this application, after the audio switching is completed, the final audio output, which is a mixture or switch of the current audio signal and the interrupted audio, is analyzed to quantify its smoothness. The number of amplitude abrupt changes between adjacent sampling points is an important objective indicator for measuring audio smoothness. The more abrupt changes, the less smooth the switching process usually is. Specifically, the sampling point data of the switched audio signal can be monitored in real time by a digital signal processor, and the amplitude difference between adjacent sampling points can be calculated to determine the abrupt changes. Secondly, it also includes collecting user feedback on the audio switching experience. A rating module can be set on the human-machine interface of the in-vehicle infotainment system. After the audio switching occurs, a rating window pops up and allows the user to rate the switching.

[0063] In one embodiment of this application, the method further includes using machine learning to improve the audio switching strategy through continuous feedback loops. This can be achieved by updating the strategy using a deep reinforcement learning framework based on collected historical subjective and objective smoothness scores as reward signals. The prediction accuracy of audio interruption events and the calculation method of audio transition processing parameters are iteratively optimized based on the current subjective and objective smoothness scores. Specifically, a reinforcement learning model is constructed with the optimization objective of maximizing the subjective smoothness score and minimizing the objective smoothness score. This model uses the prediction logic of the current audio interruption event and the calculation logic of the audio transition processing parameters as variables. The current subjective and objective smoothness scores are normalized and substituted into a preset reward function to calculate the reward value for the current audio switching. This reward function is configured to provide positive incentives for rising subjective smoothness scores and negative incentives for rising objective smoothness scores. The reinforcement learning model is trained based on the aforementioned reward values ​​to optimize the prediction logic of audio interruption events and the calculation logic of audio transition processing parameters. In some actual implementation processes, the actual output audio signal is collected in real time through the vehicle-mounted microphone and related equipment, and the number of amplitude abrupt changes between adjacent sampling points in the signal is detected as objective data to evaluate the smoothness of the switching. Secondly, the subjective satisfaction rating of the user on the listening experience of this audio switching is received through the vehicle-mounted human-machine interface as subjective data to evaluate the switching experience. Subsequently, objective and subjective data are uploaded to a cloud server via the vehicle communication module. On the cloud server, a reinforcement learning model (DRL model) is constructed with the optimization objective of maximizing historical subjective satisfaction scores and minimizing the number of historical amplitude mutations. The uploaded data is used to train the model to iteratively optimize the prediction logic for audio interruption events and the calculation logic for audio transition processing parameters. Finally, the optimized prediction logic and parameter calculation logic are packaged into an update file and distributed to the vehicle audio system via wireless communication. The application is loaded when the vehicle is in an updateable state to complete the adaptive improvement of the audio switching control method. The aforementioned distribution process also includes labeling data types, such as the number of amplitude mutations and listening satisfaction. In some feasible implementations, the distribution update cycle can be further set to 3 months for small optimizations and 6 months for large iterations to achieve continuous improvement in interaction and optimization learning.

[0064] The embodiments of this application introduce a closed-loop feedback optimization mechanism, enabling the audio switching control method to continuously learn and adapt, constantly adjusting and optimizing its internal audio interruption event prediction logic and audio transition processing parameter calculation logic. This allows it to adapt to various complex and changing driving environments and personalized user needs, and provides a continuously optimized and adaptive audio experience.

[0065] This application provides an audio switching control method that acquires the current audio signal, the vehicle's hardware computing power status, and scene event information, and performs feature fusion to obtain fused perception features. Based on these fused perception features, it predicts audio interruption events and corresponding time windows. Then, based on the audio interruption events, corresponding time windows, the current audio signal, and the vehicle's hardware computing power status, it determines corresponding audio transition processing parameters. Finally, it performs smooth transition processing on the interrupted audio and the current audio signal based on these parameters to complete the audio switching. This application achieves advance prediction of audio switching by fusing features from the current audio signal, vehicle hardware computing power status, and scene event information, and proactively predicting audio interruption events and corresponding time windows based on the fused perception features. This proactive prediction mechanism can plan and execute smooth transition processing in advance, avoiding popping sounds and auditory gaps caused by processing delays, thus improving the continuity of the user experience. Furthermore, by fusing features from multi-source heterogeneous information, a more comprehensive fusion perception feature is constructed. This allows for the comprehensive consideration of various influencing factors, enabling unified management and control of audio switching. It also establishes an integrated processing framework covering all audio sources and scenarios, improving the intelligence and adaptability of in-vehicle audio system scheduling. Finally, this application can select appropriate fade-in / fade-out durations and curves based on music type and computing load. This dynamic adjustment capability allows audio switching processing to better adapt to different scenario requirements, hardware performance changes, and user-specific habits, optimizing the overall user listening experience.

[0066] The following describes an embodiment of the apparatus described in this application, which can be used to execute the audio switching control method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the audio switching control method described above.

[0067] Figure 4 This is a schematic diagram illustrating an audio switching control device according to an exemplary embodiment of this application. The device can be applied to… Figure 2 The method implementation process shown can be based on the device Figure 1 The implementation environment shown can be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the device is applicable.

[0068] like Figure 4 As shown, the exemplary audio switching control device includes an audio signal processing module 401 and an audio smooth switching module 402.

[0069] The audio signal processing module 401 is used to acquire the current audio signal, the computing power status of the vehicle hardware, and scene event information; perform feature fusion on the current audio signal, the computing power status of the vehicle hardware, and the scene event information to obtain fused perception features, and predict audio interruption events and corresponding time windows based on the fused perception features. The audio interruption event refers to an audio signal switching event that interrupts the current audio signal with interrupting audio associated with scene event information. The audio smooth switching module 402 is used to determine the corresponding audio transition processing parameters based on the audio interruption event and the corresponding time window, as well as the current audio signal and the computing power status of the vehicle hardware; perform smooth transition processing on the interrupted audio and the current audio signal based on the audio transition processing parameters to complete the audio switching. The interrupted audio is determined according to the scene event information.

[0070] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the audio switching control method provided in the above embodiments.

[0071] Figure 5 This is a schematic diagram illustrating the structure of a computer system for an electronic device, as shown in an exemplary embodiment of this application. It should be noted that... Figure 5 The computer system 500 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0072] like Figure 5 As shown, the computer system 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in Read-Only Memory (ROM) 502 or a program loaded from storage into Random Access Memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus. An I / O interface 505 is also connected to the bus 504, where the I / O interface 505 refers to an input / output interface.

[0073] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to I / O interface 505 as needed. Removable media 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.

[0074] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs various functions defined in the system of this application.

[0075] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0076] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0077] In the corresponding figures of the above embodiments, connecting lines can represent the connection relationship between various components, indicating more constitutive signal paths and / or one or more ends of some lines having arrows to indicate the main information flow direction. Connecting lines serve as an identifier and are not a limitation on the scheme itself, but rather, using these lines in conjunction with one or more exemplary embodiments helps to more easily connect circuits or logic units. Any signal represented (determined by design requirements or preferences) can actually include one or more signals that can be transmitted in any direction and can be implemented in any suitable type of signal scheme.

[0078] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0079] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0080] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements an audio switching control method as described in any of the above embodiments.

[0081] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0082] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0083] This application can be used in a wide range of general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0084] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0085] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. An audio switching control method, characterized by, The audio switching control method comprises: obtaining a current audio signal, a vehicle-mounted hardware computing power state, and scene event information; performing feature fusion on the current audio signal, the vehicle-mounted hardware computing power state, and the scene event information to obtain fused perception features, and predicting an audio interruption event and a corresponding time window based on the fused perception features, the audio interruption event being an audio signal switching event in which interrupting audio associated with the scene event information interrupts the current audio signal; determining corresponding audio transition processing parameters according to the audio interruption event and the corresponding time window, and the current audio signal and the vehicle-mounted hardware computing power state; performing smooth transition processing on the interrupting audio and the current audio signal based on the audio transition processing parameters to complete audio switching, the interrupting audio being determined according to the scene event information.

2. The audio switching control method of claim 1, wherein, The feature fusion on the current audio signal, the vehicle-mounted hardware computing power state, and the scene event information to obtain fused perception features comprises: quantizing, aligning, and unifying the real-time amplitude, sampling rate, and channel number of the current audio signal with the computing power occupancy rate of the vehicle-mounted hardware computing power state, and splicing to generate a first feature vector, the vehicle-mounted hardware computing power state representing the computing power occupancy rate of an audio digital signal processor; converting the scene event information into an event identifier according to a preset event type encoding rule to obtain a second feature vector, the scene event information being determined by listening to system broadcasts, vehicle bus signals, or user interaction events, and being related to a sound source interruption; aligning and splicing the first feature vector and the second feature vector in the time dimension to generate the fused perception features.

3. The audio switching control method of claim 2, wherein, The prediction of the audio interruption event and the corresponding time window based on the fused perception features comprises: inputting the fused perception features into a hybrid prediction framework for processing, the hybrid prediction framework comprising a rule engine, a decision tree model, and a long short-term memory network model; based on the rule engine, matching the scene event information in the fused perception features with a preset sound source priority rule, judging high-confidence audio interruption events, and outputting a rule determination result; based on the decision tree model, performing probability inference based on the current audio signal features and historical operation data in the fused perception features, and outputting a probability value of the audio interruption event; based on the long short-term memory network model, analyzing the time-dependent relationship of the scene event information in the fused perception features, and outputting a prediction probability of the audio interruption event and a corresponding time window; based on the rule determination result, the probability value of the audio interruption event, the prediction probability of the audio interruption event, and the corresponding time window, performing weighted processing to obtain the audio interruption event and the corresponding time window.

4. The audio switching control method of claim 3, wherein, The hybrid prediction framework further comprises: performing a hierarchical response based on the prediction probability of the audio interruption event, the hierarchical response comprising: if the prediction probability of the audio interruption event is greater than or equal to a first threshold value, triggering early execution of smooth transition processing based on a preset early time length within the predicted time window. If the predicted probability of the audio interruption event is less than the first threshold value and greater than or equal to the second threshold value, the amplitude of the current audio signal is monitored, and when the amplitude of the current audio signal is greater than or equal to a preset amplitude threshold, the early execution of the smooth transition processing is triggered based on a preset early time length; If the predicted probability of the audio interruption event is less than the second threshold value, only the amplitude of the current audio signal is monitored.

5. The audio switching control method of claim 1, wherein, According to the audio interruption event and the corresponding time window, and the current audio signal and the vehicle-mounted hardware computing power state, the corresponding audio transition processing parameters are determined, including: According to the duration of the interrupted audio, the reference duration of the fade-in and fade-out processing is determined; According to the frequency of the interrupted audio, the decay curve is determined, including a logarithmic decay curve for interrupted audio with a high low-frequency proportion and an exponential decay curve for interrupted audio with a high high-frequency proportion; Based on the sampling rate of the current audio signal, the target amplitude change requirement, and the vehicle-mounted hardware computing power state, the fade-out processing step length for controlling the current audio signal and the fade-in processing slope for controlling the interrupted audio are calculated, and the target amplitude change requirement represents the amplitude change expectation value of the current audio signal from the current amplitude to the target amplitude; The reference duration, decay curve, fade-out processing step length, and fade-in processing slope are determined as the audio transition processing parameters.

6. The audio switching control method of claim 5, wherein, According to the duration of the interrupted audio, the reference duration of the fade-in and fade-out processing includes: If the duration of the interrupted audio is less than a first duration threshold, a first reference duration is used; If the duration of the interrupted audio is greater than a second duration threshold, a second reference duration is used, and the first duration threshold is less than the second duration threshold; If the duration of the interrupted audio is less than or equal to a third duration threshold, a third reference duration is used, and the third duration threshold is less than the first duration threshold.

7. The audio switching control method of claim 5, wherein, Based on the audio transition processing parameters, the smooth transition processing is performed on the interrupted audio and the current audio signal, including: Based on the fade-out processing step length and the decay curve, the current audio signal is subjected to linear fade-out processing; When the amplitude of the current audio signal decays to a preset auditory neglect threshold, the interrupted audio is subjected to linear fade-in processing according to the fade-in processing slope and the decay curve; When the interrupted audio reaches the target amplitude based on the target amplitude change requirement, the switching from the current audio signal to the interrupted audio is completed, and the audio processing resources related to the current audio signal are released.

8. The audio switching control method of claim 1, wherein, The audio switching control method further includes: When the audio switching is completed, the switching audio signal output in the audio switching process is obtained, and the number of amplitude mutations between adjacent sampling points in the switching audio signal is determined to determine the objective evaluation score of the smoothness of this time; The audio switching experience score fed back by the human-computer interaction interface is obtained, and the audio switching experience score is determined as the subjective evaluation score of the smoothness of this time; Based on the objective evaluation score of the smoothness of this time and the subjective evaluation score of the smoothness of this time, the audio switching control optimization is performed, including the optimization of the prediction logic of the audio interruption event and the optimization of the calculation logic of the audio transition processing parameters.

9. The audio switching control method of claim 8, wherein, The audio switching control optimization based on the current smoothness objective score and the current smoothness subjective score comprises: A reinforcement learning model is constructed to maximize the smoothness subjective score and minimize the smoothness objective score, and the reinforcement learning model takes the predicted logic of the current audio interruption event and the calculation logic of the audio transition processing parameter as variables; The current smoothness subjective score and the current smoothness objective score are normalized and substituted into a preset reward function to calculate the reward value of the current audio switching, and the reward function is configured to give positive incentive to the smoothness subjective score showing an upward trend and negative incentive to the smoothness objective score showing an upward trend; The reinforcement learning model is trained based on the reward value to optimize the predicted logic of the audio interruption event and the calculation logic of the audio transition processing parameter.

10. An audio switching control device, characterized by The audio switching control device comprises: An audio signal processing module is configured to obtain a current audio signal, a vehicle-mounted hardware computing power state, and scene event information; perform feature fusion on the current audio signal, the vehicle-mounted hardware computing power state, and the scene event information to obtain fused perception features, and predict an audio interruption event and a corresponding time window based on the fused perception features, wherein the audio interruption event refers to an audio signal switching event in which a breaking audio associated with the scene event information interrupts the current audio signal; An audio smooth switching module is configured to determine corresponding audio transition processing parameters according to the audio interruption event and the corresponding time window, as well as the current audio signal and the vehicle-mounted hardware computing power state; perform smooth transition processing on the breaking audio and the current audio signal based on the audio transition processing parameters to complete audio switching, and the breaking audio is determined according to the scene event information.

11. An electronic device, comprising: A processor, a memory, and a communication bus are included; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the audio switching control method of any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is used to make a computer execute the audio switching control method of any one of claims 1-9.

Citation Information

Cited By

  • An audio playing management method and an audio playing system based on sound field anchoring

    CN122219876A