A vehicle-mounted call audio processing method and device
By standardizing the format and aligning the timeline of the uplink and downlink audio data in the vehicle conferencing system, a unified hybrid audio data packet is generated, which solves the problem of inconsistent audio format and timeline in the vehicle conferencing system and improves the accuracy and efficiency of the AI model's speech processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FORYOU GENERAL ELECTRONICS
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the format standards and timelines of uplink and downlink audio data in vehicle-mounted conferencing systems are inconsistent, making it difficult to efficiently utilize large AI models, increasing development complexity and affecting the accuracy and reliability of speech recognition and semantic analysis.
By acquiring uplink and downlink audio data, reading preset audio configuration files, standardizing the format, aligning them according to the timeline, generating a unified format mixed audio data packet, merging multiple tracks using the algorithm engine module, and generating a data packet for use by the upper-layer AI model.
It shields the differences in underlying audio, provides high-quality, structured input data, and improves the accuracy and efficiency of voice-related business processing.
Smart Images

Figure CN122111367A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to a method and apparatus for processing audio during in-vehicle calls. Background Technology
[0002] With the deep integration of artificial intelligence and in-vehicle systems, the application of large-scale AI models in in-vehicle conferencing scenarios is becoming increasingly widespread. Such applications typically rely on uplink audio (local speech) and downlink audio (remote speech) generated by the conferencing system as input data. However, in existing solutions, the processing of uplink and downlink audio data is usually handled separately by individual application-layer apps, lacking system-level unified coordination. This results in significant differences in audio data format standards (such as sampling rate and bit width) and audio timelines (i.e., the chronological order of dialogue content) between different vehicle models or devices, presenting an inconsistent state. This inconsistency makes it difficult for upper-layer large-scale AI models to directly and efficiently utilize the raw audio stream for subsequent processing, increasing not only the complexity and adaptation workload of AI edge development but also affecting the accuracy and reliability of services such as speech recognition and semantic analysis. Therefore, a solution that can achieve unified acquisition, standardized processing, and time-series alignment of uplink and downlink audio at the system level is urgently needed. Summary of the Invention
[0003] The purpose of this invention is to disclose a method and apparatus for processing in-vehicle call audio, which solves the technical problem of inconsistent audio data in existing call recordings.
[0004] To achieve the above objectives, the present invention discloses a vehicle-mounted call audio processing method, comprising: Acquire downlink audio data from a remote server and uplink audio data from a local acquisition device; Read a preset audio configuration file, which contains processing parameters associated with the audio stream type; Based on the audio configuration file, the uplink and downlink audio data are standardized to generate a standard audio data packet with a unified format. According to the preset multitrack merging format, the uplink and downlink audio data in the standard audio data packet are aligned by the time axis and multitrack merged to generate a mixed audio data packet; The mixed audio data packets are stored, and the storage location is notified to the application layer service module.
[0005] As an optional implementation, the processing parameters include at least one of the following: sampling rate, number of channels, bit width, gain factor, and weighting factor.
[0006] As an optional implementation, the format standardization process includes: Based on the audio stream type, obtain the corresponding target sampling rate, target number of channels, and target bit width from the audio configuration file; Adjust at least one of the sampling rate, number of channels, and bit width of the uplink or downlink audio data to the corresponding target value.
[0007] As an optional implementation, the format normalization process further includes a gain normalization step: Based on the gain factor and weight factor corresponding to the audio stream type in the audio configuration file, as well as the gain value inherent in the audio data, the final gain value is calculated and applied. The final gain value is calculated using the formula mStandardVol = LinkVol * Multiplier + WeightFactor, where mStandardVol is the final gain value, LinkVol is the gain value inherent in the audio data, Multiplier is the gain factor, and WeightFactor is the weight factor.
[0008] As an optional implementation, the preset multitrack merging format defines the allocation relationship of each channel in the mixed audio data packet for different audio stream types.
[0009] On the other hand, in order to achieve the above objectives, the present invention discloses an in-vehicle call audio processing device, comprising: The audio data acquisition module is used to acquire downlink audio data from a remote server and uplink audio data from a local acquisition device. A configuration management module is used to store and provide audio configuration files, which contain processing parameters associated with the audio stream type; The algorithm engine module, which communicates with the audio data acquisition module and the configuration management module, is used for: Based on the audio configuration file, the uplink and downlink audio data are standardized to generate a standard audio data packet with a unified format. According to the preset multitrack merging format, the uplink and downlink audio data in the standard audio data packet are aligned by the time axis and multitrack merged to generate a mixed audio data packet; A storage module is used to store the mixed audio data packets; The interface module is used to notify the application layer business module of the storage location of the mixed audio data packet.
[0010] As an optional implementation, the algorithm engine module includes: The raw data processing unit is used to receive raw audio data and package it into a raw data packet containing an audio stream type identifier and audio parameters based on the audio configuration file. The resampling and normalization unit is used to parse audio parameters from the original data packet, perform format conversion and gain normalization, and generate the standard audio data packet; The multitrack merging unit is used to perform channel allocation and time axis alignment merging on the standard audio data packets according to the multitrack merging format to generate the mixed audio data packets.
[0011] As an optional implementation, the resampling and normalization unit is configured to: Obtain the corresponding gain factor and weight factor from the audio configuration file according to the audio stream type; Based on the gain value inherent in the audio data, the gain factor, and the weight factor, the final gain value is calculated and applied using the formula mStandardVol = LinkVol * Multiplier + WeightFactor, where mStandardVol is the final gain value, LinkVol is the gain value inherent in the audio data, Multiplier is the gain factor, and WeightFactor is the weight factor.
[0012] As an optional implementation, the device adopts an architecture that separates the framework layer and the application layer. The algorithm engine module, configuration management module, and storage module are located in the framework layer, and the application layer business module interacts with the framework layer through the interface module.
[0013] The beneficial effects of this invention are as follows: By setting up a dedicated algorithm engine module, this invention centrally receives heterogeneous uplink and downlink audio data from remote and local sources, performs format resampling and standardization based on configuration files, and performs multitrack merging processing according to predefined rules, thereby generating time-aligned, format-uniform mixed audio data packets for use by upper-layer AI models. This design effectively shields the differences in underlying audio, enabling the entire system to adopt a standardized audio processing flow, thus solving the problem of audio format and timeline chaos caused by different processing methods at each end. This provides high-quality, structured input data for large AI models, improving the accuracy and efficiency of voice-related business processing. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1This is a schematic diagram of the structure of the vehicle-mounted audio processing device of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] In this invention, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing the invention and its embodiments, and are not intended to limit the indicated devices, elements, or components to having a specific orientation, or to be constructed and operated in a specific orientation.
[0018] Furthermore, in addition to indicating direction or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in certain situations to indicate a dependency or connection. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0019] Furthermore, the terms "installation," "setup," "equipped with," "connection," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0020] Furthermore, the terms "first," "second," etc., are primarily used to distinguish different devices, components, or parts (which may be the same or different in specific type and construction), and are not intended to indicate or imply the relative importance or quantity of the indicated devices, components, or parts. Unless otherwise stated, "a plurality of" means two or more.
[0021] The technical solution of the present invention will be further described below with reference to the embodiments and accompanying drawings.
[0022] The present invention provides a method and apparatus for processing audio during in-vehicle calls. The apparatus is adapted to the method and aims to solve the problem that uplink and downlink audio streams in in-vehicle conferencing scenarios cannot be efficiently utilized by large AI models due to inconsistent formats and timelines.
[0023] like Figure 1 As shown, the device includes an algorithm engine module located in the framework layer and connected to it an audio data acquisition module, a configuration management module, a storage module, and an interface module. The storage module is also connected to the interface module, and the interface module is also connected to the business module (AI processing module) located in the application layer. The audio data acquisition module is also connected to a remote server on the external service side and a local acquisition device (such as a microphone) on the local peripheral side. The audio data acquisition module is used to acquire downlink audio data from a remote server and uplink audio data from a local acquisition device. The configuration management module is used to store and provide audio configuration files, which contain processing parameters associated with the audio stream type; The algorithm engine module is used for: Based on the audio configuration file, the uplink and downlink audio data are standardized to generate a standard audio data packet with a unified format. According to the preset multitrack merging format, the uplink and downlink audio data in the standard audio data packet are aligned by the time axis and multitrack merged to generate a mixed audio data packet; The storage module is used to store the mixed audio data packets; The interface module is used to notify the application layer business module of the storage location of the mixed audio data packet.
[0024] In this embodiment, the algorithm engine module includes: The raw data processing unit is used to receive raw audio data and package it into a raw data packet containing an audio stream type identifier and audio parameters based on the audio configuration file. The resampling and normalization unit is used to parse audio parameters from the original data packet, perform format conversion and gain normalization, and generate the standard audio data packet; The multitrack merging unit is used to perform channel allocation and time axis alignment merging on the standard audio data packets according to the multitrack merging format to generate the mixed audio data packets.
[0025] The in-vehicle call audio processing method described in this embodiment specifically includes: Step 1: Obtain uplink and downlink audio data.
[0026] The audio data acquisition module acquires remote call audio (downlink audio data) via the network and sends it to the algorithm engine module; it also acquires local call audio (uplink audio data) via a digital signal processor (DSP) and sends it to the algorithm engine module. The uplink and downlink audio data are different in format and time order.
[0027] Step 2: Read audio configuration information.
[0028] The algorithm engine module reads the uplink and downlink audio configuration files pre-installed in the configuration management module. These configuration files contain detailed processing parameters for each audio stream (corresponding to an uplink / downlink ID), such as: a. sampling rate; b. number of channels; c. bit width; d. gain factor; e. weighting factor.
[0029] Step 3: Receive and cache the raw audio data packets.
[0030] The raw data processing unit of the algorithm engine module creates a raw data packaging thread, receives raw audio stream data from remote and local sources, and combines it with the sampling rate, number of channels, and bit width information corresponding to the uplink and downlink IDs read from the configuration file to package and generate raw uplink and downlink audio data packets, which are then stored in the raw data queue. The data structure of the raw uplink and downlink audio data packets includes: uplink and downlink IDs, sampling rate, number of channels, bit width, and raw audio data.
[0031] Step 4: Resample and standardize the audio data.
[0032] The algorithm engine module's resampling and normalization unit creates a resampling thread and executes the following sub-steps: Step 401: Read the raw uplink and downlink audio data packets from the raw data queue and parse the uplink and downlink IDs, sampling rate, number of channels, bit width and raw audio data contained therein.
[0033] Step 402: Determine if the audio data format needs adjustment. Compare the parsed sampling rate, number of channels, and bit width with the preset standard uplink and downlink audio data format (e.g., sampling rate 16000Hz, number of channels 1, bit width 16bit). If they are inconsistent, resample the original audio data to standardize its format and generate a standard uplink and downlink audio data packet. The data structure of this data packet includes: uplink / downlink ID, sampling rate, number of channels, bit width, and standard audio data.
[0034] Step 403: Adjust the gain of the standardized audio data.
[0035] Based on the gain factor and weight factor corresponding to the uplink and downlink IDs in the configuration file, and combined with the gain value inherent in the audio stream, the final gain value is calculated using the gain normalization formula, and this final gain value is applied to the audio data. The gain normalization formula is: mStandardVol = LinkVol * Multiplier + WeightFactor Where mStandardVol is the final gain value, LinkVol is the built-in gain value of the audio stream, Multiplier is the gain factor, and WeightFactor is the weight factor.
[0036] Step 404: Store the processed standard uplink and downlink audio data packets into the standard uplink and downlink audio data queue.
[0037] Step 5: Merge the standardized audio data into multiple tracks.
[0038] The multitrack merging unit of the algorithm engine module creates a multitrack merging thread and executes the following sub-steps: Step 501: Read the standard uplink and downlink audio data packets from the standard uplink and downlink audio data queue, and parse the uplink and downlink IDs and standard audio data contained therein.
[0039] Step 502: According to the multitrack merging format preset in the configuration file, allocate the standard audio data corresponding to different uplink and downlink IDs to different audio tracks. The multitrack merging format defines the uplink and downlink ID types that each channel (such as the left channel and the right channel) should carry. For example, the format {02, 01} indicates that the left channel carries downlink audio (ID 02) and the right channel carries uplink audio (ID 01).
[0040] Step 503: Align and merge the audio data of each track according to the timeline to generate a standard uplink and downlink audio merged data packet. The data structure of this data packet includes: sampling rate, number of channels, bit width, and merged audio data. In the merged audio data, each channel carries a specified type of audio stream, and the chronological order of the dialogue is maintained in the time dimension.
[0041] Step 6: Store and notify.
[0042] The algorithm engine module stores the generated standard uplink and downlink audio merged data packets to a designated location in the storage module (such as a register or file system). Subsequently, the algorithm engine module notifies the application layer's AI processing module of the storage location information of the merged data packets through a predefined application programming interface (API).
[0043] Step 7: Call up the AI model.
[0044] Upon receiving the notification, the AI processing module reads a standardized uplink and downlink audio merged data packet with a uniform format and timeline alignment from the storage module based on the provided storage location information. This data is then used for business processing such as speech recognition, semantic understanding, and meeting minutes generation, without needing to consider the source, format, or time sequence differences of the original audio.
[0045] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.
Claims
1. A method for processing audio during in-vehicle calls, characterized in that, include: Acquire downlink audio data from a remote server and uplink audio data from a local acquisition device; Read a preset audio configuration file, which contains processing parameters associated with the audio stream type; Based on the audio configuration file, the uplink and downlink audio data are standardized to generate a standard audio data packet with a unified format. According to the preset multitrack merging format, the uplink and downlink audio data in the standard audio data packet are aligned by the time axis and multitrack merged to generate a mixed audio data packet; The mixed audio data packets are stored, and the storage location is notified to the application layer business module.
2. The vehicle-mounted call audio processing method according to claim 1, characterized in that, The processing parameters include at least one of the following: sampling rate, number of channels, bit width, gain factor, and weighting factor.
3. The in-vehicle call audio processing method according to claim 1 or 2, characterized in that, The format standardization process includes: Based on the audio stream type, obtain the corresponding target sampling rate, target number of channels, and target bit width from the audio configuration file; Adjust at least one of the sampling rate, number of channels, and bit width of the uplink or downlink audio data to the corresponding target value.
4. The in-vehicle call audio processing method according to claim 3, characterized in that, The format normalization process also includes a gain normalization step: Based on the gain factor and weight factor corresponding to the audio stream type in the audio configuration file, as well as the gain value inherent in the audio data, the final gain value is calculated and applied. The final gain value is calculated using the formula mStandardVol = LinkVol * Multiplier + WeightFactor, where mStandardVol is the final gain value, LinkVol is the gain value inherent in the audio data, Multiplier is the gain factor, and WeightFactor is the weight factor.
5. The in-vehicle call audio processing method according to claim 1, characterized in that, The preset multitrack merging format defines the allocation relationship of each channel in the mixed audio data packet for different audio stream types.
6. A vehicle-mounted audio processing device for voice communication, characterized in that, include: The audio data acquisition module is used to acquire downlink audio data from a remote server and uplink audio data from a local acquisition device. A configuration management module is used to store and provide audio configuration files, which contain processing parameters associated with the audio stream type; The algorithm engine module, which communicates with the audio data acquisition module and the configuration management module, is used for: Based on the audio configuration file, the uplink and downlink audio data are standardized to generate a standard audio data packet with a unified format. According to the preset multitrack merging format, the uplink and downlink audio data in the standard audio data packet are aligned by the time axis and multitrack merged to generate a mixed audio data packet; A storage module is used to store the mixed audio data packets; The interface module is used to notify the application layer business module of the storage location of the mixed audio data packet.
7. The vehicle-mounted audio processing device according to claim 6, characterized in that, The algorithm engine module includes: The raw data processing unit is used to receive raw audio data and package it into a raw data packet containing an audio stream type identifier and audio parameters based on the audio configuration file. The resampling and normalization unit is used to parse audio parameters from the original data packet, perform format conversion and gain normalization, and generate the standard audio data packet; The multitrack merging unit is used to perform channel allocation and time axis alignment merging on the standard audio data packets according to the multitrack merging format to generate the mixed audio data packets.
8. The vehicle-mounted audio processing device according to claim 7, characterized in that, The resampling and normalization unit is configured to perform gain normalization as follows: Obtain the corresponding gain factor and weight factor from the audio configuration file according to the audio stream type; Based on the gain value inherent in the audio data, the gain factor, and the weight factor, the final gain value is calculated and applied using the formula mStandardVol= LinkVol * Multiplier + WeightFactor, where mStandardVol is the final gain value, LinkVol is the gain value inherent in the audio data, Multiplier is the gain factor, and WeightFactor is the weight factor.
9. The vehicle-mounted audio processing device according to claim 6, characterized in that, The device adopts an architecture that separates the framework layer and the application layer. The algorithm engine module, configuration management module and storage module are located in the framework layer, and the application layer business module interacts with the framework layer through the interface module.