Audio processing method, audio playing method, device and control equipment

By recording and managing cache information of input and output channels in audio processing, and using available cache sequences, the problem of poor audio processing versatility in the prior art is solved, and multi-channel audio processing across platforms and across cores is realized, and the performance and versatility of audio processing are improved.

CN119724260BActive Publication Date: 2025-05-16CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510224487.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-16
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, the ALSA software architecture has compatibility problems in multi-channel audio processing, which occupies a large memory space and is only suitable for Linux operating systems, resulting in poor universality of audio processing.

Method used

By determining the input channel cache information, the input cache metadata of each audio input channel in multiple audio input channels is recorded, including status and memory addresses, and the input cache is managed using available cache sequences to achieve audio input entry. Similarly, for the audio output channel, audio playback is achieved by determining the output channel cache information and available cache sequences.

Benefits of technology

This method can be applied to any kernel and hardware device, improving the versatility of audio processing, simplifying the software architecture, reducing the waste of hardware resources, and improving the real-time and performance of audio streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724260B_ABST
    Figure CN119724260B_ABST
Patent Text Reader

Abstract

The present application relates to an audio processing method, an audio playback method, an apparatus and a control device. The audio processing method includes: determining input channel cache information, the input channel cache information is used to record input cache metadata corresponding to multiple audio input channels respectively, for each audio input channel, determining input cache metadata in an idle state, adding the determined input cache metadata to the tail of an available cache sequence corresponding to the audio input channel; obtaining an audio segment to be recorded, splitting the audio data to be recorded corresponding to each audio input channel from the audio segment to be recorded; for each audio input channel, taking out the target input cache metadata at the head from the available cache sequence corresponding to the audio input channel, and writing the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata. The use of this method can improve the versatility of audio processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to an audio processing method, an audio playing method, an apparatus and a control device. Background Art

[0002] With the continuous development of audio technology, the demand for multi-channel audio processing is increasing. For example, in the automotive field, the car system can adopt multi-channel audio input and output to provide a rich and high-quality audio experience.

[0003] In the related art, the ALSA (Advanced Linux Sound Architecture) software architecture can be used to support multi-channel audio processing. ALSA is an audio driver framework used in Linux systems.

[0004] Then, the ALSA software architecture depends on specific hardware devices. If the hardware device does not support ALSA or has compatibility issues, ALSA may not work properly. In addition, ALSA occupies a large amount of memory space and is only applicable to the Linux operating system, resulting in poor versatility in audio processing. Summary of the invention

[0005] Based on this, it is necessary to provide an audio processing method, audio playback method, device, control device, readable storage medium and program product that can improve the versatility of audio processing in response to the above technical problems.

[0006] On the one hand, the present application provides an audio processing method, including: determining input channel cache information, the input channel cache information is used to record the input cache metadata corresponding to each audio input channel in multiple audio input channels, the input cache metadata includes the status and memory address of the input cache assigned to the audio input channel; for each of the audio input channels, determining the input cache metadata whose status is an idle state, and adding the determined input cache metadata to the tail of an available cache sequence corresponding to the audio input channel; obtaining an audio segment to be recorded, and splitting the audio data to be recorded corresponding to each of the audio input channels from the audio segment to be recorded; for each of the audio input channels, taking out the target input cache metadata at the head from the available cache sequence corresponding to the audio input channel, and writing the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata, so as to realize audio recording.

[0007] On the other hand, the present application also provides an audio processing device, including: a first information determination module, used to determine input channel cache information, the input channel cache information is used to record the input cache metadata corresponding to each audio input channel in a plurality of audio input channels, the input cache metadata includes the status and memory address of the input cache assigned to the audio input channel; a first sequence update module, used to determine the input cache metadata whose status is an idle state for each of the audio input channels, and add the determined input cache metadata to the tail of the available cache sequence corresponding to the audio input channel; a first splitting module, used to obtain an audio segment to be recorded, and split the audio data to be recorded corresponding to each of the audio input channels from the audio segment to be recorded; a first data writing module, used to take out the target input cache metadata at the head from the available cache sequence corresponding to the audio input channel for each of the audio input channels, and write the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata, so as to realize audio recording.

[0008] On the other hand, the present application further provides a control device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned audio processing method when executing the computer program.

[0009] On the other hand, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above audio processing method are implemented.

[0010] On the other hand, the present application also provides a computer program product, including a computer program, which implements the steps in the above audio processing method when executed by a processor.

[0011] The above-mentioned audio processing method, apparatus, control device, storage medium and computer program product can record the input cache metadata corresponding to each audio input channel in multiple audio input channels through the input channel cache information, and the input cache metadata includes the state and memory address of the input cache assigned to the audio input channel, so that the input channel cache information can be used to manage the input cache assigned to each audio input channel, and for each audio input channel, the input cache metadata whose state is an idle state is determined, and the determined input cache metadata is added to the tail of the available cache sequence corresponding to the audio input channel to obtain the audio segment to be recorded, and the audio data to be recorded corresponding to each audio input channel is separated from the audio segment to be recorded. For each audio input channel, the target input cache metadata at the head is taken out from the available cache sequence corresponding to the audio input channel, and the audio data to be recorded corresponding to the audio input channel is written into the input cache pointed to by the memory address in the target input cache metadata, so that the input cache of each audio input channel can be used to write the audio data to be recorded in an orderly manner through the available cache sequence, so as to realize the function of audio recording. The method can be applicable to any kernel and hardware device, thereby improving the versatility of audio processing.

[0012] On the other hand, the present application provides an audio playback method, including: determining output channel cache information, wherein the output channel cache information is used to record the output cache metadata corresponding to each audio output channel in multiple audio output channels, and the output cache metadata includes the status and memory address of the output cache assigned to the audio output channel; determining an audio segment to be played, and splitting the audio data to be played corresponding to each of the audio output channels from the audio segment to be played; for each of the audio output channels, determining the output cache metadata whose status is an idle state, writing the audio data to be played corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and adding the determined output cache metadata to the tail of the available cache sequence corresponding to the audio output channel; respectively taking out the target output cache metadata at the head from the available cache sequence corresponding to each of the audio output channels, synthesizing the audio segment with the data in the output cache pointed to by the memory address in each of the target output cache metadata, and writing the synthesized audio segment into the playback cache.

[0013] On the other hand, the present application also provides an audio playback device, including: a second information determination module, used to determine output channel cache information, the output channel cache information is used to record the output cache metadata corresponding to each audio output channel in a plurality of audio output channels, the output cache metadata includes the state and memory address of the output cache assigned to the audio output channel; a second splitting module, used to determine the audio segment to be played, and split the audio data to be played corresponding to each of the audio output channels from the audio segment to be played; a second sequence updating module, used to determine the output cache metadata of the idle state for each of the audio output channels, write the audio data to be played corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and add the determined output cache metadata to the tail of the available cache sequence corresponding to the audio output channel; a second data writing module, used to take out the target output cache metadata at the head from the available cache sequence corresponding to each of the audio output channels, synthesize the audio segment with the data in the output cache pointed to by the memory address in each of the target output cache metadata, and write the synthesized audio segment into the playback cache.

[0014] On the other hand, the present application further provides a control device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned audio playback method when executing the computer program.

[0015] On the other hand, the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned audio playback method are implemented.

[0016] On the other hand, the present application also provides a computer program product, including a computer program, which implements the steps in the above-mentioned audio playback method when executed by a processor.

[0017] The above-mentioned audio playback method, device, control device, storage medium and computer program product can record the output cache metadata corresponding to each audio output channel in multiple audio output channels through the output channel cache information, and the output cache metadata contains the status and memory address of the output cache assigned to the audio output channel, so that the output channel cache information can be used to manage the output cache assigned to each audio output channel, and the audio segment to be played is determined, and the audio data to be played corresponding to each audio output channel is split from the audio segment to be played. For each audio output channel, the output cache metadata with an idle state is determined, and the audio data to be played corresponding to the audio output channel is written into the output cache pointed to by the memory address in the determined output cache metadata, and the determined output cache metadata is added to the tail of the available cache sequence corresponding to the audio output channel, and the target output cache metadata at the head is taken out from the available cache sequence corresponding to each audio output channel, and the data in the output cache pointed to by the memory address in each target output cache metadata is synthesized into an audio segment, and the synthesized audio segment is written into the playback cache. The audio segment can be synthesized in order using the data in the output cache of the audio output channel and stored in the playback cache, so as to realize the function of audio playback. The method can be applicable to any kernel and hardware device, and improves the versatility of audio processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 A diagram showing application scenarios of audio processing methods in some embodiments;

[0020] Figure 2 It is a schematic diagram of the structure of the control device in some embodiments;

[0021] Figure 3 is a flowchart of an audio processing method in some embodiments;

[0022] Figure 4 A schematic diagram of input channel buffer information in some embodiments;

[0023] Figure 5 A schematic diagram of sequence information in some embodiments;

[0024] Figure 6 A schematic diagram of recording audio in some embodiments;

[0025] Figure 7 A schematic diagram of a flow chart of an audio playback method in some embodiments;

[0026] Figure 8 is an architectural diagram of a control device in some embodiments;

[0027] Fig. 9 is a timing diagram of data transmission in I2S mode in some embodiments;

[0028] Fig.10 is a timing diagram of data transmission in short frame mode in the PCM protocol in some embodiments;

[0029] Fig.11 is a timing diagram of data transmission in long frame mode in PCM protocol in some embodiments;

[0030] Fig.12 is a structural block diagram of an audio processing device in some embodiments;

[0031] Fig.13 is a structural block diagram of an audio playback device in some embodiments;

[0032] Fig.14 1 is a diagram of the internal structure of a control device in some embodiments. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0034] The audio processing method provided in the embodiment of the present application can be applied to Figure 1 In an application scenario, the application scenario includes a vehicle 100, and the vehicle 100 includes a control device 102. The control device 102 is connected to multiple audio acquisition devices and multiple audio playback devices, such as audio acquisition device 1~audio acquisition device m, audio playback device 1~audio playback device n, 2≤m, 2≤n, and m and n can be the same or different. The audio acquisition device can be but is not limited to a microphone. The audio playback device can be but is not limited to a speaker. The vehicle 100 can be any vehicle, which can be a passenger car, a bus, a truck or a tractor, etc. It can be a car, a sports utility vehicle, a truck or a bus, it can be a fuel vehicle or a new energy vehicle, it can be an unmanned vehicle, an automatic driving vehicle, a semi-automatic driving vehicle or a non-automatic driving vehicle.

[0035] In some embodiments, Figure 2As shown, a schematic diagram of the structure of the control device 102 is shown, in which IN represents input, each input can be understood as an audio acquisition device, and OUT represents output, each output can be understood as an audio playback device. The control device 102 includes CODEC (Coder-Decoder), SOC (System on Chip) and FLASH (Flash Memory). CODEC is responsible for the conversion of digital and analog signals and supports the compression and decompression of video and audio. SOC includes two IIS (I2S, Inter-IC Sound, integrated circuit built-in audio bus) controllers, namely IIS0 controller and IIS1 controller, DMA (Direct Memory Access) controller, SPI controller, DDR controller and CPU (Central Processing Unit). IIS controller is usually used for the transmission and management of audio data. IIS controller synchronizes data transmission through clock signal to ensure the accuracy and stability of data. During the data transmission process, IIS controller is responsible for generating clock signal, sending and receiving data, and processing the format conversion of audio data. A DMA (Direct Memory Access) controller is a unique peripheral device that transfers data within a system. It can be thought of as a controller that can connect internal and external memory to each DMA-capable peripheral device through a set of dedicated buses. The DMA controller performs transfer tasks under the programmed control of the processor. It usually includes an address bus, a data bus, and control registers. When performing data transfer, the DMA controller can read data directly from the source address and write it to the destination address. The SPI controller is a hardware device or software module used to manage SPI (Serial Peripheral Interface) communications. SPI is a high-speed, full-duplex, synchronous communication bus that is commonly used to connect various peripheral devices. The SPI controller communicates with the FLASH through the SPI bus. The DDR (Double Data Rate Controller) controller is a hardware device or software module responsible for managing and controlling DDR memory access. The DDR controller communicates with the DDR memory through the DDR bus. Based on Figure 2The control device can realize audio recording and playback. For example, during the recording process: the analog signal is collected through four audio IN channels and sent to CODEC for encoding and decoding, and then converted into a digital signal and transmitted to the IIS0 controller. Then, when the RX FIFO (receive buffer) of the IIS0 controller is full, an interrupt will be triggered. At this time, the CPU will schedule the DMA controller to move the audio data from the RX FIFO to the DDR memory space of the DDR controller. If you want to avoid losing the audio data in the event of a local power failure, you can write the audio data into the SPI controller's FLASH. In addition, you can also upload the audio data to the cloud via Ethernet for storage. During playback: the audio playback file is obtained from FLASH or Ethernet cloud to the DDR memory space of the DDR controller, and then the CPU schedules the DDR controller to move the DDR memory audio data to the TX FIFO (transmit buffer) of the IIS1 controller. The IIS1 controller sends the audio data (audio digital signal) to the CODEC codec through the IIS bus protocol. The CODEC codec converts the audio digital signal into four-channel analog signals for playback to achieve surround sound effects. The hardware only requires a SOC chip with two IIS controllers and a multi-channel CODEC codec, which reduces hardware costs, has a simple spatial structure, is energy-saving and environmentally friendly, and has a simple software architecture, which is conducive to the rapid iteration and update of automotive electronic products.

[0036] The IIS controller and CODEC communicate through the IIS bus. The IIS bus, also known as the integrated circuit built-in audio bus, is a bus standard developed by Philips for audio data transmission between digital audio devices. The IIS bus supports half-duplex and full-duplex, and supports master-slave mode. The I2S variant also supports multi-channel time division multiplexing, so it can support multiple channels. The IIS bus contains three main signal lines: serial clock SCK / SCLK (Continuous Serial Clock), serial data SD (Serial Data) (or DOUT / DIN), and left and right channel selection WS (Word Select) signals. SCK is also called bit clock BCK (Bit Clock) or BCLK (Bit Clock). This clock has a pulse for each bit of digital audio data. IIS includes data for two channels (Left / Right), and the left and right channel data are switched under the control of the channel selection / word selection (WS) issued by the master device (SOC). WS is also called LRCK (Left-Right Clock, frame clock), which is used to switch the left and right channel data. LRCK is 1 for right channel data transmission, and 0 for left channel data transmission. The frequency of LRCK, lrck, is usually set to be equal to the sampling rate, fs, i.e., lrck = fs. The input data DIN and the output data DOUT refer to the audio data represented by binary complement. The master clock MCLK (Master Clock) is the reference clock for the bit clock and the serial clock. A system should use the same MCLK to ensure clock synchronization requirements. The common frequencies of the master clock MCLK are 256×fs or 384×fs, where fs (frequency of sample) is the sampling frequency of the audio signal. Sometimes, in order to achieve better synchronization between systems, this clock is used when I2S is configured in master mode. The clock output frequency is 256×fs.

[0037] An audio processing method can be implemented based on a control device. Specifically, the control device determines input channel cache information, which is used to record input cache metadata corresponding to each audio input channel in multiple audio input channels. The input cache metadata contains the state and memory address of the input cache assigned to the audio input channel. For each audio input channel, the control device determines the input cache metadata whose state is an idle state, and adds the determined input cache metadata to the tail of the available cache sequence corresponding to the audio input channel. The control device obtains the audio segment to be recorded, and separates the audio data to be recorded corresponding to each audio input channel from the audio segment to be recorded. For each audio input channel, the control device takes out the target input cache metadata at the head from the available cache sequence corresponding to the audio input channel, and writes the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata. This method can be applied to any kernel and hardware device, which improves the versatility of audio processing.

[0038] In some embodiments, Figure 3 As shown, an audio processing method is provided, which is executed by a control device in a vehicle, and includes steps 302 to 308. Among them:

[0039] Step 302: determine input channel cache information, where the input channel cache information is used to record input cache metadata corresponding to each audio input channel in a plurality of audio input channels, and the input cache metadata includes a state and a memory address of the input cache allocated to the audio input channel.

[0040] The multiple audio input channels are audio input channels supported by the control device, for example, they can be audio input channels supported by the codec in the control device, that is, the number of audio input channels in the multiple audio input channels is the maximum number of audio input channels supported by the control device, for example, the control device supports up to 4 audio input channels, then the multiple audio input channels are the 4 audio input channels supported by the control device. The input buffer is a part of the memory, for example, the input buffer is a part of the DDR memory.

[0041] One or more input buffers can be allocated to each audio input channel. The input buffer allocated to the audio input channel is used to store the digital audio data corresponding to the audio input channel. The digital audio data is the data obtained by analog-to-digital conversion of the input analog audio data. Each audio input channel corresponds to multiple input buffer metadata, where multiple means at least two. The input buffer metadata contains the status and memory address of the input buffer. The memory address points to the input buffer, that is, the memory address represents the location of the input buffer in the memory. The input channel buffer information is generated by the control device through initialization. For each audio input channel, the number of input buffers allocated to the audio input channel is consistent with the maximum number of input buffers supported by the audio input channel.

[0042] In some embodiments, the input channel cache information is in the form of a two-dimensional structure array, each element in the input channel cache information represents an input cache metadata, and the same row in the input channel cache information records each input cache metadata corresponding to the same audio input channel. For example, the input channel cache information is a two-dimensional structure array chn_info[MAX_CHN][MAX_BUF_CNT], where MAX_CHN represents the maximum number of audio input channels supported by the control device, and MAX_BUF_CNT represents the maximum number of input buffers supported by each audio input channel. For example, the control device supports 4 audio input channels, and each audio input channel supports 8 input buffers. Figure 4 As shown, a schematic diagram of the input channel cache information in the form of a two-dimensional structure array is shown. Each small square in the figure represents an element, namely the input cache metadata, which includes the status and memory address of the input cache.

[0043] Step 304: for each audio input channel, determine the input buffer metadata in the idle state, and add the determined input buffer metadata to the end of the available buffer sequence corresponding to the audio input channel.

[0044] The idle state can be represented as an idle state. Each audio input channel corresponds to its own available buffer sequence, and the available buffer sequence can be in the form of a linked list or a queue, for example, a linked list queue. The data in the available buffer sequence is taken from the head and added from the tail.

[0045] Specifically, for each audio input channel, the control device can traverse the input cache metadata corresponding to the audio input channel in the input channel cache information. If the state of the traversed input cache metadata is idle, the traversed input cache metadata is added to the end of the available cache sequence.

[0046] In some embodiments, after the control device determines that the input cache metadata is in an idle state, in the input channel cache information, the state in the determined input cache metadata is changed to a busy state, such as a busy state, and then the determined input cache metadata is added to the available cache sequence. For example, the determined input cache metadata is changed from (idle state, memory address) to (busy state, memory address), and then (busy state, memory address) is added to the available cache sequence.

[0047] Step 306: obtain the audio segment to be recorded, and separate the audio data to be recorded corresponding to each audio input channel from the audio segment to be recorded.

[0048] The audio clip to be recorded may include one or more audio frames, and the audio frames include digital audio data corresponding to each audio input channel, and each digital audio data in the audio frame is data obtained at the same sampling point. Since only part of the multiple audio input channels may be used or in working state, the digital audio data of the unused or non-working audio input channels may be set to 0, for example, if each digital audio data occupies 4 bits, it is set to 0000.

[0049] In some embodiments, for each audio frame in the audio segment to be recorded, the control device may separate the digital audio data corresponding to each audio input channel. For each audio input channel, the separated digital audio data is arranged as the audio data to be recorded corresponding to the audio input channel, wherein the arrangement order of the digital audio data in the audio data to be recorded is consistent with the arrangement order of the audio frames to which the digital audio data belongs in the audio segment to be recorded, for example, the digital audio data corresponding to channel A in the first audio frame in the audio segment to be recorded is the first digital audio data in the audio data to be recorded corresponding to channel A.

[0050] In some embodiments, the control device writes the generated audio frames into the recording buffer of the DMA in sequence, and the control device can obtain the audio segments stored in the recording buffer to obtain the audio segments to be recorded. Step 306 can be executed by the CPU in the control device.

[0051] Step 308, for each audio input channel, take out the target input cache metadata of the head from the available cache sequence corresponding to the audio input channel, and write the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata to realize audio recording.

[0052] Wherein, step 308 may be executed by a CPU in a control device. The available cache sequence is a queue or a linked list, and when data is taken from the available cache sequence, data is taken from the head of the available cache sequence, and the target input cache metadata is the input cache metadata ranked first in the available cache sequence.

[0053] Specifically, taking four audio input channels as an example, the four audio input channels are the first audio input channel, the second audio input channel, the third audio input channel and the fourth audio input channel, the audio data to be recorded corresponding to the first audio input channel is A, the audio data to be recorded corresponding to the second audio input channel is B, the audio data to be recorded corresponding to the third audio input channel is C, and the audio data to be recorded corresponding to the fourth audio input channel is D. The memory address in the target input cache metadata taken out from the available cache sequence corresponding to the first audio input channel points to the first input cache, the memory address in the target input cache metadata taken out from the available cache sequence corresponding to the second audio input channel points to the second input cache, the memory address in the target input cache metadata taken out from the available cache sequence corresponding to the third audio input channel points to the third input cache, and the memory address in the target input cache metadata taken out from the available cache sequence corresponding to the fourth audio input channel points to the fourth input cache. A is written into the first input cache, B is written into the second input cache, C is written into the third input cache, and D is written into the fourth input cache. Thus, the audio segment to be recorded is split into 4 audio data to be recorded and stored in the input buffers of the 4 audio input channels respectively, thus realizing multi-channel audio recording.

[0054] In some embodiments, the data stored in the input cache can be used for playback, or the control device can store the data in the input cache in a flash memory, a disk, or upload it to the cloud via a network. For example, in order to ensure that the data is not lost during a local power outage, the data in the input cache can be written to the FLASH of the SPI controller.

[0055] In some embodiments, each audio input channel corresponds to its own release cache sequence, and the release cache sequence can be in the form of a linked list or a queue, for example, a linked list queue. The data in the release cache sequence is taken from the head and added from the tail. After writing the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata, the control device can add the target input cache metadata to the release cache sequence corresponding to the audio input channel. The control device can traverse the input cache metadata in the release cache sequence, take the traversed input cache metadata out of the release cache sequence, and restore the state of the traversed input cache metadata in the input channel cache information to an idle state, so that audio recording can continue using the input cache metadata.

[0056] In the above audio processing method, since the input buffer metadata corresponding to each audio input channel in multiple audio input channels can be recorded through the input channel buffer information, and the input buffer metadata includes the status and memory address of the input buffer assigned to the audio input channel, the input buffer assigned to each audio input channel can be managed by using the input channel buffer information, and for each audio input channel, the input buffer metadata whose status is idle is determined, and the determined input buffer metadata is added to the end of the available buffer sequence corresponding to the audio input channel to obtain the audio segment to be recorded, and the audio data to be recorded corresponding to each audio input channel is split from the audio segment to be recorded, and for For each audio input channel, the target input cache metadata of the head is taken out from the available cache sequence corresponding to the audio input channel, and the audio data to be recorded corresponding to the audio input channel is written into the input cache pointed to by the memory address in the target input cache metadata. Thus, through the available cache sequence, the input cache of each audio input channel can be used to write the audio data to be recorded in an orderly manner, thereby realizing the function of audio recording. Compared with the ALSA software architecture, this method realizes audio recording by configuring an available cache sequence and releasing a cache sequence for each audio input channel. It is not only applicable to operating systems such as Linux, but also to any kernel and hardware device, thereby improving the versatility of audio processing.

[0057] In addition, this method is also suitable for bare-metal development and requires a small amount of code, so the software architecture corresponding to this method is simple and easy to learn (while ALSA is a complex framework, it takes a lot of time and effort to understand and master its API (Application Programming Interface) and internal working principles, so the learning curve is steep). This method has fewer call levels, less stack space overhead, high execution efficiency, and high real-time audio streaming, so it has better performance (while ALSA needs to interact with hardware devices, it will occupy certain system resources, and when processing more audio data, it is easy to affect system performance).

[0058] In some embodiments, the determined input cache metadata is added to the tail of the available cache sequence corresponding to the audio input channel, including: updating the idle state to a busy state, and adding the input cache metadata to the tail of the available cache sequence corresponding to the audio input channel; after writing the audio data to be recorded corresponding to the audio input channel to the input cache pointed to by the memory address in the target input cache metadata, it also includes: adding the target input cache metadata to the tail of the release cache sequence corresponding to the audio input channel; when the target input cache metadata is taken out from the release cache sequence, the state in the target input cache metadata in the input channel cache information is restored to an idle state to release the input cache pointed to by the memory address in the target input cache metadata.

[0059] Specifically, after updating the state in the input buffer metadata to a busy state, the control device may add the input buffer metadata to the end of the available buffer sequence corresponding to the audio input channel.

[0060] In some embodiments, a corresponding release thread can be assigned to each release cache sequence in the control device. The release thread corresponding to the release cache sequence can traverse the input cache metadata in the release cache sequence, for example, traverse from the head to the tail, take the traversed input cache metadata out of the release cache sequence, and change the status of the traversed input cache metadata in the input channel cache information from a busy state to an idle state.

[0061] In this embodiment, by releasing the cache sequence, storing the input cache metadata in a busy state, and restoring the corresponding input cache metadata in the input channel cache information to an idle state when taking out the input cache metadata, the input cache pointed to by the memory address in the input cache metadata can be released, so that the input cache can be used for audio recording again.

[0062] In some embodiments, the audio processing method also includes: determining sequence information, the sequence information including sequence elements corresponding to each audio input channel, the sequence elements corresponding to the audio input channels including a first sequence identifier and a second sequence identifier, the first sequence identifier points to an available cache sequence, and the second sequence identifier points to a released cache sequence.

[0063] Specifically, the control device may determine the available buffer sequence corresponding to the audio input channel according to the first sequence identifier in the identification element corresponding to the audio input channel, and determine the release buffer sequence corresponding to the audio input channel according to the second sequence identifier in the identification element corresponding to the audio input channel.

[0064] The sequence information may be in array form, for example, the sequence information is Chn_buf_list[MAX_CHN]. The elements in the array are identification elements corresponding to the audio input channels, and the elements include a first sequence identifier and a second sequence identifier. When the released buffer sequence and the available buffer sequence are linked list queues, the first sequence identifier and the second sequence identifier may be head pointers of the linked list queues.

[0065] Specifically, Figure 5 As shown, taking four audio input channels as an example, the sequence information in array form (four-channel Chn_buf_list[MAX_CHN]) includes four identification elements, each identification element corresponds to a different audio input channel, and the first sequence identifier in the identification element points to an available cache sequence in the form of a linked list queue (Idle_buf linked list queue), and the second sequence identifier points to a released cache sequence in the form of a linked list queue (Unmap_buf linked list queue).

[0066] In this embodiment, the sequence elements corresponding to each audio input channel are recorded through sequence information, and the sequence elements include a first sequence identifier and a second sequence identifier, the first sequence identifier points to an available cache sequence, and the second sequence identifier points to a released cache sequence, so that the available cache sequence and the released cache sequence can be quickly located through the sequence information, thereby improving the efficiency of locating the available cache sequence and the released cache sequence.

[0067] In some embodiments, obtaining an audio segment to be recorded includes: collecting audio through each of the audio input channels and synthesizing audio frames; writing the audio frames into a receiving buffer of a first I2S controller, and transferring the audio frames in the receiving buffer to a recording buffer of the DMA controller through a DMA controller; when a recording buffer of the DMA controller changes from being not full to being full, obtaining the audio segment in the full recording buffer to obtain the audio segment to be recorded.

[0068] The first I2S controller is, for example, Figure 2 The IIS0 controller in the . The audio frame is generated by the codec, and the codec can transmit the audio frame to the first I2S controller through the I2S bus, so that the audio frame is stored in the receive buffer (RXFIFO) of the first I2S controller. The DMA controller can transfer the audio frame in the receive buffer of the first I2S controller to the input buffer of the DMA controller. The DMA controller can have multiple input buffers, and the multiple input buffers can be continuous in the memory. For example, the control device can apply for a memory space for the DMA, and the memory space satisfies the address alignment and the size is an integer multiple of an input buffer, thereby generating multiple input buffers for the DMA controller. The ratio of the size of each recording buffer to the size of the input buffer is equal to the number of audio input channels. For example, if there are 4 audio input channels, the size of the recording buffer is 4 times the size of the input buffer. Therefore, when one recording buffer is full, the audio segment (including one or more audio frames) in the recording buffer is determined as the audio segment to be recorded, and the audio segment to be recorded is split into audio data to be recorded corresponding to the 4 audio input channels. Then, the storage space occupied by the audio data to be recorded is equal to the size of one input buffer. Therefore, the audio data to be recorded can be written into the input buffer, which can just fill the input buffer.

[0069] Specifically, the control device may obtain analog audio data obtained by respectively collecting audio from each audio input channel at a sampling point, convert each analog audio data into digital audio data, and synthesize each converted digital audio data into an audio frame.

[0070] In some embodiments, when an input buffer of a DMA controller changes from being not full to being full, a first DMA interrupt is triggered, and the interrupt task corresponding to the first DMA interrupt is to write the audio clips in the full input buffer into the input buffer. The CPU of the control device can execute the interrupt task corresponding to the first DMA interrupt, that is, to obtain the audio clips in the full input buffer, obtain the audio clips to be recorded, split the audio data to be recorded corresponding to each audio input channel from the audio clips to be recorded, and for each audio input channel, take out the target input buffer metadata of the head from the available cache sequence corresponding to the audio input channel, and write the audio data to be recorded corresponding to the audio input channel into the input buffer pointed to by the memory address in the target input buffer metadata.

[0071] In this embodiment, audio data transmission can be achieved through a DMA controller and an I2S controller (ie, a first I2S controller), and the hardware complexity is low. Compared with the hardware requirements of the ALSA framework, the hardware requirements are simplified and the versatility is improved.

[0072] like Figure 6 As shown, a schematic diagram of recording audio is shown, including: 1. Traverse the buf (input buffer) of the corresponding audio input channel to find the buf in idle state, configure the buf state to busy and hang it in the idle_buf linked list queue, that is, add the input buffer metadata of the buf to the idle_buf linked list queue. 2. Audio acquisition: The DMA controller transfers data from the RXFIFO (receive buffer) of the first IIS controller to the DMA recording buffer. When a recording buffer of the DMA controller is filled from not full, it indicates that an audio data of the recording buffer has been collected, thereby triggering the first DMA interrupt and entering the interrupt processing function scheduling task. 3. Execute the interrupt processing function scheduling task: take a buf node of the input buffer from the idle_buf linked list queue. The buf node refers to the input buffer metadata, and then copy the data of the DMA memory to the buf to complete the audio recording function. 4. After copying the data of DMA memory to the buf in step 3, hang the buf node in the unmap_buf linked list queue to release it (the release process will configure the member buf state of the buf node to idle). Among them, the four-channel structure array chn_info[MAX_CHN][MAX_BUF_CNT] refers to the input channel buffer information, and the four-channel Chn_buf_list[MAX_CHN] refers to the sequence information. DMA memory refers to the continuous memory occupied by each DMA input buffer.

[0073] In some embodiments, the audio processing method further includes: determining output channel cache information, where the output channel cache information is used to record output cache metadata corresponding to each audio output channel in multiple audio output channels, and the output cache metadata includes the status and memory address of the output cache assigned to the audio output channel; determining the audio segment to be played, and splitting the audio data to be played corresponding to each audio output channel from the audio segment to be played; for each audio output channel, determining the output cache metadata with an idle status from each output cache metadata corresponding to the audio output channel in the output channel cache information, writing the audio data to be played corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and adding the determined output cache metadata to the tail of the available cache sequence corresponding to the audio output channel; respectively taking out the target output cache metadata at the head from the available cache sequence corresponding to each audio output channel, synthesizing the audio segment with the data in the output cache pointed to by the memory address in each target output cache metadata, and writing the synthesized audio segment into the playback cache to play the synthesized audio segment.

[0074] In this embodiment, the recording and playing functions of audio processing can be realized through two I2S controllers, the hardware structure complexity is low, and the method can be applied to any kernel and hardware device, thereby improving the versatility of audio processing.

[0075] In some embodiments, Figure 7 As shown, an audio playback method is provided, which is executed by a control device in a vehicle, and includes steps 702 to 708. Among them:

[0076] Step 702, determine output channel cache information, where the output channel cache information is used to record output cache metadata corresponding to each audio output channel in a plurality of audio output channels, and the output cache metadata includes a state and a memory address of an output cache allocated to the audio output channel.

[0077] Among them, the multiple audio output channels are the audio output channels supported by the control device, for example, they can be the audio output channels supported by the codec in the control device. Each audio output channel can be allocated one or more output buffers. The output buffer allocated to the audio output channel is used to store audio data that needs to be played from the audio output channel. Each audio output channel corresponds to multiple output buffer metadata, and the output buffer metadata contains the status and memory address of the output buffer. The memory address points to the output buffer, that is, the memory address represents the location of the output buffer in the memory. The output channel cache information is generated by the control device through initialization. For each audio output channel, the number of output buffers allocated to the audio output channel is consistent with the maximum number of output buffers supported by the audio output channel.

[0078] In some embodiments, the output channel cache information is in the form of a two-dimensional structure array, each element in the output channel cache information represents an output cache metadata, and the same row in the output channel cache information records each output cache metadata corresponding to the same audio output channel.

[0079] Step 704: determine the audio segment to be played, and separate the audio data to be played corresponding to each audio output channel from the audio segment to be played.

[0080] The audio clip to be played includes one or more audio frames. The storage space occupied by the audio data to be played is consistent with the size of the output buffer. The ratio of the storage space occupied by the audio clip to be played to the size of the output buffer can be the number of audio output channels.

[0081] Specifically, the control device may obtain the audio clip to be played locally or in the cloud, for example, from the FLASH of the SPI controller, or from the cloud via a network.

[0082] In some embodiments, for each audio frame in the audio segment to be played, the control device can separate the digital audio data corresponding to each audio output channel. For each audio output channel, the separated digital audio data is arranged as the audio data to be played corresponding to the audio output channel, wherein the arrangement order of the digital audio data in the audio data to be played is consistent with the arrangement order of the audio frames to which the digital audio data belongs in the audio segment to be played. For example, the digital audio data corresponding to channel E in the first audio frame in the audio segment to be played is the first digital audio data in the audio data to be played corresponding to channel E. The separation to obtain the audio data to be played in step 704 can be executed by the CPU in the control device.

[0083] Step 706: For each audio output channel, determine the output cache metadata that is in an idle state from the output cache metadata corresponding to the audio output channel in the output channel cache information, write the to-be-played audio data corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and add the determined output cache metadata to the end of the available cache sequence corresponding to the audio output channel.

[0084] Each audio output channel corresponds to its own available buffer sequence, which may be in the form of a linked list or a queue, for example, a linked list queue. Data in the available buffer sequence is taken from the head and added from the tail.

[0085] Specifically, for each audio output channel, the controller may determine output buffer metadata whose state is an idle state from each output buffer metadata corresponding to the audio output channel in the output channel buffer information.

[0086] In some embodiments, for each audio output channel, after the control device determines that the output cache metadata is in an idle state, the to-be-played audio data corresponding to the audio output channel is written into the output cache pointed to by the memory address in the determined output cache metadata, and in the output channel cache information, the state in the determined output cache metadata is changed to a busy state, such as a busy state, and then the determined output cache metadata is added to the available cache sequence. For example, the determined output cache metadata is changed from (idle state, memory address) to (busy state, memory address), and then (busy state, memory address) is added to the available cache sequence.

[0087] Step 708, respectively take out the target output cache metadata of the header from the available cache sequence corresponding to each audio output channel, synthesize the audio clip with the data in the output cache pointed to by the memory address in each target output cache metadata, and write the synthesized audio clip to the playback cache to play the synthesized audio clip.

[0088] Among them, the playback cache is the cache of the DMA controller. The DMA controller can transfer the audio frames in the playback cache to the sending cache of the second I2S controller. The second I2S controller can transmit the audio frames in the sending cache to the codec. The codec decodes the audio frames to obtain the analog audio data corresponding to each audio output channel. Outputting the analog audio data is playing the analog audio data.

[0089] Specifically, since the audio data to be played corresponding to the audio output channel is written into the output cache pointed to by the memory address in the determined output cache metadata, and the determined output cache metadata is added to the tail of the available cache sequence corresponding to the audio output channel, when the target output cache metadata at the head is taken out from the available cache sequence corresponding to each audio output channel, the data stored in the taken out target output cache metadata is originally the audio data to be played obtained after splitting an audio clip to be played, so that the data in the output cache pointed to by the memory address in each target output cache metadata can be restored to an audio clip before the split.

[0090] In the above audio playback method, since the output buffer metadata corresponding to each audio output channel in the multiple audio output channels can be recorded through the output channel buffer information, and the output buffer metadata includes the status and memory address of the output buffer assigned to the audio output channel, the output channel buffer information can be used to manage the output buffer assigned to each audio output channel, and the audio segment to be played is determined, and the audio data to be played corresponding to each audio output channel is split from the audio segment to be played. For each audio output channel, the output buffer metadata with an idle state is determined, and the audio data to be played corresponding to the audio output channel is written into the output buffer pointed to by the memory address in the determined output buffer metadata, and the determined output buffer metadata is added to the tail of the available buffer sequence corresponding to the audio output channel. From the available buffer sequence corresponding to each audio output channel, the target output buffer metadata at the head is taken out respectively, and the data in the output buffer pointed to by the memory address in each target output buffer metadata is synthesized into an audio segment, and the synthesized audio segment is written into the playback buffer. The audio segment can be synthesized in order using the data in the output buffer of the audio output channel and stored in the playback buffer, so as to realize the function of audio playback. The method can be applied to any kernel and hardware device, and improves the versatility of audio processing.

[0091] In some embodiments, the determined output cache metadata is added to the end of the available cache sequence corresponding to the audio output channel, including: updating the idle state to the busy state, and adding the output cache metadata to the end of the available cache sequence corresponding to the audio output channel; after writing the synthesized audio clip to the playback cache, it also includes: for each audio output channel, adding the retrieved target output cache metadata to the end of the release cache sequence corresponding to the audio output channel; when the target output cache metadata is retrieved from the release cache sequence, restoring the state in the target output cache metadata in the output channel cache information to the idle state to release the output cache pointed to by the memory address in the target output cache metadata.

[0092] Specifically, after updating the state in the output buffer metadata to a busy state, the control device may add the output buffer metadata to the end of the available buffer sequence corresponding to the audio output channel.

[0093] In some embodiments, each audio output channel corresponds to its own release buffer sequence, and the release buffer sequence can be in the form of a linked list or a queue, for example, a linked list queue. The data in the release buffer sequence is taken from the head and added from the tail. The method for determining the available buffer sequence and the release buffer sequence can refer to the relevant description in the above-mentioned audio processing method, which will not be repeated here.

[0094] In this embodiment, after the synthesized audio clip is written into the playback cache, the retrieved target output cache metadata is added to the end of the release cache sequence corresponding to the audio output channel, so that the output cache pointed to by the memory address in the output cache metadata can be released, so that the output cache can be used for audio playback again.

[0095] In some embodiments, the play buffer is a buffer of a DMA controller, and the method further includes: transferring, through the DMA controller, the audio frames in the play buffer of the DMA controller to the send buffer of the second I2S controller.

[0096] In some embodiments, the target output buffer metadata of the header is respectively taken out from the available buffer sequence corresponding to each audio output channel, the data in the output buffer pointed to by the memory address in each target output buffer metadata is synthesized into an audio clip, and the synthesized audio clip is written to the playback buffer, including: when the direct memory access DMA controller transfers the data in a playback buffer to the sending buffer of the second I2S controller, the target output buffer metadata of the header is respectively taken out from the available buffer sequence corresponding to each audio output channel, the data in the output buffer pointed to by the memory address in each target output buffer metadata is synthesized into an audio clip, and the synthesized audio clip is written to the playback cache of the DMA controller.

[0097] The second I2S controller is, for example, Figure 2 IIS1 controller in. The DMA controller can have multiple play buffers, and the multiple play buffers can be continuous in the memory. For example, the control device can apply for a piece of memory space for DMA, and the memory space satisfies address alignment and the size is an integer multiple of a play buffer, thereby generating multiple play buffers for the DMA controller. The ratio of the size of each play buffer to the size of the output buffer is equal to the number of audio output channels. For example, if there are 4 audio output channels, the size of the play buffer is 4 times the size of the output buffer. Therefore, when the DMA controller transfers data in a play buffer to the send buffer of the second I2S controller, it means that an audio clip to be played has been written into the send buffer of the second I2S controller for playing out. Therefore, a new audio clip to be played can be written into a play buffer of the DMA controller, that is, when the DMA controller transfers data in a play buffer to the send buffer of the second I2S controller, the target output buffer metadata of the head is respectively taken out from the available buffer sequence corresponding to each audio output channel, and the data in the output buffer pointed to by the memory address in each target output buffer metadata is synthesized into an audio clip, and the synthesized audio clip is written into the play buffer of the DMA controller.

[0098] In some embodiments, when the DMA controller transfers data in a playback buffer to a sending buffer of a second I2S controller, a second DMA interrupt is triggered, and the CPU of the control device can execute the interrupt task corresponding to the second DMA interrupt, that is, to perform: from the available buffer sequence corresponding to each audio output channel, respectively take out the target output buffer metadata of the head, synthesize the audio segment with the data in the output buffer pointed to by the memory address in the memory address of each target output buffer metadata, and write the synthesized audio segment to the playback buffer of the DMA controller.

[0099] In this embodiment, the audio data playback function can be implemented through a DMA controller and an I2S controller (ie, a second I2S controller), and the hardware complexity is low. Compared with the hardware requirements of the ALSA framework, the hardware requirements are simplified and the versatility is improved.

[0100] From a hardware perspective, in the related art, one IIS controller is used for one audio channel. For multi-channel audio, multiple IIS controllers are required, which undoubtedly wastes hardware resources. In this application, two IIS controllers can be used to achieve multi-channel acquisition and playback, saving hardware resources, with a simple hardware structure, energy saving and environmental protection (less IIS controller resources are enabled, energy saving), and due to the simple software architecture, it is conducive to the rapid development and iterative release of automotive electronic products.

[0101] In some embodiments, the specific implementation of the audio playback method can be: 1. Traverse the buf (output buffer) of the corresponding audio output channel to find the buf in idle state, copy the audio data to be played to the buf, and configure the buf state to busy and hang it in the idle_buf linked list queue. 2. Audio playback: When the DMA controller moves the data in a DMA play buffer to the TXFIFO (transmit buffer) of the second IIS controller, it indicates that an audio data of a play buffer has been played, thereby triggering the second DMA interrupt and entering the interrupt processing function scheduling task. 3. Execute the interrupt processing function scheduling task: take a buf node from the idle_buf linked list queue, and then copy the buf data to the DMA memory (complete the audio playback function). Then, hang the buf node in the unmap_buf linked list queue to release it (the release process will configure the member buf state of the buf node to idle).

[0102] In some embodiments, Figure 8As shown, an architecture diagram of the control device is shown, including an application layer, a kernel layer (such as Linux kernel), a driver layer, a SOC control layer and a hardware layer. The audio software architecture (program) corresponding to the audio processing method and the audio playback method provided in the present application can be in, that is, deployed in, the driver layer. Taking the Linux kernel as an example, the driver layer operates the SOC IIS controller downward to receive and send audio data streams from the hardware, and mounts it on the Linux kernel platform bus upward, providing a method for reading and writing audio streams for the Linux kernel system call interface. Platform in the Linux kernel is a bus type used to manage devices in embedded systems. The application layer realizes the collection and playback of audio streams by calling Linux system calls.

[0103] In some embodiments, in the I2S bus, the MSB (Most Significant Bit) of the left and right channel data is valid on the second SCK / BCLK rising edge after the WS changes. The WCLK (Word Clock) / LRCLK (Left / Right Clock) signal is used to indicate the channel to which the data currently being sent belongs. When it is 0, it indicates the left channel data. The LRCLK signal is valid from one clock before the first bit (MSB) of the current channel data. The LRCLK signal changes on the falling edge of BCLK. The sender changes the data on the falling edge of the clock signal BCLK, and the receiver reads the data on the rising edge of the clock signal BCLK. The LRCLK frequency is equal to the sampling frequency fs, and one LRCLK cycle (1 / fs) includes sending the left channel and right channel data.

[0104] For signals in the standard I2S format, no matter how many bits of valid data there are, the highest bit of the data always appears at the second BCLK / SCLK pulse after the WCLK / LRCK changes (that is, the start of a frame). This allows the number of valid bits at the receiving end to be different from that at the transmitting end. If the receiving end can process fewer valid bits than the transmitting end, it can discard the excess low-bit data in the data frame; if the receiving end can process more valid bits than the transmitting end, it can make up the remaining bits by itself. This synchronization mechanism makes the interconnection of digital audio devices more convenient and will not cause data misalignment. I2S has three operating modes, namely I2S mode, left-aligned mode, and right-aligned mode. Fig. 9As shown in the figure, it shows the timing diagram of data transmission in I2S mode. In the figure, the English name of the least significant bit is Least Significant Bit, abbreviated as LSB. One clock before the most significant bit is "1 clock before MSB", the left channel is "Left Channel", the right channel is "Right Channel", the serial data input is SDIN (Serial Data Input), and the serial data output is SDOUT (Serial Data Output).

[0105] In some embodiments, the CODEC generates a digital signal through the PCM (Pulse Code Modulation) protocol and transmits the generated digital signal to the I2S controller through the I2S bus. The PCM protocol mainly has two modes: short frame and long frame. Fig.10 As shown in Figure 1, the timing diagram of data transmission in the short frame mode of the PCM protocol is shown. Fig.11 The figure shows the timing diagram of data transmission in the long frame mode of the PCM protocol. In the figure, a short frame is "short frame" and a long frame is "longframe", which means that a frame of data in a certain channel can be called a slot. Due to the lack of a unified standard, different manufacturers may have different settings for the FSYNC (Frame Synchronization) pulse width and trigger edge. For example, in the short frame mode, the frame synchronization clock width is one bit clock cycle, and the second rising edge of BCLK is valid after the data is valid in FSYNC; in the long frame mode, the frame synchronization clock width is two bit clock cycles, and the first rising edge of BCLK is valid after the data is valid in FSYNC. The FSYNC pulse width is a key parameter in the PCM frame synchronization clock mode, which determines the duration of the frame synchronization signal. In the short frame mode: the clock frequency of BCLK bclk = lrck * (slot_num * slot_width + 1), lrck is the clock frequency of LRCK; in the long frame mode: the clock frequency of BCLK bclk = lrck * slot_num * slot_width. Slot_width refers to the number of bits in a frame of data. slot_num refers to the number of channels. m in the figure refers to the bit width of a channel. Sample is sample.

[0106] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0107] Based on the same inventive concept, the embodiment of the present application also provides an audio processing device for implementing the audio processing method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more audio processing device embodiments provided below can refer to the limitations on the audio processing method above, and will not be repeated here.

[0108] In some embodiments, Fig.12 In the embodiment, an audio processing device is provided, comprising: a first information determining module 1202, a first sequence updating module 1204, a first splitting module 1206 and a first data writing module 1208, wherein:

[0109] The first information determination module 1202 is used to determine input channel cache information, where the input channel cache information is used to record input cache metadata corresponding to each audio input channel in a plurality of audio input channels, where the input cache metadata includes a status and a memory address of the input cache allocated to the audio input channel.

[0110] The first sequence updating module 1204 is configured to determine, for each audio input channel, input buffer metadata in an idle state, and add the determined input buffer metadata to the end of an available buffer sequence corresponding to the audio input channel.

[0111] The first splitting module 1206 is used to obtain the audio segment to be recorded, and split the audio segment to be recorded into the audio data to be recorded corresponding to each audio input channel.

[0112] The first data writing module 1208 is used to take out the target input cache metadata of the head from the available cache sequence corresponding to the audio input channel for each audio input channel, and write the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata to realize audio recording.

[0113] In some embodiments, the first sequence update module 1204 is also used to update the idle state to a busy state, and add the input cache metadata to the end of the available cache sequence corresponding to the audio input channel; the device also includes a first information update module, the first information update module is used to add the target input cache metadata to the end of the release cache sequence corresponding to the audio input channel; when the target input cache metadata is taken out from the release cache sequence, the state in the target input cache metadata in the input channel cache information is restored to an idle state to release the input cache pointed to by the memory address in the target input cache metadata.

[0114] In some embodiments, the audio processing device also includes a sequence determination module, which is used to determine sequence information, wherein the sequence information includes sequence elements corresponding to each audio input channel, and the sequence elements corresponding to the audio input channels include a first sequence identifier and a second sequence identifier, wherein the first sequence identifier points to an available cache sequence, and the second sequence identifier points to a released cache sequence.

[0115] In some embodiments, the device also includes an audio frame synthesis module, which is used to obtain analog audio data obtained by collecting audio for each audio input channel at a sampling point, convert each analog audio data into digital audio data, and synthesize the converted digital audio data into an audio frame; a first audio frame transmission module, which is used to write the audio frame into a receiving cache of a first I2S controller, and transfer the audio frame in the receiving cache to the recording cache of the DMA controller through the DMA controller; the first splitting module 1206 is also used to obtain the audio segment in the full recording cache when a recording cache of the DMA controller is filled from not full, so as to obtain the audio segment to be recorded.

[0116] In some embodiments, the device further includes: a second information determination module, which is used to determine output channel cache information, the output channel cache information is used to record output cache metadata corresponding to each audio output channel in a plurality of audio output channels, the output cache metadata corresponding to the audio output channel includes the state and memory address of the output cache assigned to the audio output channel; a second splitting module, which is used to determine the audio segment to be played, and split the audio data to be played corresponding to each audio output channel from the audio segment to be played; a second sequence updating module, which is used to determine the output cache metadata with an idle state from each output cache metadata corresponding to the audio output channel in the output channel cache information for each audio output channel, write the audio data to be played corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and add the determined output cache metadata to the tail of the available cache sequence corresponding to the audio output channel; a second data writing module, which is used to take out the target output cache metadata at the head from the available cache sequence corresponding to each audio output channel, synthesize the audio segment with the data in the output cache pointed to by the memory address in each target output cache metadata, and write the synthesized audio segment into the playback cache.

[0117] In some embodiments, Fig.13 In the embodiment, an audio playback device is provided, comprising: a second information determination module 1302, a second splitting module 1304, a second sequence updating module 1306, and a second data writing module 1308, wherein:

[0118] The second information determination module 1302 is used to determine the output channel cache information, and the output channel cache information is used to record the output cache metadata corresponding to each audio output channel in the multiple audio output channels. The output cache metadata corresponding to the audio output channel includes the status and memory address of the output cache assigned to the audio output channel.

[0119] The second splitting module 1304 is used to determine the audio segment to be played, and split the audio segment to be played into the audio data to be played corresponding to each audio output channel.

[0120] The second sequence updating module 1306 is used to determine, for each audio output channel, output cache metadata in an idle state from the output cache metadata corresponding to the audio output channel in the output channel cache information, write the to-be-played audio data corresponding to the audio output channel into the output cache pointed to by the memory address in the determined output cache metadata, and add the determined output cache metadata to the end of the available cache sequence corresponding to the audio output channel.

[0121] The second data writing module 1308 is used to take out the target output cache metadata of the header from the available cache sequence corresponding to each audio output channel, synthesize the audio segment with the data in the output cache pointed to by the memory address in each target output cache metadata, and write the synthesized audio segment to the playback cache to play the synthesized audio segment.

[0122] In some embodiments, the second sequence update module 1306 is also used to update the idle state to a busy state, and add the output cache metadata to the end of the available cache sequence corresponding to the audio output channel; the device also includes a second information update module, the second information update module is used to, for each audio output channel, add the retrieved target output cache metadata to the end of the release cache sequence corresponding to the audio output channel; when the target output cache metadata is retrieved from the release cache sequence, the state in the target output cache metadata in the output channel cache information is restored to an idle state to release the output cache pointed to by the memory address in the target output cache metadata.

[0123] In some embodiments, the playback cache is a cache of a DMA controller, and the device also includes a first audio frame transmission module, which is used to transfer the audio frames in the playback cache of the DMA controller to the sending cache of the second I2S controller through the DMA controller; the second data writing module 1308 is also used to, when a sending cache of the second I2S controller is filled from underfilled, respectively retrieve the target output cache metadata of the header from the available cache sequence corresponding to each audio output channel, synthesize the audio segment with the data in the output cache pointed to by the memory address in the metadata of each target output cache, and write the synthesized audio segment to the playback cache of the DMA controller.

[0124] Each module in the above audio processing device can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of the processor in the control device in the form of hardware, or can be stored in the memory in the control device in the form of software, so that the processor can call and execute the operations corresponding to each module.

[0125] In some embodiments, a control device is provided, whose internal structure diagram can be as follows: Fig.14As shown. The control device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The system bus can be but not limited to a CAN (Controller Area Network) bus, an SPI bus or an IIS bus. Among them, the processor of the control device is used to provide computing and control capabilities. The memory of the control device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the control device is used to store data involved in the audio processing method. The input / output interface of the control device is used to exchange information between the processor and an external device. The input / output interface can be but not limited to a codec. The communication interface of the control device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an audio processing method is implemented.

[0126] Those skilled in the art will understand that Fig.14 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the control device to which the scheme of the present application is applied. The specific control device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0127] In some embodiments, a control device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above audio processing method when executing the computer program.

[0128] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned audio processing method are implemented.

[0129] In some embodiments, a computer program product is provided, including a computer program, which implements the steps in the above audio processing method when executed by a processor.

[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0131] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0132] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0133] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. An audio processing method, characterized in that: The method comprises: Determine input channel cache information, where the input channel cache information is used to record input cache metadata corresponding to each audio input channel of the plurality of audio input channels, where the input cache metadata includes a state and a memory address of an input cache allocated to the audio input channel; For each of the audio input channels, determine the input buffer metadata in the idle state, and add the determined input buffer metadata to the end of the available buffer sequence corresponding to the audio input channel; Obtaining an audio segment to be recorded, and separating the audio data to be recorded corresponding to each of the audio input channels from the audio segment to be recorded; For each of the audio input channels, the target input cache metadata of the header is taken out from the available cache sequence corresponding to the audio input channel, and the audio data to be recorded corresponding to the audio input channel is written into the input cache pointed to by the memory address in the target input cache metadata to realize audio recording.

2. The method according to claim 1, characterized in that Adding the determined input buffer metadata to the end of the available buffer sequence corresponding to the audio input channel includes: Updating the idle state to a busy state, and adding the input buffer metadata to the end of the available buffer sequence corresponding to the audio input channel; After writing the to-be-recorded audio data corresponding to the audio input channel into the input buffer pointed to by the memory address in the target input buffer metadata, the method further includes: Adding the target input buffer metadata to the end of the release buffer sequence corresponding to the audio input channel; When the target input buffer metadata is taken out from the release buffer sequence, the state of the target input buffer metadata in the input channel buffer information is restored to an idle state to release the input buffer pointed to by the memory address in the target input buffer metadata.

3. The method according to claim 1, characterized in that The method further comprises: Determine sequence information, where the sequence information includes sequence elements corresponding to each of the audio input channels, where the sequence elements corresponding to the audio input channels include a first sequence identifier and a second sequence identifier, where the first sequence identifier points to an available cache sequence, and the second sequence identifier points to a released cache sequence.

4. The method according to any one of claims 1 to 3, characterized in that: Get the audio clip to be recorded, including: Collecting audio through each of the audio input channels and synthesizing audio frames; Writing the audio frame into a receiving buffer of a first I2S controller, and transferring the audio frame in the receiving buffer to an input buffer of the DMA controller through a direct memory access (DMA) controller; When a recording buffer of the DMA controller is filled from being not full, an audio segment in the full recording buffer is acquired to obtain an audio segment to be recorded.

5. An audio playing method, characterized in that: The method comprises: Determine output channel cache information, where the output channel cache information is used to record output cache metadata corresponding to each audio output channel in a plurality of audio output channels, where the output cache metadata includes a state and a memory address of an output cache allocated to the audio output channel; Determine an audio segment to be played, and separate the audio data to be played corresponding to each of the audio output channels from the audio segment to be played; For each of the audio output channels, determine the output buffer metadata that is in an idle state, write the to-be-played audio data corresponding to the audio output channel into the output buffer pointed to by the memory address in the determined output buffer metadata, and add the determined output buffer metadata to the end of the available buffer sequence corresponding to the audio output channel; From the available cache sequence corresponding to each of the audio output channels, the target output cache metadata of the header is respectively taken out, the data in the output cache pointed to by the memory address in each of the target output cache metadata is synthesized into an audio clip, and the synthesized audio clip is written into the playback cache.

6. The method according to claim 5, characterized in that Adding the determined output buffer metadata to the end of the available buffer sequence corresponding to the audio output channel includes: Updating the idle state to a busy state, and adding the output buffer metadata to the end of the available buffer sequence corresponding to the audio output channel; After writing the synthesized audio clip to the playback buffer, it also includes: For each of the audio output channels, adding the retrieved target output buffer metadata to the tail of a release buffer sequence corresponding to the audio output channel; When the target output buffer metadata is taken out from the buffer release sequence, the state of the target output buffer metadata in the output channel buffer information is restored to an idle state to release the output buffer pointed to by the memory address in the target output buffer metadata.

7. The method according to claim 5 or 6, characterized in that: The method further comprises: taking out target output buffer metadata of the header from the available buffer sequence corresponding to each of the audio output channels, synthesizing audio segments with the data in the output buffer pointed to by the memory address in each of the target output buffer metadata, and writing the synthesized audio segments to the playback buffer, including: In the case where the direct memory access DMA controller transfers data in a playback buffer to the sending buffer of the second I2S controller, the target output buffer metadata of the header is respectively taken out from the available buffer sequence corresponding to each of the audio output channels, the data in the output buffer pointed to by the memory address in each of the target output buffer metadata is synthesized into an audio segment, and the synthesized audio segment is written into the playback buffer of the DMA controller.

8. An audio processing device, characterized in that: The device comprises: A first information determination module is used to determine input channel cache information, where the input channel cache information is used to record input cache metadata corresponding to each audio input channel in a plurality of audio input channels, where the input cache metadata includes a state and a memory address of an input cache allocated to the audio input channel; A first sequence updating module is used to determine, for each of the audio input channels, input buffer metadata in an idle state, and add the determined input buffer metadata to the end of an available buffer sequence corresponding to the audio input channel; A first splitting module is used to obtain an audio segment to be recorded, and split the audio segment to be recorded into audio data to be recorded corresponding to each of the audio input channels; The first data writing module is used to take out the target input cache metadata of the head from the available cache sequence corresponding to the audio input channel for each of the audio input channels, and write the audio data to be recorded corresponding to the audio input channel into the input cache pointed to by the memory address in the target input cache metadata to realize audio recording.

9. An audio playback device, characterized in that: The device comprises: A second information determination module is used to determine output channel cache information, where the output channel cache information is used to record output cache metadata corresponding to each audio output channel in a plurality of audio output channels, where the output cache metadata includes a state and a memory address of an output cache allocated to the audio output channel; A second splitting module is used to determine the audio segment to be played, and split the audio data to be played corresponding to each of the audio output channels from the audio segment to be played; A second sequence updating module is used to determine, for each of the audio output channels, the output buffer metadata in the idle state, write the to-be-played audio data corresponding to the audio output channel into the output buffer pointed to by the memory address in the determined output buffer metadata, and add the determined output buffer metadata to the end of the available buffer sequence corresponding to the audio output channel; The second data writing module is used to respectively take out the target output cache metadata of the head from the available cache sequence corresponding to each of the audio output channels, synthesize the audio clip with the data in the output cache pointed to by the memory address in each of the target output cache metadata, and write the synthesized audio clip into the playback cache.

10. A control device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Audio stream processing method and device, electronic equipment, storage medium and vehicle

    CN116600234A

  • Audio interface device, system, electronic component, electronic equipment and method

    CN118426728A