Real-time audio transmission method and system based on AVAudio Engine audio engine, medium, program product and terminal
Through the AVAudioEngine audio engine and dynamic buffer segmentation mechanism, the problems of real-time acquisition and format adaptation of audio streams in iOS system are solved, and low-latency and efficient audio transmission is achieved, which is suitable for real-time voice transfer and other services.
Patent Information
- Application Number
- CN202510508568.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the existing iOS system, the AVAudioRecorder class cannot obtain audio stream data in real time, the format adaptation is difficult and the resource occupancy is high, so it cannot meet the needs of real-time voice transfer and other services.
It adopts the AVAudioEngine audio engine, combined with the dynamic buffer fragmentation mechanism, obtains audio data in real time and transmits it to the target server, supports customized audio formats, uses the WebSocket protocol for data transmission, and introduces an exception handling mechanism.
It realizes millisecond audio stream capture and transmission, ensures low latency and high real-time, reduces memory consumption, improves the operating efficiency of iOS devices, and operates stably under abnormal conditions.
Smart Images

Figure CN120447859A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of photon counting CT technology and self-supervised noise reduction technology, and in particular to a real-time audio transmission method, system, medium, program product and terminal based on the AVAudioEngine audio engine. Background Art
[0002] Currently, audio recording is implemented in the existing iOS system through the AVAudioRecorder class. The AVAudioRecorder class provides an audio metering function (averagePowerForChannel) for obtaining audio volume information, which can be used to determine whether the audio is overloaded or muted. This function of the AVAudioRecorder class can be used in simple audio monitoring scenarios, but it is obviously insufficient for application scenarios that require the real-time transmission of audio streams in specific formats. In other words, the AVAudioRecorder class cannot directly obtain real-time audio stream data during the audio recording process.
[0003] For example, in the application scenario of iFlytek's real-time speech transcription (ASR) service, developers need to transmit a continuous audio stream in real time to obtain the corresponding text stream. However, traditional audio processing methods have the following drawbacks:
[0004] (1) Unable to capture audio stream in real time: Using the AVAudioRecorder class cannot provide real-time data callback.
[0005] (2) Difficulty in format adaptation: Audio data must meet specific format requirements. Manual processing of audio format conversion and fragmented transmission is inefficient.
[0006] (3) High resource usage: Frequent audio data processing will occupy a large amount of memory and may cause performance bottlenecks. Summary of the Invention
[0007] In view of the above-mentioned shortcomings of the prior art, the present invention provides a real-time audio transmission method, system, medium, program product and terminal based on the AVAudioEngine audio engine, which is used to solve the problems that the prior art using the AVAudioRecorder class cannot meet the needs of real-time acquisition of audio streams, format adaptation and efficient transmission of services such as real-time speech transcription.
[0008] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a real-time audio transmission method based on the AVAudioEngine audio engine, which is characterized by including: obtaining audio format information of the target server; initializing the AVAudioEngine audio engine based on the audio format information, and configuring the AVAudioEngine audio engine; obtaining audio data according to the configured AVAudioEngine audio engine, and processing the obtained audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
[0009] In some embodiments of the first aspect of the present application, an AVAudioEngine audio engine is initialized based on audio format information, and the AVAudioEngine audio engine is configured, including: configuring audio parameters of audio recording of the AVAudioEngine audio engine based on the audio format information; defining an audio unit, setting the input format and buffer size of the audio unit according to the audio parameters of the audio recording, and registering an audio callback function on the audio unit to obtain a target audio unit.
[0010] In some embodiments of the first aspect of the present application, audio data is obtained according to the configured AVAudioEngine audio engine, and the obtained audio data is processed using a dynamic buffer slicing mechanism, including: the target audio unit obtains audio data in real time through the audio callback function, and sends the audio data to the buffer of the target audio unit; the buffer of the target audio unit receives the audio data and slices the audio data according to a preset slicing threshold.
[0011] In some embodiments of the first aspect of the present application, the audio format information of the target server includes: a PCM format with a sampling rate of 16kHz, a bit depth of 16 bits, and a mono channel number.
[0012] In some embodiments of the first aspect of the present application, initializing the AVAudioEngine audio engine further includes creating a WAV file; the WAV file is used to locally store the recorded audio data.
[0013] In some embodiments of the first aspect of the present application, the processed audio data is transmitted to the target server using the WebSocket protocol.
[0014] To achieve the above-mentioned purpose and other related purposes, the second aspect of the present application provides a real-time audio transmission system based on the AVAudioEngine audio engine, including: an audio format information acquisition module for acquiring the audio format information of the target server; an audio engine configuration module for initializing the AVAudioEngine audio engine based on the audio format information and configuring the AVAudioEngine audio engine; a real-time transmission module for acquiring audio data according to the configured AVAudioEngine audio engine, and processing the acquired audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
[0015] To achieve the above-mentioned purpose and other related purposes, the third aspect of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the real-time audio transmission method based on the AVAudioEngine audio engine.
[0016] To achieve the above-mentioned objectives and other related objectives, the fourth aspect of the present application provides a computer program product, which includes computer program code. When the computer program code is run on a computer, the computer implements the real-time audio transmission method based on the AVAudioEngine audio engine.
[0017] To achieve the above-mentioned purpose and other related purposes, the fifth aspect of the present application provides an electronic terminal, including a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the real-time audio transmission method based on the AVAudioEngine audio engine.
[0018] As described above, the real-time audio transmission method, system, medium, program product, and terminal based on the AVAudioEngine audio engine provided by this application have the following beneficial effects:
[0019] This application uses the AVAudioEngine audio engine to achieve millisecond-level audio stream capture and transmission, ensuring low latency and high real-time performance of audio transmission. This application supports customized audio formats, can flexibly adapt to mainstream service interfaces such as voice recognition, and has strong compatibility. This application uses a dynamic buffer sharding mechanism to effectively reduce memory consumption and significantly improve the operating efficiency of iOS devices. This application introduces an exception handling mechanism to ensure that audio processing can continue to operate stably in the face of various abnormal situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1Shown is a flowchart of a real-time audio transmission method based on the AVAudioEngine audio engine in one embodiment of the present application.
[0021] Figure 2 Shown is a schematic diagram of the working principle of a real-time audio transmission method based on the AVAudioEngine audio engine in one embodiment of the present application.
[0022] Figure 3 Shown is a structural diagram of a real-time audio transmission system based on the AVAudioEngine audio engine in one embodiment of the present application.
[0023] Figure 4 Shown is a structural schematic diagram of an electronic terminal in one embodiment of the present application. DETAILED DESCRIPTION
[0024] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict. Before the present invention is described in further detail, the nouns and terms involved in the embodiments of the present invention are explained, and the nouns and terms involved in the embodiments of the present invention are subject to the following explanations:
[0025] <1> AVAudioEngine: An audio processing framework provided by Apple. It is an audio engine framework provided by Apple for audio processing and audio synthesis. It is a powerful audio tool on the iOS and macOS platforms that supports real-time audio effect processing, audio mixing, audio playback, audio recording, and audio synthesis. It is used for audio playback and processing in iOS, macOS, and tvOS applications. It provides low-latency audio processing capabilities and supports real-time audio effects and mixing. AVAudioEngine is suitable for a variety of application scenarios, including music players, game development, speech recognition, speech synthesis, etc. It provides flexible audio processing capabilities to meet the needs of different applications.
[0026] <2> PCM (Pulse Code Modulation) format: A coding method used for digital audio and video signals. The PCM data format converts analog signals into digital signals by sampling, quantizing, and encoding them for storage, transmission, and processing on digital media. The PCM data format is widely used in digital audio processing, such as music production, audio editing, and speech recognition. In these applications, the PCM data format provides high-quality audio, meeting the needs of professional users.
[0027] The real-time audio transmission method, system, medium, program product and terminal based on the AVAudioEngine audio engine provided in this application capture audio data in real time through AVAudioEngine and combine it with a dynamic buffer slicing mechanism to achieve efficient and low-latency audio stream processing and transmission.
[0028] To facilitate understanding of the embodiments of this application, first Figure 1 Detailed description. Figure 1 The following is a flow chart showing a method for real-time audio transmission based on the AVAudioEngine audio engine according to an embodiment of the present invention. The method in this embodiment includes:
[0029] Step S11: Acquire the audio format information of the target server, which includes a PCM format with a sampling rate of 16 kHz, a bit depth of 16 bits, and a mono channel.
[0030] It's important to note that in the current iOS system, the AVAudioRecorder class is an audio recording class provided by the iOS system, primarily used to record audio and save it to a file. The AVAudioRecorder class doesn't directly access real-time audio stream data. That is, it can't extract audio data in real time during the recording process and transmit it to other systems or servers. However, scenarios such as speech recognition, online education, and real-time translation require the complete audio data to be transmitted to the server for processing, not just the volume information. For example, speech recognition converts speech signals into text messages, and scenarios like intelligent voice assistants and voice input methods require real-time audio data acquisition and processing. In online education, the voice communication between teachers and students in online classrooms must be accurately transmitted and recognized to enable interactive teaching and voice Q&A. Real-time translation involves translating speech from one language into another in real time, which requires the ability to quickly acquire and process audio data to output the translation results.
[0031] In this embodiment, audio data is transmitted in real time to a target server, which is an iFlytek server. The corresponding audio format information is a 16kHz sampling rate, 16-bit bit depth, mono PCM format, and 1280 bytes of data must be sent every 40ms. The target server can also be another server that needs to receive and process audio data in real time. Different servers can be selected and the corresponding audio format information can be obtained based on actual needs, and this embodiment does not limit this.
[0032] It should be explained that audio data in PCM format is uncompressed and can better restore the original audio signal. Sampling rate is an important parameter, which indicates the number of samples per second. The higher the sampling rate, the better the audio quality that can be represented, but at the same time, the larger the storage space occupied. The sampling rate of 16kHz is usually used for voice recording and can meet the needs of general voice recognition and calls. Bit depth refers to the quantization accuracy of each sample, measured in bits. Mono audio has only one channel and is suitable for scenarios such as voice recording that do not require stereo effects. Compared with stereo (two-channel), mono audio has a smaller amount of data.
[0033] Step S12: Initialize the AVAudioEngine audio engine based on the audio format information, and configure the AVAudioEngine audio engine.
[0034] It should be noted that this embodiment utilizes the AVAudioEngine audio engine for real-time audio data transmission. Specifically, the AVAudioEngine audio engine is initialized based on the target server's audio format information and appropriately configured to obtain an initialized and configured AVAudioEngine audio engine. This AVAudioEngine audio engine can meet the needs of real-time audio data processing and transmission in scenarios such as speech recognition, online education, and real-time translation.
[0035] In some examples, the process of initializing an AVAudioEngine audio engine based on audio format information and configuring the AVAudioEngine audio engine includes: configuring audio parameters for audio recording of the AVAudioEngine audio engine based on the audio format information; defining an audio unit, setting the input format and buffer size of the audio unit according to the audio parameters of the audio recording, and registering an audio callback function on the audio unit to obtain a target audio unit.
[0036] Specifically, the audio parameters of the audio recording of the AVAudioEngine audio engine are configured according to the audio format information. The audio parameters include a sampling rate of 16kHz and a number of bytes per frame of 2 (16 bits). Then the audio session type is configured to recording mode (AVAudioSessionCategoryRecord class). The AVAudioSessionCategoryRecord class indicates that only audio recording is supported, and playback is not supported. That is, this category is mainly used for applications that need to record, such as recorder applications. On iOS devices, after setting the AVAudioSessionCategoryRecord category, the application can record normally, but other system sounds will not be played. That is, the system will give priority to the recording function, and other system sounds (such as incoming call ringtones, alarms, etc.) will not be played to ensure that the recording quality is not disturbed.
[0037] Furthermore, define an audio unit, which is used to capture audio data collected by the microphone. Set the properties of the audio unit to enable the audio input function, and define the audio format according to the audio parameters of the audio recording through AudioStreamBasicDescription, that is, the audio format is PCM format, the sampling rate is 16kHz, the bit depth is 16bit, and it is mono. Set the defined audio format as the input format of the audio unit, and configure the buffer size in the audio unit to 1280 bytes, that is, 1280 bytes of audio data are sent every 40ms. Register the audio callback function to the audio unit, and specify the audio callback function as AURenderCallback. The audio unit receives audio data in real time through the audio callback function, that is, the audio callback function ensures that the audio data can be correctly collected and passed to the subsequent processing logic.
[0038] It should be understood that after performing the above configuration, a target audio unit can be obtained, and the input format of the target audio unit is consistent with the audio format information of the target server, which can ensure compatibility with the interface of the target server. For example, when the target server is an iFlytek server, the input format of the target audio unit is set to PCM format, with a sampling rate of 16kHz, a bit depth of 16 bits, and mono. The target audio unit is compatible with the iFlytek interface, and audio data can be transmitted to the iFlytek server in real time through the target audio unit.
[0039] It should be noted that the PCM format offers advantages such as good reproducibility, strong compatibility, and high flexibility. Specifically, the PCM data format is a lossless compression method that accurately represents the original analog signal. Through sampling, quantization, and encoding, the PCM data format can reproduce the same audio and video effects as the original analog signal on digital media. The PCM data format has excellent compatibility and can be widely used in various audio and video processing software. Furthermore, the PCM data format can be easily combined with other digital signal processing technologies, such as audio effects processing and video encoding. The PCM data format is highly flexible, allowing users to adjust parameters such as sampling rate and quantization bit number as needed. This makes the PCM data format highly adaptable to different application scenarios.
[0040] Step S13: Acquire audio data according to the configured AVAudioEngine audio engine, and process the acquired audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
[0041] In this embodiment, the audio data is processed by the configured AVAudioEngine audio engine. The specific processing process is as follows: during the audio recording process of the AVAudioEngine audio engine, the recorded audio data is appended to the buffer in real time, the audio data is dynamically cached in the buffer, and the size of the buffer is checked cyclically; if it is detected that the accumulated audio data in the buffer reaches the preset fragmentation threshold, the audio data is intercepted and processed according to the preset fragmentation threshold to obtain the intercepted fragmented data, and the fragmented data is sent to the target server. In this embodiment, millisecond-level audio stream capture and transmission can be achieved through AVAudioEngine, that is, the AVAudioEngine audio engine is combined with the dynamic buffer fragmentation mechanism to achieve efficient and low-latency audio stream processing and transmission.
[0042] The dynamic buffer segmentation mechanism segments the audio data in the buffer according to a preset segmentation threshold. For example, if the preset segmentation threshold is 1280 bytes, the audio data is segmented every 40ms and 1280 bytes. The processed audio data is released immediately after the processing is completed, that is, the processed audio data in the buffer is cleared to avoid memory overflow.
[0043] In some examples, audio data is obtained according to a configured AVAudioEngine audio engine, and the obtained audio data is processed using a dynamic buffer slicing mechanism, including: the target audio unit obtains audio data in real time through the audio callback function, and sends the audio data to the buffer of the target audio unit; the buffer of the target audio unit receives the audio data and slices the audio data according to a preset slicing threshold.
[0044] It's important to explain that when the target audio unit of the AVAudioEngine audio engine is started and audio recording begins, it captures the audio data collected by the microphone device in real time. The target audio unit obtains the audio data in real time through the audio callback function, renders the audio data in the callback, and sends the audio data to the buffer after successful rendering. The obtained audio data is processed using a dynamic buffer slicing mechanism.
[0045] Among them, the target audio unit obtains audio data in real time through the audio callback function and sends the audio data to the buffer of the target audio unit. The process is: initialize the buffer of the target audio unit, the buffer size is 1280 bytes, and the buffer is used to store audio data; detect the validity of the audio unit, if invalid, release the buffer memory and return an error status, if valid, the audio unit obtains audio data and renders the audio data and sends it to the buffer, and performs subsequent processing.
[0046] The processing process in the buffer is: append the acquired audio data to the buffer of the target audio unit, and cyclically determine the amount of audio data accumulated in the buffer; if the amount of audio data in the buffer is greater than or equal to the preset fragmentation threshold, the audio data in the buffer is intercepted and fragmented according to the preset fragmentation threshold to obtain fragmented data, and the fragmented data is sent to the target server; if the audio data bytes in the buffer are less than the preset fragmentation threshold, the audio data continues to be accumulated.
[0047] Furthermore, this embodiment can avoid memory overflow by dynamically caching audio data in a buffer, fragmenting the audio data according to a preset fragmentation threshold, and then clearing the processed audio data from the buffer. Dynamic buffer management reduces memory consumption and improves the operating efficiency of iOS devices.
[0048] In one embodiment, initializing the AVAudioEngine audio engine further includes creating a WAV file; the WAV file is used to locally store the recorded audio data.
[0049] It should be noted that in this embodiment, after the processed audio data is sent to the target server, a WAV file is also created locally for backup storage. WAV files store audio data in PCM format, a lossless audio format that can fully preserve the original audio data. The WAV file format is widely supported and can be used in a variety of audio players and editing software. The recorded audio data is saved as a backup to prevent data loss or problems during transmission, and can also be used for subsequent debugging and analysis.
[0050] In one embodiment, the processed audio data is transmitted to the target server using the WebSocket protocol. In this embodiment, the preset fragmentation threshold size and the transmission protocol are adjusted according to the interface of the target server, which is not limited here.
[0051] It needs to be explained that after the acquired audio data is processed using the dynamic buffer slicing mechanism, slicing data is obtained and transmitted to the target server in real time, wherein the slicing data is sent to the target server using the WebSocket protocol. WebSocket is a network communication protocol used to establish a full-duplex communication channel between the client and the server. It allows two-way, real-time data transmission between the client and the server without the need to frequently establish and close connections like the traditional HTTP protocol. Using the WebSocket protocol to establish a connection between the AVAudioEngine audio engine and the target server can achieve continuous sending of processed audio data (slicing data) to the target server without the need to re-establish the connection, reducing latency. The full-duplex communication mechanism of the WebSocket protocol ensures that audio data can be transmitted in real time, achieving low-latency network transmission, and is suitable for application scenarios that require real-time audio transmission.
[0052] In some examples, the process of processing the acquired audio data using the dynamic buffer slicing mechanism further includes: detecting the validity of the target audio unit using an exception handling mechanism.
[0053] It's important to note that checking the validity of audio units is a crucial step in ensuring the proper functioning of the audio processing pipeline. By introducing an exception handling mechanism to check the validity of audio units, we can prevent null pointer exceptions and free invalid memory resources when exceptions occur. Checking the validity of target audio units ensures that the audio processing pipeline can run reliably despite errors and exceptions.
[0054] In order to facilitate understanding of the real-time audio transmission method based on the AVAudioEngine audio engine provided in this embodiment, the following specific embodiments are provided for illustration. Figure 2 shown.
[0055] Example 1: Real-time speech-to-text system
[0056] Device configuration includes iOS device (iPhone / iPad), system version ≥ iOS12, and integrated AVAudioEngine audio engine.
[0057] Initializes and configures the AVAudioEngine audio engine, including the audio format and registered callback functions. It uses a dynamic buffer sharding mechanism to fragment and transmit audio data in the buffer, and sends the fragmented data to the iFLYTEK server via WebSocket.
[0058] The specific processing process is as follows Figure 2 As shown, the original input of the iOS device microphone is 44.1kHz. The configured AVAudioEngine obtains the audio data collected by the microphone in real time. The audio format is 16kHz, 16bit, and PCM format. In the AVAudioEngine audio engine, the acquired audio data is appended to the buffer. When the amount of data in the buffer is ≥1280 bytes, the audio data is fragmented and sent to the target server (iFlytek server). If not, the buffer continues to accumulate. After the audio data is fragmented and sent to the target server, the remaining audio data in the buffer is updated to determine whether to continue recording. If recording is allowed, the above steps of obtaining audio data, appending it to the buffer, fragmenting it, and sending it are repeated. If not, the recording ends.
[0059] The test used the AVAudioEngine audio engine to transmit audio data to the iFlytek server in real time. The test results showed that the audio stream transmission delay was ≤40ms, the recognition accuracy was improved by 15%, and the memory usage was stabilized within 10MB.
[0060] It should be emphasized that the real-time audio transmission method based on the AVAudioEngine audio engine provided by this application realizes millisecond-level audio stream capture and transmission by adopting the AVAudioEngine audio engine, ensuring low latency and high real-time performance of audio transmission. This application supports customized audio formats, can flexibly adapt to mainstream service interfaces such as speech recognition, and has strong compatibility. This application effectively reduces memory consumption and significantly improves the operating efficiency of iOS devices by adopting a dynamic buffer sharding mechanism management. This application introduces an exception handling mechanism to ensure that audio processing can still run stably in the face of various abnormal situations.
[0061] In the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects, and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different.
[0062] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" represent examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0063] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc or abc, where a, b, c can be single or multiple.
[0064] Figure 3 The embodiment of the present application provides a schematic block diagram of a real-time audio transmission system based on the AVAudioEngine audio engine. Figure 3 As shown, the system 300 includes:
[0065] The audio format information acquisition module 310 is used to obtain the audio format information of the target server;
[0066] The audio engine configuration module 320 is used to initialize the AVAudioEngine audio engine based on the audio format information and configure the AVAudioEngine audio engine;
[0067] The real-time transmission module 330 is used to obtain audio data according to the configured AVAudioEngine audio engine, and process the obtained audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
[0068] It should be understood that the specific process of each module executing the above corresponding steps has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0069] It should also be understood that the division of modules in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.
[0070] Figure 4 : is a schematic block diagram of an electronic terminal provided in an embodiment of the present application. Figure 4 As shown, the electronic terminal includes: at least one processor 401, a memory 402, at least one network interface 403 and a user interface 405. The various components in the device are coupled together via a bus system 404. It is understood that the bus system 404 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 404 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 4 In the text, various buses are labeled as bus systems.
[0071] The user interface 405 may include a display, a keyboard, a mouse, a trackball, a click gun, keys, buttons, a touch pad or a touch screen.
[0072] It will be appreciated that the memory 402 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM) or a programmable read-only memory (PROM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memory described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0073] The memory 402 in the embodiment of the present invention is used to store various categories of data to support the operation of the electronic terminal 400. Examples of such data include: any executable program for operating on the electronic terminal 400, such as an operating system 4021 and an application 4022; the operating system 4021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application 4022 can include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The real-time audio transmission method based on the AVAudioEngine audio engine provided by the embodiment of the present invention can be included in the application 4022.
[0074] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 401. Processor 401 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 401 or by software instructions. The above processor 401 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 401 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 401 may be a microprocessor or any conventional processor. The steps of the accessory optimization method provided in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium located in a memory. The processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0075] In an exemplary embodiment, the electronic terminal 400 may be configured to execute the aforementioned method using one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).
[0076] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, which, when running on a computer, enables the computer to execute the real-time audio transmission method based on the AVAudioEngine audio engine of any embodiment of the illustrated embodiments.
[0077] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores program code. When the program code is run on a computer, the computer executes the real-time audio transmission method based on the AVAudioEngine audio engine of any embodiment of the illustrated embodiments.
[0078] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0079] Those skilled in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0080] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0081] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0082] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0083] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0084] In the above embodiments, the functions of each functional unit can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (program) are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. Available media may be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid state disks (SSDs)).
[0085] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program codes.
[0086] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0087] In summary, the present application provides a real-time audio transmission method, system, medium, program product, and terminal based on the AVAudioEngine audio engine, comprising: obtaining audio format information from a target server; initializing and configuring the AVAudioEngine audio engine based on the audio format information; obtaining audio data according to the configured AVAudioEngine audio engine, and processing the obtained audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time. By utilizing the AVAudioEngine audio engine, the present application achieves millisecond-level audio stream capture and transmission, ensuring low latency and high real-time performance for audio transmission. The present application supports customized audio formats, can flexibly adapt to mainstream service interfaces such as speech recognition, and possesses strong compatibility. By employing a dynamic buffer slicing mechanism, the present application effectively reduces memory consumption and significantly improves the operating efficiency of iOS devices. The present application introduces an exception handling mechanism to ensure stable audio processing in the face of various abnormal situations. Therefore, the present application effectively overcomes the shortcomings of the existing technology and has high industrial application value.
[0088] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A real-time audio transmission method based on AVAudioEngine audio engine, characterized in that: include: Get the audio format information of the target server; Initialize an AVAudioEngine audio engine based on the audio format information and configure the AVAudioEngine audio engine; The audio data is acquired according to the configured AVAudioEngine audio engine, and the acquired audio data is processed using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
2. The real-time audio transmission method based on the AVAudioEngine audio engine according to claim 1, wherein Initialize the AVAudioEngine audio engine based on the audio format information and configure the AVAudioEngine audio engine, including: Configure audio parameters for audio recording of the AVAudioEngine audio engine based on the audio format information; An audio unit is defined, an input format and a buffer size of the audio unit are set according to audio parameters of the audio recording, and an audio callback function is registered on the audio unit to obtain a target audio unit.
3. The real-time audio transmission method based on the AVAudioEngine audio engine according to claim 2, wherein Get audio data based on the configured AVAudioEngine audio engine and process the acquired audio data using a dynamic buffer slicing mechanism, including: The target audio unit obtains audio data in real time through the audio callback function, and sends the audio data to the buffer of the target audio unit; The buffer of the target audio unit receives the audio data and performs fragmentation processing on the audio data according to a preset fragmentation threshold.
4. The real-time audio transmission method based on the AVAudioEngine audio engine according to claim 1, wherein The audio format information of the target server includes: a PCM format with a sampling rate of 16kHz, a bit depth of 16 bits, and a mono channel.
5. The real-time audio transmission method based on the AVAudioEngine audio engine according to claim 1, characterized in that, Initializing the AVAudioEngine audio engine also includes creating a WAV file; the WAV file is used to store the recorded audio data locally.
6. The real-time audio transmission method based on the AVAudioEngine audio engine according to claim 1, characterized in that, The processed audio data is transmitted to the target server using the WebSocket protocol.
7. A real-time audio transmission system based on the AVAudioEngine audio engine, characterized in that: include: An audio format information acquisition module is used to obtain the audio format information of the target server; An audio engine configuration module, configured to initialize an AVAudioEngine audio engine based on audio format information and configure the AVAudioEngine audio engine; The real-time transmission module is used to obtain audio data according to the configured AVAudioEngine audio engine, and process the obtained audio data using a dynamic buffer slicing mechanism to transmit the processed audio data to the target server in real time.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the real-time audio transmission method based on the AVAudioEngine audio engine according to any one of claims 1 to 6 is implemented.
9. A computer program product, characterized in that The computer program product includes computer program code, and when the computer program code is run on a computer, the computer is enabled to implement the real-time audio transmission method based on the AVAudioEngine audio engine according to any one of claims 1 to 6.
10. An electronic terminal comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the real-time audio transmission method based on the AVAudioEngine audio engine as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Audio format conversion method and apparatus
CN105551512A
HTML5-based streaming media processing method, system and related components
CN109151570A
Distributed data processing method, device and system and electronic device
CN110704536A
Audio and video network transmission fragmentation shaping method and system
CN112350986A
Audio and video playing method and system for HTML5 browser
CN114745361A