Audio transcriptor

By connecting multiple microphones to an audio transcriber and using different audio signal frequency bands, the compatibility problem of smart terminals using microphones in multiple applications at the same time is solved, enabling clear recording and transmission of multi-person calls and improving audio quality.

CN224154341UActive Publication Date: 2026-04-21SHENZHEN FLYING CELEBRATION ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
SHENZHEN FLYING CELEBRATION ELECTRONIC TECH CO LTD
Filing Date
2025-05-21
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing smart terminals cannot handle multi-person call scenarios when multiple applications use the microphone simultaneously, and the microphone can only clearly capture close-range sound, while the sound quality is poor at a distance.

Method used

An audio transcriber is used, and multiple microphones are connected through a DSP digital audio module and an SPI recording audio processing module. Different audio signal transmission frequency bands are used to achieve signal non-interference, and a recording file with a unified frequency band is generated through a mixing processing unit.

Benefits of technology

It enables the simultaneous use of multiple microphones in various application scenarios, clearly recording and transmitting the voices of each person in a multi-person call, thus improving the audio quality of multi-person remote conferences and live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN224154341U_ABST
    Figure CN224154341U_ABST
Patent Text Reader

Abstract

The embodiment of the utility model discloses an audio transcriptor, which is used for an intelligent terminal. The system comprises a DSP digital audio module, an SPI recording audio processing module, a TF storage module and at least one external microphone. The DSP digital audio module is provided with a first DSP interface, a second DSP interface, a third DSP interface and a fourth DSP interface; the SPI recording audio processing module is provided with a first SPI interface, a second SPI interface and at least two microphone input interfaces; the first DSP interface is used for being connected with an intelligent terminal, the second DSP interface is used for being connected with an earphone and / or a microphone, the third DSP interface is connected with the first SPI interface, and the second SPI interface is connected with the TF storage module; the at least two microphone input interfaces are used for connecting microphones; wherein the signal transmission frequency bands of the at least two microphone input interfaces and the third DSP interface are different.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This utility model relates to the field of audio processing technology, and in particular to an audio transcription device. Background Technology

[0002] Current smart devices, such as smartphones and tablets, cannot support multiple applications using the microphone simultaneously when recording audio or screen recording. For example, when playing games on a phone with the microphone on, you cannot simultaneously make voice and video calls. Similarly, recording multi-person conferences cannot accurately record both internal and external audio from the phone, or because only one microphone is used, it can only clearly capture the voices of those closest to the microphone, while the voices of those further away appear as background noise. In short, a single smart device cannot adequately handle scenarios where multiple applications require microphone access simultaneously, or multi-person calls involving a certain distance. Utility Model Content

[0003] In view of this, the present invention provides an audio transcription device to solve the problems in the background art described above.

[0004] To achieve one or more of the above objectives or other objectives, this utility model proposes an audio transcriber for a smart terminal, comprising: a DSP digital audio module, an SPI recording audio processing module, a TF storage module, and at least one external microphone;

[0005] The DSP digital audio module has a first DSP interface, a second DSP interface, a third DSP interface, and a fourth DSP interface;

[0006] The SPI recording audio processing module has a first SPI interface, a second SPI interface, and at least two microphone input interfaces;

[0007] The first DSP interface is used to connect to a smart terminal, the second DSP interface is used to connect to headphones and / or microphones, the third DSP interface is connected to the first SPI interface, and the second SPI interface is connected to the TF storage module;

[0008] Both of the at least two microphone input interfaces are used to connect microphones; wherein...

[0009] The signal transmission frequency bands of at least two microphone input interfaces and the third DSP interface are different.

[0010] Furthermore, the fourth DSP interface is used to connect to an external data interface so that external devices can read data from the TF storage module.

[0011] Furthermore, the SPI recording audio processing module includes an audio acquisition unit and a mixing processing unit, wherein the audio acquisition unit is connected to at least two microphone input interfaces and a third DSP interface through audio transmission interfaces of different frequency bands.

[0012] To achieve one, some, or all of the above objectives, or other objectives, this utility model also proposes an audio processing method based on the aforementioned audio transcriber, comprising the following steps:

[0013] Adjust the method of obtaining communication files in the app to obtain the latest generated recording files from the TF storage module;

[0014] The audio acquisition unit acquires audio signals from the microphone and all external microphones in real time at different preset frequency bands and sends them to the mixing processing unit.

[0015] The audio signals from all the microphones and external microphones of different preset frequency bands are mixed synchronously by the mixing processing unit to generate a recording file.

[0016] The audio file is sent to the recipient in real time via the app to be communicated with.

[0017] To achieve one, some, or all of the above objectives, or other objectives, this utility model further proposes that the step of acquiring audio signals from microphones and all external microphones in real time through the audio acquisition unit and sending them to the mixing processing unit further includes:

[0018] The audio acquisition unit acquires the background sound audio signal of the smart terminal in real time through the same preset frequency band as the microphone, and sends the real-time acquired background sound audio signal and the real-time acquired microphone audio signal to the mixing processing unit.

[0019] Furthermore, the step of synchronously mixing audio signals from different preset frequency bands through the mixing processing unit to generate a recording file also includes:

[0020] The audio signals of the background sound acquired in real time are synchronously mixed with the audio signals of the microphone acquired in real time and the audio signals of the external microphone acquired in real time to generate an audio recording file.

[0021] To achieve one or more of the above objectives or other objectives, this utility model also proposes a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above-described audio processing methods.

[0022] To achieve one or more of the above objectives or other objectives, this utility model also proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described audio processing methods.

[0023] Implementing the embodiments of this utility model will have the following beneficial effects:

[0024] After adopting the aforementioned audio transcriber, because the audio transcriber connects to multiple microphones via the SPI recording audio processing module through different audio signal transmission frequency bands, the transmission signals between different microphones do not interfere with each other. It can simultaneously collect audio signals through different microphones. That is, in scenarios where multiple applications need to use microphones at the same time, the audio signals collected by the above-mentioned different microphones are mixed and processed to form a signal of a unified frequency band and stored in the TF storage module. At the same time, the smart terminal extracts the mixed audio signal from the TF storage module and outputs it, thus realizing the situation where multiple applications can use microphones simultaneously. For example, in remote conferences with at least one party consisting of multiple people, the existing technology requires one party to use multiple terminal devices or the speaker to share a microphone. People far from the microphone cannot transmit their voices clearly to the other party. However, the audio transcriber of this utility model supports multiple microphones communicating simultaneously, and each microphone uses a different audio frequency band, so there is no interference between them. Only one terminal device and multiple microphones are needed to transmit the voice of each person in one party to the other party at the same time in real time. In this case, the meeting recording or screen recording can clearly hear everyone's speech throughout. The sound effect is comparable to a multi-person live conference. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this utility model or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this utility model. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] in:

[0027] Figure 1 This is a schematic diagram of the module structure of the audio transcriber in one embodiment of the present invention;

[0028] Figure 2 This is a schematic diagram of the SPI recording audio processing module in one embodiment of the present invention. Detailed Implementation

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this invention are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or accompanying drawings of this invention are used to distinguish different objects, not to describe a particular order.

[0030] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the present invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0032] Reference Figure 1 and Figure 2 This invention proposes an audio transcriber for a smart terminal. The audio transcriber includes a DSP digital audio module, an SPI recording audio processing module, a TF storage module, and at least one external microphone. The DSP digital audio module has a first DSP interface, a second DSP interface, a third DSP interface, and a fourth DSP interface. The SPI recording audio processing module has a first SPI interface, a second SPI interface, and at least two microphone input interfaces. The first DSP interface is used to connect to the smart terminal, the second DSP interface is used to connect headphones and / or a microphone, the third DSP interface is connected to the first SPI interface, and the second SPI interface is connected to the TF storage module. The at least two microphone input interfaces are each used to connect a microphone. The signal transmission frequency bands of the at least two microphone input interfaces and the third DSP interface are different. For example, two microphones or three microphones.

[0033] In some embodiments, the fourth DSP interface is used to connect to an external data interface so that external devices can read data from the TF storage module.

[0034] In some embodiments, the SPI recording audio processing module includes an audio acquisition unit and a mixing processing unit, wherein the audio acquisition unit is connected to at least two microphone input interfaces and a third DSP interface through audio transmission interfaces of different frequency bands.

[0035] Because this audio transcriber connects to multiple microphones via an SPI recording audio processing module that uses different audio signal transmission frequency bands, the transmission signals between different microphones do not interfere with each other. It can simultaneously collect audio signals from different microphones. In scenarios where multiple applications require microphone use, the audio signals collected from these different microphones are mixed and processed to form a signal with a unified frequency band, which is then stored in the TF storage module. Simultaneously, the smart terminal extracts the mixed audio signal from the TF storage module and outputs it, thus achieving the goal of accommodating microphone use in multiple application scenarios. For example, if a user is screen recording and playing a game with others using their microphone, and a WeChat call comes in, the user can only use the microphone in the game or in WeChat. That is, if the microphone is used in the game, the WeChat friend cannot hear the user, and vice versa. After using the aforementioned external audio transcriber, since the audio transcriber is a physical external device and is connected to multiple microphones of different frequency bands, the SPI recording audio processing module picks up the audio input of the corresponding microphone through different audio signal transmission frequency bands. After the picked-up audio signal is mixed synchronously by the mixing processing unit, the mixed audio is transmitted and played through the smart terminal, thus enabling the simultaneous use of microphones in multiple application scenarios.

[0036] Refer again Figures 1 to 2 The working principle of this audio transcriber is as follows:

[0037] S1. Adjust the method of obtaining communication files in the APP to obtain the latest generated recording files from the TF storage module;

[0038] S2. The audio acquisition unit acquires audio signals from the microphone and all external microphones in real time at different preset frequency bands and sends them to the mixing processing unit.

[0039] S3. The audio signals from all different preset frequency band microphones and external microphones are mixed synchronously through the mixing processing unit to generate a recording file.

[0040] S4. The audio file is sent to the recipient in real time via the communication APP.

[0041] In step S1 above, the communication app can be any app software with real-time calling capabilities on a smart terminal, such as social software like WeChat and QQ that can make voice or video calls, or game software like Honor of Kings and Peacekeeper Elite that can play games together in real time. In this step, the method of obtaining the communication file for such communication apps that require the audio processing method of this invention is adjusted to obtain the latest generated recording file from the TF storage module of the audio transcriber.

[0042] In step S2 above, each microphone or external microphone is equipped with a separate audio frequency band in the audio acquisition unit. Each microphone or external microphone acquires audio signals through an independent audio frequency band, which ensures that the sound acquired by each microphone does not interfere with each other.

[0043] In step S3 above, the audio signals from all microphones and external microphones of different preset frequency bands are synchronously mixed by the mixing processing unit to generate an audio file. In other words, after synchronously mixing the different audio signals, an audio file in a specific frequency band recording format is generated. Taking a multi-person remote conference as an example, since each person is equipped with an independent microphone, the frequency bands of the signals transmitted between these independent microphones do not interfere with each other, and the acquired sound information is very close to the real scene sound.

[0044] In step S4 above, the audio file is sent to the receiver in real time through the communication APP. The sound played by the receiver's smart terminal is a synthesis of all the audio signals collected by the microphone or external microphone. Taking a multi-person remote conference as an example, the sound played by the receiver is close to the real scene sound of the sender.

[0045] In some embodiments, the step of acquiring audio signals from the microphone and all external microphones in real time through different preset frequency bands via the audio acquisition unit and sending them to the mixing processing unit further includes:

[0046] The audio acquisition unit acquires the background sound audio signal of the smart terminal in real time through a preset frequency band identical to that of the microphone, and sends the acquired background sound audio signal and the acquired microphone audio signal to the mixing processing unit. For example... Figure 1 As shown, the audio signal of the background sound of the smart terminal can be transmitted to the SPI recording audio processing module through the L / R channels.

[0047] In this embodiment, compared to the previous embodiment, the background sound currently being played on the smart terminal is recorded. For example, if the smart terminal is playing background music and multiple people are singing together, this step records both the background music and the singing together.

[0048] In some embodiments, the step of synchronously mixing audio signals of different preset frequency bands through a mixing processing unit to generate a recording file further includes:

[0049] The audio signals of the background sound acquired in real time are synchronously mixed with the audio signals of the microphone acquired in real time and the audio signals of the external microphone acquired in real time to generate an audio recording file.

[0050] In this embodiment, building upon the previous embodiment, taking the example of a smart terminal playing background music while multiple people sing along, the mixed audio incorporates the background music played by the terminal itself. This usage scenario is more suitable for live streaming, such as multiple people using the same terminal in the same location to conduct a live stream.

[0051] In some embodiments, the present invention also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the audio processing method of claim above.

[0052] In some embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described audio processing methods.

[0053] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), and enhanced SDRAM.

[0054] ESDRAM, Synchlink DRAM (SLDRAM), Rambus Direct RAM (RDRAM), Direct Memory Bus Dynamic RAM (DRDRAM), and Memory Bus Dynamic RAM (RDRAM), etc.

[0055] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0056] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An audio transcriber for a smart terminal, characterized by, Includes: a DSP digital audio module, an SPI recording and audio processing module, a TF storage module, and at least one external microphone; The DSP digital audio module has a first DSP interface, a second DSP interface, a third DSP interface, and a fourth DSP interface; The SPI recording audio processing module has a first SPI interface, a second SPI interface, and at least two microphone input interfaces; The first DSP interface is used to connect to a smart terminal, the second DSP interface is used to connect to headphones and / or microphones, the third DSP interface is connected to the first SPI interface, and the second SPI interface is connected to the TF storage module; Both of the at least two microphone input interfaces are used to connect microphones; wherein... The signal transmission frequency bands of at least two microphone input interfaces and the third DSP interface are different.

2. The audio transcriber of claim 1, wherein, The fourth DSP interface is used to connect to an external data interface so that external devices can read data from the TF storage module.

3. The audio transcriber of claim 2, wherein, The SPI recording audio processing module includes an audio acquisition unit and a mixing processing unit, wherein the audio acquisition unit is connected to at least two microphone input interfaces and a third DSP interface through audio transmission interfaces of different frequency bands.