Voice transfer system text editing method based on voice control and related device

By setting up a voice chip module in the voice transfer system and using voice signal analysis to generate system control instructions, the problem of users frequently switching operations during audio playback and text editing is solved, and a more efficient text editing experience is achieved.

CN120496532APending Publication Date: 2025-08-15电视电声研究所(中国电子科技集团公司第三研究所)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510463912.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

During the audio playback and text editing process of existing voice transfer products, users need to frequently switch typing and playback operations, resulting in inefficiency.

Method used

Set up the voice chip module to connect to the voice transfer system, generate system control instructions through voice signal analysis, realize audio playback control of the voice transfer system, and the user completes the operation through voice control.

Benefits of technology

Users can control audio playback through voice, while their hands only need to be responsible for keyboard typing, which improves text editing efficiency and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496532A_ABST
    Figure CN120496532A_ABST
Patent Text Reader

Abstract

The invention discloses a voice control-based text editing method for a voice transfer system and a related device, and the method is specially provided with a voice chip module, so that a user can control the audio playback of the voice transfer system through voice, the two hands are only responsible for the knocking of a keyboard, and the user experience is improved. Through the operation, the text editing efficiency of the user can be effectively improved, and then the user experience is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a text editing method and related device for a speech transcription system based on voice control. Background Art

[0002] At present, common speech transcription products themselves have the function of supporting audio playback and manual editing of transcribed text, so that users can easily make manual corrections to the transcribed text. During the playback process, users need to repeatedly drag the audio progress bar, or click buttons such as "Fast forward 5 seconds", "Rewind 5 seconds", "Play", and "Pause" to locate the audio clip. During playback, users also need to type with both hands to edit the text. In other words, when operating existing speech transcription products, users need to take into account both typing and playback operations, so they need to frequently switch between typing and playback, which leads to low work efficiency of users. Summary of the Invention

[0003] The present invention provides a text editing method and related devices for a speech transcription system based on voice control, so as to solve the problem that the existing speech transcription products cannot be controlled efficiently and simply.

[0004] In a first aspect, the present invention provides a text editing method for a speech transcription system based on voice control, comprising: setting a voice chip module and connecting the voice chip module to a speech transcription system; parsing the received voice signal through the voice chip module and converting the voice signal into a system control instruction corresponding to the speech transcription system; triggering the speech transcription system to complete the corresponding playback operation according to the converted system control instruction; wherein the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal to jump to a certain hour, minute and second, a play signal to jump to a certain hour, minute and second, a start signal to jump to a certain hour, minute and second, a play signal to jump to a certain hour, minute and second, a start signal to jump to a certain hour, minute and second, a play signal to jump to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

[0005] Optionally, the voice chip module is connected to a microphone and receives the voice signal through the microphone; and the voice chip module includes a battery unit, which is used to power the voice chip module and charge the battery unit through a USB plug on the voice chip module.

[0006] Optionally, the voice chip module is an intelligent voice chip, which is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and parses the received voice signal in an offline state.

[0007] Optionally, the multi-layer reconfigurable spatial computing architecture is used to dynamically adjust hardware resources to achieve efficient analysis and recognition of the voice signal; the digital signal processing core DSP is used to preprocess the received voice signal and post-process the voice signal after feature extraction and classification, and the neural network processing core NPU is used to run a deep learning model and perform feature extraction and classification on the preprocessed voice signal through the deep learning model to achieve accurate analysis and recognition of the voice recognition.

[0008] Optionally, parsing the received voice signal by the voice chip module and converting the voice signal into a system control instruction corresponding to the voice transcription system includes: When the voice chip module receives a voice signal, the digital signal processing core DSP performs analog-to-digital conversion, buffering, voice activity detection (VAD), and recognition processing on the voice signal to obtain the voice content of the voice signal; The neural network processing core NPU is used to determine whether the voice content is a correct voice instruction. If so, the voice content is converted into a system control instruction corresponding to the voice transcription system according to a preset communication protocol.

[0009] Optionally, the system control instruction includes: a packet header, an instruction type bit, a data packet length bit, several instruction content bits and a packet tail, wherein: The packet header is used to identify the beginning of the data packet of the system control instruction; The instruction type bit is used to identify the type of instruction in the data packet, which includes two types: fixed instructions and non-fixed instructions; The data packet length bit is used to indicate the length of the instruction content in the data packet so that the speech transcription system can correctly parse the data packet; The instruction content bits include data of the system control instruction; The packet tail is used to mark the end of the data packet.

[0010] In a second aspect, the present invention provides a voice chip module, the voice chip module comprising: A receiving unit, configured to receive a voice signal; a processing unit, configured to parse the received voice signal, convert the voice signal into a system control instruction corresponding to the voice transcription system, and trigger the voice transcription system to perform a corresponding playback operation according to the converted system control instruction; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

[0011] Optionally, the voice chip module is an intelligent voice chip, which is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and parses the received voice signal in an offline state; The voice chip module includes a battery unit, which is used to power the voice chip module, and the battery unit is charged through a USB plug on the voice chip module.

[0012] In a third aspect, the present invention provides a system for text editing of a speech transcription system based on voice control, the system comprising: a voice chip module and a speech transcription system connected to each other; The voice chip module is used to analyze the received voice signal and convert the voice signal into a system control instruction corresponding to the voice transcription system; The speech transcription system is used to complete the corresponding playback operation according to the system control instructions converted by the speech chip module; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned text editing methods for a speech transcription system based on voice control.

[0014] The beneficial effects of the present invention are as follows: The present invention specifically sets up a voice chip module, allowing users to control the audio playback of the voice transcription system through voice, while their hands are only responsible for typing on the keyboard. This operation can effectively improve the user's text editing efficiency, thereby greatly improving the user experience.

[0015] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 This is a flow chart of a text editing method for a speech transcription system based on voice control provided by an embodiment of the present invention; Figure 2 Schematic diagram of hardware connection provided by an embodiment of the present invention; Figure 3 1 is a flow chart of another text editing method for a speech transcription system based on voice control provided by an embodiment of the present invention; Figure 4 This is a schematic diagram of the workflow of the voice chip module provided by an embodiment of the present invention; Figure 5 This is a schematic diagram of the workflow of the voice command recognition algorithm provided by the embodiment of the present invention; Figure 6 It is a structural diagram of the voice chip module provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0018] A speech transcription system is a system that automatically converts speech into text based on technologies such as speech recognition. It can be a simple software or a combination of software and hardware. Speech transcription systems are usually used in scenarios such as smart conferences, press conferences, and court trials, and can replace manual labor (such as stenographers, conference secretaries, etc.) to a certain extent. Since speech recognition technology cannot currently achieve 100% recognition accuracy, there are often interferences such as noise and reverberation at the speaking scene, so the results of the speech transcription system usually need to be manually corrected. During the correction process, the staff is required to play back the audio containing the speech and manually modify the transcribed text. If the text accuracy is required to be high, the staff is often required to correct it sentence by sentence. When the audio is unclear, or the speech content is too professional to be understood, the staff usually needs to play back a certain audio segment repeatedly. In other words, the user needs to constantly adjust the audio position, and then frequently switch back and forth between typing and playback operations, resulting in low work efficiency for the user. In order to solve the above problems, an embodiment of the present invention provides a text editing method for a speech transcription system based on voice control, see. Figure 1 and Figure 2 , the method comprising: S101, setting a voice chip module and connecting the voice chip module to a voice transcription system; Among them, in order to improve the accuracy of voice acquisition, the embodiment of the present invention connects the voice chip module with the microphone, and receives the voice signal through the microphone. In specific implementation, the microphone and the voice chip module can be connected through a line or through Bluetooth or other methods. Specific technical personnel in this field can make settings according to actual needs, and the present invention does not make specific limitations on this.

[0019] In addition, in a specific implementation, an embodiment of the present invention is to provide a battery unit in the voice chip module, and the voice chip module is powered by the battery unit, and then the battery unit is charged through the USB plug on the voice chip module.

[0020] By setting the voice chip module to operate in battery mode, the configuration of the voice chip module can be simplified, making it easy to carry and store, thereby improving user experience. In specific implementation, those skilled in the art can carry the voice chip module of the present invention with them and connect it to the voice transcription system when needed to perform tasks such as text editing and proofreading in the voice transcription system.

[0021] In specific implementation, the voice chip module in the embodiment of the present invention is an intelligent voice chip. The intelligent voice chip is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and analyzes the received voice signal in an offline state.

[0022] The multi-layer reconfigurable spatial computing architecture of the present invention is used to improve the energy efficiency and flexibility of intelligent voice chips, and achieve efficient computing processing by dynamically adjusting hardware resources according to different computing requirements. The multi-layer reconfigurable spatial computing architecture of the present invention has: Dynamic Reconfigurability: The multi-layer reconfigurable spatial computing architecture of the present invention allows hardware resources to be dynamically configured according to different computing tasks. For example, in speech recognition, the allocation of computing resources can be adjusted according to the needs of different neural network layers, thereby improving resource utilization. High energy efficiency: By dynamically adjusting hardware resources, the multi-layer reconfigurable spatial computing architecture of the present invention can significantly reduce power consumption while maintaining high performance; Flexibility: The multi-layer reconfigurable spatial computing architecture of the present invention supports calculations with multiple data precisions and can use different precision representations according to different neural network layers, thereby achieving more efficient calculations; Heterogeneous computing: The multi-layer reconfigurable spatial computing architecture of the present invention is based on a data flow graph and is oriented towards heterogeneous spatial computing. A fixed circuit structure is formed by a one-time configuration and repeatedly executed with near-ASIC efficiency, thereby improving resource utilization and data reuse. Support for mixed-precision computing: The multi-layer reconfigurable spatial domain computing architecture of the present invention can support mixed-precision computing from 1 bit to 16 bits. Different neural network layers can use different precision representations, thereby improving computing efficiency. Optimizing the efficiency of non-neural network calculations: The multi-layer reconfigurable spatial domain computing architecture of the present invention optimizes the efficiency of non-neural network calculations, such as FFT and MEL FILTER.

[0023] Furthermore, in the embodiment of the present invention, the NPU is mainly responsible for running deep learning models, such as convolutional neural networks (CNN) or recurrent neural networks (RNN), for feature extraction and classification of voice signals, thereby achieving high-precision voice recognition; before the voice signal enters the NPU, the DSP can pre-process the signal, such as noise reduction and echo cancellation, to improve the quality of the voice signal.

[0024] Among them, the functions and effects of the NPU in the embodiment of the present invention include: Efficient neural network computing: As a neural network computing optimization, the NPU can efficiently handle deep learning tasks. In the voice chip module, the NPU is mainly used to calculate the speech recognition model; Floating-point computing capability: The NPU supports 16-bit floating-point convolution operations and can perform 32 16-bit floating-point MAC operations per cycle, providing powerful computing power; Support for multiple deep learning frameworks: NPU supports multiple deep learning frameworks such as TensorFlow, Caffe, Tflite, Pytorch, Onnx NN, Android NN, etc., enabling it to be flexibly applied to different AI tasks.

[0025] In the embodiment of the present invention, the DSP is mainly used for processing audio signals, including functions such as noise suppression and echo cancellation. The DSP in the embodiment of the present invention also has efficient audio processing: the DSP can efficiently process audio signals, realize functions such as audio noise reduction, echo cancellation, and sound beautification, thereby improving audio quality; the DSP supports multiple audio formats, such as MP3, WAV, etc., and can be flexibly applied to different audio processing scenarios; the DSP realizes efficient signal processing through hardware acceleration and can quickly process audio data.

[0026] The multi-layer reconfigurable spatial computing architecture described in the embodiment of the present invention is used to dynamically adjust hardware resources to achieve efficient analysis and recognition of the voice signal; the digital signal processing core DSP is used to preprocess the received voice signal and post-process the voice signal after feature extraction and classification, and the neural network processing core NPU is used to run a deep learning model and perform feature extraction and classification on the preprocessed voice signal through the deep learning model to achieve accurate analysis and recognition of the voice recognition.

[0027] S102, parsing the received voice signal through the voice chip module, and converting the voice signal into a system control instruction corresponding to the voice transcription system; Specifically, in the embodiment of the present invention, the digital signal processing core DSP is used to pre-process the received voice signal and post-process the voice signal after feature extraction and classification, and the neural network processing core NPU is used to run a deep learning model and perform feature extraction and classification on the pre-processed voice signal through the deep learning model to achieve accurate analysis and recognition of the voice recognition.

[0028] In specific implementation, the embodiment of the present invention is that after the voice chip module receives a voice signal, the voice signal is subjected to analog-to-digital conversion, caching, voice activity detection VAD and recognition processing by the digital signal processing core DSP to obtain the voice content of the voice signal; the neural network processing core NPU is used to determine whether the voice content is a correct voice instruction, and if so, the voice content is converted into a system control instruction corresponding to the voice transcription system according to a preset communication protocol.

[0029] In simple terms, the NPU and DSP in the intelligent voice chip of the embodiment of the present invention each have different functions, but the NPU and DSP work together to achieve efficient voice processing. The NPU focuses on deep learning tasks, while the DSP is responsible for the pre-processing and post-processing of audio signals. The coordination of the NPU and DSP can greatly improve the audio quality. Therefore, the multi-core heterogeneous architecture of the present invention enables the intelligent voice chip to flexibly cope with various complex voice processing tasks.

[0030] S103: triggering the speech transcription system to complete corresponding playback operations according to the converted system control instructions.

[0031] That is, after receiving the system control instruction, the voice chip module sends the system control instruction to the voice transcription system, and the voice transcription system completes the corresponding playback operation according to the converted system control instruction.

[0032] The system control instruction in the embodiment of the present invention includes: a packet header, an instruction type bit, a data packet length bit, several instruction content bits and a packet tail, wherein: The packet header is used to mark the beginning of the data packet to help the receiver, ie, the speech transcription system, identify the starting position of the data packet. The packet header is specifically a fixed byte sequence, such as 0xA5 and 0x5A, and is used to ensure the synchronization and integrity of the data packet.

[0033] The instruction type bit is used to identify the type of instruction in the data packet, which can be a fixed instruction, a non-fixed instruction, etc. The instruction type bit field is one byte, and different values represent different instruction types, for example, 0x80 represents a fixed instruction, and 0x81 represents a non-fixed instruction. Those skilled in the art can set it arbitrarily, and the present invention does not specifically limit this.

[0034] The packet length field is used to indicate the length of the instruction content in the packet so that the receiver can correctly parse the packet. This field is one byte and indicates the number of bytes of the subsequent data content.

[0035] The command content bit is used to contain actual command data, such as play, start, stop, etc. The length of this part of the content is specified by the data packet length bit and can contain multiple bytes.

[0036] The packet trailer is used to mark the end of the data packet and help the receiver identify the end position of the data packet. The packet trailer is a fixed byte sequence and its purpose is to ensure the synchronization and integrity of the data packet.

[0037] It should be noted that the various fields of the system control instruction in the embodiment of the present invention can be set arbitrarily, as long as the speech transcription system can complete the corresponding playback operation according to the converted system control instruction.

[0038] In specific implementation, the voice signal described in the embodiment of the present invention includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

[0039] The corresponding system control instructions may include one or more of the following: play instruction, stop instruction, pause instruction, fast forward instruction, rewind instruction, mute instruction, unmute instruction, loop playback instruction, play at a certain point and speed instruction, position to a certain hour, minute and second to start instruction, position to a certain hour, minute and second to play instruction, jump to a certain hour, minute and second to start instruction, jump to a certain hour, minute and second to play instruction, jump to a certain hour, minute and second to start instruction, jump to a certain hour, minute and second to play instruction, start from a certain hour, minute and second instruction and play from a certain hour, minute and second instruction.

[0040] It should be noted that the voice signals and system control instructions in the embodiment of the present invention are set in a one-to-one correspondence. The above is only an example of a voice signal and a system control instruction. During specific implementation, technical personnel in this field can arbitrarily set specific voice signals and system control instructions according to actual needs. The present invention does not make any arbitrary settings for this.

[0041] In summary, the method described in the embodiment of the present invention is to specially set up a voice chip module, so that the user can control the audio playback of the voice transcription system through voice, while both hands are only responsible for typing on the keyboard. This operation can effectively improve the user's text editing efficiency, thereby greatly improving the user experience.

[0042] The following will be combined Figures 2 to 6 The method described in the embodiment of the present invention is explained and illustrated in detail by taking a specific example: See also Figure 2In the embodiment of the present invention, the voice is transmitted to the voice chip module through the microphone. The voice chip module can also be called a voice command recognition module or a voice recognition module, etc., and of course it can also be simply called a module. When in use, the microphone and the voice command recognition module are connected through a common 3.5mm plug, and the voice command recognition module is connected to the terminal computer of the voice transcription system through USB or other means. Among them, the microphone can be a common product and does not need to be specially customized. The voice command recognition module is used to complete the recognition of voice commands. It has a reserved 3.5mm jack at the input end and a USB plug at the output end. A battery unit is provided in the voice chip module, which is powered by USB and can work after connection. The core part of the module is the intelligent voice chip. The chip is designed based on a multi-layer reconfigurable spatial domain computing architecture, integrating a neural network processing core (NPU) and a digital signal processing core (DSP). It can complete the recognition of fixed command words and non-fixed command words based on grammatical rules in an offline state (not connected to the Internet).

[0043] See also Figure 3 During use, when the voice transcription system starts working, it will actively detect whether there is a voice command recognition module connected to the system, and then establish a USB port connection. If the transcription system detects that a module is connected, it will establish a USB port connection with the module. After the connection is successfully established, the voice transcription system will receive the data packets sent by the voice command recognition module in real time; Detect whether it is a communication protocol data packet: The specific transcription system receives the data packet sent by the voice command recognition module, and parses the data packet type and content to determine whether the current data packet is a voice control command data packet. If it is a voice control command data packet, it parses the specific control command content of the current data packet according to the protocol and enters the next software execution operation. If it is not a voice control command data packet, it returns to continue data packet collection.

[0044] When the voice transcription system receives the data packet sent by the module, it parses the contents and performs the corresponding operation on the function on the page based on the analysis results. Because the majority of the voice command recognition calculations are performed on the voice command recognition module side and connected via the USB port, the voice transcription system responds to the user's voice control commands in real time. Therefore, the phenomenon of the action being executed long after the user speaks the command will not occur.

[0045] See also Figure 4 The workflow of the voice command recognition module in the embodiment of the present invention specifically includes: Initialize the USB communication port. After powering on, the voice command recognition module first initializes the USB port for communication with the speech transcription system, waiting for a connection to be established. Once the module begins operating, the microphone associated with the module continuously picks up audio, which is then received by the module.

[0046] The module processes the microphone audio stream through analog-to-digital conversion, caching, VAD (Voice Activity Detection), and recognition, ultimately outputting the collected speech recognition content. The module does not require voice wake-up; simply speak a command word within the specified range. This eliminates the wake-up interaction process and allows users' voice commands to take effect more quickly.

[0047] It should be noted that in actual applications, users of this function usually work in a relatively quiet environment, and the pickup range of the head-mounted microphone is limited. Therefore, even if there is no wake-up operation process, the probability of a valid voice command being mistakenly triggered is very small.

[0048] Determine whether the voice recognition result is a correct command word. Based on the recognition result from the previous step, the module determines whether the current voice command is a valid voice control command and outputs the specific command word. If it is a command word, it proceeds to the next step of generating the protocol data packet. If it is not a command word, it returns to continue voice content recognition.

[0049] Setting the communication protocol data packet specifically includes: when the module recognizes the instruction word, the recognition result is packaged into a communication protocol data packet with a fixed format according to the preset communication protocol. The present invention supports instruction words and corresponding transcription system operations as shown in the following table, which are divided into 10 fixed instruction words and 10 random number instruction words, see Table 1. These instruction words can meet all the functional requirements of the transcription system for audio playback control during the text correction process. The data sent is in a fixed communication protocol format, and each packet of data contains a 2-bit header, an instruction type bit, a data packet length bit, several instruction content bits, and a 2-bit tail.

[0050] Table 1 Comparison table of voice command words and voice transcription system operations Send protocol data packets to the speech transcription system. The module sends the generated communication protocol data packets to the speech transcription system through the initialized USB communication port.

[0051] See also Figure 5 The voice command recognition module recognition algorithm proposed in the present invention specifically includes: After the voice command recognition module is powered on, it initializes the command word recognition network model and prepares to receive audio streams for command word recognition. When the module starts working, the microphone continuously picks up audio, and the module performs analog-to-digital conversion, gain amplification, and preliminary buffering on the microphone audio stream.

[0052] The module performs VAD (Voice Activity Detection) on the initially cached audio stream data to determine whether it is the start of a valid speech event. If not, it continues with real-time audio stream capture and VAD testing. If it is, it intercepts the valid speech segment.

[0053] The active speech segment captured by the VAD is subjected to noise reduction processing and fed into the command word recognition model for speech content recognition. This noise reduction effectively removes noise introduced by the headset microphone placed near the mouth, such as exhalation, sneezing, and coughing. This facilitates recognition by the command word recognition model, thereby improving the accuracy of the recognition results.

[0054] Based on the recognition results of the command word recognition network model, it can eventually be determined whether there is a valid command word in the current audio stream. If there is a valid command word, the specific command word recognition result is output. If there is no valid command word, the real-time audio acquisition is continued to be returned for recognition.

[0055] The audio playback of the transcription system is controlled by voice control, replacing the existing manual control method. The recognition of voice commands is realized by a dedicated module, which does not occupy the computing resources of the voice transcription system, has low recognition latency, and the user's voice is not uploaded to other devices except the chip module, so it can also protect user privacy. In general, the present invention provides an application mode that combines a speech transcription system with a speech chip module with speech recognition, which moves the recognition process related to speech control to the hardware end, and can realize voice control of the software without excessive modification of the software. This application mode can also be applied to other software besides the speech transcription system to realize other software control functions besides audio playback control. Ultimately, it provides users of software such as the transcription system with an operation method based on voice control in addition to manual control, avoiding the user's hands repeatedly switching between the keyboard and the mouse when correcting the transcribed text, improving the user's work efficiency in audio playback and text modification, and enhancing the automation and intelligence of the transcription system.

[0056] Accordingly, the embodiment of the present invention also provides a voice chip module, see Figure 6 , the voice chip module includes: A receiving unit, configured to receive a voice signal; a processing unit, configured to parse the received voice signal, convert the voice signal into a system control instruction corresponding to the voice transcription system, and trigger the voice transcription system to perform a corresponding playback operation according to the converted system control instruction; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

[0057] In specific implementation, the voice chip module in the embodiment of the present invention is an intelligent voice chip, which is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and parses the received voice signal in an offline state; and the voice chip module includes a battery unit, which is used to power the voice chip module, and the battery unit is charged through the USB plug on the voice chip module.

[0058] Furthermore, an embodiment of the present invention further provides a system for text editing of a speech transcription system based on speech control, wherein the system comprises: a speech chip module and a speech transcription system connected to each other; The voice chip module is used to analyze the received voice signal and convert the voice signal into a system control instruction corresponding to the voice transcription system; The speech transcription system is used to complete the corresponding playback operation according to the system control instructions converted by the speech chip module; The text editing system of the speech transcription system based on voice control of the present invention is specially equipped with a voice chip module, so that the user can control the audio playback of the speech transcription system through voice, while both hands are only responsible for typing on the keyboard. This operation can effectively improve the user's text editing efficiency and greatly enhance the user experience.

[0059] In addition, embodiments of the present invention further provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for automatically sorting table fields in a PostgreSQL database as described in any one of the method embodiments of the present invention. For details, please refer to the method embodiments of the present invention and will not be discussed in detail here.

[0060] The relevant contents of the module embodiment, system embodiment and storage medium embodiment of the present invention can be understood by referring to the method embodiment of the present invention, and will not be discussed in detail here.

[0061] Although the preferred embodiments of the present invention have been disclosed for illustrative purposes, those skilled in the art will appreciate that various modifications, additions and substitutions are possible, and thus, the scope of the present invention should not be limited to the above embodiments.

Claims

1. A text editing method for a speech transcription system based on voice control, characterized in that: include: Setting up a voice chip module and connecting the voice chip module to a voice transcription system; The voice chip module analyzes the received voice signal and converts the voice signal into a system control instruction corresponding to the voice transcription system; Triggering the speech transcription system to complete corresponding playback operations according to the converted system control instructions; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

2. The method according to claim 1, characterized in that The voice chip module is connected to a microphone and receives the voice signal through the microphone; The voice chip module includes a battery unit, which is used to power the voice chip module and charge the battery unit through the USB plug on the voice chip module.

3. The method according to claim 1, characterized in that The voice chip module is an intelligent voice chip. The intelligent voice chip is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and analyzes the received voice signal in an offline state.

4. The method according to claim 3, characterized in that The multi-layer reconfigurable spatial domain computing architecture is used to dynamically adjust hardware resources to achieve efficient analysis and recognition of the voice signal; the digital signal processing core DSP is used to preprocess the received voice signal and post-process the voice signal after feature extraction and classification, and the neural network processing core NPU is used to run a deep learning model and perform feature extraction and classification on the preprocessed voice signal through the deep learning model to achieve accurate analysis and recognition of the voice recognition.

5. The method according to claim 3, characterized in that The voice chip module analyzes the received voice signal and converts the voice signal into a system control instruction corresponding to the voice transcription system, including: When the voice chip module receives a voice signal, the digital signal processing core DSP performs analog-to-digital conversion, buffering, voice activity detection (VAD), and recognition processing on the voice signal to obtain the voice content of the voice signal; The neural network processing core NPU is used to determine whether the voice content is a correct voice instruction. If so, the voice content is converted into a system control instruction corresponding to the voice transcription system according to a preset communication protocol.

6. The method according to any one of claims 1 to 5, characterized in that The system control instruction includes: a packet header, an instruction type bit, a data packet length bit, several instruction content bits and a packet tail, wherein: The packet header is used to identify the beginning of the data packet of the system control instruction; The instruction type bit is used to identify the type of instruction in the data packet, and the instruction type includes fixed instructions and non-fixed instructions; The data packet length bit is used to indicate the length of the instruction content in the data packet so that the speech transcription system can correctly parse the data packet; The instruction content bit is the data of the system control instruction; The packet tail is used to mark the end of the data packet.

7. A voice chip module, characterized in that: The voice chip module includes: A receiving unit, configured to receive a voice signal; a processing unit, configured to parse the received voice signal, convert the voice signal into a system control instruction corresponding to the voice transcription system, and trigger the voice transcription system to perform a corresponding playback operation according to the converted system control instruction; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

8. The voice chip module according to claim 7, characterized in that: The voice chip module is an intelligent voice chip. The intelligent voice chip is designed based on a multi-layer reconfigurable spatial computing architecture, integrates a neural network processing core NPU and a digital signal processing core DSP, and analyzes the received voice signal in an offline state; The voice chip module includes a battery unit, which is used to power the voice chip module, and the battery unit is charged through a USB plug on the voice chip module.

9. A system for text editing of speech transcription system based on voice control, characterized in that: The system includes: a voice chip module and a voice transcription system connected to each other; The voice chip module is used to analyze the received voice signal and convert the voice signal into a system control instruction corresponding to the voice transcription system; The speech transcription system is used to complete the corresponding playback operation according to the system control instructions converted by the speech chip module; Among them, the voice signal includes one or more of a play signal, a stop signal, a pause signal, a fast forward signal, a rewind signal, a mute signal, an unmute signal, a loop play signal, a play signal at a certain point and a certain speed, a start signal positioned at a certain hour, minute and second, a play signal positioned at a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal jumped to a certain hour, minute and second, a play signal jumped to a certain hour, minute and second, a start signal from a certain hour, minute and second, and a play signal from a certain hour, minute and second.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for editing text in a speech transcription system based on voice control according to any one of claims 1 to 6 is implemented.