Speech processing method, apparatus, device, storage medium, and computer program product
By creating target threads and instances using a speech processing chain description file, the problem of low efficiency in microphone array module function adjustment is solved, enabling more efficient speech processing function adjustment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT DIGITAL TIANJIN
- Filing Date
- 2021-12-15
- Publication Date
- 2026-07-21
AI Technical Summary
When existing microphone array modules need to implement new voice processing functions or adjust voice processing functions, the fixed code needs to be manually modified, which is inefficient and slow.
By creating target threads and speech processing instances based on a speech processing chain description file, speech processing functionality can be achieved by modifying the description file rather than the fixed code.
It improves the efficiency and speed of voice processing functions and reduces maintenance pressure.
Smart Images

Figure CN116264080B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to speech processing methods, speech processing apparatus, computer equipment, computer-readable storage media, and computer program products. Background Technology
[0002] With the popularization of online meetings and the refinement of meeting equipment, microphone array modules, with their professional features, microphone array algorithms, no consumption of host computing power, and hot-swappability, are widely used in desktop and large-screen meeting systems, which can greatly improve the purity and intelligibility of the voice in the meeting.
[0003] Currently, microphone array modules implement corresponding voice processing functions based on fixed code. When a new voice processing function needs to be implemented in the microphone array module, or when the implementation method of the voice processing function needs to be adjusted, the fixed code needs to be modified manually. However, manually modifying the fixed code is inefficient and slow. Summary of the Invention
[0004] This application provides a speech processing method, apparatus, device, storage medium, and computer program product, which can quickly implement the speech processing function of the speech processing device based on the speech processing chain description file, with higher efficiency and faster speed.
[0005] This application provides a speech processing method, including:
[0006] Acquire the voice data to be processed;
[0007] The target thread is invoked to perform calculations on the speech data to be processed according to the speech processing instance, and the target speech processing result of the speech data to be processed is obtained.
[0008] The target thread is created based on the thread-related information included in the speech processing chain description file, and the speech processing instance is created based on the speech processing function-related information included in the speech processing chain description file.
[0009] One embodiment of this application provides a voice processing device, including:
[0010] The acquisition module is used to acquire the voice data to be processed;
[0011] The processing module is used to call the target thread to perform calculations on the speech data to be processed according to the speech processing instance, and obtain the target speech processing result of the speech data to be processed.
[0012] The target thread is created based on the thread-related information included in the speech processing chain description file, and the speech processing instance is created based on the speech processing function-related information included in the speech processing chain description file.
[0013] One embodiment of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the voice processing method provided in this application embodiment.
[0014] One aspect of this application provides a computer storage medium storing a computer program, which includes program instructions. When the program instructions are executed by a processor, they perform the speech processing method provided in this application.
[0015] One aspect of this application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. When the computer instructions are executed by the processor of a computer device, the voice processing method provided in this application is executed.
[0016] This application embodiment creates a target thread based on the thread-related information included in the voice processing chain description file, and calls the target thread to create a voice processing instance based on the voice processing function-related information included in the voice processing chain description file. This allows for the rapid implementation of voice processing functions of voice processing devices (such as microphone array modules) based on the voice processing chain description file. Thus, when a new voice processing function needs to be implemented in the voice processing device, or when the implementation method of the voice processing function needs to be adjusted, only the voice processing chain description file needs to be modified, which is more efficient and faster than modifying the fixed code. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a network architecture to which the speech processing method provided in the embodiments of this application is applicable;
[0019] Figure 2 This is a schematic flowchart of a speech processing method provided in an embodiment of this application;
[0020] Figure 3This is a schematic diagram of a sequentially arranged functional description data group provided in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of a sequentially arranged speech processing example provided in an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of another sequentially arranged speech processing example provided in the embodiments of this application;
[0023] Figure 6 This is a flowchart illustrating another speech processing method provided in an embodiment of this application;
[0024] Figure 7 This is a schematic diagram illustrating an example of running voice processing provided in an embodiment of this application;
[0025] Figure 8 This is a schematic diagram of another example of running voice processing provided in the embodiments of this application;
[0026] Figure 9 This is a schematic diagram illustrating another example of voice processing provided in this application embodiment;
[0027] Figure 10 This is a schematic diagram illustrating another example of voice processing provided in this application embodiment;
[0028] Figure 11 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application;
[0029] Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0031] This application provides a speech processing method that can quickly implement the speech processing function of a speech processing device based on a speech processing chain description file, resulting in higher efficiency and faster speed. In feasible embodiments, the speech processing method provided in this application can be implemented based on cloud technology.
[0032] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0033] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.
[0034] Cloud computing is a computing model that distributes computing tasks across a large pool of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" appear infinitely scalable, readily available, on-demand, and expandable, with payment based on usage.
[0035] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of storage devices of various types (storage devices are also called storage nodes) in the network to work together through application software or application interfaces to provide data storage and business access functions to the outside world.
[0036] A database, simply put, can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, with minimal redundancy, and independent of application programs.
[0037] The embodiments of this application may specifically involve one or more of cloud technologies, such as cloud computing, cloud storage, and cloud database. For example, cloud computing is used to process the voice data to be processed, or data required to execute the voice processing method is obtained from a cloud database, or the target voice processing result obtained from the processing is stored in cloud storage.
[0038] The speech processing method provided in this application embodiment can be applied to... Figure 1 The network architecture shown. Figure 1 The voice processing device 10 shown can be a server or terminal device with data processing capabilities. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal device can be a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device, smart home appliance, vehicle terminal, etc., but is not limited to these.
[0039] Optionally, the voice processing device 10 can be a microphone array module. This microphone array module can be understood as a sound acquisition system. It should be noted that a lightweight operating system such as Linux or a Real-Time Operating System (RTOS) can run inside the microphone array module to support the implementation of various functions within it. The microphone array module can also be equipped with one or more microphone terminals for acquiring sound from different spatial directions. The microphone array module typically carries multiple voice function modules, which are arranged according to certain requirements and coupled with corresponding algorithms to construct a voice processing link capable of performing various functions. For example, assuming that the microphone array module sequentially carries an echo cancellation module, a noise reduction module, and an angle determination module, a voice processing link with echo cancellation, noise reduction, and angle determination functions can be constructed using these three voice function modules. The microphone array module can be externally connected to a computer or a large-screen conference device via a Universal Serial Bus (USB) to achieve functions such as voice acquisition and voice preprocessing, and to provide high-quality voice processing for terminal devices.
[0040] exist Figure 1In this context, database 11 can be a local database of the voice processing device 10, or a cloud database accessible to the voice processing device 10. The voice processing method provided in this embodiment can be executed by the voice processing device 10, specifically:
[0041] The system retrieves the voice data to be processed from database 11 and calls the target thread created based on the thread-related information included in the voice processing chain description file. It then creates voice processing instances based on the voice processing function-related information included in the voice processing chain description file, and connects these instances into a voice processing chain according to the thread-related information in the file. This chain processes the retrieved voice data to be processed, yielding the target voice processing result. This allows each functional module involved in the voice processing chain to quickly create corresponding voice processing instances using the aforementioned voice processing chain description file. Furthermore, subsequent maintenance can be performed on this voice processing chain description file, reducing the burden of maintaining multiple sets of description files.
[0042] Figure 2 This is a schematic flowchart of a speech processing method provided in an embodiment of this application. The speech processing method can be... Figure 1 The speech processing method is executed by the speech processing device shown, specifically by its processor. The speech processing method includes the following steps:
[0043] S201, Obtain the speech processing chain description file. The speech processing chain description file includes one or more object structures. Each object structure includes thread-related information and speech processing function-related information. The thread-related information includes the processing thread identifier and the processor core identifier.
[0044] The aforementioned speech processing chain description file can be a file described using JavaScript Object Notation (JSON). Specifically, the speech processing device can obtain this speech processing chain description file by parsing the basic code implementing the speech processing device's functions using a JSON parser.
[0045] It should be noted that the basic code for implementing the above functions can be a collection of code written in a voice processing device to implement various functions. Based on this basic code, the voice processing device can determine the thread-related information, processor core information, and the serial relationships between the various functional modules, and then implement various functions sequentially according to the above information and serial relationships.
[0046] It's also worth noting that JSON is a lightweight data-interchange format. JSON is easy to read and write, as well as easy to parse and generate. Similar to Extensible Markup Language (XML), JSON is plain text and possesses self-descriptive and hierarchical structure. However, unlike XML, JSON lacks closing tags, occupies fewer bytes, is faster to read and write, and can incorporate data group structures. JSON has two representation structures: objects and data groups. The following section will briefly describe these two JSON representation structures: object structure and data group structure.
[0047] An object structure is an unordered collection of key / value pairs, typically beginning with curly braces "{" and ending with curly braces "}". The middle part of the object structure can consist of zero or more key / value pairs separated by commas; keys and values can be separated by colons ":". The following example illustrates a specific way to describe an object structure:
[0048]
[0049] As can be seen from the above description of the object structure, the object structure can include two "key / value" pairs, namely key1:value1 and key2:value2; where the key can be a string, and the value can be any of the following: string, number, true, false, null, object, or array.
[0050] Data group structures typically begin with square brackets "[" and end with square brackets "]". The middle part of a data group structure can consist of zero or more comma-separated lists of values. Data group structures and object structures can be nested within each other. The following example illustrates a specific way to describe a data group structure:
[0051]
[0052] As can be seen from the above example description of the data group structure, this data group structure contains two nested object structures, each of which includes two "keyword / value" pairs. It is understood that the inclusion of "keyword / value" pairs in the object structures in the above example description is merely illustrative; data group structures can also be nested within these two object structures, and this application does not impose any restrictions on this.
[0053] It should be noted that the aforementioned speech processing chain description file can be a data group structure, which can nest one or more object structures. Optionally, one or more data groups can be nested within these object structures; this application does not impose any restrictions on this.
[0054] In this context, each of the aforementioned object structures can correspond to a target thread, and the aforementioned thread-related information can be information describing the thread corresponding to the object structure, such as the processing thread identifier of the target thread and the processor core identifier of the target thread.
[0055] The aforementioned speech processing function-related information can be descriptive information describing the speech processing function corresponding to the target thread; specifically, it can include information about the function description data group (i.e., speech processing function module) within the target thread, as well as the function parameter description information within the function module. This speech processing function-related information can include one or more function description data groups, meaning that one target thread can correspond to one or more speech processing function modules; each of these multiple function description data groups can include: function name, function identifier, and function parameter description information. The relevant content of this speech processing function-related information will be described in detail below, and will not be repeated here.
[0056] S202, based on the processing thread identifier and processor core identifier included in each object structure, create a target thread that matches each object structure.
[0057] After the voice processing device obtains the aforementioned voice processing chain description file, it can determine the processing thread identifier and processor core identifier corresponding to each object structure based on one or more object structures in the voice processing chain description file, and create a target thread that matches each object structure according to the processing thread identifier and processor core identifier corresponding to each object structure.
[0058] For example, suppose the speech processing chain description file includes an object structure, the processing thread of which corresponds to 1 and the processor core identifier corresponds to 2; then the target thread can be determined as thread 1, and the target thread can be bound to processor core 2, that is, thread 1 can be run on processor core 2.
[0059] In one possible implementation, when the aforementioned speech processing chain description file includes multiple object structures, the multiple target threads corresponding to these multiple object structures belong to one or more processor cores.
[0060] In this context, multiple target threads can belong to multiple processor cores; that is, each target thread can be bound to a different processor core. It's understandable that when the number of processor cores is greater than or equal to the number of target threads, to avoid insufficient computing power on any single processor core, each target thread can be assigned to a separate processor core to improve the efficiency of subsequent processing of the audio data. For example, assuming there are 3 processor cores and 3 target threads, then these 3 object structures can be assigned to 3 processor cores respectively.
[0061] Optionally, multiple target threads can belong to a single processor core. That is, each or some of the target threads can be bound to the same processor core. It is understood that when the number of processor cores is less than the number of target threads, each of the target threads can be assigned to a single processor core, or some of the target threads (two or more target threads) can be assigned to a single processor core.
[0062] For example, assuming there is one processor core and two object structures, both object structures can be assigned to the same processor core. Alternatively, assuming there are two processor cores and three object structures, the object structure with the highest computing power can be assigned to one processor core, and the other two object structures with lower computing power can be assigned to the other processor core, thus avoiding insufficient computing power on any single processor core. For example, for a dual-core processor, if object structure 1 occupies 45% of the computing power, object structure 2 occupies 70% of the computing power, and object structure 3 occupies 20% of the computing power, then object structure 2 can be assigned to one processor core, and object structures 1 and 3 can be assigned to the other processor core.
[0063] In one possible implementation, the aforementioned thread-related information also includes downstream thread identifiers; based on the downstream thread identifiers included in the aforementioned object structures and the processing thread identifiers corresponding to the aforementioned target threads, the aforementioned target threads are connected.
[0064] The downstream thread identifier mentioned above can be used to indicate the identifier of the next target thread to run after the current target thread has finished running. The voice processing device can determine the processing order of each target thread based on the downstream thread identifier and the processing thread identifier of each target thread, thereby connecting the aforementioned target threads.
[0065] It should be noted that the connection between the aforementioned target threads can be a series connection, a parallel connection, or a combination of series and parallel connections. This application does not impose any restrictions on this. The following description uses a series connection as an example and does not limit this application.
[0066] For example, suppose there are three target threads, and the processing thread identifier of the first target thread is 2, and the downstream thread identifier is end; the processing thread identifier of the second target thread is 3, and the downstream thread identifier is 2; the processing thread identifier of the third target thread is 1, and the downstream thread identifier is 3. Then, based on the processing thread identifiers and downstream thread identifiers of the three target threads, these three target threads can be connected sequentially, meaning the execution order of the three target threads can be: third target thread, second target thread, first target thread. Where the current target thread is the last target thread, i.e., the downstream thread is empty, the downstream thread identifier can be set to end or set to 0; this application does not impose any restrictions on this.
[0067] S203, calls each target thread to create a voice processing instance that matches each target thread based on the corresponding voice processing function information.
[0068] After the voice processing device creates each target thread, it can call each target thread so that each target thread can create a voice processing instance that matches the target thread based on the voice processing function information included in the object structure.
[0069] In one possible implementation, the above-mentioned speech processing instance may include one or more of the following: an instance corresponding to noise reduction processing, an instance corresponding to echo cancellation processing, and an instance corresponding to angle decision processing.
[0070] Among them, the noise reduction processing instance can be used to reduce noise, the echo cancellation processing instance can be used to eliminate echo, and the angle decision processing instance can be used to make a decision on the speaker's angle.
[0071] In one possible implementation, for any target thread, the target thread is invoked to create a speech processing instance that matches each of the corresponding one or more function description data groups.
[0072] As can be seen from the foregoing, the aforementioned voice processing function-related information may include one or more function description data groups. Each of these function description data groups may include: function name, function identifier, and function parameter description information.
[0073] For speech processing instances that can achieve the same function (such as instances corresponding to noise reduction processing), the same function name can be used, meaning the description file corresponding to the function name can be the same. However, it is understood that the function identifier in the speech processing chain description file is unique to ensure that each function that needs to be implemented can correspond to a unique identifier. That is to say, even if the function names corresponding to the speech processing chain description files are the same, factors such as different bound processor cores, different target threads, or different corresponding function parameter description information may exist, thus making the function identifier a unique identifier to distinguish the above-mentioned different situations.
[0074] When the voice processing function information includes multiple function description data groups, the voice processing device can call the target thread that includes multiple function description data groups, so that the target thread can create voice processing instances that are respectively matched with the multiple function description data groups.
[0075] For example, assuming that the speech processing function information of the target thread includes two function description data groups, such as the function description data group corresponding to noise reduction processing and the function description data group corresponding to angle decision processing, the target thread can create an instance corresponding to noise reduction processing and an instance corresponding to angle decision processing based on the two function description data groups respectively.
[0076] In one possible implementation, when there are multiple speech processing instances corresponding to any of the above target threads, the execution order among the multiple speech processing instances corresponding to any of the target threads is determined.
[0077] Because each speech processing instance has certain restrictions on the input format, output format, and data characteristics, there is a sequential order among the various functional description data groups. For example, ... Figure 3 As shown, the data input format for the echo cancellation processing instance can be a multi-channel speech signal, while the data output format after echo cancellation processing can be a single-channel speech signal. Similarly, the data input format for the noise reduction processing instance can be a single-channel speech signal, while the data output format after noise reduction processing can also be a single-channel speech signal. In other words, the noise reduction processing instance can be placed after the echo cancellation processing instance. Furthermore, the success rate of angle decision processing is highly sensitive to the purity of the speech signal. Therefore, to obtain a more accurate speaker angle decision, the angle decision processing instance can be placed after the noise reduction processing instance.
[0078] It should be noted that the speech processing device can obtain the execution order among multiple speech processing instances corresponding to any given target thread based on the acquired speech processing chain description file. For example, such as Figure 4 As shown, the multiple speech processing instances obtained by the speech processing device are as follows: the instance corresponding to noise reduction processing ( Figure 4 (represented by NS in Chinese), and instances corresponding to echo cancellation processing ( Figure 4 (referred to as AEC in Chinese) and the corresponding instance of angle decision processing ( Figure 4 (represented by DOA in Chinese), then based on the obtained speech processing chain description file, the execution order of the three speech processing instances can be determined as follows: the instance corresponding to echo cancellation, the instance corresponding to noise reduction, and the instance corresponding to angle decision processing. Among these, in... Figure 4 The execution order of the three speech processing instances is for illustrative purposes only and does not limit the scope of this application. Optionally, in Figure 4 In this example, the three worker threads correspond to three different processor cores (i.e., core 1, core 2, and core 3 in the diagram), which is only used as an example and does not impose any limitations on this application.
[0079] Optionally, if the voice processing device acquires multiple identical voice processing instances, the order between these instances can be adjusted to some extent. For example... Figure 5 As shown, assuming the speech processing device obtains four speech processing instances: two instances corresponding to noise reduction processing, one instance corresponding to echo cancellation processing, and one instance corresponding to angle decision processing, the execution order of these four speech processing instances can be determined based on the obtained speech processing chain description file as follows: the instance corresponding to echo cancellation processing, the instance corresponding to noise reduction processing, the instance corresponding to angle decision processing, and the instance corresponding to noise reduction processing. Figure 5 In this application, the instance corresponding to the second noise reduction process can be located after the instance corresponding to the angle decision process, and this application does not impose any restrictions on this.
[0080] For example, this application provides the following exemplary speech processing chain description files. Specifically, speech processing chain description file 1 is as follows:
[0081]
[0082] The aforementioned speech processing chain description file 1 includes an object structure corresponding to a target thread. From the thread-related information corresponding to this object structure, it can be seen that the processing thread identifier of this target thread is 1 (also referred to as target thread 1); the processor core identifier corresponding to this object structure is 1; the downstream thread identifier corresponding to this object structure is 2 (also referred to as target thread 2), meaning that the next target thread after target thread 1 is target thread 2. From the speech processing function-related information corresponding to this object structure, it can be seen that target thread 1 includes a function description data group, the function name corresponding to this function description data group is AEC (i.e., echo cancellation function module), the function identifier corresponding to this function description data group is 1, and the function parameter description information corresponding to this function description data group includes parameters such as enable signal (aec), gain, policy (hpf_policy), and filter length (aec_filter_len_ms). The parameters included in the above function parameter description information are for illustrative purposes only and do not limit this application.
[0083] Specifically, the speech processing chain description file 2 is as follows:
[0084]
[0085] The aforementioned speech processing chain description file 2 includes two object structures, each corresponding to a target thread (target thread 1 and target thread 2). Since the downstream thread identifier of target thread 1 is 2, target thread 2 can be the next in line after target thread 1. Since the downstream thread identifier of target thread 2 is "end", target thread 2 can be the last target thread. Since the processor core identifier for target thread 1 is 1 and the processor core identifier for target thread 2 is 2, target threads 1 and 2 are bound to different processor cores. Based on the speech processing function information of target thread 1, it includes a function description data group with the function name AEC and a function identifier of 1. Similarly, based on the speech processing function information of target thread 2, it includes a function description data group with the function name NS (i.e., noise reduction module) and a function identifier of 2. The order of these two function description data groups is: AEC, NS.
[0086] Specifically, the speech processing chain description file 3 is as follows:
[0087]
[0088]
[0089] The aforementioned speech processing chain description file 3 includes three object structures, which correspond to three target threads (i.e., target thread 1, target thread 2, and target thread 3). Based on the processing thread identifiers and downstream thread identifiers of target thread 1, target thread 2, and target thread 3, the order of the three target threads is: target thread 1, target thread 2, and target thread 3. Based on the processor core identifiers corresponding to the three target threads, the three target threads are bound to different processor cores (i.e., processor core 1, processor core 2, and processor core 3). Based on the speech processing function information of the three target threads, each of the three target threads includes a function description data group, and the order and function name of each function description data group are: AEC, NS, and DOA (i.e., angle decision function module).
[0090] It should be noted that the aforementioned speech processing chain description files 1, 2, and 3 are examples where one target thread corresponds to one functional description data set, and each target thread is bound to a different processor core. The difference lies in the data set: in speech processing chain description file 1, the data set corresponds to one object structure, meaning one speech processing instance can be created; in speech processing chain description file 2, the data set corresponds to two object structures, meaning two speech processing instances can be created; and in speech processing chain description file 3, the data set corresponds to three object structures, meaning three speech processing instances can be created. As can be seen from the aforementioned speech processing chain description files 1, 2, and 3, one data set structure can correspond to one or more object structures.
[0091] Specifically, the speech processing chain description file 4 is as follows:
[0092]
[0093]
[0094] The aforementioned speech processing chain description file 4 includes two object structures, each corresponding to a target thread (target thread 1 and target thread 2). The order of these two target threads is: target thread 1, target thread 2. These two target threads are bound to two processor cores (processor core 1 and processor core 2). Based on the speech processing function information of target thread 2, it is known that target thread 2 includes two function description data groups. The order and function names of these two function description data groups are: NS and DOA. It can be understood that, considering one function description data group corresponding to target thread 1, the order and function names of the three function description data groups included in this data group are: AEC, NS, and DOA.
[0095] It should be noted that the target thread 2 of the aforementioned speech processing chain description file 4 corresponds to two functional description data groups, namely NS and DOA, which share one thread and one core. For dual-core processors, the configuration method provided in speech processing chain description file 4 allows two functional description data groups to be placed in one thread and bound to one processor core. This avoids placing three functional description data groups in the same thread and binding them to the same processor core, thus preventing a situation where one processor core has insufficient computing power while the other processor core has idle computing power.
[0096] Specifically, the speech processing chain description file 5 is as follows:
[0097]
[0098]
[0099] The aforementioned speech processing chain description file 5 includes three object structures, which correspond to three target threads (i.e., target thread 1, target thread 2, and target thread 3). As can be seen from the processor core identifiers corresponding to the three target threads, target thread 1 and target thread 3 are bound to the same processor core (i.e., processor core 1); target thread 2 is bound to another processor core (i.e., processor core 2).
[0100] In real-world scenarios, the complexity of an instance corresponding to echo cancellation processing is 45% of that of a single core, the complexity of an instance corresponding to noise reduction processing is 70% of that of a single core, and the complexity of an instance corresponding to angle decision processing is 25% of that of a single core. Therefore, for a dual-core processor, the instance corresponding to noise reduction processing, which has the highest complexity, can be bound to one processor core, while the instances corresponding to echo cancellation processing and angle decision processing, which have lower complexity, can be bound to the other processor core. This ensures a more balanced computing power distribution between the two cores in the dual-core processor while meeting the real-time requirements of each speech processing instance.
[0101] In this embodiment of the application, by obtaining a speech processing chain description file, it can be found that the speech processing chain description file includes one or more object structures. Based on the processing thread identifier and processor core identifier included in each object structure, a target thread matching each object structure is created. Each target thread is then called to create a speech processing instance matching each target thread based on the corresponding speech processing function information. This allows the speech processing device to call the target thread to perform calculations on the speech data to be processed according to the speech processing instance when processing the speech data to be processed in the future.
[0102] Figure 6 This is a schematic flowchart of a speech processing method provided in an embodiment of this application. The speech processing method can be... Figure 1The speech processing method is executed by the speech processing device shown, specifically by its processor. The speech processing method includes the following steps:
[0103] S601, acquire the voice data to be processed;
[0104] The aforementioned voice data to be processed can be any voice input data acquired by the voice processing device. For example, the speaker's voice input data when speaking through the voice processing device in a meeting; or the voice input data of the caller when making a video call on a personal computer through the voice processing device.
[0105] S602, the target thread is invoked to perform calculations on the speech data to be processed according to the speech processing instance, and the target speech processing result of the speech data to be processed is obtained; wherein, the target thread is created based on the thread-related information included in the speech processing chain description file, and the speech processing instance is created by the target thread based on the speech processing function-related information included in the speech processing chain description file.
[0106] It should be noted that the content regarding the creation of the target thread based on the thread-related information included in the speech processing chain description file, and the content regarding the creation of a speech processing instance based on the speech processing function-related information included in the speech processing chain description file, can be found in the aforementioned documents. Figure 2 The detailed descriptions in the corresponding embodiments are not repeated here.
[0107] After acquiring the aforementioned voice data to be processed, the voice processing device can call the created target thread and perform calculations on the aforementioned voice data to be processed according to the created voice processing instance, so as to obtain the target voice processing result corresponding to the voice data to be processed.
[0108] In one possible implementation, based on the thread connection result, each of the aforementioned target threads is invoked according to the corresponding speech processing instance to sequentially process the speech data to be processed, or the intermediate speech processing result of the speech data to be processed, to obtain the target speech processing result of the speech data to be processed; wherein, the input of the second thread is the second intermediate speech processing result obtained by the first thread processing the speech data to be processed or the first intermediate speech processing result of the speech data to be processed according to the corresponding speech processing instance, the first thread being any target thread, and the second thread being the target thread ranked after the first thread in the thread connection result.
[0109] As described above, when creating each target thread, the speech processing device can connect the target threads based on the downstream thread identifiers included in each object structure and the processing thread identifiers corresponding to each target thread. It can be understood that the speech processing device can, based on the aforementioned thread connection results, call each of the target threads and, according to the speech processing instance corresponding to each target thread, sequentially process the acquired speech data to be processed, or the intermediate speech processing results of the speech data to be processed, to obtain the target speech processing result of the speech data to be processed.
[0110] For example, suppose there are 3 target threads, and the connection order of the 3 target threads is: target thread 1, target thread 3 and target thread 2; then the voice processing device can, based on the connection order, perform calculations on the above-mentioned voice data to be processed, or the intermediate voice processing results of the above-mentioned voice data to be processed, according to the voice processing instance 1 corresponding to target thread 1, the voice processing instance 3 corresponding to target thread 3 and the voice processing instance 2 corresponding to target thread 2, and obtain the target voice processing result.
[0111] It is understandable that if the first thread is any target thread and the second thread is the target thread that is ranked after the first thread in the above thread connection result, then the second intermediate speech processing result obtained by the first thread through the above calculation and processing of the speech data to be processed, or the first intermediate speech processing result of the speech data to be processed, can be the input of the second thread.
[0112] It should be noted that when the first thread is the first target thread, it can directly process the aforementioned speech data to be processed to obtain the second intermediate speech processing result. Optionally, when the first thread is not the first target thread, it can process the output of the target thread that ranks before the first thread in the aforementioned thread connection result (i.e., the first intermediate speech processing result of the aforementioned speech data to be processed) to obtain the second intermediate speech processing result.
[0113] For example, suppose there are three target threads, and the connection order of these three target threads is: target thread 1, target thread 3, and target thread 2; if the first thread is target thread 3, then the second thread is target thread 2. It can be understood that the second intermediate speech processing result obtained by target thread 3 after processing the first intermediate speech processing result of the speech data to be processed can be the input of target thread 2; wherein, the first intermediate speech processing result of the speech data to be processed can be the output obtained by target thread 1 after processing the speech data to be processed.
[0114] It should be noted that, when the second thread is the last target thread, the second intermediate speech processing result can be the target speech processing result, that is, the final target speech processing result obtained by the speech processing device after processing the speech data to be processed. Optionally, when the second thread is not the last target thread, the speech processing device can continue to process the second intermediate speech processing result.
[0115] In one possible implementation, when there are multiple speech processing instances corresponding to any of the aforementioned target threads, when any target thread is called to perform calculations on the speech data to be processed, the target thread performs calculations on the speech data to be processed or the intermediate speech processing results of the speech data to be processed in sequence according to the aforementioned multiple speech processing instances and the aforementioned execution order.
[0116] As described above, when a voice processing device invokes any of the aforementioned target threads to create multiple voice processing instances corresponding to those target threads, it can create these multiple voice processing instances sequentially according to the execution order of each voice processing instance in the voice processing chain description file. It can be understood that when the voice processing device invokes any of the target threads to perform calculations on the voice data to be processed, it can also perform calculations on the voice data to be processed, or the intermediate voice processing results of the voice data to be processed, sequentially according to the aforementioned multiple voice processing instances and their corresponding execution order.
[0117] For example, suppose there are two target threads, and the connection order of the two target threads is: target thread 1 and target thread 2; where target thread 1 corresponds to speech processing instance 1, and target thread 2 corresponds to speech processing instance 2 and speech processing instance 3 in sequence; then when the speech processing device calls the above two target threads to perform calculations on the speech data to be processed, it can perform calculations on the speech data to be processed in the order of speech processing instance 1, speech processing instance 2 and speech processing instance 3 in sequence.
[0118] For example, please see Figure 7 , Figure 7 A speech processing chain 1 can be created for speech based on the aforementioned speech processing chain description file 1. This speech processing chain 1 includes a speech processing instance, specifically the instance corresponding to echo cancellation processing. After acquiring the speech data to be processed, the speech processing device can invoke this speech processing chain 1 to perform calculations on the speech data and obtain the target speech processing result after echo cancellation processing.
[0119] Please see Figure 8 , Figure 8A speech processing chain 2 can be created for speech based on the aforementioned speech processing chain description file 2. This speech processing chain 2 includes two speech processing instances: one for echo cancellation and one for noise reduction. After acquiring the speech data to be processed, the speech processing device can invoke this speech processing chain 2 to perform calculations on the speech data and obtain the target speech processing result after echo cancellation and noise reduction processing.
[0120] Please see Figure 4 , Figure 4 A speech processing chain (also referred to as speech processing chain 3) can be created based on the speech created using the aforementioned speech processing chain description file 3. This speech processing chain 3 includes three speech processing instances: an instance corresponding to echo cancellation processing, an instance corresponding to noise reduction processing, and an instance corresponding to angle decision processing. After acquiring the speech data to be processed, the speech processing device can call this speech processing chain 3 to perform calculations on the speech data and obtain the target speech processing result after echo cancellation processing, noise reduction processing, and angle decision processing have been performed.
[0121] Please see Figure 9 , Figure 9 A speech processing chain 4 can be created for speech based on the aforementioned speech processing chain description file 4. This speech processing chain 4 includes three speech processing instances: an instance for echo cancellation, an instance for noise reduction, and an instance for angle decision processing. Specifically, the instances for noise reduction and angle decision processing can share worker thread 2 and processor core 2. After acquiring the speech data to be processed, the speech processing device can call this speech processing chain 4 to perform calculations on the speech data and obtain the target speech processing result after echo cancellation, noise reduction, and angle decision processing.
[0122] Please see Figure 10 , Figure 10 A speech processing chain 5 can be created for speech based on the aforementioned speech processing chain description file 5. This speech processing chain 5 includes three speech processing instances: an instance for echo cancellation, an instance for noise reduction, and an instance for angle decision processing. Specifically, the instance for noise reduction is bound to processing core 2, while the instances for echo cancellation and angle decision processing can be assigned to different work threads (i.e., work thread 1 and work thread 2 respectively) and share processor core 1. After acquiring the speech data to be processed, the speech processing device can call this speech processing chain 5 to perform calculations on the speech data and obtain the target speech processing result after echo cancellation, noise reduction, and angle decision processing.
[0123] In practical applications, due to increasingly diverse needs, the functional requirements of microphone array modules are also constantly increasing, leading to a wide variety of microphone array module solutions. The differences between these various microphone array modules often lie in the processor cores and clock speeds, allowing different module solutions to correspond to different voice processing chains. During microphone array module initialization, its different algorithm modules can be "combined" and "ordered," and based on the number of processor cores and computing power, these algorithm modules can be rationally allocated to the corresponding processor cores for computation and processing.
[0124] It's important to note that in the early stages of microphone array module projects, due to limited business requirements, fixed code is typically used to specify the algorithm modules within the speech processing chain and the sequential relationships between them. For example, when constructing an "echo cancellation-noise reduction-angle decision" processing chain, a worker thread can be created first, followed by the sequential creation of three processing instances: echo cancellation, noise reduction, and angle decision. These instances are then attached to the worker thread. When processing the speech data to be processed through this worker thread, the three processing instances are executed sequentially, thus achieving the aforementioned three functions and outputting the processed speech data. However, when using fixed code to generate the speech processing chain, the processor core scheduled during computation is entirely determined by the operating system, and the algorithm modules within the worker thread are also fixed. If a particular algorithm module requires high computing power, a single-threaded speech processing chain cannot handle it.
[0125] To adapt to a dual-core processor module, you can first create two processing threads, then create the necessary voice algorithm modules sequentially, and then attach each voice algorithm module to its corresponding thread. Finally, attach the two threads to their respective processor cores, thus completing the "core binding" of the algorithm modules. It's important to note that this process actually uses fixed code; therefore, any modifications required to adapt to different projects, such as changes to any part of the code, will necessitate code modifications.
[0126] It is understandable that, unlike the aforementioned method of creating a speech processing chain through fixed code, the present application embodiment uses a JSON description file to generate a speech processing chain within a microphone array module. Leveraging the strong readability, natural hierarchy, nested object structure, and array structure of JSON, the functional algorithm modules, upstream and downstream relationships (sequentiality), and bound processor core information (core binding) required within the speech processing chain can be described using the JSON language.
[0127] During microphone array module initialization, a JSON parser can be used to parse the aforementioned JSON description file to obtain information related to the speech processing chain. This information is then used to create each algorithm module, connect threads and algorithm modules, and bind threads to specified processor cores using core binding information, thus generating a complete speech processing chain. If subsequent business requirements require changes to the core binding information or expansions of algorithm modules, only the JSON description file needs to be modified, avoiding extensive code changes.
[0128] It's worth noting that by using a single JSON language to create multiple JSON description files, different functional modules can be implemented. This avoids the need for multiple description languages for microphone array modules with different functional algorithms, thus greatly reducing the pressure of maintaining the described speech. Due to the high readability of the JSON language, the computational resources required for reading and parsing it are minimal. Furthermore, the JSON language has good compatibility across different hardware platforms, and can be used on both desktop and large-screen conferencing devices.
[0129] In this embodiment, by acquiring the voice data to be processed, a target thread created based on the thread-related information included in the voice processing chain description file can be invoked. The voice processing instance created by the target thread based on the voice processing function-related information included in the voice processing chain description file is then used to process the voice data to be processed, so as to obtain the target voice processing result of the voice data to be processed. This allows each functional module to create a corresponding voice processing instance through a set of description files, thereby reducing the pressure of maintaining the description files and facilitating subsequent functional expansion.
[0130] Furthermore, in the specific embodiments of this application, data related to the voice data to be processed, the target voice processing result, the voice processing chain description file, the target thread, and the voice processing instance are involved, and all data used is authorized by the user. When the above embodiments of this application are applied to specific products or technologies, the data used must be authorized or agreed to by the user, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0131] For further details, please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a voice processing device provided in an embodiment of this application. Figure 11 As shown, the voice processing device 1100 can be applied to the above-mentioned Figure 1The corresponding embodiment describes a voice processing device. Specifically, the voice processing device 1100 can be a computer program (including program code) running within the voice processing device; for example, the voice processing device 1100 is an application software. The voice processing device 1100 can be used to execute... Figure 1 The corresponding steps in the method provided in the corresponding embodiment.
[0132] The voice processing device 1100 may include an acquisition unit 1101 and a processing unit 1102.
[0133] Acquisition unit 1101 is used to acquire the voice data to be processed;
[0134] Processing unit 1102 is used to call the target thread to perform calculations on the speech data to be processed according to the speech processing instance, and obtain the target speech processing result of the speech data to be processed.
[0135] The target thread is created based on the thread-related information included in the speech processing chain description file, and the speech processing instance is created based on the speech processing function-related information included in the speech processing chain description file.
[0136] In one possible implementation, the acquisition unit 1101 is further configured to acquire a speech processing chain description file, which includes one or more object structures. Each object structure includes thread-related information and speech processing function-related information. The thread-related information includes a processing thread identifier and a processor core identifier. The processing unit 1102 is further configured to create a target thread matching each object structure based on the processing thread identifier and processor core identifier included in each object structure. The processing unit 1102 is further configured to call each target thread to create a speech processing instance matching each target thread based on the corresponding speech processing function-related information.
[0137] In one possible implementation, the aforementioned thread-related information also includes a downstream thread identifier; the aforementioned processing unit 1102 is further configured to connect the aforementioned target threads based on the downstream thread identifiers included in the aforementioned object structures and the processing thread identifiers corresponding to the aforementioned target threads.
[0138] In one possible implementation, the processing unit 1102 is further configured to call each of the target threads according to the corresponding speech processing instance based on the thread connection result, and sequentially perform calculations on the speech data to be processed or the intermediate speech processing result of the speech data to be processed to obtain the target speech processing result of the speech data to be processed; wherein, the input of the second thread is the second intermediate speech processing result obtained by the first thread according to the corresponding speech processing instance, and the first thread is any target thread, and the second thread is the target thread ranked after the first thread in the thread connection result.
[0139] In one possible implementation, the aforementioned speech processing function-related information includes one or more function description data groups, each of which includes a function name, a function identifier, and function parameter description information; for any target thread, the aforementioned processing unit 1102 is further configured to call that target thread to create a speech processing instance that matches each function description data group based on the corresponding one or more function description data groups.
[0140] In one possible implementation, when there are multiple speech processing instances corresponding to any of the target threads, the processing unit 1102 is further configured to determine the execution order among the multiple speech processing instances corresponding to the target thread; wherein, when the target thread is called to perform calculation processing on the speech data to be processed, the target thread performs calculation processing on the speech data to be processed or the intermediate speech processing result of the speech data to be processed in sequence according to the multiple speech processing instances and the execution order.
[0141] In one possible implementation, when the aforementioned speech processing chain description file includes multiple object structures, the multiple target threads corresponding to these multiple object structures belong to one or more processor cores.
[0142] In one possible implementation, the above-mentioned speech processing instance may include one or more of the following: an instance corresponding to noise reduction processing, an instance corresponding to echo cancellation processing, and an instance corresponding to angle decision processing.
[0143] According to the embodiments of this application, Figure 11The various units in the illustrated voice processing device can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the voice processing device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0144] According to embodiments of this application, it is possible to execute functions such as those described above by running on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 2 and Figure 6 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 11 The speech processing apparatus shown herein, and the speech processing method implemented according to the embodiments of this application. The aforementioned computer program may be recorded on, for example, a computer storage medium, loaded onto the aforementioned computing device via the computer storage medium, and run therein.
[0145] Further, please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 1200 can also be used to implement the server functions described in the above method embodiments. For example... Figure 12 As shown, the computer device 1200 may include at least: a processor 1201, a communication interface 1202, and a computer storage medium 1203. The processor 1201, communication interface 1202, and computer storage medium 1203 may be connected via a bus or other means.
[0146] Computer storage medium 1203 can be stored in memory 1204 of computer device 1200. This computer storage medium 1203 is used to store computer programs, which include program instructions. The processor 1201 is used to execute the program instructions stored in the computer storage medium 1203. The processor 1201 (or CPU (Central Processing Unit)) is the computing and control core of computer device 1200, suitable for implementing one or more instructions, specifically suitable for loading and executing:
[0147] Acquire the voice data to be processed;
[0148] The target thread is invoked to perform calculations on the speech data to be processed according to the speech processing instance, and the target speech processing result of the speech data to be processed is obtained.
[0149] The target thread is created based on the thread-related information included in the speech processing chain description file, and the speech processing instance is created based on the speech processing function-related information included in the speech processing chain description file.
[0150] In one possible implementation, the processor 1201 is further configured to obtain a speech processing chain description file, which includes one or more object structures. Each object structure includes thread-related information and speech processing function-related information. The thread-related information includes a processing thread identifier and a processor core identifier. The processor 1201 is further configured to create a target thread matching each object structure based on the processing thread identifier and processor core identifier included in each object structure. The processor 1201 is further configured to call each target thread to create a speech processing instance matching each target thread based on the corresponding speech processing function-related information.
[0151] In one possible implementation, the aforementioned thread-related information also includes a downstream thread identifier; the processor 1201 is further configured to connect the target threads based on the downstream thread identifiers included in the respective object structures and the processing thread identifiers corresponding to the respective target threads.
[0152] In one possible implementation, the processor 1201 is further configured to call each of the target threads based on the thread connection result, according to the corresponding speech processing instance, to sequentially process the speech data to be processed or the intermediate speech processing result of the speech data to be processed, to obtain the target speech processing result of the speech data to be processed; wherein, the input of the second thread is the second intermediate speech processing result obtained by the first thread according to the corresponding speech processing instance, which is the speech data to be processed or the first intermediate speech processing result of the speech data to be processed, and the first thread is any target thread, and the second thread is the target thread that is ranked after the first thread in the thread connection result.
[0153] In one possible implementation, the aforementioned speech processing function-related information includes one or more function description data groups, each of which includes a function name, a function identifier, and function parameter description information; for any target thread, the aforementioned processor 1201 is further configured to call that target thread to create a speech processing instance that matches each function description data group based on the corresponding one or more function description data groups.
[0154] In one possible implementation, when there are multiple speech processing instances corresponding to any of the above target threads, the processor 1201 is further configured to determine the execution order among the multiple speech processing instances corresponding to the any of the target threads; wherein, when the any of the target threads is called to perform calculations on the speech data to be processed, the any of the target threads performs calculations on the speech data to be processed or the intermediate speech processing results of the speech data to be processed in sequence according to the multiple speech processing instances and the above execution order.
[0155] In one possible implementation, when the aforementioned speech processing chain description file includes multiple object structures, the multiple target threads corresponding to these multiple object structures belong to one or more processor cores.
[0156] In one possible implementation, the above-mentioned speech processing instance may include one or more of the following: an instance corresponding to noise reduction processing, an instance corresponding to echo cancellation processing, and an instance corresponding to angle decision processing.
[0157] It should be understood that the computer device 1200 described in the embodiments of this application can perform the foregoing... Figure 2 and Figure 6 The description of the speech processing method in the corresponding embodiments can also be performed as described above. Figure 11 The speech processing device 1100 described in this embodiment will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated here.
[0158] Furthermore, it should be noted that this application embodiment also provides a computer storage medium, which stores the computer program executed by the aforementioned voice processing device 1100. The computer program includes program instructions, and when the processor executes the program instructions, it can execute the aforementioned... Figure 2 and Figure 6 The description of the speech processing method in the corresponding embodiments is already provided and will not be repeated here. Similarly, the beneficial effects of using the same method will not be repeated here either. For technical details not disclosed in the embodiments concerning the computer storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed and executed on a single computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed across multiple locations and interconnected via a communication network. Multiple computer devices distributed across multiple locations and interconnected via a communication network can be combined to form a blockchain network.
[0159] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned... Figure 2 and Figure 6 The methods described in the corresponding embodiments are therefore not repeated here.
[0160] Those skilled in the art will recognize that the units and steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0161] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0162] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this invention should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A speech processing method, characterized in that, The method includes: A speech processing chain description file is obtained, which includes one or more object structures. Each object structure includes thread-related information and speech processing function-related information. The thread-related information includes a processing thread identifier, a processor core identifier, and a downstream thread identifier. When the number of processor cores is less than the number of target threads, the processor cores are allocated according to the computing power occupancy of each object structure for the one or more object structures. Based on the processing thread identifier and processor core identifier included in each object structure, a target thread matching the respective object structure is created; Based on the downstream thread identifiers included in each object structure and the processing thread identifiers corresponding to each target thread, the target threads are connected. Each target thread is invoked to create a voice processing instance that matches the corresponding voice processing function information. Acquire the voice data to be processed; Based on the thread connection result, each target thread is invoked to perform calculations on the speech data to be processed or the intermediate speech processing result of the speech data to be processed in sequence according to the corresponding speech processing instance, so as to obtain the target speech processing result of the speech data to be processed.
2. The method according to claim 1, characterized in that, The input of the second thread is the second intermediate speech processing result obtained by the first thread through the calculation and processing of the speech data to be processed or the first intermediate speech processing result of the speech data to be processed according to the corresponding speech processing instance. The first thread is any target thread, and the second thread is the target thread that is ranked after the first thread in the thread connection result.
3. The method according to claim 1 or 2, characterized in that, The voice processing function-related information includes one or more function description data groups, each function description data group including function name, function identifier and function parameter description information; The step of calling each target thread to create a voice processing instance matching each target thread based on the corresponding voice processing function information includes: For any target thread, the target thread is invoked to create a speech processing instance that matches each of the corresponding one or more function description data groups.
4. The method according to claim 3, characterized in that, The method further includes: When there are multiple speech processing instances corresponding to any target thread, determine the execution order among the multiple speech processing instances corresponding to any target thread; Specifically, when any of the target threads is invoked to perform calculations on the voice data to be processed, the target thread performs calculations on the voice data to be processed or the intermediate voice processing results of the voice data to be processed in sequence according to the plurality of voice processing instances and the execution order.
5. The method according to claim 1, characterized in that, When the speech processing chain description file includes multiple object structures, the multiple target threads corresponding to the multiple object structures belong to one or more processor cores.
6. The method according to claim 1, characterized in that, The speech processing examples include one or more of the following: examples of noise reduction processing, examples of echo cancellation processing, and examples of angle decision processing.
7. A voice processing device, characterized in that, include: The acquisition module is used to acquire a speech processing chain description file. The speech processing chain description file includes one or more object structures. Each object structure includes thread-related information and speech processing function-related information. The thread-related information includes a processing thread identifier, a processor core identifier, and a downstream thread identifier. When the number of processor cores is less than the number of target threads, processor cores are allocated for the one or more object structures according to the computing power occupied by each object structure. The processing module is used to create a target thread that matches each object structure based on the processing thread identifier and processor core identifier included in each object structure, connect each target thread based on the downstream thread identifier included in each object structure and the processing thread identifier corresponding to each target thread, and call each target thread to create a voice processing instance that matches each target thread based on the corresponding voice processing function related information. The acquisition module is also used to acquire the voice data to be processed; The processing module is further configured to call each target thread based on the thread connection result to perform calculations on the speech data to be processed or the intermediate speech processing result of the speech data to be processed in sequence according to the corresponding speech processing instance, so as to obtain the target speech processing result of the speech data to be processed.
8. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause a computer device having the processor to perform the steps of the method according to any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.