Offline voice recognition method and system, electronic equipment and storage medium
By introducing gateways and multiple middleware service nodes into the voice recognition system and processing multiple voice requests in parallel, the problem of limited synchronization processing capabilities of existing systems in high concurrency scenarios is solved, and efficient concurrent processing and response speed improvement is achieved.
Patent Information
- Application Number
- CN202510347556.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
When existing speech recognition systems face large-scale, highly concurrent offline speech recognition tasks, their synchronization processing capabilities are limited, resulting in performance bottlenecks.
By introducing gateways and multiple middleware service nodes into the offline voice recognition system, multiple voice requests are processed in parallel, and the system's concurrent processing capability and response speed are improved. The specific implementation includes the client sending a request to obtain the target task ID to the gateway, the gateway allocates the request to the target middleware service node, and the middleware service node performs voice recognition and returns the result.
It realizes the system's efficient processing capabilities in high concurrency scenarios, makes full use of multi-core CPU resources, and improves response speed and overall performance.
Smart Images

Figure CN120183399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and particularly to an offline speech recognition method, system, electronic device, and storage medium. Background Art
[0002] Currently, when facing large-scale and high-concurrency offline speech recognition tasks, speech recognition systems have the following main problems: limited synchronous processing ability and performance bottleneck. Since speech is continuous, traditional speech recognition systems usually process speech data in a synchronous manner. This synchronous processing method cannot fully utilize the processing power of multi-core CPUs, resulting in the system being unable to provide sufficient processing power in high-concurrency scenarios and becoming a performance bottleneck. Summary of the Invention
[0003] Embodiments of this application provide an offline speech recognition method, system, electronic device, and storage medium, which can process multiple speech requests in parallel, improving the concurrent processing ability and response speed of the speech recognition system.
[0004] In a first aspect embodiment of this application, an offline speech recognition method is provided, which is applied to an offline speech recognition system. The offline speech recognition system includes: a client, a gateway, and multiple middleware service nodes. The gateway establishes communication connections with the client and each middleware service node among the multiple middleware service nodes respectively. The method includes:
[0005] When the client detects a speech recognition operation, it sends a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation. At least the following are included in the request for obtaining a target task ID: the audio ID of the audio to be recognized, and a target string, where the target string is the audio to be recognized or is used to indicate the audio to be recognized;
[0006] The gateway assigns the request for obtaining a target task ID to a corresponding target middleware service node;
[0007] After receiving the request for obtaining a task ID, the target middleware service node generates a target task ID, returns the target task ID to the client through the gateway, and obtains the audio to be recognized according to the target string, performs speech recognition on the audio to be recognized, and obtains a target recognition result;
[0008] After obtaining the target task ID, the client sends a request for obtaining a target recognition result to the target middleware service node through the gateway. At least the following are included in the request for obtaining a recognition result: the audio ID of the audio to be recognized and the target task ID;
[0009] After receiving the request for obtaining the target recognition result, the target middleware service node obtains the target recognition result and returns the target recognition result to the client through the gateway.
[0010] In some possible embodiments, different intermediate service nodes correspond to different preset hash value ranges. The gateway allocates the request for obtaining the target task ID to the corresponding target middleware service node, including:
[0011] Calculate a target hash value according to the audio ID of the audio to be recognized and the hash algorithm;
[0012] Determine the intermediate service node corresponding to the target hash value range where the target hash value is located as the target intermediate service node, and the target hash value range is one of the multiple preset hash value ranges;
[0013] Allocate the request for the target task ID to the target intermediate service node.
[0014] In some possible embodiments, the method further includes:
[0015] Map the multiple middleware service nodes to a consistent hash ring to obtain the initial hash value range corresponding to each middleware service node, and the initial hash value ranges corresponding to different middleware service nodes are the same;
[0016] Obtain the weight of each middleware service node according to the performance of each middleware service node, where the better the performance, the greater the weight;
[0017] Adjust the range of partial rings corresponding to different middleware service nodes on the consistent hash ring according to the weight of each middleware service node, and determine the preset hash value range of each middleware service node according to the range of partial rings corresponding to each middleware service node.
[0018] In some possible embodiments, after receiving the request for obtaining the task ID, the target middleware service node generates a target task ID, including:
[0019] After receiving the request for obtaining the task ID, the target middleware service node determines whether the request for obtaining the task ID includes a task ID. If the request for obtaining the task ID does not include a task ID, the target task ID is generated for the request for obtaining the task ID.
[0020] In some possible embodiments, the information type of the target information is further included in the request for obtaining the target task ID. The information type includes URL type, base64 type, and binary type. The target middleware service node obtains the audio to be recognized according to the target string, including:
[0021] The target middleware service node obtains the audio to be recognized according to the information type and the target string;
[0022] Wherein, when the information type is the URL type, the audio to be recognized is obtained by downloading the audio in the link corresponding to the target string;
[0023] When the information type is the base64 type, the target string is decoded by the base64 decoding method to obtain the audio to be recognized;
[0024] When the information type is the binary type, the target string is determined as the audio to be recognized.
[0025] In some possible embodiments, the middleware service node further includes a local database, which is used to store the speech recognition status and recognition results corresponding to each existing task ID. The speech recognition status includes one of in recognition, recognition completed, and error. The target middleware service node performs speech recognition on the audio to be recognized to obtain a target recognition result, including:
[0026] The target middleware service node adds a target recognition task to the speech recognition queue. The target recognition task includes the target task ID and the audio to be recognized;
[0027] Read the execution result of the speech recognition queue and update the local database;
[0028] Wherein, when the execution result does not include the execution result of the target task, update the speech recognition status of the target task ID in the local database to the recognition status; when the execution result includes the execution result of the target task ID, update the speech recognition status of the target task ID to recognition completed and store the execution result, and the execution result is the target recognition result; when the execution result includes the error reason of the target task, update the speech recognition status of the target task ID to error and store the error reason.
[0029] In some possible embodiments, after receiving the request for obtaining the target recognition result, the target middleware service node obtains the target recognition result, including:
[0030] The target middleware service node searches the local database for the speech recognition status corresponding to the target task ID;
[0031] If the speech recognition status is the recognition status, return a status identifier indicating that the speech recognition status is in recognition;
[0032] If the speech recognition status is recognition completed, return a status identifier indicating that the speech recognition status is recognition completed and the target recognition result;
[0033] If the speech recognition status is an error, return a status identifier indicating that the speech recognition status is an error and the error reason.
[0034] An embodiment of the second aspect of the present application provides an offline speech recognition system, the offline speech recognition system includes: a client, a gateway, and a plurality of middleware service nodes, the gateway respectively establishes a communication connection with the client and each middleware service node among the plurality of middleware service nodes, wherein:
[0035] The client is configured to, when detecting a speech recognition operation, send a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation; at least the following are included in the request for obtaining a target task ID: the audio ID of the audio to be recognized, and a target string, the target string being the audio to be recognized or used to indicate the audio to be recognized;
[0036] The gateway is configured to allocate the request for obtaining a target task ID to a corresponding target middleware service node;
[0037] The target middleware service node is configured to, after receiving the request for obtaining a task ID, generate a target task ID, return the target task ID to the client through the gateway, obtain the audio to be recognized according to the target string, and perform speech recognition on the audio to be recognized to obtain a target recognition result;
[0038] The client is configured to, after obtaining the target task ID, send a request for obtaining a target recognition result to the target middleware service node through the gateway, at least the following are included in the request for obtaining a recognition result: the audio ID of the audio to be recognized and the target task ID;
[0039] The target middleware service node is configured to, after receiving the request for obtaining a target recognition result, obtain the target recognition result, and return the target recognition result to the client through the gateway.
[0040] An embodiment of the third aspect of the present application provides an electronic device, including:
[0041] A processor;
[0042] A memory for storing instructions executable by the processor;
[0043] Wherein, the processor is configured to execute the instructions to implement the method described in any one of the embodiments of the first aspect of this application.
[0044] The embodiments of the fourth aspect of this application provide a storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the method described in any one of the embodiments of the first aspect of this application.
[0045] The technical solutions provided by the embodiments of this application at least include the following beneficial effects:
[0046] The embodiments of this application propose an offline speech recognition method, which is applied to an offline speech recognition system. The offline speech recognition system includes: a client, a gateway, and multiple middleware service nodes. The gateway establishes communication connections with the client and each middleware service node among the multiple middleware service nodes respectively. The method includes: when the client detects a speech recognition operation, it sends a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation; at least included in the request for obtaining a target task ID are: the audio ID of the audio to be recognized, and a target string, where the target string is the audio to be recognized or is used to indicate the audio to be recognized; the gateway allocates the request for obtaining a target task ID to a corresponding target middleware service node; after receiving the request for obtaining a task ID, the target middleware service node generates a target task ID, returns the target task ID to the client through the gateway, and obtains the audio to be recognized according to the target string, performs speech recognition on the audio to be recognized, and obtains a target recognition result; after obtaining the target task ID, the client sends a request for obtaining a target recognition result to the target middleware service node through the gateway. At least included in the request for obtaining a recognition result are: the audio ID of the audio to be recognized and the target task ID; after receiving the request for obtaining a target recognition result, the target middleware service node obtains the target recognition result and returns the target recognition result to the client through the gateway. In this way, the system can process multiple speech requests in parallel, make full use of the multi-core CPU resources, and improve the concurrent processing ability and response speed of the system. The asynchronous processing of the request for obtaining a target task ID and the request for obtaining a target recognition result not only reduces the response time of each request, but also avoids the blocking problem that may be brought by synchronous calls, thereby effectively improving the overall performance of the system. Description of the Drawings
[0047] Figure 1 Schematic diagram of an application scenario of an offline speech recognition method proposed in an embodiment of this application;
[0048] Figure 2 Schematic flowchart of an offline speech recognition method proposed in an embodiment of this application;
[0049] Figure 3 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0050] Figure 4 Schematic diagram of a consistent hashing ring proposed in an embodiment of this application;
[0051] Figure 5 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0052] Figure 6 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0053] Figure 7 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0054] Figure 8 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0055] Figure 9 Schematic flowchart of another offline speech recognition method proposed in an embodiment of this application;
[0056] Figure 10 Schematic diagram of the structure of an offline speech recognition system proposed in an embodiment of this application;
[0057] Figure 11 Schematic diagram of the structure of an electronic device proposed in an embodiment of this application. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of this application will be described with reference to the accompanying drawings.
[0059] To facilitate a clear description of the technical solutions in the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions and effects. For example, the first instruction and the second instruction are used to distinguish different user instructions, and their order is not limited. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0060] It should be noted that in the embodiments of the present application, words such as "exemplarily" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0061] In addition, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0062] It should be noted that in the embodiments of the present application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element.
[0063] The technical implementation of offline speech recognition systems has long faced performance bottleneck problems brought about by the synchronous processing mechanism. Traditional systems generally adopt a linear pipeline architecture, and speech data needs to sequentially complete serialized processing processes such as audio preprocessing, feature extraction, acoustic model inference, and language model decoding. There is a strict sequential dependency relationship between each link, resulting in low utilization of multi-core computing resources. Especially when processing long audio or high-concurrency tasks, the phenomenon of idle waiting for computing resources is significant. Although some solutions attempt to introduce asynchronous interfaces to relieve the real-time pressure, due to the lack of a unified asynchronous task management standard and cross-platform adaptation capabilities, there are significant differences in protocol compatibility, status feedback mechanisms, etc. among the asynchronous interfaces of different manufacturers, making it difficult to form a general-purpose technology ecosystem, which severely restricts the horizontal expansion ability of the system.
[0064] In terms of scalability and session maintenance, the existing architecture faces more complex engineering challenges. Since offline speech recognition tasks often require maintaining session-level context consistency (such as in multi-turn dialogue scenarios), traditional load balancing strategies are difficult to meet the rigid requirement that the same task request must be continuously routed to a fixed node, resulting in service interruptions easily caused by state migration and session reconstruction operations involved in the node expansion process. In addition, the system expansion mode based on static resource allocation cannot dynamically adapt to task volume fluctuations. After adding new nodes, local resource idle or overload often occurs due to uneven task migration, further exacerbating the uncontrollability of the overall service quality. Traditional speech recognition systems usually process speech data in a synchronous manner, which means that the entire process from audio input to text output is completed in a continuous and linear process. In such a system, each step must wait for the previous step to complete before it can start, and the whole process is blocking until the final result is generated and returned to the user. This synchronous processing method cannot fully utilize the processing power of multi-core CPUs, resulting in the system being unable to provide sufficient processing power in high-concurrency scenarios and becoming a performance bottleneck. Although some manufacturers provide asynchronous interfaces, they have not formed standardized interfaces and cannot be adapted to the interfaces of other manufacturers.
[0065] In view of this, an embodiment of the present application proposes an offline speech recognition method, which is applied to an offline speech recognition system, wherein the offline speech recognition system includes: a client, a gateway, and multiple middleware service nodes, wherein the gateway establishes a communication connection with the client and each of the multiple middleware service nodes respectively, and the method includes: when the client detects a speech recognition operation, sending a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation; the request for obtaining the target task ID includes at least: an audio ID of the audio to be recognized, and a target string, wherein the target string is the audio to be recognized or is used to indicate the audio to be recognized; the gateway assigns the request for obtaining the target task ID to the corresponding The target middleware service node; after receiving the request to obtain the task ID, the target middleware service node generates a target task ID, returns the target task ID to the client through the gateway, obtains the audio to be recognized according to the target string, performs speech recognition on the audio to be recognized, and obtains the target recognition result; after obtaining the target task ID, the client sends the request to obtain the target recognition result to the target middleware service node through the gateway, and the request to obtain the recognition result includes at least: the audio ID of the audio to be recognized and the target task ID; after receiving the request to obtain the target recognition result, the target middleware service node obtains the target recognition result and returns the target recognition result to the client through the gateway. In this way, the system can process multiple voice requests in parallel, make full use of multi-core CPU resources, and improve the concurrent processing capability and response speed of the system. The asynchronous processing of the request to obtain the target task ID and the request to obtain the target recognition result not only reduces the response time of each request, but also avoids the blocking problem that may be caused by synchronous calls, thereby effectively improving the overall performance of the system.
[0066] The following is a detailed explanation of an application scenario example of an embodiment of the present application. For example, Figure 1 As shown, it includes: a client 101, a gateway 102, and multiple middleware service nodes 103.
[0067] Exemplarily, the client serves as a user interaction portal and may include a mobile terminal (such as an e-commerce APP built with React Native), a Web terminal (a management backend developed with Vue), an IoT device (smart home controller), and other forms. Exemplarily, the client initiates a request through http, https, or WebSocket protocol.
[0068] Exemplarily, in the embodiments of the present application, the gateway may be an Application Programming Interface (API) gateway. The API gateway is a dedicated application layer gateway that specifically manages application programming interface traffic, such as request routing and protocol conversion in a microservices architecture. The API gateway is an intermediate layer between the client and the backend service and serves as the access point for all external requests. Its functions include authentication, request routing, traffic limiting, caching, etc., to ensure that requests can be accurately forwarded to the appropriate backend service. The API gateway realizes load balancing of requests through a load balancing module to ensure the high availability of the backend service. Since the API gateway needs to undertake the load balancing task in a high-concurrency scenario to utilize all middleware service nodes, it is crucial to adopt a suitable load balancing algorithm. The gateway establishes communication connections with the client and each of the multiple middleware service nodes respectively.
[0069] Exemplarily, the middleware service node is a core component for processing speech recognition tasks. Multiple middleware service nodes are used for load balancing, improving concurrency capabilities, and system reliability. At the same time, requests are reasonably allocated through a weighting strategy to optimize resource usage. Among them, since speech recognition is a computationally intensive task (such as feature extraction and decoding), a single node is limited by CPU / memory resources and cannot handle large-scale offline speech processing (such as audio of tens of thousands of hours). Therefore, in the embodiments of the present application, there are multiple middleware service nodes, including middleware service node 1, middleware service node 2,..., middleware service node n.
[0070] After understanding an application scenario and related functions of the embodiments of the present application, please refer to Figure 2 and the following will detail the execution steps of the offline speech recognition method proposed in the embodiments of the present application.
[0071] Step 201: The client detects a speech recognition operation.
[0072] Exemplarily, the speech recognition operation may be an active operation by the user, such as clicking the "Start Recognition" button, uploading an audio file, or the end of voice input (such as releasing the recording button). It may also be an automatically detected speech activity, for example, detecting human voices in a real-time speech stream and automatically triggering recognition. It may also be triggered by a system event, such as automatically initiating speech transcription after a call ends or a meeting recording is saved.
[0073] Step 202: The client sends a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation.
[0074] The target task ID acquisition request at least includes: the audio ID of the audio to be recognized, and a target string, where the target string is the audio to be recognized or is used to indicate the audio to be recognized. Exemplarily, the target string can be of the network address URL type or base64 string, etc., that indicates the audio to be recognized.
[0075] Exemplarily, the target task ID acquisition request may further include audio recognition request parameters, including but not limited to: whether to add punctuation, whether to turn on inverse text normalization, whether to turn on the language model for smoothing processing, and whether to turn on speaker recognition. The audio recognition request parameters are used to control the specific processing method of speech-to-text conversion and the optimization of the target recognition result in subsequent steps.
[0076] Exemplarily, automatically adding punctuation marks (such as full stops, commas, question marks, etc.) to the speech recognition result can significantly improve the readability of the text and make it more in line with the expression habits of natural language. For example, for the original target recognition result without punctuation "Today the weather is nice let's go for a walk", the target recognition result after adding punctuation is "Today the weather is nice, let's go for a walk." This function is applicable to scenarios that require structured text, such as meeting records, customer service recording transcription, subtitle generation, etc., and can effectively improve the clarity and practicality of the text. Exemplarily, whether to turn on inverse text normalization is to convert symbols such as numbers, dates, and currencies in the speech recognition result into a preset standard text format to make the text more in line with the written expression habits. Whether to turn on the language model for smoothing processing is to use a language model (such as N-gram, neural network model) to optimize the recognition result in context and correct semantic or grammar errors. Whether to turn on speaker recognition is to identify different speakers in the audio and label the speaker identities (such as "Speaker 1", "Speaker 2"). Turning on speaker recognition helps to distinguish multi-role conversations and is applicable to multi-person meeting or interview recordings.
[0077] When the user performs the speech recognition operation, the client will obtain the audio to be recognized and the audio ID of the audio to be recognized according to the speech recognition operation, and generate a target task ID acquisition request.
[0078] Exemplarily, if the audio to be recognized is actively uploaded by the user, directly generate the audio ID of the audio to be recognized according to the audio to be recognized and the pre-set audio ID generation rule, and then generate a target task ID acquisition request.
[0079] Exemplarily, if the voice recognition operation indicates the network address URL type of the audio to be recognized, it is necessary to first obtain the online link of the audio file provided by the user, download the audio to be recognized through this online link, generate the audio ID of the audio to be recognized according to the audio to be recognized and the preset audio ID generation rule, and then generate a request to obtain the target task ID.
[0080] Exemplarily, the preset audio ID generation rule can be any one of selecting a random number, (Message-Digest Algorithm 5, md5) value, hash value of the audio file name, or the audio file name. Among them, md5 is a widely used hash function that is used to convert data of any length (such as files, texts, audios, etc.) into a fixed-length 128-bit (16-byte) hash value (usually represented as a 32-bit hexadecimal string).
[0081] After the client generates a request to obtain the target task ID, it sends the request to obtain the target task ID to the gateway through a preset communication connection. Exemplarily, the communication connection can be established based on protocols such as http, https, and websocket, which depends on the design requirements of the system. The client will pass the data to the gateway by calling the corresponding API or sending a network request and wait for the response from the gateway.
[0082] Step 203: The gateway allocates the request to obtain the target task ID to the corresponding target middleware service node.
[0083] When a large number of requests (such as audio recognition tasks) pour in simultaneously, a single middleware service node cannot withstand the high concurrency pressure, and the requests need to be distributed to multiple middleware service nodes.
[0084] After receiving the request to obtain the target task ID sent by the client, the gateway will route the request to the corresponding target middleware service node according to the preset allocation strategy. This allocation process can be achieved in various ways, depending on the system architecture and requirements. The following are several common allocation methods and their examples:
[0085] Exemplarily, the gateway can dynamically select a node with a lower load according to the load situation of the current middleware service node. For example, if there are three middleware service nodes in the system, the gateway will monitor the CPU usage rate, memory occupancy, or request queue length of each node in real time, and then allocate the request to the most idle node to ensure the reasonable utilization of system resources and the efficient processing of requests.
[0086] Exemplarily, the gateway can also use the consistent hashing algorithm to distribute requests to the target middleware service nodes. For example, the gateway can calculate the hash value based on the audio ID in the request for obtaining the target task ID, and then distribute the request for obtaining the target task ID to the middleware service node corresponding to the hash value.
[0087] Step 204: After receiving the request for obtaining the task ID, the target middleware service node generates a target task ID and returns the target task ID to the client through the gateway.
[0088] Exemplarily, after receiving the request for obtaining the task ID, the target middleware service node determines whether the request for obtaining the task ID includes a task ID. If the request for obtaining the task ID does not include a task ID, the target task ID is generated for the request for obtaining the task ID.
[0089] To ensure the uniqueness and traceability of the target task ID, a specific algorithm or mechanism is usually adopted to generate the target task ID. The following are exemplary generation methods and their implementation logics:
[0090] Exemplarily, the middleware service node can combine the current timestamp and its own unique identifier to generate the target task ID. For example, using a high-precision timestamp (such as millisecond or microsecond level) as the basis, then concatenating the middleware service node ID, and finally appending a random number or an incrementing sequence to ensure the global uniqueness of the task ID in a distributed environment.
[0091] Exemplarily, the middleware service node can use the Universally Unique Identifier (UUID) algorithm to generate the task ID. UUID is a standardized 128-bit identifier, and its generation method is based on elements such as timestamps, random numbers, and hardware addresses, which can ensure that the target task IDs generated in a distributed system are almost non-repetitive.
[0092] After generating the target task ID, the middleware service node binds the target task ID to the audio ID. If the http protocol is used for communication between the gateway and the middleware service node, the middleware service node will return the task ID to the gateway in the form of JSON or other structured data formats through the http response, and return the target task ID to the client through the gateway. The client can continuously track the task status (such as in progress / completed / error) based on the task ID, while the middleware service node can quickly locate the task details through the task ID to achieve efficient status management and result feedback.
[0093] Step 205: The target middleware service node obtains the audio to be recognized according to the target string, performs speech recognition on the audio to be recognized, and obtains a target recognition result.
[0094] The target string is the audio to be recognized or used to indicate the audio to be recognized. After obtaining the content or acquisition method of the audio to be recognized, the underlying Automatic Speech Recognition (ASR) interface is called to perform speech recognition on the speech to be recognized, and a target recognition result is obtained. The ASR interface is a programming interface provided by speech recognition technology. By calling the service through code, the audio to be recognized (such as voice recording, real-time stream) is converted into text, and the target recognition result is stored in the local database. The recognition status and the corresponding target recognition result can also be stored.
[0095] Step 206: After obtaining the target task ID, the client sends a request for obtaining the target recognition result to the target middleware service node through the gateway. The request for obtaining the target recognition result at least includes: the audio ID of the audio to be recognized and the target task ID.
[0096] In this step, the client generates a request for obtaining the target recognition result according to the target task ID and the audio ID of the audio to be recognized corresponding to the target task ID, and polls to send the request for obtaining the target recognition result. The request for obtaining the target recognition result is allocated to the same middleware service node through the gateway according to a preset allocation policy. Since the middleware service node is matched by calculating the hash value of the audio ID during the process of allocating the request for obtaining the target recognition result, different requests for the same speech recognition task can always be allocated to the same middleware service node for processing, thus maintaining the continuity and consistency of the session.
[0097] Step 207: After receiving the request for obtaining the target recognition result, the target middleware service node obtains the target recognition result and returns the target recognition result to the client through the gateway.
[0098] Among them, the target recognition result at least includes: the speech recognition status corresponding to the task ID and the recognition result.
[0099] After the target middleware service node receives the request for obtaining the target recognition result, it will first retrieve the corresponding target recognition result from the local database according to the task ID. This process usually involves querying the mapping relationship between the task ID and the recognition result to ensure that the correct data can be quickly located.
[0100] Once the target recognition result is obtained, the middleware service node will send the target recognition result to the gateway through a preset communication connection. After receiving the target recognition result, the gateway will verify and forward it. For example, the gateway may check the integrity of the data and the legitimacy of the source, and then return the recognition result to the client through the same communication protocol. After receiving the recognition result, the client will parse and process it, such as presenting the recognition text to the user or storing it locally for subsequent use.
[0101] In the embodiments of the present application, the method further includes: standardizing the synchronous interface, asynchronous interface, and streaming interface into two types of asynchronous interfaces through a standardized request interface and multiple middleware services. This standardized processing method enables the system to handle different types of requests more flexibly, while optimizing resource utilization and improving the processing capacity and response speed of the system. The streaming interface is a type of interface used to process continuous data streams.
[0102] Exemplarily, the implementation of the middleware combining asynchronous programming and multi-threaded programming is an efficient solution for handling mixed tasks (both IO-intensive and CPU-intensive operations). For IO-intensive tasks (such as network interfaces), asynchronous programming is adopted because IO operations usually involve waiting for responses from external resources (such as requests), and the asynchronous method can avoid blocking the main thread and make full use of the waiting time to process other tasks. For CPU-intensive tasks (such as audio conversion), multi-threaded programming is adopted because CPU-intensive tasks require a large amount of computing resources, and multi-threading can make full use of the parallel computing power of multi-core CPUs to speed up task processing.
[0103] In the above embodiments, by adopting the asynchronous processing mechanism and multi-threaded processing technology, it is possible to make full use of multi-core CPU resources, improve the concurrent processing capacity and processing efficiency of the system. Different from the traditional synchronous call method, the asynchronous processing mechanism can effectively reduce the response time of each request and avoid the system bottleneck caused by blocking.
[0104] To solve the problems of session interruption or recognition failure that may occur during the multi-node expansion process, the embodiments of the present application combine the weighted consistent hashing algorithm to evenly and only distribute the requests of the speech recognition task to different nodes. Exemplarily, the weighted consistent hashing algorithm is integrated into the API gateway in the form of a plugin, and can dynamically route requests to middleware service nodes with different capabilities. The API gateway includes but is not limited to openresty, apisix, and kong gateway.
[0105] The weighted consistent hashing algorithm, while fully considering the differences in server hardware, ensures that different requests for the same speech recognition task are always routed to the same node for processing, thus maintaining the continuity and consistency of the session. This mechanism significantly improves the stability of the system in the case of large-scale concurrency and avoids the problems of session loss or interruption in traditional load balancing methods.
[0106] The following details how the gateway distributes and obtains target task requests to intermediate service nodes through the weighted consistent hashing algorithm.
[0107] In some possible embodiments, different intermediate service nodes correspond to different preset hash value ranges, and the gateway distributes the obtained target task ID request to the corresponding target middleware service node. Please refer to Figure 3 , including the following steps:
[0108] Step 301: Calculate the target hash value according to the audio ID of the audio to be recognized and the hashing algorithm.
[0109] Exemplarily, the calculation of the target hash value can be to convert input data (audio ID) of any length into a fixed-length value (usually hexadecimal or decimal digits) through a hash function. Among them, the audio ID ensures that different requests for the same audio are routed to the same middleware service node, maintaining session consistency and avoiding task status confusion.
[0110] Step 302: Determine the intermediate service node corresponding to the target hash value range where the target hash value is located as the target intermediate service node.
[0111] Exemplarily, calculate the target hash value corresponding to the audio ID, and find the target hash value range corresponding to the target hash value on the pre-created consistent hash ring. The intermediate service node corresponding to the target hash value range is the target middleware service node. The target hash value range is one of multiple preset hash value ranges.
[0112] Step 303: Distribute the target task ID request to the target intermediate service node.
[0113] Exemplarily, in the embodiments of the present application, a consistent hash ring is pre-constructed to map intermediate service nodes and their virtual nodes to the hash space, forming a ring structure to achieve node distribution and dynamic expansion. The consistent hash ring is a ring-shaped space with a range from 0 to 2^ 32 -1. Each middleware service node occupies a certain range on the hash ring, and the size of the range is proportional to the weight of the node. In the case of no weight, assuming there are n computing nodes, each computing node can be assigned a hash range of 2 32 / n, and the schematic diagram effect is as Figure 4 shown:
[0114] Exemplarily, assume there are 3 middleware service nodes, namely node1, node2, and node3. Among them, the performance of node1 is better, and the performance of node2 and node3 is half of that of node1. The corresponding hash ring is as Figure 4 shown. The hash range divided by node1 is larger, and the subsequent requests assigned to it will also increase accordingly.
[0115] Exemplarily, the different preset hash value ranges corresponding to the different intermediate service nodes are related to the weights of each intermediate service node. The greater the weight, the larger the mapped range on the consistent hash ring.
[0116] Optionally, referring to Figure 5 , the method further includes the following steps:
[0117] Step 501: Map the multiple middleware service nodes onto the consistent hash ring to obtain the initial hash value range corresponding to each middleware service node. The initial hash value ranges corresponding to different middleware service nodes are the same.
[0118] Exemplarily, map the multiple middleware service nodes evenly in the clockwise direction onto the consistent hash ring. Assume there are n computing nodes, and the initial hash value range corresponding to each middleware service node is 2 32 / n.
[0119] Step 502: Obtain the weight of each middleware service node according to the performance of each middleware service node.
[0120] After obtaining the initial hash value range, calculate the weight of each middleware service node, where the better the performance, the greater the weight.
[0121] Exemplarily, the performance of each middleware service node can be reflected by the hardware characteristics of the middleware service node. The hardware characteristics can at least include the following parameters: CPU main frequency (cpu freq ), the support degree of the avx instruction set of the CPU (cpu avx ), memory size (mem gb ), memory main frequency (mem freq ), network bandwidth (network bd ). Then the calculation formula for the weight of each middleware service node is as follows:
[0122] weight = w1 * cpu freg + w2 * cpu avx + w3 * mem gb + w4 * mem freq + w5 * networkbd (1)
[0123] Among them, w1 - w5 represent weights, which can be configured according to manual experience or calculated through machine learning and deep learning methods. The CPU main frequency (cpu freq ) reflects the computing speed of the CPU. The higher the main frequency, the more instructions can be processed per unit time. The degree of support for the avx instruction set of the CPU (cpu avx ) The avx instruction set can improve the performance of the CPU when executing specific vector operations. Nodes with larger memory will be assigned more tasks because they can provide a more stable and efficient storage environment for speech recognition tasks, avoiding performance bottlenecks caused by insufficient memory and ensuring the smooth operation of the system. The memory main frequency (mem freq ) determines the read - write speed of memory data. A high memory main frequency means that data can be read and written to memory by the CPU faster. Sufficient network bandwidth can ensure that audio data is quickly uploaded to the middleware service node and the recognition results are promptly returned to the client.
[0124] It should be noted that the weights of the middleware service nodes do not need to be calculated every time, and only need to be recalculated when adding, deleting, or modifying.
[0125] Step 503: Adjust the ranges of the partial rings corresponding to different middleware service nodes on the consistent hashing ring according to the weights of each middleware service node.
[0126] According to the calculated weights of each middleware service node, adjust the ranges of the partial rings corresponding to different middleware service nodes on the consistent hashing ring weight * 2 32 / n. Among them, weight = [weight1, weight2,..., weightn], and weight is the weight vector corresponding to the weight values of the 1st to the nth middleware service nodes.
[0127] Step 504: Determine the preset hash value ranges of each middleware service node according to the ranges of the partial rings corresponding to each middleware service node.
[0128] After obtaining the ranges of the partial rings corresponding to each middleware service node weight * 2 32 / n, determine the starting hash value of each middleware service node, and determine the preset hash value range of each middleware service node according to the range of the partial ring corresponding to each middleware service node weight * 2 32 / n and the starting hash value.
[0129] Exemplarily, after the gateway assigns the request to the middleware service node, the middleware service node needs to process the content of the request. Then, the audio to be recognized in the request for obtaining the target task ID needs to be obtained in advance. The following details how the middleware service node obtains the audio to be processed in the request for obtaining the target task ID.
[0130] In some possible embodiments, the request for obtaining the target task ID further includes the information type of the target information. The information type includes URL type, base64 type, and binary type. In step 205, the target middleware service node obtains the audio to be recognized according to the target string, as Figure 6 shown, including the following steps:
[0131] Step 601: The target middleware service node obtains the audio to be recognized according to the information type and the target string.
[0132] Step 602: Determine the information type.
[0133] Exemplarily, the information type includes URL type, base64 type, and binary type. The methods for obtaining the audio to be recognized are different for different information types. The following details how to obtain the audio to be recognized for different information types.
[0134] Step 603: The information type is the URL type;
[0135] Step 604: Obtain the audio to be recognized by downloading the audio in the link corresponding to the target string.
[0136] Exemplarily, the target string is a network address URL type pointing to the audio to be recognized. The user provides an online link to the audio file, and it needs to be downloaded from a remote location, for example: http: / / example.com / audio.wav.
[0137] Exemplarily, obtaining the audio to be recognized by downloading the audio in the link corresponding to the target string includes the following steps: Verify whether the target string conforms to the URL format (such as including the http: / / or https: / / prefix). Initiate an http request: Use an http client to download the audio file. Process the response: Verify the status code (such as 200 OK). Read the binary data in the response body as the audio content.
[0138] Step 605: The information type is the base64 type.
[0139] Step 606: Decode the target string by the base64 decoding method to obtain the audio to be recognized.
[0140] Exemplarily, the target string is audio binary data encoded in Base64. To transmit the audio via a text protocol (such as JSON), direct transmission of binary data is avoided to prevent format errors.
[0141] Exemplarily, decoding the target string to obtain the audio to be recognized through the base64 decoding method includes the following steps:
[0142] 1. Verify the Base64 format: Check whether the target string complies with the Base64 encoding rules (such as being a multiple of 4 in length and having a legitimate character set).
[0143] 2. Decoding operation: Restore the Base64 string to the original binary data.
[0144] 3. Verify data validity: Verify whether the decoded data is in a valid audio format (such as a WAV header identifier).
[0145] Step 607: The information type is binary.
[0146] Step 608: Determine the target string as the audio to be recognized.
[0147] Exemplarily, the target string is the original audio binary byte stream. The client has directly read the audio file into memory or transmitted the data via a binary protocol. This method can directly verify whether the byte stream conforms to the audio format (for example, by detecting the file header through libmagic).
[0148] After obtaining the audio to be recognized in the request, the detailed steps for the middleware service node to perform speech recognition on the audio to be recognized and obtain the target recognition result are described below.
[0149] In some possible embodiments, the middleware service node further includes a local database, which is used to store the speech recognition status and recognition results corresponding to each existing task ID. The speech recognition status includes one of in recognition, recognition completed, and error. The target middleware service node performs speech recognition on the audio to be recognized to obtain the target recognition result, as Figure 7 shown, including the following steps:
[0150] Step 701: The target middleware service node adds the target recognition task to the speech recognition queue.
[0151] The target recognition task includes the target task ID and the audio to be recognized.
[0152] Exemplarily, a speech recognition queue is a buffering and scheduling system for managing speech processing tasks. Its core function is to process high-concurrency speech requests in an orderly manner, balance system load, and ensure service stability. When a user submits a speech recognition request, the task is stored in the queue in sequence, and the background worker nodes consume and process it step by step according to strategies (such as first-in-first-out, priority scheduling) to avoid crashes caused by instantaneous traffic overload.
[0153] Step 702: Read the execution result of the speech recognition queue and update the local database.
[0154] Among them, in the case where the execution result does not include the execution result of the target task, update the speech recognition status of the target task ID in the local database to the recognition status; in the case where the execution result includes the execution result of the target task ID, update the speech recognition status of the target task ID to recognition completed and store the execution result, and the execution result is the target recognition result; in the case where the execution result includes the error reason of the target task, update the speech recognition status of the target task ID to error and store the error reason.
[0155] After receiving a request from the client to obtain the target recognition result, the middleware service node obtains the target recognition result. There are various situations, and the result may be in the process of recognition, or recognition error, or recognition completed. In different states, different information is returned to the client. As follows:
[0156] Exemplarily, please refer to Figure 8 , after receiving the request to obtain the target recognition result, the target middleware service node obtains the target recognition result, including:
[0157] Step 801: The target middleware service node searches the local database for the speech recognition status corresponding to the target task ID.
[0158] In a speech recognition system, the middleware service node realizes full-process control through a task status management mechanism. Three states are defined: processing, completed, and failed. Each state corresponds to a clear status code and processing logic. When the service node receives a task query request, it first retrieves the status record of the task ID from the local database. The database adopts a structured storage scheme, including fields such as task ID, status enumeration, result text, error code, etc., and records the task life cycle through a timestamp.
[0159] Step 802: If the speech recognition status is the recognition status.
[0160] Step 803: Return a status identifier indicating that the speech recognition status is in progress.
[0161] Exemplarily, for the in-progress status, the service node returns a 202 Accepted status code and a 1001 service code, prompting the client that the task is being processed. Exemplarily, an asynchronous processing mode is adopted to avoid resource occupation caused by long connections. The client can obtain the final result through a polling mechanism or a Webhook callback.
[0162] Exemplarily, a typical response includes an estimated remaining time field (such as "estimated_time": 30), guiding the client to set a reasonable retry interval. The server implements a timeout fuse mechanism at this stage. If the task processing exceeds a preset threshold (such as 300 seconds), it is automatically marked as a failed state to prevent the system resources from being occupied for a long time.
[0163] Step 804: If the speech recognition status is completed.
[0164] Step 805: Return a status identifier indicating that the speech recognition status is completed and the target recognition result.
[0165] Exemplarily, when the task enters the completed state, the system returns a 200 OK status code and a 2000 service code, and carries the structured recognition result in the response body. The result data includes the recognized text, confidence score, and optional timestamp segmented text, meeting the refined requirements of different scenarios.
[0166] For example, "segments": [{"start": 0.0, "end": 1.5, "text": "Hello"}].
[0167] Exemplarily, the server adopts a hierarchical storage strategy, storing short-term results in a high-performance database (such as Redis) and archiving long-term data to an object storage (such as AWS S3), taking into account both query efficiency and storage cost.
[0168] Step 806: If the speech recognition status is an error.
[0169] Step 807: Return a status identifier indicating that the speech recognition status is an error and the error reason.
[0170] Exemplarily, the error status accurately locates the root cause of the problem through the 5xxx series of service codes.
[0171] Exemplarily, the error code design follows the classification and grading principle: 5001 indicates abnormal input data (such as an invalid audio format), 5002 identifies processing timeouts, 5003 corresponds to internal system failures, and 5004 is used for scenarios such as invalid task IDs.
[0172] Exemplarily, the error response integrates self-explanatory error details, including error type, readable description, and a link to technical documentation, significantly improving the efficiency of problem troubleshooting.
[0173] For example:
[0174] "error_detail":{"type":"invalid_audio_format","doc_url":"https: / / api.example.co m / errors / 5001"}.
[0175] Exemplarily, the server can implement automatic log capture and exception tracing for error tasks, and combine monitoring tools such as Sentry to achieve error clustering analysis.
[0176] The above method provides real-time status feedback, further improving the user experience.
[0177] In some possible embodiments, the flowchart of the middleware service node obtaining the recognition result according to the request sent by the client is as Figure 9 shown.
[0178] Step 901: The middleware service node receives the request sent by the client. The request can be a request to obtain a target task ID or a request to obtain a target result.
[0179] Step 902: Determine whether there is a request in the local database. If there is a request, execute Step 903: Whether there is a target task ID; if there is no target task ID, execute Step 904: Generate a target task ID, and execute Step 905: Store the status corresponding to the target task ID in the local database, execute Step 906: Return the local database to the gateway. If there is a target task ID, execute Step 907: Query the status corresponding to the target task ID from the local database, and then execute Step 908: Whether the target task ID is queried. If not, execute Step 909: Return an error code and the reason for the error. If the target task ID is queried, return the status code corresponding to the target task ID and the target recognition result.
[0180] It should be understood that in Figure 9 the above steps described above, the detailed process of the steps has been described in detail in the above embodiments and will not be repeated here.
[0181] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of these steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0182] In some possible embodiments, an offline speech recognition system is proposed, as Figure 10 shown, the offline speech recognition system includes: a client 1001, a gateway 1002, and multiple middleware service nodes 1003. The gateway establishes communication connections with the client and each of the multiple middleware service nodes respectively, where:
[0183] The client is configured to, when detecting a speech recognition operation, send a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation; at least included in the request for obtaining a target task ID are: the audio ID of the audio to be recognized, and a target string, where the target string is the audio to be recognized or used to indicate the audio to be recognized;
[0184] The gateway is configured to allocate the request for obtaining a target task ID to a corresponding target middleware service node;
[0185] The target middleware service node is configured to, after receiving the request for obtaining a task ID, generate a target task ID, return the target task ID to the client through the gateway, obtain the audio to be recognized according to the target string, and perform speech recognition on the audio to be recognized to obtain a target recognition result;
[0186] The client is configured to, after obtaining the target task ID, send a request for obtaining a target recognition result to the target middleware service node through the gateway. At least included in the request for obtaining a recognition result are: the audio ID of the audio to be recognized and the target task ID;
[0187] The target middleware service node is configured to, after receiving the request for obtaining a target recognition result, obtain the target recognition result, and return the target recognition result to the client through the gateway.
[0188] Another embodiment provides a storage medium for storing a computer program. The computer program includes instructions for implementing the method described in the embodiments of the present application. By installing this computer program on a computer, the computer can execute the corresponding method.
[0189] Another embodiment provides a computer program product that includes computer program code. When the computer program code runs on a computer, it causes the computer to implement the method proposed in the embodiments of the present application. In this way, the user can achieve this by using this computer program product.
[0190] Exemplarily, Figure 11 is a schematic block diagram of an electronic device provided in the embodiments of the present application. The electronic device 1100 may include: a memory 1101 storing executable program code and a processor 1102 coupled to the memory 1101.
[0191] Among them, the processor calls the executable program code stored in the memory and executes any one of the methods disclosed in the embodiments of the present application. Those skilled in the art can understand that Figure 11 the structure of the electronic device shown in
[0192] does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0193] The processor is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory, and by calling the data stored in the memory, the processor executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor may include one or more processing units; preferably, the processor may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor.
[0194] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0195] In the implementation process, each step of the above method may be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0196] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0197] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0198] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces. The indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0199] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, in each embodiment of this application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0201] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0202] The above is only the specific implementation manner of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the embodiments of this application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be subject to the protection scope of the claims.
Claims
1. An off-line speech recognition method, characterized in that: Applied to an offline speech recognition system, the offline speech recognition system comprises: a client, a gateway, and a plurality of middleware service nodes, the gateway respectively establishes a communication connection with the client and each of the plurality of middleware service nodes, the method comprises: When the client detects a voice recognition operation, the client sends a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the voice recognition operation; the request for obtaining a target task ID includes at least: an audio ID of the audio to be recognized, and a target string, where the target string is the audio to be recognized or is used to indicate the audio to be recognized; The gateway distributes the request to obtain the target task ID to the corresponding target middleware service node; After receiving the request to obtain the target task ID, the target middleware service node generates a target task ID, returns the target task ID to the client through the gateway, obtains the audio to be recognized according to the target character string, performs speech recognition on the audio to be recognized, and obtains a target recognition result; After obtaining the target task ID, the client sends a target recognition result acquisition request to the target middleware service node through the gateway, wherein the target recognition result acquisition request at least includes: the audio ID of the audio to be recognized and the target task ID; After receiving the request for obtaining the target recognition result, the target middleware service node obtains the target recognition result and returns the target recognition result to the client through the gateway.
2. The method according to claim 1, characterized in that: Different intermediate service nodes correspond to different preset hash value ranges, and the gateway distributes the request to obtain the target task ID to the corresponding target middleware service node, including: Calculate a target hash value based on the audio ID of the audio to be identified and a hash algorithm; Determine that an intermediate service node corresponding to a target hash value range where the target hash value is located is the target intermediate service node, and the target hash value range is one of a plurality of preset hash value ranges; Allocate the target task ID request to the target intermediate service node.
3. The method according to claim 2, characterized in that The method further comprises: Mapping the multiple middleware service nodes to a consistent hash ring to obtain an initial hash value range corresponding to each middleware service node, where the initial hash value ranges corresponding to different middleware service nodes are the same; According to the performance of each middleware service node, obtaining the weight of each middleware service node, wherein the better the performance, the greater the weight; According to the weight of each middleware service node, the range of the partial ring corresponding to different middleware service nodes on the consistent hash ring is adjusted, and according to the range of the partial ring corresponding to each middleware service node, the preset hash value range of each middleware service node is determined.
4. The method according to claim 1, characterized in that: After receiving the request to obtain the target task ID, the target middleware service node generates a target task ID, including: After receiving the request to obtain the target task ID, the target middleware service node determines whether the request to obtain the task ID includes the task ID. If the request to obtain the task ID does not include the task ID, the target task ID is generated for the request to obtain the task ID.
5. The method according to claim 1, characterized in that The request for obtaining the target task ID also includes the information type of the target information, which includes a URL type, a base64 type, and a binary type. The target middleware service node obtains the audio to be recognized according to the target character string, including: The target middleware service node acquires the audio to be recognized according to the information type and the target character string; Wherein, when the information type is the URL type, the audio to be recognized is obtained by downloading the audio in the link corresponding to the target character string; In the case where the information type is the base64 type, decoding the target character string by a base64 decoding method to obtain the audio to be recognized; In a case where the information type is the binary type, the target character string is determined as the audio to be recognized.
6. The method according to claim 1, characterized in that The middleware service node also includes a local database, and the local database is used to store the speech recognition status and recognition results corresponding to each existing task ID, and the speech recognition status includes one of the recognition in progress, recognition completed and error. The target middleware service node performs speech recognition on the audio to be recognized and obtains the target recognition result, including: The target middleware service node adds a target recognition task to a speech recognition queue, wherein the target recognition task includes the target task ID and the audio to be recognized; Reading the execution result of the speech recognition queue and updating the local database; Among them, when the execution result does not include the execution result of the target task, the speech recognition status of the target task ID in the local database is updated to the recognition status; when the execution result includes the execution result of the target task ID, the speech recognition status of the target task ID is updated to recognition completion and the execution result is stored, and the execution result is the target recognition result; when the execution result includes the error cause of the target task, the speech recognition status of the target task ID is updated to error and the error cause is stored.
7. The method according to claim 1, characterized in that After receiving the request for obtaining the target recognition result, the target middleware service node obtains the target recognition result, including: The target middleware service node searches the local database for a speech recognition state corresponding to the target task ID; If the speech recognition state is the recognition state, returning a state flag indicating that the speech recognition state is the recognition state; If the speech recognition status is recognition completed, return a status identifier indicating that the speech recognition status is recognition completed and the target recognition result; If the speech recognition status is an error, a status identifier indicating that the speech recognition status is an error and the cause of the error are returned.
8. An offline speech recognition system, characterized in that: The offline speech recognition system comprises: a client, a gateway, and a plurality of middleware service nodes, wherein the gateway establishes a communication connection with the client and each of the plurality of middleware service nodes respectively, wherein: The client is configured to, when a speech recognition operation is detected, send a request for obtaining a target task ID to the gateway according to the audio to be recognized corresponding to the speech recognition operation; the request for obtaining the target task ID includes at least: an audio ID of the audio to be recognized, and a target character string, wherein the target character string is the audio to be recognized or is used to indicate the audio to be recognized; The gateway is used to distribute the request to obtain the target task ID to the corresponding target middleware service node; The target middleware service node is used to generate a target task ID after receiving the request to obtain the target task ID, return the target task ID to the client through the gateway, obtain the audio to be recognized according to the target character string, and obtain the target recognition result by performing speech recognition on the audio to be recognized; The client is used to send a target recognition result acquisition request to the target middleware service node through the gateway after acquiring the target task ID, wherein the target recognition result acquisition request at least includes: the audio ID of the audio to be recognized and the target task ID; The target middleware service node is used to obtain the target recognition result after receiving the request to obtain the target recognition result, and return the target recognition result to the client through the gateway.
9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as claimed in any one of claims 1 to 7.