Speech recognition method and system based on AIOT edge device and cloud LLM large model
By analyzing voice command data from edge devices in the Industrial Internet of Things (IIoT), the dynamic optimization module performs personalized adaptation and cloud feedback loops, solving the recognition inconsistency problem caused by differences in edge device hardware and algorithms, and improving the accuracy and adaptability of voice recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-04
- Publication Date
- 2026-04-03
AI Technical Summary
In the Industrial Internet of Things (IIoT), differences in the hardware capabilities and algorithm versions of edge devices lead to inconsistent speech recognition results, affecting cloud data integration and the accurate reproduction of worker intentions, making it difficult to maintain the stability and consistency of recognition performance in complex environments.
By acquiring voice command data from edge devices, analyzing algorithm versions and hardware performance parameters, dynamically selecting optimization modules for personalized adaptation and configuration, applying unified standard protocols at the edge device layer, and combining cloud analysis and feedback loop mechanisms, the recognition deviation control scheme is optimized, and ultimately the recognition accuracy is improved through algorithm updates.
It significantly improves the accuracy and adaptability of voice command recognition in industrial scenarios, ensures efficient and unified voice processing capabilities of edge devices, and reduces recognition deviation through cloud verification.
Smart Images

Figure CN121789677A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a speech recognition method and system based on AIoT edge devices and cloud-based LLM large models. Background Technology
[0002] In the field of Industrial Internet of Things (IIoT), voice recognition technology serves as a crucial bridge connecting people and devices.
[0003] It not only improves production efficiency, but also optimizes the operation process within the factory through intelligent interaction.
[0004] However, with the widespread application of edge devices in industrial scenarios, ensuring the accuracy and adaptability of speech recognition has become a major issue that urgently needs to be addressed.
[0005] Research in this field is directly related to the future development of intelligent manufacturing and has an undeniable value in enhancing the overall system's collaborative capabilities.
[0006] Currently, although many speech recognition methods have been applied in industrial environments, they often struggle to cope with complex real-world scenarios.
[0007] Many solutions overlook the differences between devices during the design process. Since the equipment in a factory may be purchased in batches at different times, its hardware capabilities and software algorithm versions may be different. If the inconsistencies in hardware capabilities or algorithm versions of different devices are not fully considered, the recognition effect will fluctuate greatly in different scenarios.
[0008] This problem is not simply a technical limitation, but rather a lack of comprehensive consideration of the diversity of equipment and the complexity of the environment, which in turn affects the stability of the overall system.
[0009] The deeper technical challenge lies in the contradiction between algorithm consistency and personalized adaptation at the edge device level.
[0010] Due to differences in hardware computing power and algorithm versions among edge devices, the same voice command may produce vastly different results on different devices.
[0011] This contradiction directly leads to difficulties in subsequent cloud-based analysis, because the cloud needs to integrate data from multiple devices, and data inconsistency can interfere with the overall judgment.
[0012] For example, in a factory, when workers give instructions in dialects, some equipment may misrecognize them because the algorithm is not adapted to the specific accent, while other equipment may miss key information in noisy environments due to insufficient noise resistance.
[0013] In particular, for example, when using a multi-device collaborative processing scheme for the same audio, if the same person says a sentence, the sentence is divided into three segments, which are received and processed by devices A, B, and C respectively and then sent to the cloud. The cloud then combines these three segments into a complete sentence. At this time, due to the different dialect recognition capabilities and noise resistance capabilities of the three devices, the final synthesized content will have deviations.
[0014] This discrepancy in the identification results between devices ultimately makes it difficult for the cloud to accurately reconstruct the worker's complete intentions, affecting the accuracy of instruction execution.
[0015] Therefore, how to balance the consistency and personalized adaptation of algorithms at the edge device level, and ensure that the recognition results of the same voice command on different devices are as similar as possible, has become a key problem that this research needs to overcome.
[0016] Solving this problem will directly affect the reliability and efficiency of voice interaction in industrial scenarios. Summary of the Invention
[0017] This invention provides a speech recognition method based on AIoT edge devices and cloud-based LLM large models, mainly including: Acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method, and determine the degree of recognition deviation of the voice command data; Based on the degree of recognition deviation, a matching dynamic tuning module is selected from the preset algorithm library to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. Through the personalized adaptation configuration, a unified standard protocol is applied at the edge device layer. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result. Consistency indicators are extracted from the preliminary speech recognition results of the revised version, and the consistency indicators are transmitted using a cloud analysis interface to determine the overall processing path of the speech content. Based on the comprehensive processing path, algorithm consistency feedback data between edge devices is obtained, and the unified standard protocol is adjusted through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme. Noise environment adaptation parameters are extracted from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description. The enhanced edge device voice processing capabilities are described, and the feedback verification data from cloud analysis is obtained. A consistency verification method is used to confirm that the recognition deviation has been reduced, and the final algorithm version update list is determined. Based on the final algorithm version update list, update packages are distributed from the industrial IoT system to edge devices. If the update is complete, the worker dialect recognition rate is tested to obtain a complete record of the unified standard implementation.
[0018] This invention provides a speech recognition system based on AIoT edge devices and cloud-based LLM large models, mainly comprising: The data acquisition and analysis module is used to acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, and to analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method to determine the degree of recognition deviation of the voice command data. The dynamic optimization module selection module is used to select a matching dynamic optimization module from a preset algorithm library according to the degree of recognition deviation, so as to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. The edge processing module is used to apply a unified standard protocol at the edge device layer through the personalized adaptation configuration. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result. The consistency index extraction and transmission module is used to extract consistency indexes from the preliminary speech recognition results of the revised version, transmit the consistency indexes using a cloud analysis interface, and determine the comprehensive processing path of the overall speech content. The feedback loop adjustment module is used to obtain algorithm consistency feedback data between edge devices according to the comprehensive processing path, and adjust the unified standard protocol through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme. The noise parameter integration module is used to extract noise environment adaptation parameters from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description. The verification and confirmation module is used to obtain the verification data returned from cloud analysis through the enhanced edge device voice processing capabilities, use a consistency check method to confirm that the recognition deviation has been reduced, and determine the final algorithm version update list. The update distribution and testing module is used to distribute update packages from the industrial IoT system to edge devices according to the final algorithm version update list. If the update is completed, the dialect recognition rate of workers is tested to obtain a complete record of the implementation of the unified standard.
[0019] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses an optimized method for recognizing worker voice commands in industrial IoT factories, aiming to solve the recognition deviation problem of edge devices under complex noise environments and dialect differences. By analyzing algorithm version differences and hardware performance parameters, the degree of deviation in voice command data is determined, and an optimization module is dynamically selected based on the deviation to generate personalized adaptation configurations to cope with dialect and noise challenges. This invention applies a unified standard protocol at the edge device layer, activates the dynamic optimization module to correct recognition results, and optimizes the deviation control scheme through cloud analysis and feedback loop mechanisms. Finally, it integrates noise adaptation parameters to enhance voice processing capabilities. Simultaneously, this invention ensures reduced deviation through cloud verification and consistency checks, and distributes algorithm update packages to improve the recognition rate. This invention significantly improves the accuracy and adaptability of voice commands in industrial scenarios, providing edge devices with efficient and unified voice processing capabilities. Attached Figure Description
[0020] The above and other objects, features and advantages of the present invention will become more apparent from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a flowchart of the speech recognition method based on AIoT edge devices and cloud-based LLM large models of the present invention.
[0021] Figure 2 This is a schematic diagram of a speech recognition method based on AIoT edge devices and cloud-based LLM large models according to the present invention.
[0022] Figure 3 This is another schematic diagram of the speech recognition method based on AIoT edge devices and cloud-based LLM large models of the present invention.
[0023] Figure 4 This is a schematic diagram of the structure of a speech recognition system based on AIoT edge devices and cloud-based LLM large models according to the present invention.
[0024] Figure 5 This is a flowchart illustrating the identification deviation analysis and threshold comparison process of this invention.
[0025] Figure 6 This is a flowchart of the selection and personalized configuration of the dynamic optimization module in this invention.
[0026] Figure 7 This is a flowchart illustrating the data interaction process between the edge device and the cloud in this invention.
[0027] Figure 8 This is a flowchart showing the adjustment of the feedback loop mechanism of the present invention.
[0028] Figure 9 This is a graph comparing the speech recognition performance of the present invention.
[0029] Figure 10This is a flowchart of the algorithm version update and deployment process of this invention. Detailed Implementation
[0030] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0031] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0032] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0033] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0034] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0035] like Figure 1As shown, the speech recognition method based on AIoT edge devices and cloud-based LLM large model provided by this invention includes the following steps: First, in step S101, worker voice command data collected from multiple edge devices in an industrial IoT factory is acquired. Algorithm version differences and hardware performance parameters are analyzed using a preset threshold comparison method to determine the degree of recognition deviation in the voice command data. In step S102, based on the degree of recognition deviation, a matching dynamic optimization module is selected from a preset algorithm library to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. In step S103, through the personalized adaptation configuration, a unified standard protocol is applied at the edge device layer. If the degree of recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the preliminary speech recognition result. In step S104, a consistency index is extracted from the corrected version of the preliminary speech recognition result, and the consistency index is transmitted using a cloud analysis interface. The flowchart clearly illustrates the complete technical path from voice acquisition, deviation analysis, dynamic optimization, cloud collaboration to algorithm updates. In step S105, based on the comprehensive processing path, algorithm consistency feedback data between edge devices is obtained. A preset feedback loop mechanism is used to adjust the unified standard protocol, resulting in an optimized recognition deviation control scheme. In step S106, noise environment adaptation parameters are extracted from the optimized recognition deviation control scheme. If the parameters meet personalized adaptation requirements, they are integrated into the dynamic optimization module to obtain an enhanced description of the edge device's voice processing capabilities. In step S107, based on the enhanced edge device voice processing capability description, feedback verification data from cloud analysis is obtained. A consistency verification method is used to confirm the reduction in recognition deviation, and the final algorithm version update list is determined. Finally, in step S108, based on the final algorithm version update list, update packages are distributed from the industrial IoT system to the edge devices. If the update is complete, the worker's dialect recognition rate is tested, resulting in a complete unified standard implementation record. This flowchart clearly demonstrates the complete technical path from voice acquisition, deviation analysis, dynamic optimization, cloud collaboration to algorithm updates, reflecting the technical solution of edge computing and cloud-based large-scale model collaboration.
[0036] like Figure 2As shown, the speech recognition method of this invention adopts a three-layer architecture design, including an edge device layer, a dynamic optimization layer, and a cloud analysis layer. In the edge device layer, multiple edge devices (edge device A 101, edge device B 102, and edge device C 103) are responsible for collecting worker voice commands 104 and performing preliminary speech recognition processing. Each edge device uploads the recognition results to the recognition deviation analysis module 204 of the dynamic optimization layer, which performs deviation analysis on the recognition results from different devices. The core of the dynamic optimization layer is the dynamic optimization module 201, which interacts bidirectionally with the personalized adaptation configuration 202 and the unified standard protocol 203, adjusting and optimizing parameters based on the recognition deviation analysis results. The dynamically optimized data is uploaded to the cloud LLM large model 301 of the cloud analysis layer for in-depth analysis. The cloud LLM large model 301 sends the analysis results to the consistency verification module 302 and the algorithm version management module 303 for verification and version control. The outputs of the consistency verification module 302 and the algorithm version management module 303 converge to the feedback optimization module 304. This module generates optimization parameters and sends the optimization results to the dynamic tuning layer and the edge device layer through the feedback path indicated by the dashed arrow, achieving closed-loop optimization. The entire system, through a hierarchical data flow and feedback mechanism, realizes a complete speech recognition process from edge acquisition to cloud analysis and then to optimization feedback, ensuring continuous improvement in recognition accuracy and system adaptability.
[0037] like Figure 3 As shown, this invention demonstrates a scenario of multi-device collaborative speech processing. A worker's voice command is divided into three segments, received by edge devices A, B, and C, respectively. Each edge device performs preliminary recognition of the received speech segments. Due to potential differences in dialect recognition and noise immunity among the devices, the recognition results will vary. The recognition results from the three devices are uploaded to a cloud processing module. The cloud first performs speech concatenation, integrating the three segments into a complete sentence, and then performs consistency analysis to detect deviations between the recognition results of each device. When a deviation is detected, the cloud sends optimization instructions to the edge devices through a feedback loop mechanism, guiding the devices to adjust recognition parameters and improve subsequent recognition accuracy. After comprehensive processing and analysis by the cloud, a complete speech recognition result is finally output. This multi-device collaborative processing mechanism effectively solves the problem of insufficient recognition capability of a single device. Through unified coordination and feedback optimization by the cloud, it significantly improves the overall recognition accuracy and robustness of the system.
[0038] like Figure 4As shown, the speech recognition system of this invention comprises two main parts: an edge terminal and a cloud terminal, totaling eight functional modules. The edge terminal includes a data acquisition and analysis module 101, a dynamic optimization module selection module 102, an edge processing module 103, and a noise parameter integration module 106. The data acquisition and analysis module 101 is responsible for acquiring speech command data, analyzing algorithm version differences and hardware performance parameters, and determining the degree of recognition deviation. The dynamic optimization module selection module 102 selects a dynamic optimization module from the algorithm library based on the degree of deviation and generates a personalized adaptation configuration. The edge processing module 103 applies a unified standard protocol to activate dynamic optimization and obtain a corrected version. The noise parameter integration module 106 is responsible for extracting noise adaptation parameters and integrating them into the dynamic optimization module. The cloud terminal includes a consistency index extraction and transmission module 104, a feedback loop adjustment module 105, a verification and confirmation module 107, and an update distribution and testing module 108. The consistency index extraction and transmission module 104 extracts the consistency indexes transmitted by the edge processing module 103, transmits them through the cloud interface, and determines the comprehensive processing path; the feedback loop adjustment module 105 obtains consistency feedback data, adjusts the unified standard protocol, obtains the optimization scheme, and feeds it back to the edge end; the verification and confirmation module 107 obtains the returned verification data, confirms the reduction of deviation, and determines the update list; the update distribution and testing module 108 is responsible for distributing the update package, testing the dialect recognition rate, and feeding the test results back to the edge end. The entire system achieves dynamic optimization and continuous improvement of speech recognition through the collaborative work of the edge end and the cloud. The edge processing module 103 transmits the consistency indexes to the cloud module, the cloud module 105 transmits the feedback data back to the edge module 102 via the dotted arrow, the module 107 transmits the noise parameters back to the module 106, and the module 108 transmits the test results back to the module 101, forming a complete closed-loop optimization process.
[0039] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0040] This embodiment of a speech recognition method and system based on AIoT edge devices and cloud-based LLM large models may specifically include: S101. Acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, and analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method to determine the degree of recognition deviation of the voice command data.
[0041] By acquiring worker voice command data from edge devices in an Industrial Internet of Things (IIoT) environment, and processing the data using a preset threshold comparison method, the impact of algorithm version differences and hardware performance parameters on recognition is determined, yielding preliminary deviation assessment results. Based on these preliminary deviation assessment results, the voice command data is processed in layers. If a recognition deviation exceeds a preset threshold, the affected data segments are marked, identifying key data ranges requiring further analysis. The marked key data ranges are then used to classify the voice command data using a Support Vector Machine (SVM) algorithm, determining the performance differences of different algorithm versions under specific hardware performance conditions, resulting in categorized deviation distribution data. This categorized deviation distribution data is used to analyze the fluctuations in hardware performance parameters of edge devices in the IIoT environment. If the fluctuation amplitude exceeds a preset range, the relevant device data is prioritized to identify a high-risk device list. Based on the high-risk device list, the voice command data acquisition process is optimized and adjusted, obtaining the adjusted acquisition frequency and data integrity indicators to determine if they meet preset stability standards. Finally, by comparing the adjusted acquisition frequency and data integrity indicators with historical data records, the long-term trend of recognition deviation is analyzed to determine the final deviation control range. Based on the final deviation control range, a voice command data processing strategy for edge devices in the industrial IoT environment is generated, resulting in an optimized configuration scheme suitable for different hardware performance and algorithm versions.
[0042] like Figure 5 As shown, the identification deviation analysis and threshold comparison process includes the following steps: First, S501 acquires voice command data from edge devices, which comes from multiple edge devices deployed in the industrial IoT environment; second, S502 analyzes algorithm version differences, comparing version A HMM model (accuracy 85.3%) with version B RNN model (accuracy 92.7%), calculating an algorithm version difference of 7.4%; then, S503 analyzes the impact of hardware performance parameters, statistically showing that the recognition accuracy of the high-performance group is 93.1%, while that of the low-performance group is 88.9%, with a hardware performance difference of 4.2%; next, S504 calculates... The overall deviation value is calculated using a weighted calculation method, with an algorithm difference weight of 0.6 and a hardware difference weight of 0.4, resulting in an overall deviation of 7.4% × 0.6 + 4.2% × 0.4 = 6.12%. Next, it is determined whether the deviation exceeds a preset tolerance threshold of 5%. If it does, step S505 marks the data as needing optimization, designating that batch of data as requiring algorithm optimization or hardware upgrades. If it does not exceed the threshold, the process proceeds directly to step S506. Finally, step S506 generates a deviation analysis report, summarizing algorithm differences, hardware differences, the overall deviation value, and optimization suggestions, providing data support for subsequent model iterations and equipment upgrades. This process, through quantitative analysis and threshold comparison mechanisms, achieves accurate assessment and timely early warning of performance deviations in the identification system.
[0043] Specifically, in an Industrial Internet of Things (IIoT) factory, voice command data is first collected from multiple worker locations using voice acquisition modules deployed on edge devices. Assuming each device collects 100 voice segments per minute, a total of 10 devices collect approximately 1.44 million data points daily. This data is stored in WAV format with a sampling rate of 16kHz and transmitted in real-time to a cloud-based analytics platform via the MQTT protocol. Next, a preset threshold comparison method is used to analyze algorithm version differences. Assuming two versions of speech recognition algorithms, A and B, are used: version A is based on a traditional Hidden Markov Model (HMM), and version B uses a deep learning Recurrent Neural Network (RNN) model. By comparing the recognition accuracy of the two versions on the same dataset, version A achieves 85.3%, while version B achieves 92.7%, resulting in a difference of 7.4%. A threshold of 5% is set; exceeding this threshold indicates that version B is significantly superior to version A. Simultaneously, the processing latency for each version is recorded: 0.2 seconds for A and 0.3 seconds for B. This comprehensive evaluation assesses performance and efficiency. Subsequently, the impact of hardware performance parameters on the recognition results was analyzed. Assuming the devices were divided into a high-performance group (CPU clock speed 3.5GHz, memory 16GB) and a low-performance group (CPU clock speed 2.0GHz, memory 8GB), the recognition accuracy of the two groups under the same algorithm B was statistically analyzed. The high-performance group achieved 93.1%, while the low-performance group achieved 88.9%, a difference of 4.2%. A hardware performance impact threshold of 3% was set; exceeding this threshold was considered a significant impact of hardware on recognition. Finally, the degree of recognition deviation of the voice command data was determined. Combining algorithm differences and hardware impact, a comprehensive deviation value was calculated. The algorithm difference weight was set at 0.6, and the hardware difference weight at 0.4. The comprehensive deviation = 7.4% * 0.6 + 4.2% * 0.4 = 6.12%. A deviation tolerance threshold of 5% was set; exceeding this threshold required algorithm optimization or hardware upgrades, generating a deviation analysis report, which was automatically pushed to the factory management system, forming a closed-loop optimization logic. To supplement business relevance, if the deviation exceeds the standard, the system can link with the equipment maintenance module to automatically schedule hardware upgrade tasks, ensuring improved recognition accuracy and maintaining the accuracy of production commands.
[0044] S102. Based on the degree of recognition deviation, select a matching dynamic optimization module from the preset algorithm library to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment.
[0045] An initial audio dataset is constructed by collecting worker voice data and factory background noise data, resulting in a preliminary set of voice and noise samples. Based on this preliminary set, a pre-defined feature extraction tool is used to separate dialect voice features and noise background features, determining the independent distribution range of the two types of features. If there are overlapping regions between the separated dialect voice features and noise background features, a noise reduction module is used to separate the signals in the overlapping regions, obtaining a clean dialect voice signal. For the clean dialect voice signal, a support vector machine algorithm is applied for dialect classification training, and the accuracy of the classification model in recognizing different dialects is assessed. Based on the recognition accuracy of the classification model, noise environment adaptation parameters are dynamically adjusted to generate a preliminary environment adaptation scheme for factory background noise. Using the preliminary environment adaptation scheme, the collected audio dataset is processed a second time to obtain an optimized audio processing result. If the optimized audio processing result does not match the preset recognition accuracy threshold, the parameters of the dynamic tuning module are fine-tuned to determine the final personalized configuration scheme.
[0046] like Figure 6 As shown, the dynamic optimization module selection and personalized configuration process includes the following steps: First, in step S601, the recognition deviation score (e.g., 78 points out of 100) is received; then, the score is evaluated by the diamond judgment node to see if it is greater than 70 points; when the score is greater than 70 points, the system executes steps S602 and S603 in parallel, selecting the LSTM adaptive noise suppression algorithm (based on 100,000 training data points, with an accuracy of 92%) and the HMM dialect speech recognition model (covering 5 dialects, with a recognition rate of 85%) respectively; next, in step S604, the noise suppression parameters are configured, setting the noise suppression threshold to -30dB and filtering out interference signals in the range of 500Hz to 2000Hz; then, in step S605, the dialect priority is configured, increasing the target dialect matching priority by 20%; finally, in step S606, the personalized adaptation configuration is output, increasing the speech recognition rate from the original 60% to over 85%. This process achieves intelligent algorithm selection by judging nodes and fine-tuning through parameterized configuration, thereby significantly improving speech recognition performance in different application scenarios.
[0047] Specifically, when implementing personalized configurations for worker dialect recognition and factory noise environments, the system first identifies the problem by analyzing the degree of deviation in the input speech data. For example, it assumes that dialect pronunciation causes a 35% error rate in the collected speech data, while factory noise interference causes a 40% error rate. The system automatically extracts characteristic parameters of the speech signal, such as spectral distribution and signal-to-noise ratio, calculating that the noise frequency is mainly concentrated in the 500Hz to 2000Hz range. It then combines this with the pitch characteristics of the dialect speech to quantify the deviation and obtain a comprehensive deviation score. For example, a score of 78 (out of 100) indicates that a higher degree of optimization is needed. Next, the system matches a dynamic optimization module from a pre-set algorithm library based on the deviation score. Assuming the library contains various noise reduction algorithms and dialect adaptation models, when the score exceeds 70 points, it automatically selects a deep learning-based adaptive noise suppression algorithm (such as an LSTM network model with a training dataset containing 100,000 noise samples and an accuracy of 92%) and a dialect speech enhancement model (such as an HMM-based dialect speech recognition model covering five major dialects and improving the recognition rate to 85%). Subsequently, the system generates a personalized adaptation configuration. For noisy environments, it automatically adjusts the parameters of the noise reduction algorithm, for example, setting the noise suppression threshold to -30dB to ensure the filtering of interference signals from 500Hz to 2000Hz while preserving the main speech frequency band. For dialect recognition, the system loads an acoustic model specific to the dialect, optimizes the recognition weights, for example, increasing the matching priority of the target dialect by 20%, and combines contextual semantic analysis to reduce ambiguity, ultimately forming an adaptation configuration scheme. The entire process is completed automatically through data analysis and algorithm matching, forming a complete logical chain from deviation identification to configuration generation, ensuring that the voice recognition rate in the factory environment increases from the initial 60% to over 85%, meeting actual business needs.
[0048] S103. Through the personalized adaptation configuration, a unified standard protocol is applied at the edge device layer. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the preliminary speech recognition result.
[0049] Based on the business content and extracted relevant attributes, the following business solution is generated. Focusing on the optimization goal of speech recognition results, and combining attributes such as personalized configuration, adaptation settings, edge devices, device level, unified standards, standard protocols, recognition deviation, deviation degree, preset threshold, dynamic optimization, optimization module, and correction version, the technical process steps are as follows: Raw speech data is collected through the edge device layer and initially formatted according to personalized configuration and adaptation settings to obtain standardized speech input. Based on the standardized speech input, data transmission and preliminary parsing are performed at the device level using unified standards and standard protocols to determine the initial recognition result. For the initial recognition result, the recognition deviation and deviation degree are calculated. If the deviation degree exceeds the preset threshold, subsequent processing is triggered to obtain deviation analysis data. Based on the deviation analysis data, the dynamic optimization function is activated, and the optimization module adjusts the parameters of the initial recognition result to generate an optimized intermediate result. The intermediate result is iteratively processed multiple times through the optimization module, and combined with feedback data on the recognition deviation, a further refined speech recognition output is obtained. Based on the refined speech recognition output, a final corrected version is generated. It is then determined whether the corrected version meets the preset quality standards. If not, the process reverts to the dynamic optimization step for further processing. The final corrected version completes the storage and distribution of the speech recognition results, ensuring synchronous updates of data within the edge device layer and determining the final output.
[0050] Specifically, in the process of applying a unified standard protocol at the edge device layer, let's assume we use the MQTT protocol as the communication standard to ensure low latency and high reliability of data transmission between devices.
[0051] For example, in a smart home scenario, an edge device such as a voice gateway processes 1000 voice data packets per second, each packet being 512 bytes in size. These packets are sent to the cloud analysis module via the MQTT protocol at 5-millisecond intervals, while the system locally caches the most recent 10 minutes of data for dynamic optimization. Next, if the recognition error exceeds a preset threshold—for example, if the word error rate (WER) is calculated to be 15% while the preset threshold is 10%—the system automatically triggers the dynamic optimization module. This module uses a deep learning-based adaptive algorithm (such as an LSTM model) and combines it with 5000 locally cached historical voice data packets to fine-tune the model parameters. The learning rate is adjusted to 0.001, and the number of iterations is set to 50. Analysis shows that the WER decreases to 8% after optimization, achieving the expected result. Subsequently, the dynamic optimization module outputs a corrected version of the initial voice recognition result. For example, the original recognized text "open the living room, etc." is corrected to "turn on the living room lights," and the semantic similarity score of the text before and after the correction is calculated using a semantic analysis algorithm, increasing it from 0.75 to 0.95, ensuring the result better matches the user's intent. To form a tight logical chain, the correction results will be further linked with the smart home control system to automatically generate control commands, send them to the corresponding devices, and complete the light-on operation. At the same time, the deviation data of this optimization will be recorded in the log for reference in subsequent system optimization. The log storage capacity is limited to 1GB, and the oldest record will be automatically overwritten after it is exceeded. Through the above process, a complete technical closed loop is formed from protocol application to result correction and business linkage.
[0052] S104. Extract consistency indicators from the preliminary speech recognition results of the revised version, transmit the consistency indicators using the cloud analysis interface, and determine the comprehensive processing path for the overall speech content.
[0053] Preliminary results are obtained from the input audio data using speech recognition technology. Consistency indicators are extracted from these preliminary results to obtain initial indicator data. Based on the extracted consistency indicator data, a cloud analysis interface is used for transmission processing, and the indicator data is uploaded to the cloud system to ensure data integrity. If the uploaded data integrity meets a preset threshold, the consistency indicators are deeply analyzed using the cloud analysis module to obtain the analyzed feature data. For the analyzed feature data, a pre-established decision tree model is used for classification processing to determine the corresponding speech content category. If the classified speech content category matches the overall speech target category, a content processing path is planned based on the category information to obtain a specific processing path scheme. Using the planned content processing path, the overall speech data is segmented to obtain segmented speech fragment data, and the final processing order is determined. Based on the segmented speech fragment data and the processing order, content optimization operations are performed segment by segment to obtain optimized speech content data.
[0054] Specifically, such as Figure 7 As shown, in the process of processing speech recognition results to extract consistency indicators and determine the comprehensive processing path, the initial speech recognition results are first analyzed by text segmentation. The recognized speech text is divided into segments of 5 seconds each. Assuming a total duration of 60 seconds, 12 segments are obtained. Then, the semantic consistency between segments is calculated using cosine similarity in natural language processing algorithms. Specifically, each segment is converted into a word vector, and the cosine value between each pair of segments is calculated. A threshold of 0.8 is set. If the similarity between any two segments is higher than 0.8, they are considered consistent. The percentage of consistent segments is used as the consistency indicator. Assuming that 30 pairs of consistent segments are calculated, and the total number of pairs is 66, the consistency indicator is 45.45%. Next, this indicator is transmitted to the server in JSON format through a cloud analysis interface. The interface adopts a RESTful style, and the data packet size is controlled within 2KB to ensure transmission efficiency. After receiving the data, the server records the indicator value and generates a timestamp log for subsequent traceability. Ultimately, based on the consistency index, a comprehensive processing path for the overall speech content is determined. If the index is greater than 40%, the system automatically selects the deep semantic analysis path, calling a deep learning model such as BERT for content refinement. Assuming the processing time is 3 seconds, the refined semantic structure is output. If the index is less than 40%, the fast keyword extraction path is selected, using the TF-IDF algorithm to extract high-frequency words. Assuming 5 keywords are extracted, the processing time is 1.5 seconds, generating a brief content summary. The above process forms a closed-loop logic, with data extraction and path selection all completed automatically by the system. Simultaneously, it is linked to the speech content quality assessment business, ensuring that the consistency index is not only used for path selection but also serves as a basis for quality feedback, improving subsequent optimizations.
[0055] S105. Based on the comprehensive processing path, obtain algorithm consistency feedback data between edge devices, and adjust the unified standard protocol through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme.
[0056] By integrating processing paths, algorithm consistency-related data is obtained from edge devices to determine the data interaction status between devices. Based on the obtained data interaction status, differences in algorithm consistency among edge devices are analyzed to determine if there are identification deviation issues. If a deviation exceeds a preset threshold, the deviation data is recorded. For the recorded deviation data, a feedback loop mechanism is used to dynamically adjust the unified standard protocol to obtain the adjusted protocol content. Based on the adjusted protocol content, a control scheme for identification deviation is generated, and the applicable scope and execution conditions of the scheme are determined. Based on the generated control scheme, consistency verification tools are deployed among edge devices to obtain execution feedback data between devices. After obtaining the execution feedback data, if the feedback data indicates that the deviation still exists, the control scheme is recalibrated through a preset loop mechanism to obtain the final optimized scheme. Based on the final optimized scheme, the standard protocol between edge devices is updated, completing the identification deviation control process.
[0057] like Figure 8 As shown, the feedback loop mechanism adjustment process includes the following steps: First, in S801, algorithm consistency feedback data of each device is collected through the edge device communication interface, assuming that 10 edge devices are running the same image recognition algorithm; in S802, for a test dataset containing 1000 images, the recognition accuracy of each edge device is calculated; in S803, the average accuracy of all devices is calculated to be 84.55%, with a standard deviation of 1.12, and the performance deviation between devices is analyzed; in S804, the algorithm parameters are adjusted using a weighted average algorithm, with the weights based on the historical performance of each device, and the recognition threshold is adjusted from 0.75 to 0.78, thereby reducing the false recognition rate by 2.5%; in S805, the adjustment effect is verified through simulation testing. Next, an accuracy benchmark is assessed. If the benchmark is met, the unified standard protocol is updated in S806, reducing the error tolerance range from ±5% to ±3%, and a dynamic adjustment module is added. This module triggers an adaptive parameter update when three consecutive deviations exceed 2%. If the benchmark is not met, the process returns to S804 to readjust the parameters, forming a feedback loop mechanism. Finally, the optimized solution is output in S807. Verification shows that the overall accuracy has improved from 84.55% to 85.2%. This process ensures consistency and reliability of algorithm execution across multiple devices through iterative optimization.
[0058] Specifically, based on the comprehensive processing path, algorithm consistency feedback data is first collected through the communication interface between edge devices. Assuming there are 10 edge devices, each running the same image recognition algorithm, their recognition results on the same test dataset are collected. The dataset contains 1000 images, and the recognition accuracy of each device is recorded. Assuming the accuracy rates of devices 1 to 5 are 85.3%, 86.1%, 84.9%, 85.7%, and 86.0%, and the accuracy rates of devices 6 to 10 are 83.2%, 82.9%, 83.5%, 84.1%, and 83.8%, the calculated average accuracy rate is 84.55%, with a standard deviation of 1.12. The analysis shows that there is a certain deviation between devices, which requires further adjustment. Next, using a pre-defined feedback loop mechanism, the above data is input into the consistency optimization model. A weighted average algorithm is used to adjust the parameters between devices. The weights are based on the historical performance of the devices. Assuming that device 1 has a weight of 0.12 and the weights of the remaining devices are evenly distributed, a new parameter configuration scheme is calculated, adjusting the recognition threshold from 0.75 to 0.78, reducing the false recognition rate by approximately 2.5%. Simulation tests verify that the overall accuracy under the new threshold is improved to 85.2%. Finally, based on the adjusted parameters, a unified standard protocol is optimized to form a recognition deviation control scheme. Specifically, the fault tolerance range in the protocol is reduced from ±5% to ±3%, and a dynamic adjustment module is added in combination with business scenarios (such as intelligent monitoring). If the recognition deviation exceeds 2% for three consecutive times, the parameter adaptive update mechanism is triggered. By analyzing historical data, the deviation trend is predicted, and a deviation control curve is generated, keeping the error within 1.8%, ensuring the stability of the system in different scenarios. Logically, this forms a complete closed loop from data collection to parameter optimization to protocol adjustment.
[0059] S106. Extract noise environment adaptation parameters from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, integrate them into the dynamic tuning module to obtain an enhanced edge device voice processing capability description.
[0060] Step 1: Obtain noise environment-related adaptation parameters from the control scheme. Perform preliminary analysis on the collected environmental audio data to determine the initial range of these adaptation parameters. Step 2: Based on the initial range of adaptation parameters determined in Step 1, use a pre-established classification model to perform personalized adaptation screening of the parameters. If the screened parameters meet preset adaptation conditions, a personalized parameter set is obtained. Step 3: Based on the personalized adaptation parameter set obtained in Step 2, integrate it with the dynamic optimization module. Use a layer-by-layer matching method to determine the degree of fit between the parameter set and the preset rules within the module, obtaining parameter configurations with a high degree of fit. Step 4: Based on the parameter configuration obtained in Step 3, load the speech processing module in real-time on the edge device. Process the input speech signal in segments to obtain an optimized speech processing path. Step 5: Based on the speech processing path obtained in Step 4, enhance the speech processing capabilities of the edge device. Determine the enhanced processing effect by real-time monitoring of key nodes in the processing path. Step Six: If the processing effect determined in Step Five does not reach the preset threshold, the dynamic optimization module is fine-tuned. By analyzing the processing latency and accuracy of the voice signal, the adjusted optimization strategy is obtained. Step Seven: Based on the optimization strategy obtained in Step Six, it is applied to the voice processing flow of the edge device. By continuously updating parameter configurations and processing paths, it is determined whether the final voice processing capability meets business requirements.
[0061] Specifically, in optimizing the recognition deviation control scheme, the noise environment adaptation parameters are first extracted. The input audio signal is then processed using a spectrum analysis algorithm. Assuming a sampling rate of 44.1kHz, the audio signal is decomposed into multiple frequency bands. The characteristic values of the dominant frequency range of environmental noise from 200Hz to 800Hz are extracted, and the signal-to-noise ratio (SNR) is calculated to be 12.5dB. By comparing this with a preset threshold of 15dB, the impact of the current environmental noise on speech recognition is determined. The analysis results indicate that parameter adjustments are necessary. Next, the extracted parameters are matched with personalized adaptation requirements. An adaptive algorithm based on user history is used, combined with the user's voice input data from the past 30 days, to calculate a personalized noise suppression coefficient of 0.75. This coefficient is compared with the standard coefficient of 0.6 to confirm that the parameters meet personalized requirements. Simultaneously, a machine learning model is used to perform cluster analysis on the user's voice features to ensure that the accuracy of the adaptation parameters reaches over 90%. Subsequently, the parameters meeting the requirements were integrated into the dynamic optimization module. Specifically, a real-time feedback mechanism was used, inputting the noise suppression coefficient of 0.75 and the main frequency range data into the module's neural network optimizer. The network weights were adjusted using a gradient descent algorithm with a learning rate of 0.01. After 10 iterations, the model converged, and the speech recognition accuracy improved from 85% to 92%, thereby enhancing the speech processing capabilities of the edge device. To form a robust logical chain and supplement relevant business content, the processed speech data was combined with the low-power mode of the edge device, dynamically adjusting the allocation of computing resources to ensure that power consumption only increased by 5% in noisy environments. By analyzing the balance between power consumption and recognition rate, the optimal resource allocation scheme was derived, further improving the overall performance of the device.
[0062] S107. Based on the enhanced edge device voice processing capability description, obtain the feedback verification data from cloud analysis, use the consistency verification method to confirm the reduction of the recognition deviation, and determine the final algorithm version update list.
[0063] Raw speech signals are acquired through edge devices and pre-established filtering mechanisms to initially clean the signals, resulting in processed speech segments. These processed speech segments are then transmitted to a cloud-based analysis module for deep feature extraction, yielding a feature vector set. If the feature vector set's match with a pre-defined benchmark feature library is below a threshold, a verification process is triggered, retrieving verification data from the cloud. A consistency check method is used to compare the verification data with local recognition results to determine if any recognition deviation exists. If the deviation exceeds a preset range, the algorithm version is iteratively adjusted to determine candidate versions. Based on the test performance of the candidate versions, batch comparative analysis is employed to determine the final algorithm version update list.
[0064] Specifically, in enhancing the voice processing capabilities of edge devices, the process begins by deploying an optimized speech recognition model on the edge devices. This increases the sampling rate of the original speech signal to 16kHz, and uses the Mel-frequency cepstral coefficients (MFCC) feature extraction algorithm combined with a Long Short-Term Memory (LSTM) network for initial recognition, improving the recognition accuracy from 75% to 82.3%. Subsequently, the processed speech feature data from the edge devices is uploaded to the cloud for in-depth analysis. The cloud uses a convolutional neural network (CNN) combined with an attention mechanism to perform secondary modeling of the data, generating feedback verification data. Analysis results show that the speech recognition error rate decreased from 12.5% to 8.7%, and a detailed deviation distribution report is generated. Next, a consistency verification method is used. By comparing the differences between the recognition results from the edge devices and the cloud, the matching rate on 5000 test samples is calculated, yielding a consistency index of 90.2%, confirming a reduction in recognition deviation of approximately 3.8 percentage points. The verification process uses a cosine similarity algorithm to quantify feature vector differences, ensuring the reliability of the results. Finally, based on the verification results and deviation analysis, a list of algorithm version updates was determined. Priority was given to updating the LSTM model parameters on edge devices, increasing the number of hidden layer nodes from 128 to 256. Simultaneously, the kernel size of the cloud-based CNN model was optimized from 3x3 to 5x5, improving its adaptability to complex speech scenarios. The updated model achieved an overall accuracy of 85.6% on the test set. To form a robust logical chain, the system automatically matched the update list with the hardware resource usage of the edge devices, ensuring that the CPU utilization of the updated devices did not exceed 80% and the memory usage was controlled below 60%. Version push and verification were completed through automated deployment tools, forming a closed-loop optimization process.
[0065] S108. Based on the final algorithm version update list, distribute the update package from the industrial IoT system to the edge devices. If the update is completed, test the worker dialect recognition rate to obtain a complete unified standard implementation record.
[0066] The latest update list is retrieved from a pre-established algorithm version database. For target edge devices in the Industrial Internet of Things (IIoT), the content of the update package to be distributed is determined. An automated distribution tool pushes the update package to the edge devices, monitoring the transmission status during the push process. If transmission is interrupted, the distribution request is re-initiated, and confirmation of successful reception of the update package is obtained. The update package is installed using the edge device's built-in update module. If an anomaly is detected during installation, it is rolled back to the previous version to determine if the update is complete. The running status of the updated edge devices is obtained. For the worker dialect recognition module, a pre-set test script is initiated to obtain dialect recognition test results. Based on the comparison and analysis of the test results with a unified standard, if the recognition rate is lower than a preset threshold, a secondary testing process is triggered to determine the final recognition test data. By matching the test data with the unified standard, corresponding implementation records are generated and saved to the central database of the IIoT using structured storage, completing the record generation process.
[0067] like Figure 10 As shown, the algorithm version update and deployment process includes the following steps: First, the cloud platform distributes the update package (S1001), preparing the 50MB algorithm update package for fragmented transmission; then, fragmented transmission is executed (S1002), dividing the update package into 5MB data blocks, transmitting them via the MQTT protocol at a bandwidth of 10MB / s, with the entire transmission expected to be completed within 5 seconds; after transmission, MD5 verification is performed (S1003), verifying data integrity piece by piece. If the verification fails, the failed fragments are retransmitted until all fragments pass verification; after successful verification, the device is automatically restarted (S1004), initiating the loading process for the new algorithm version; after the device restarts, version verification is performed (S1005), checking whether the current algorithm version number is the latest V2. Version 3.1 is used; if the version is inconsistent, it will automatically roll back to the stable version V2.3.0. After the version verification is passed, dialect recognition testing is conducted (S1006), using 1000 sample data to test the recognition effect of 5 dialects, with a processing time of 2 seconds per sample. The system judges whether the recognition rate has reached the threshold standard of 95%. If it has not reached the standard, the model parameters are adjusted, and the learning rate is changed from 0.01 to 0.005 for optimization, and then the test is re-executed. After the recognition rate reaches the standard, an implementation record is generated (S1007), and the detailed data of the update process is stored in a distributed database to ensure traceability. Finally, linear regression prediction analysis is performed (S1008) to predict the future recognition rate trend based on historical data. When the predicted value is lower than 90%, the algorithm optimization process is triggered in advance. The entire update and deployment process ensures the security and reliability of the algorithm version update through multiple verification mechanisms and automatic rollback function, while performance testing and predictive analysis ensure the continuous and stable operation of the system.
[0068] Specifically, in the process of distributing update packages to edge devices and completing related tests and recordings within the industrial IoT system, the latest algorithm version update package is first distributed to the edge devices via the IoT cloud platform. Assuming the update package size is 50MB, containing an optimized speech recognition algorithm model, it uses a fragmented transmission technology, with each fragment being 5MB in size, transmitted via the MQTT protocol at a bandwidth of 10MB per second, and is expected to complete distribution within 5 seconds. The system automatically verifies the MD5 value of each fragment to ensure data integrity; if verification fails, it automatically retransmits until 100% success. Next, after the update is complete, the system automatically triggers the edge device restart mechanism. After restarting, a preset script verifies whether the algorithm version is the latest version V2.3.1. If the version matches, the system enters the testing phase; otherwise, it rolls back to V2.3.0 and records the error log. Subsequently, when testing the dialect recognition rate, the system used a built-in test dataset containing 1000 dialect speech samples covering five major dialects. Each sample was 2 seconds long and used a deep learning-based LSTM model for recognition. The recognition rate was calculated by dividing the number of correctly recognized samples by the total number of samples. Assuming the test result was 920 correctly recognized samples, the recognition rate was 92%. If the recognition rate fell below the preset threshold of 95%, the system automatically adjusted the model parameters, such as reducing the learning rate from 0.01 to 0.005, and retrained until the standard was met. Finally, the system generated a unified standard implementation record, including update time, version number, and recognition rate data, stored in a distributed database using a timestamp format of YYYY-MM-DD HH:MM:SS. The total number of records was expected to be 10,000. By analyzing the recognition rate trend in the records, the system used a linear regression algorithm to predict the recognition rate change for the next week. If the predicted value was below 90%, an optimization task was triggered in advance to ensure system stability. Through the above process, the data at each stage was closely linked, forming a complete technical chain from update to testing to recording.
[0069] like Figure 9As shown, the method of this invention exhibits significant advantages over the traditional HMM method in multiple performance metrics. Regarding speech recognition accuracy, after six iterations of training, the method of this invention achieves a recognition accuracy of 93%, an improvement of 8 percentage points compared to the traditional method's 85%. In terms of dialect recognition accuracy, the method of this invention ultimately reaches 88%, an improvement of 18 percentage points compared to the traditional method's 70%, fully demonstrating the superiority of this invention in processing dialect speech. Regarding noise environment recognition accuracy, the method of this invention reaches 85%, an improvement of 20 percentage points compared to the traditional method's 65%, indicating that this invention has stronger resistance to noise interference. The curve trend shows that the method of this invention exhibits a faster performance improvement rate in the early stages of iteration, and the performance improvement trend remains stable as the number of iterations increases, while the performance improvement of the traditional method gradually flattens out. The above experimental data verifies the effectiveness and advancement of the technical solution of this invention, especially the more significant performance improvement in complex scenarios such as dialect recognition and noise environment recognition.
[0070] This invention also provides a speech recognition system based on AIoT edge devices and cloud-based LLM large models, mainly comprising: The data acquisition and analysis module is used to acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, and to analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method to determine the degree of recognition deviation of the voice command data. The dynamic optimization module selection module is used to select a matching dynamic optimization module from a preset algorithm library according to the degree of recognition deviation, so as to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. The edge processing module is used to apply a unified standard protocol at the edge device layer through the personalized adaptation configuration. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result. The consistency index extraction and transmission module is used to extract consistency indexes from the preliminary speech recognition results of the revised version, transmit the consistency indexes using a cloud analysis interface, and determine the comprehensive processing path of the overall speech content. The feedback loop adjustment module is used to obtain algorithm consistency feedback data between edge devices according to the comprehensive processing path, and adjust the unified standard protocol through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme. The noise parameter integration module is used to extract noise environment adaptation parameters from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description. The verification and confirmation module is used to obtain the verification data returned from cloud analysis through the enhanced edge device voice processing capabilities, use a consistency check method to confirm that the recognition deviation has been reduced, and determine the final algorithm version update list. The update distribution and testing module is used to distribute update packages from the industrial IoT system to edge devices according to the final algorithm version update list. If the update is completed, the dialect recognition rate of workers is tested to obtain a complete record of the implementation of the unified standard.
[0071] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A speech recognition method based on AIoT edge devices and cloud-based LLM large models, characterized in that, The method includes: Acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method, and determine the degree of recognition deviation of the voice command data; Based on the degree of recognition deviation, a matching dynamic tuning module is selected from the preset algorithm library to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. Through the personalized adaptation configuration, a unified standard protocol is applied at the edge device layer. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result. Consistency indicators are extracted from the preliminary speech recognition results of the revised version, and the consistency indicators are transmitted using a cloud analysis interface to determine the overall processing path of the speech content. Based on the comprehensive processing path, algorithm consistency feedback data between edge devices is obtained, and the unified standard protocol is adjusted through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme. Noise environment adaptation parameters are extracted from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description. The enhanced edge device voice processing capabilities are described, and the feedback verification data from cloud analysis is obtained. A consistency verification method is used to confirm that the recognition deviation has been reduced, and the final algorithm version update list is determined. Based on the final algorithm version update list, update packages are distributed from the industrial IoT system to edge devices. If the update is complete, the worker dialect recognition rate is tested to obtain a complete record of the unified standard implementation.
2. The method according to claim 1, characterized in that, The process of acquiring worker voice command data collected from multiple edge devices in an industrial IoT factory, analyzing algorithm version differences and hardware device performance parameters using a preset threshold comparison method, and determining the degree of recognition deviation of the voice command data includes: The system acquires worker voice command data collected by industrial IoT edge devices, compares and analyzes algorithm version differences and hardware performance parameters by setting preset thresholds to obtain preliminary deviation assessment results; marks data segments with deviations exceeding the threshold, classifies voice data using support vector machine algorithm, determines the deviation distribution of different algorithm-hardware combinations; and generates optimized configuration schemes adapted to different hardware and algorithms.
3. The method according to claim 1, characterized in that, The step of selecting a matching dynamic optimization module from a preset algorithm library based on the degree of recognition deviation to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment includes: A dataset is constructed by collecting worker voices and factory noise. Dialect voice and noise background features are extracted and their distribution range is determined. Noise reduction is performed on overlapping feature regions to obtain clean dialect signals. A dialect classification model is trained using a support vector machine algorithm. The noise adaptation parameters are dynamically adjusted based on the model's recognition accuracy to generate an environment adaptation scheme. If the optimized audio processing results do not reach the accuracy threshold, the parameters of the dynamic tuning module are fine-tuned to determine the personalized adaptation configuration.
4. The method according to claim 1, characterized in that, The personalized adaptation configuration applies a unified standard protocol at the edge device layer. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result, including: The edge device layer formats and processes the raw voice data, and applies a unified standard protocol to parse it to obtain the initial recognition result; it calculates the recognition deviation, and if it exceeds the preset threshold, it activates the dynamic optimization module to iteratively adjust the initial result to generate intermediate results; it refines the intermediate result to obtain a corrected version, and after confirming that it meets the quality standards, it completes the storage and distribution of the corrected version.
5. The method according to claim 1, characterized in that, The step of extracting consistency indicators from the preliminary speech recognition results of the revised version, transmitting the consistency indicators using a cloud analysis interface, and determining the overall processing path for the speech content includes: Consistency indicators are extracted from the revised recognition results, uploaded and verified for data integrity via the cloud analysis interface; the cloud analysis module parses the indicator data and classifies the speech content using a decision tree model; if the classification result is consistent with the target, a comprehensive processing path is planned and the speech data is processed in segments; the speech content is optimized segment by segment according to the processing order to obtain the final optimized data.
6. The method according to claim 1, characterized in that, The step of obtaining algorithm consistency feedback data between edge devices based on the integrated processing path, and adjusting the unified standard protocol through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme includes: Based on the comprehensive processing path, algorithm consistency data between edge devices is obtained, differences are analyzed and deviations exceeding the threshold are recorded; the unified standard protocol is adjusted through a feedback loop mechanism to generate a deviation control scheme; consistency verification tools are deployed on edge devices, and if the feedback shows that the deviation still exists, the scheme is calibrated a second time; the device standard protocol is updated to complete the deviation control.
7. The method according to claim 1, characterized in that, The noise environment adaptation parameters are extracted from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description, including: Noise environment adaptation parameters are extracted from the optimized control scheme, and the parameter range is determined by combining environmental audio data; a parameter set that meets the personalized requirements is selected through a classification model, integrated with the dynamic tuning module, and matched with preset rules; parameter configuration is loaded on the edge device and the voice processing path is optimized, and the enhanced processing effect is monitored; if the effect does not meet the standard, the module is fine-tuned, and finally it is confirmed whether the voice processing capability meets the business requirements.
8. The method according to claim 1, characterized in that, The description of obtaining cloud-based verification data through the enhanced edge device voice processing capabilities, using a consistency check method to confirm the reduction in recognition deviation, and determining the final algorithm version update list includes: Edge devices collect raw speech, filter and clean it, and transmit it to the cloud to extract deep feature vector sets. If the matching degree between the vector set and the benchmark library is insufficient, the cloud is triggered to send back verification data. The verification data is compared with the local results through consistency verification to determine the recognition deviation. The algorithm version is iteratively adjusted to generate candidate solutions, and the final updated list is determined after batch comparison and testing.
9. A speech recognition system based on AIoT edge devices and cloud-based LLM large models, characterized in that, The system includes: The data acquisition and analysis module is used to acquire worker voice command data collected from multiple edge devices in an industrial IoT factory, and to analyze algorithm version differences and hardware device performance parameters through a preset threshold comparison method to determine the degree of recognition deviation of the voice command data. The dynamic optimization module selection module is used to select a matching dynamic optimization module from a preset algorithm library according to the degree of recognition deviation, so as to obtain a personalized adaptation configuration for worker dialect recognition and factory noise environment. The edge processing module is used to apply a unified standard protocol at the edge device layer through the personalized adaptation configuration. If the recognition deviation exceeds a preset threshold, the dynamic optimization module is activated to obtain a corrected version of the initial speech recognition result. The consistency index extraction and transmission module is used to extract consistency indexes from the preliminary speech recognition results of the revised version, transmit the consistency indexes using a cloud analysis interface, and determine the comprehensive processing path of the overall speech content. The feedback loop adjustment module is used to obtain algorithm consistency feedback data between edge devices according to the comprehensive processing path, and adjust the unified standard protocol through a preset feedback loop mechanism to obtain an optimized identification deviation control scheme. The noise parameter integration module is used to extract noise environment adaptation parameters from the optimized recognition deviation control scheme. If the parameters meet the personalized adaptation requirements, they are integrated into the dynamic tuning module to obtain an enhanced edge device voice processing capability description. The verification and confirmation module is used to obtain the verification data returned from cloud analysis through the enhanced edge device voice processing capabilities, use a consistency check method to confirm that the recognition deviation has been reduced, and determine the final algorithm version update list. The update distribution and testing module is used to distribute update packages from the industrial IoT system to edge devices according to the final algorithm version update list. If the update is completed, the dialect recognition rate of workers is tested to obtain a complete record of the implementation of the unified standard.
Citation Information
Patent Citations
Speech recognition processing method based on accent, electronic device and storage medium
CN109074804A
Voice control method, central control device and storage medium
CN109074808A
Equipment calibration method and device, storage medium and electronic device
CN113314098A
Conference voice recognition method and device, electronic equipment and storage medium
CN120472883A
Work AR intelligent auxiliary management system based on voice AI interaction driving
CN121096334A