Voice recognition software interaction implementation method based on workshop production operation scene

By constructing an intelligent production speech recognition application, low-pass filtering and Fbank method are used for signal preprocessing. Combined with BiGRU and attention mechanism speech recognition algorithms, a full-duplex communication channel is established, and software operation components are built. This solves the problem of poor adaptability of speech recognition technology in the production workshop, realizes efficient speech operation mapping in high-noise environments, and improves production efficiency and system usability.

CN120998183AInactive Publication Date: 2025-11-21INSPUR HONGQI (SHANDONG) DIGITAL TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202511517216.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing speech recognition technology is difficult to adapt to the high noise, multiple terms and strong real-time requirements in the production workshop environment. It cannot achieve closed-loop control of the entire process of voice commands and software operations, resulting in high operating threshold, low efficiency and high error rate.

Method used

To build an intelligent production speech recognition application, low-pass filtering and Fbank method are used for signal preprocessing. A speech recognition algorithm with BiGRU and attention mechanism is combined to establish a full-duplex communication channel and build software operation components to realize real-time transmission of speech recognition results and software operation mapping.

Benefits of technology

Improving speech recognition accuracy in high-noise environments, achieving deep integration of speech and software systems, lowering the operational threshold, and increasing production efficiency are particularly suitable for special groups with weak skill foundations or operational difficulties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998183A_ABST
    Figure CN120998183A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition software interaction implementation method based on a workshop production operation scene, and relates to the technical field of man-machine interaction and software engineering technologies, and the method comprises the steps: constructing an intelligent production voice recognition application; constructing a full duplex communication channel between the intelligent production voice recognition application and an intelligent production software service; constructing a software operation component; recognizing voice data in a workshop production operation scene through the intelligent production voice recognition application to obtain a voice recognition result; transmitting the voice recognition result to the software operation component through the full duplex communication channel; and analyzing the intention of the user based on the voice recognition result through the software operation component and executing corresponding software operation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of human-computer interaction and software engineering technology, and particularly relates to a voice recognition software interaction implementation method based on a workshop production operation scene. BACKGROUND

[0002] With the continuous deepening of the transformation of manufacturing industry to digitization and intelligentization, various production management software systems are generally used in the workshop production environment to improve operational efficiency. However, in the actual operation process, traditional production personnel are not familiar with the operation of computers and software systems, and there are problems such as high operation threshold, high learning cost, and slow response speed. At the same time, special operators (such as people with physical disabilities) have obvious obstacles when facing traditional interaction methods such as keyboards and mice, resulting in low efficiency and high error rate in production task execution.

[0003] Although existing voice recognition technology has been applied in some scenarios, such as customer service telephone semantic classification and vehicle-mounted voice interaction, it is usually oriented to general fields or specific closed scenes, and lacks the ability to adapt to special environments such as high noise, multiple terminologies, and strong real-time performance in production workshops. In addition, existing voice systems mainly focus on semantic understanding and content feedback, and cannot be deeply combined with the operation logic of production software, so they cannot realize the whole-process closed-loop control of "voice instruction - system operation - execution feedback", and thus cannot realize "barrier-free interaction" in the production site in the true sense. SUMMARY

[0004] The application provides a voice recognition software interaction implementation method based on a workshop production operation scene to solve one of the above technical problems.

[0005] The technical solution adopted by the application is as follows: The application embodiment provides a voice recognition software interaction implementation method based on a workshop production operation scene, which comprises: constructing an intelligent production voice recognition application; constructing a full-duplex communication channel between the intelligent production voice recognition application and an intelligent production software service; constructing a software operation component; recognizing voice data in the workshop production operation scene through the intelligent production voice recognition application to obtain a voice recognition result; transmitting the voice recognition result to the software operation component through the full-duplex communication channel; analyzing a user's intention based on the voice recognition result through the software operation component and performing corresponding software operation.

[0006] According to one embodiment of the application, the construction of the intelligent production voice recognition application comprises: Constructing a voice data preprocessing module and an algorithm recognition public component; Constructing a BiGRU voice recognition algorithm module component with a fusion attention mechanism; Constructing a multi-dimensional high-quality corpus; Constructing a model training and optimization service.

[0007] According to an embodiment of the present application, the voice data preprocessing module and the algorithm recognition public component include: Low-pass filtering is used to remove background noise and other non-target sound interference; The Fbank method is used to extract parameters reflecting the essential characteristics of the sound from the preprocessed digital signal.

[0008] According to an embodiment of the present application, the BiGRU voice recognition algorithm module component with a fusion attention mechanism includes: A bidirectional gated recurrent unit is used as the basis of the voice recognition model; The attention mechanism is fused to strengthen the model's attention to important parts of the input voice sequence; Output the corresponding text representation.

[0009] According to an embodiment of the present application, the construction of a multi-dimensional high-quality corpus includes: Intelligent production industry standard kernel corpus is constructed, covering key terms and instructions of typical business processes in the intelligent production industry; Intelligent production enterprise extensible corpus is constructed, combining with enterprise individualized needs, and incorporating enterprise-specific process parameters, quality standards, and production terminology; The corpus is deeply optimized, including multi-channel data collection, data cleaning, and deep labeling.

[0010] According to an embodiment of the present application, the construction of a full-duplex communication channel between the intelligent production voice recognition application and the intelligent production software service includes: Intelligent production voice recognition application is deployed on each independent PC workstation; Obtain the unique identification code of the device workstation terminal and send it as a signal source to the software server; The server verifies the legality of the identification code and establishes a persistent two-way communication connection.

[0011] According to an embodiment of the present application, the software operation component includes: Build an instruction parsing and intent mapping and instruction confirmation module; Build an operation execution module.

[0012] According to an embodiment of the present application, the instruction parsing and intent mapping and instruction confirmation module includes: The multi-dimensional automatic recognition is based on a classification algorithm, and involves a feature extraction judgment index system and a multi-dimensional integrated algorithm. The difference analysis is performed on various scene system software operation items, operation modes, system software mapping interfaces and system software confirmation instructions of a specific industry.

[0013] According to an embodiment of the present application, the building operation execution module comprises: The voice recognition result after the user's intention is parsed is received through the full-duplex communication channel; The specific button or operation item to be executed in the system interface is mapped based on the instruction parsing and intention mapping component; The software operation is executed, and log recording and voice prompting are performed according to the execution result.

[0014] The second aspect embodiment of the present application provides an electronic device, comprising a memory, a processor and a program stored in the memory and executable on the processor, wherein the processor executes the program to realize the steps in the method as described.

[0015] Due to the adoption of the above technical solutions, the present application has the following beneficial effects: The present application constructs an intelligent production voice recognition application, realizes accurate recognition and text conversion of voice signals in a workshop environment, improves the recognition accuracy and robustness in a high-noise, multi-terminology scene, realizes real-time, stable and bidirectional data interaction between the voice recognition end and the software service end through the construction of a full-duplex communication channel, ensures low delay and high reliability of instruction transmission, parses the voice recognition result into a specific software operation intention and maps it to the actual operation item (such as button clicking, form submission, etc.) in the system interface through the construction of a software operation component, and completes the closed-loop control of "voice-intention-operation-feedback". Overall, the voice recognition and the software system are deeply integrated, the operation threshold of the production personnel on the software system is significantly reduced, the operation efficiency and system usability are improved, and it is particularly suitable for special groups with weak skills or operation obstacles, enhances system inclusiveness, and reduces enterprise training and operation cost. BRIEF DESCRIPTION OF DRAWINGS

[0016] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a voice recognition software interaction implementation method based on a workshop production operation scene is provided for the embodiments of the present application; Figure 2 A structural schematic diagram of an electronic device is provided for the embodiments of the present application.

[0017] Reference Signs: 810, processor; 820, communication interface; 830, memory; 840, communication bus. DETAILED DESCRIPTION

[0018] In order to more clearly illustrate the overall concept of the present application, the following detailed description will be made in conjunction with the accompanying drawings.

[0019] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details and other implementations can be employed. In other instances, well-known methods, procedures, components, and circuits have not been described in detail as not to unnecessarily obscure aspects of the present application. Embodiments of the present application can be implemented in various ways without departing from the spirit or scope of the present application. Embodiments of the present application and features of the embodiments can be combined with each other, if not in conflict.

[0020] In the present application, unless specifically stated and limited otherwise, the first feature is "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.

[0021] Embodiment 1 As shown in Figure 1 A voice recognition software interaction implementation method based on a workshop production operation scene, comprising: Constructing an intelligent production voice recognition application.

[0022] As described above, constructing an intelligent production voice recognition application means creating a special software system specifically for understanding and processing voice commands in a workshop production environment. This application is not a general voice recognition tool, but a customized design from signal processing, algorithm model to data foundation for the particularity of the industrial scene, aiming to achieve high-precision and high-robust voice recognition in a high-noise environment.

[0023] Specifically, the construction of the application is a system engineering. First, at the signal input end, it purifies the collected original audio signal through a voice data preprocessing module, such as using a low-pass filtering algorithm to suppress the persistent mechanical noise and other background interference in the workshop, and using methods such as Fbank (Filter Bank) to extract acoustic feature parameters that can reflect the essential characteristics of the voice, converting the original, mixed waveform signal into a pure and key information-rich feature vector suitable for computer deep analysis. Subsequently, the processed feature data is sent to the core recognition algorithm module. This module uses a bidirectional gated recurrent unit (BiGRU) designed for sequence data as the basic network structure, allowing it to combine the context information of the voice signal for comprehensive judgment; further, by integrating the attention mechanism, the model can autonomously and dynamically focus more computing "attention" on the critical segments or phonemes in the voice stream that are crucial to the recognition result, such as the starting part of a command keyword, thereby effectively improving the recognition accuracy of the key command. The final output of this module is the standardized text representation corresponding to the voice.

[0024] For example, on a noisy assembly line, an operator says the command "start reporting work". Its voice will be collected by the workstation microphone, at this time the signal is mixed with the running sound of the conveyor belt and the operation sound of the tool. The preprocessing module will first work to filter out most of the low-frequency mechanical noise and extract the key frequency band features representing human voice. Subsequently, the BiGRU model integrated with the attention mechanism will process these feature sequences, which may automatically identify that the two syllables "start" and "work" have a decisive role in the current context, thereby giving them higher weights, and finally accurately output the text "start reporting work", although the intermediate syllables may have been disturbed by the environment. In order to train and optimize such a special model, a high-quality corpus must be constructed. This includes an industry kernel corpus covering standard processes such as "start work", "complete work", "quality inspection", "call material", etc., as well as an enterprise-specific corpus that can be extended according to the specific needs of the enterprise, containing enterprise-specific terms such as "rough turning speed", "weld detection", etc. Finally, through model training and optimization services (such as using knowledge distillation, model pruning, etc.), the model size is compressed while ensuring high accuracy, allowing it to be efficiently deployed on the computing resources of the workshop workstation computer.

[0025] It should be noted that in specific implementation scenarios, in addition to low-pass filtering, more advanced speech enhancement algorithms such as spectral subtraction, Wiener filtering, etc. can also be introduced at the preprocessing level, or special suppression modules can be designed for specific types of impact noise (such as hammering sound).

[0026] In specific implementation scenarios, on the basis of the above scheme, the basic network structure at the algorithm model level can be replaced by a long short-term memory network (LSTM), a Transformer, or other new sequence modeling architectures that may appear in the future. The attention mechanism can be implemented as self-attention (Self-Attention) or multi-head attention (Multi-Head Attention).

[0027] In specific implementation scenarios, on the basis of the above scheme, at the corpus construction level, the source can be expanded to include multi-modal data such as oral records, training recordings, and operation video soundtracks. Data augmentation (Data Augmentation) techniques can be used to simulate speech samples in different distances and different noise environments to further improve the diversity and robustness of the corpus.

[0028] In specific implementation scenarios, on the basis of the above scheme, at the model optimization level, in addition to knowledge distillation and pruning, parameter quantization, neural architecture search (NAS), and other techniques can be used to achieve the best balance between accuracy and efficiency to adapt to different deployment environments such as cloud servers and edge computing devices.

[0029] A full-duplex communication channel is established between the intelligent production voice recognition application and the intelligent production software service.

[0030] As mentioned above, the full-duplex communication channel between the intelligent production voice recognition application and the intelligent production software service refers to the establishment of a persistent communication link between the voice recognition application of the front-end workstation and the intelligent production software service of the back-end server that can simultaneously and independently transmit data in both directions. This channel is the core hub of the real-time interactive loop from voice command recognition to execution and then to operation feedback. It ensures that the voice recognition results can be delivered to the software service side in real time and reliably, and the status or confirmation information after software execution can be returned to the workstation side in real time, thereby providing immediate feedback to the user.

[0031] Specifically, the construction of this communication channel starts with the identity authentication and link initialization of each individual PC workstation where the voice recognition application is deployed. The system acquires the unique hardware identification information of the workstation device, such as the MAC address or a hardware fingerprint generated by multiple hardware features, and sends this identification code as a "signal source" to the intelligent production software server during communication initialization. The server side acts as a "pre-filter" and verifies the legality of the identification code in its pre-set list of legal devices. After verification, instead of using the traditional "request-response" short connection, a persistent, bidirectional communication connection based on protocols such as WebSocket is established. This means that once the connection is established, the voice recognition application can actively push the recognized text instructions at any time, and the software service can also actively send operation confirmation, execution results or system notifications to the designated workstation at any time, without repeatedly establishing and disconnecting the connection, greatly reducing communication delay and ensuring the real-time nature of the interaction.

[0032] For example, in a specific workshop scenario, the operator speaks the instruction "submit current quality inspection documents" at workstation A03. The voice recognition application on this workstation first converts the voice into text "submit current quality inspection documents". Then, the application sends this text instruction along with the unique hardware MAC address (such as 00-1B-44-11-3A-B7) of workstation A03 to the background server in real time through the established full-duplex communication channel. After receiving the data packet, the server first checks whether the MAC address is an authorized production device, and if it is correct, it forwards the instruction text to the subsequent software operation component for processing. When the software system successfully performs the document submission operation, the processing result (such as "submission successful") will be sent back to workstation A03 by the server through the same communication channel. After receiving this feedback, the application on the workstation side immediately broadcasts "documents have been submitted" to the operator through the speech synthesis module, thus completing a complete, bidirectional interaction loop.

[0033] It should be noted that in a specific implementation scenario, on the basis of the above scheme, the unique identification code is not limited to the MAC address, and can also be extended to the hardware fingerprint generated by the CPU serial number, hard disk serial number or combination thereof, or other software tokens capable of uniquely identifying terminal devices. The communication protocol is not limited to WebSocket, and can also be extended to HTTP / 2 Server Push, WebRTC data channel or any other network protocol capable of supporting full-duplex communication. In terms of security verification, in addition to simple identification code whitelist verification, more advanced security mechanisms such as digital signature, dynamic token or two-way certificate authentication based on SSL / TLS can also be introduced to ensure the confidentiality and integrity of the communication link. In addition, the management strategy of the channel can also be extended, for example, to include a heartbeat keep-alive mechanism to detect connection health and an automatic reconnection mechanism to attempt to recover when the connection is unexpectedly interrupted, thereby ensuring high availability of the communication channel in a harsh industrial network environment.

[0034] Building a software operation component.

[0035] As described above, building a software operation component refers to creating a function module located inside the intelligent production software service, which is specifically responsible for parsing the voice text instruction recognized by the front end into an explicit user operation intent, and finally driving the software system to perform the corresponding specific operation. The component is the "brain" and "execution arm" that implements the conversion from "voice" to "operation", and its core role is to understand the instruction semantics and accurately map it to the visual elements of the software interface or the back-end business logic, thereby completing a complete closed loop from instruction issuance to result feedback.

[0036] Specifically, the construction of the software operation component mainly includes two core modules: instruction analysis and intention mapping module, and operation execution module. First, the instruction analysis and intention mapping module receives the speech recognition text from the full-duplex communication channel, and internally analyzes the text based on the preset classification algorithm, feature extraction judgment index system and multi-dimensional integrated algorithm. This module can differentially identify the operation items (such as triggering a "business button", generating a "material requisition form" or "submitting a form") contained in the text, the operation method (such as simulating a "button click event" or triggering a complex "operation item process"), and the corresponding system software mapping interface (such as calling an "automation script interface", accessing a "system native API" or a pre-configured "flexible configuration interface"). On this basis, the module may also trigger an instruction confirmation mechanism, such as confirming "whether to confirm the submission of the current work report data?" through voice before executing high-risk operations, and executing after the user confirms again. Subsequently, the operation execution module starts working, which calls the mapping interface according to the resolved explicit intention, actually manipulates the software interface, such as automatically positioning and clicking the "submit" button, or writing document data to the database. After execution is completed, the module will capture the execution result and record the log, and generate feedback information (such as "work report success"), which is returned to the workstation through the full-duplex communication channel, and finally informs the user in the form of voice broadcast, thus forming a complete closed loop of "instruction, analysis, mapping, execution, feedback".

[0037] For example, when the operator says "apply for quality inspection for order A1001", the speech recognition application converts it into text and sends it to the software operation component through the communication channel. The instruction analysis and intention mapping module first performs word segmentation and semantic analysis on the text, and identifies that the core operation item is "apply for quality inspection" and the key business parameter is "order A1001". Then, it is mapped in the system configuration that the instruction corresponds to the "new quality inspection form" button under the "quality inspection management" module, and a "flexible configuration interface" needs to be called to pre-fill the information of order A1001. Before formal execution, the module may ask "a quality inspection form will be created for order A1001, please confirm" through voice. After getting the voice reply of "confirmation" from the user, the operation execution module is triggered, which first generates a quality inspection form based on the information of A1001 through the interface in the background, and then automatically triggers the operation of clicking the "new quality inspection form" button on the graphical interface, so that the pre-filled form pops up on the screen. After the operation is successful, the system sends the feedback "quality inspection form has been generated" to the workstation through the communication channel, and finally the user hears the voice prompt, so as to complete the complex operation without manually searching for the order and clicking multiple buttons.

[0038] It should be noted that in a specific implementation scenario, on the basis of the above scheme, the instruction analysis and intention mapping manner is not limited to rule-based feature extraction and classification algorithm, and can be extended to utilize a natural language processing model to perform semantic similarity calculation or based on a deep learning model to perform end-to-end intention classification. The mapping interface is not limited to an automated script or a system API, and can be extended to any technical means capable of driving software to perform automated operations, including but not limited to simulating keyboard and mouse messages, operating a UI automated test framework, or triggering a microservice through a remote procedure call. The operation confirmation mechanism can be flexibly configured according to the operation risk level, for example, a key operation is forced to be confirmed, and a regular operation can be set to skip confirmation. The feedback mechanism of the operation execution module can also be extended and is not limited to voice prompts, and can simultaneously or alternatively include interface visual special effects (such as highlighting the operated button), a text prompt box pop-up, or triggering a physical signal (such as an indicator light) in linkage with a workshop execution system, and the like, to adapt to different workshop interaction needs.

[0039] The voice data in the workshop production operation scene is recognized by the intelligent production voice recognition application to obtain a voice recognition result.

[0040] As described above, the voice data in the workshop production operation scene is recognized by the intelligent production voice recognition application to obtain a voice recognition result, which refers to a process of converting the original audio signal collected by the special voice recognition application deployed on the workshop workstation machine and containing the operation instruction into a standardized text instruction that can be understood and processed by subsequent software components through a series of special signal processing and intelligent analysis. This process is not simply a general voice-to-text conversion, but a special information extraction process deeply integrated with the characteristics of the workshop environment, industry terms, and business logic.

[0041] Specifically, the recognition process starts with the capture of raw speech data. The microphone on the workstation picks up the voice command issued by the operator, which is usually mixed with inherent background noise in the workshop, such as equipment operation sound, tool impact sound, or environmental human voice. Then, the data is sent to the pre-built intelligent production speech recognition application. Inside the application, the audio is first purified by the speech data preprocessing module, for example, using a low-pass filter designed for industrial low-frequency noise to reduce noise, and using feature extraction methods such as Fbank to extract acoustic feature parameters that can represent the essential characteristics of the speech from the purified audio, thereby converting the original, mixed analog signal into a pure, structured digital feature sequence. Then, the feature sequence is input into the core recognition algorithm model (such as a BiGRU-based model) that integrates attention mechanism. The model uses its bidirectional sequence modeling capability to make comprehensive judgments in combination with the context information of the speech, and dynamically focuses on key syllables or words in the command through the attention mechanism, and finally outputs the corresponding text sequence with high accuracy, i.e. the speech recognition result, laying the foundation for subsequent intent analysis and software interaction.

[0042] For example, on a noisy engine assembly line, an operator holding a tool cannot operate the keyboard and mouse, so he says the command: "Record serial number AB123XY and submit to the next station". The microphone on the workstation captures this voice, but the recording also contains the "click" sound of the pneumatic wrench and the "hum" sound of the conveyor belt. The intelligent production speech recognition application starts the processing flow: the preprocessing module first suppresses most of the stable background noise and extracts the key spectral features representing human voice; then the core recognition model analyzes these features, and its attention mechanism may pay special attention to the keywords "record", "serial number", "AB123XY", "submit" which have high weight in the context of the workshop, although the middle part of the speech may be briefly disturbed, the model can still accurately infer and output the complete text recognition result: "Record serial number AB123XY and submit to the next station". This accurate text result, rather than the original noisy audio, will be used to drive subsequent software operations.

[0043] It should be noted that in a specific implementation scenario, on the basis of the above scheme, the preprocessing means is not limited to low-pass filtering and Fbank feature extraction, and can be extended to more advanced purification and feature extraction schemes such as spectral subtraction, Wiener filtering, and speech enhancement models based on deep learning. The core recognition model is not limited to the combination of BiGRU and attention mechanism, and its basic network structure can be equivalent to replace long short-term memory network, Transformer or other neural network architectures specialized for sequence modeling. In addition, the recognition process can be personalized for specific workshops, specific workstations or specific operators, and the recognition rate of professional terms can be improved by loading corresponding industry corpus and enterprise extended corpus. The output form of the recognition result can also be extended and is not limited to pure text, and can include timestamp, speaker identification, confidence and other metadata for subsequent modules to make more complex logical judgments and error handling. The entire recognition process can be performed in real time or in batch processing for recorded voice data to adapt to different application requirements.

[0044] The voice recognition result is transmitted to the software operation component through the full-duplex communication channel.

[0045] As described above, transmitting the voice recognition result to the software operation component through the full-duplex communication channel means that the text instructions obtained by successfully converting the front-end intelligent production voice recognition application are sent in the form of data packets to the back-end software operation component responsible for analysis and execution in real time, reliably and in order through the established, persistent and bidirectional communication link. This step is the bridge connecting the two core functions of "voice recognition" and "software operation", and the key is to use the low-latency, high-reliability and bidirectional interaction characteristics of the full-duplex channel to ensure the immediacy and context coherence of instruction transmission, providing a data foundation for subsequent precise intent analysis and closed-loop interaction.

[0046] Specifically, the transmission process begins with the voice recognition application generating a structured recognition result data packet locally. This data packet contains not only the core instruction text (such as "submit report"), but also key metadata such as the unique identification code of the workstation machine (such as the MAC address) and the timestamp, etc. Subsequently, the application does not need to wait or re-establish a connection, but directly "pushes" this data packet to the remote intelligent production software server through the already persistent full-duplex communication channel (such as a connection based on the WebSocket protocol). After the server-side communication listening service receives the data packet, it first parses it and verifies its legitimacy, and then according to the predetermined data routing rules, accurately distributes the core instruction text and its metadata in the packet to the input interface of the "software operation component" responsible for it. This process is a one-way data flow, but thanks to the characteristics of the full-duplex channel, the transmission behavior and the confirmation or feedback information returned by the operation component are logically parallel and do not interfere with each other, thus achieving efficient data flow at the system level.

[0047] For example, on a laser cutting workstation in a sheet metal workshop, after the operator finishes cutting a piece of sheet metal, he says the instruction "current task completed, perform special inspection". The voice recognition application on the workstation machine accurately recognizes it as the text "current task completed, perform special inspection". Immediately, the application encapsulates the text, the workstation machine ID "WS-Laser-05", and the current time into a data packet, and sends it to the central server through the already stable WebSocket channel. After the server unpacks it, it confirms that "WS-Laser-05" is a legal device, and then accurately routes "current task completed, perform special inspection" to the software operation component responsible for processing the production process. Almost at the same time as the instruction is delivered, the workstation machine may have received the voice feedback "instruction received, processing" sent by the server through the same channel in the opposite direction, while the operation component begins to analyze the intent and execute the operation in parallel.

[0048] It should be noted that in a specific implementation scenario, on the basis of the above scheme, the transmitted data content can be expanded to include rich context data such as voice recognition confidence, voiceprint recognition information, and environmental noise level, in addition to pure text instructions, to enable the software operation component to make more intelligent decisions. The transmission protocol and packaging format can also be diversified, and is not limited to WebSocket and specific JSON structure, but can be expanded to use other communication protocols suitable for Internet of Things or microservice architecture, such as MQTT, gRPC, and efficient data serialization formats such as Protocol Buffers. The data routing strategy can be flexibly configured according to business logic, for example, different types of instructions (such as "report work" and "quality inspection") can be distributed to different business microservice components for processing according to keywords in the instructions. In addition, the transmission process can introduce a priority queue mechanism to give higher transmission and processing priority to urgent instructions (such as "emergency stop"), and can combine data encryption and integrity verification techniques to ensure the security and reliability of production instructions during transmission.

[0049] The software operation component analyzes the user's intention based on the voice recognition result and performs corresponding software operations.

[0050] As described above, the software operation component analyzes the user's intention based on the voice recognition result and performs corresponding software operations, which means that after receiving the voice recognition text transmitted via the full-duplex communication channel, the software operation component performs deep semantic understanding and structured analysis to determine the final operation target that the user wants the software system to perform, and then automatically calls the corresponding system resources or interfaces to drive the software to complete specific business actions from interface interaction to data update, ultimately forming an intelligent closed loop of "voice instruction-intention understanding-system execution".

[0051] Specifically, the process is completed by the instruction parsing and intent mapping module and the operation execution module within the software operation component. First, the instruction parsing and intent mapping module performs natural language parsing on the received voice recognition text, which extracts key operation verbs, business objects and parameters from the text based on pre-set classification algorithms, feature extraction rules and integrated analysis models. For example, it identifies the core operation items in the instruction (such as "generate", "submit", "query"), the operation targets (such as "material requisition", "quality inspection report") and the associated business entities (such as order number, material code). Then, the module matches the parsed semantic elements with the pre-configured "intent-operation" mapping rules, thereby converting the ambiguous user instruction into one or more explicit and executable software operation commands, and determining the specific system interface to be called, such as the automation script interface, the system application programming interface or the pre-defined business logic interface through flexible configuration. Subsequently, the operation execution module takes over the process, which calls the determined interface according to the mapping result, actually manipulates the software interface elements (such as simulating clicking buttons, filling in data in input boxes) or directly triggers the backend business logic (such as updating database status, calling service functions). After execution, the module captures the operation result (success or failure) and performs log recording, and generates corresponding feedback information, which is returned to the front end via the same full-duplex channel, and finally informs the user in the form of voice broadcast, thereby realizing end-to-end automation from user instruction issuance to system response feedback.

[0052] For example, when the software operation component receives the voice recognition result "report 50 pieces of work order P2024052001", the instruction parsing module first performs word segmentation and semantic analysis, identifies that the core intent is "report work", the operation object is "work order P2024052001", and the key parameter is the quantity "50 pieces". Then, it matches in the mapping rule library and determines that the intent corresponds to the "work time and production reporting" function in the "production execution system", which needs to call a specific "flexible configuration interface" to execute, which can automatically navigate to the work reporting interface, locate to work order P2024052001, and fill in "50" in the quantity field. The operation execution module immediately calls this interface to drive the software to automatically complete a series of interface operations and data submission. After successful execution, the module records the log "work order P2024052001 reports 50 pieces of work successfully", and generates feedback information. The user will then hear the voice prompt "report work has been completed", so as not to manually find the work order and input data, greatly improving the operation efficiency and accuracy.

[0053] It should be noted that in a specific implementation scenario, on the basis of the above scheme, the intent analysis method is not limited to rule-based feature extraction and classification, but can be extended to semantic similarity calculation using a natural language processing model, or end-to-end intent classification and slot filling based on a deep learning model. The establishment and management of the "intent-operation" mapping relationship can be extended to a flexible mapping platform supporting graphical drag-and-drop configuration, allowing administrators to flexibly adjust the correspondence between voice instructions and software operations according to changes in enterprise business processes, without the need to modify program code. The interface type called by the operation execution is not limited to automation scripts and system APIs, but can be extended to any technical means that can drive software to perform automated operations, including but not limited to simulating keyboard and mouse messages, operating UI automation test frameworks, triggering microservices through remote procedure calls, or directly interacting with databases. In addition, the execution feedback mechanism can also be extended and is not limited to voice broadcast, but can also include interface visual changes (such as operation button highlighting), system notification pop-ups, or linkage with physical devices such as workshop board systems and indicator lights, thereby forming multi-dimensional interactive feedback to adapt to complex and diverse workshop production environments.

[0054] According to an embodiment of the present application, the construction of the intelligent production voice recognition application comprises: constructing a voice data preprocessing module and an algorithm recognition public component; constructing a BiGRU voice recognition algorithm module component with a fusion attention mechanism; constructing a multi-dimensional high-quality corpus; constructing a model training and optimization service.

[0055] As described above, constructing a voice data preprocessing module and an algorithm recognition public component refers to creating a general functional unit for processing raw voice signals. This module first performs noise reduction processing on the collected workshop environment voice, using a low-pass filtering method to eliminate background noise interference such as equipment operation. Subsequently, the Fbank feature extraction method is used to extract acoustic feature parameters reflecting the essential characteristics of sound from the filtered voice signal, converting the analog voice signal into a digital feature sequence suitable for computer analysis.

[0056] Constructing a BiGRU voice recognition algorithm module component with a fusion attention mechanism refers to establishing a core voice recognition model. This component uses a bidirectional gated recurrent unit as the basic network structure, which can consider both the preceding and following context information of the voice sequence. By introducing an attention mechanism, the model can dynamically adjust the attention weight for different parts of the input sequence, focusing on key voice segments, and finally outputting the corresponding text recognition result.

[0057] Building a multi-dimensional high-quality corpus refers to establishing a special language database for intelligent production scenarios. First, an industry standard core corpus is built, which includes standard terminologies and operation instructions for typical production links such as start-up, report work, and quality inspection. On this basis, an enterprise extensible corpus is built, which includes personalized content such as enterprise-specific process parameters and quality standards. Through multi-modal data collection and deep labeling optimization, the corpus is ensured to cover different production environments and speech characteristics.

[0058] Building model training and optimization services refers to establishing a model continuous improvement mechanism. First, a teacher model is trained using labeled corpus, and a lightweight student model is generated through knowledge distillation technology. Redundant parameters are removed using model pruning algorithm, and model storage requirements are reduced through parameter quantization compression technology. Finally, the optimized model is fine-tuned and performance evaluated to ensure its accuracy and computational efficiency meet the requirements of production environment.

[0059] According to an embodiment of the present application, the speech data preprocessing module and algorithm recognition public components include: Low-pass filtering is used to remove background noise and other non-target sound interference; Fbank method is used to extract parameters reflecting the essential characteristics of sound from preprocessed digital signals.

[0060] As mentioned above, low-pass filtering is used to remove background noise and other non-target sound interference, which refers to setting an electronic filter with a cutoff frequency lower than the main frequency range of human voice, filtering out the low-frequency continuous noise generated by common mechanical equipment in the workshop environment, and retaining the high-frequency components containing speech information, to achieve preliminary purification of the original speech signal.

[0061] Using Fbank method to extract parameters reflecting the essential characteristics of sound from preprocessed digital signals refers to simulating human auditory characteristics, passing the noise-reduced speech signal through a set of mel-scale distributed triangular band-pass filters, calculating the log energy of each filter output, and forming a mel-frequency cepstral coefficient feature vector that can represent the key features of the speech spectrum, providing standardized input features for subsequent speech recognition algorithms.

[0062] According to an embodiment of the present application, the BiGRU speech recognition algorithm module component with fusion attention mechanism includes: A bidirectional gated recurrent unit is used as the basis for the speech recognition model; The attention mechanism is fused to enhance the model's attention to important parts of the input speech sequence; The corresponding text representation is output.

[0063] As described above, the speech recognition model using the bidirectional gated recurrent unit as the basis refers to constructing a neural network structure composed of a forward GRU layer and a backward GRU layer. The forward GRU processes the speech feature sequence in chronological order, and the backward GRU processes the same sequence in reverse chronological order. The hidden state of the two directions is combined to obtain a sequence representation containing complete context information.

[0064] Fusing the attention mechanism to strengthen the model's attention to important parts of the input speech sequence refers to assigning different weight coefficients to each time step of the encoder output in the decoding stage. This mechanism calculates the attention probability distribution, enabling the model to automatically focus on the most relevant speech frames at the current decoding time, highlighting the contribution of key speech features to the recognition result.

[0065] Outputting the corresponding text representation refers to inputting the attention-weighted encoding features into a fully connected layer and a Softmax classifier, calculating the probability distribution of each character corresponding to each time step, and finally converting the probability sequence into the final text recognition result through the connectionist temporal classification decoding strategy.

[0066] According to an embodiment of the present application, the construction of the multi-dimensional high-quality corpus includes: Constructing an intelligent production industry standard kernel corpus covering key terms and instructions for typical business processes in the intelligent production industry; Constructing an intelligent production enterprise extensible corpus, incorporating enterprise-specific process parameters, quality standards, and production terminology based on enterprise individual needs; Optimizing the corpus in depth, including multi-channel data collection, data cleaning, and deep labeling.

[0067] As described above, constructing an intelligent production industry standard kernel corpus refers to systematically collecting and organizing general business process-related terminology in the field of intelligent manufacturing. This corpus comprehensively includes standard operating instructions from production preparation to final inspection, including core business scenarios such as starting work, reporting work, quality inspection, and material distribution, forming a standardized expression of industry-wide basic language resources.

[0068] Constructing an intelligent production enterprise extensible corpus refers to expanding the industry standard corpus based on the specific production process characteristics of a particular enterprise. This corpus focuses on collecting unique device parameters, process indicators, product specifications, and other proprietary terminology of the enterprise, such as the processing parameter range of specific processes, product precision requirements, and special inspection standards, ensuring that the corpus is highly matched with the actual production needs of the enterprise.

[0069] The deep optimization of the corpus refers to improving the quality of the corpus through a systematic data processing procedure. First, raw voice data including field recording, dictation recording, etc. is obtained through multi-device and multi-environment data collection; then, data cleaning is performed to remove duplicate, invalid and error data entries; finally, text transcription, intent classification and entity labeling are performed on the valid corpus according to professional labeling specifications to form a structured training data set.

[0070] According to an embodiment of the present application, the building of the full-duplex communication channel between the intelligent production voice recognition application and the intelligent production software service comprises: deploying the intelligent production voice recognition application on each independent PC workstation; obtaining a unique identification code of the device workstation terminal and sending it as a signal source to the software server; establishing a persistent two-way communication connection after the server end verifies the legality of the identification code.

[0071] As described above, the deployment of the intelligent production voice recognition application on each independent PC workstation refers to installing and running a dedicated voice recognition client program on the computer devices at each production workstation in the workshop, so that it has local voice collection, processing and communication capabilities.

[0072] Obtaining a unique identification code of the device workstation terminal and sending it as a signal source to the software server refers to reading the specific hardware feature information of each workstation, including the MAC address or a unique fingerprint code generated by hardware configuration, through a system interface, and sending these identification information to the background server for device authentication in the communication initialization stage.

[0073] After the server end verifies the legality of the identification code, a persistent two-way communication connection is established, which refers to that after the server receives the identification code, it compares and verifies it with the pre-registered device whitelist. After authentication, a two-way data channel in an active state is established between the client and the server based on the WebSocket communication protocol, supporting real-time parallel data transmission and reception.

[0074] According to an embodiment of the present application, the software operation component comprises: building an instruction parsing and intent mapping and instruction confirmation module; building an operation execution module.

[0075] As described above, building an instruction parsing and intent mapping and instruction confirmation module refers to establishing a functional unit that can analyze voice recognition text and convert it into system operation instructions. This module identifies the operation object, action type and business parameters in the text through feature extraction and classification algorithms, maps natural language instructions to specific software function interfaces. At the same time, a confirmation mechanism is set up to generate a secondary confirmation request for key operations, ensuring the accuracy of instruction execution.

[0076] The building operation execution module refers to creating a functional unit for driving the software system to complete the actual operation. The module receives the parsed and confirmed operation instruction, realizes interface element control and business logic triggering by calling the corresponding system API interface or executing the predefined automation script. After execution, operation log is automatically recorded, and execution state feedback information is generated and returned to the user end.

[0077] According to an embodiment of the present application, the building instruction parsing and intention mapping and instruction confirmation module comprises: Based on the classification algorithm, multi-dimensional automatic identification is performed, involving a feature extraction and judgment index system and a multi-dimensional integrated algorithm. Differential analysis of various scene system software operation items, operation methods, system software mapping interfaces and system software confirmation instructions of specific industries.

[0078] As described above, multi-dimensional automatic identification based on the classification algorithm refers to using the text classification technology in natural language processing to automatically identify the key elements of the operation instruction from the voice recognition text by establishing a feature extraction and judgment index system. This process uses a multi-dimensional integrated algorithm to comprehensively analyze the lexical features, grammatical structures and semantic information of the instruction text, and realizes accurate judgment of the operation intention.

[0079] Differential analysis of various scene system software operation items, operation methods, system software mapping interfaces and system software confirmation instructions of specific industries refers to according to the characteristics of intelligent production industries, respectively processing the software operation requirements under different business scenarios. Specifically, it includes identifying the characteristics of various system software operation items, distinguishing the technical requirements of different operation methods, establishing the corresponding relationship of system software mapping interface, and configuring the trigger conditions of system software confirmation instruction, forming a complete instruction parsing and execution rule system.

[0080] According to an embodiment of the present application, the building operation execution module comprises: Receiving the voice recognition result after parsing the user's intention through the full-duplex communication channel; Mapping the specific button or operation item to be executed in the system interface based on the instruction parsing and intention mapping component; Executing software operation and recording log and voice prompt according to the execution result.

[0081] As described above, receiving the voice recognition result after parsing the user's intention through the full-duplex communication channel refers to that the module continuously monitors the established two-way communication link to obtain the structured operation instruction data processed by the instruction parsing and intention mapping module in real time.

[0082] The instruction analysis and intention mapping component maps to a specific button or operation item to be executed in the system interface, which means that the received structured operation instruction is converted into an operation command for a specific control in the graphical user interface according to a pre-established mapping relationship configuration, including but not limited to button clicking, menu selection, data entry, and other interface interaction actions.

[0083] Performing software operations and logging and voice prompting according to the execution result means that after completing the interface interaction operation, the system automatically captures the operation execution state information, writes the success or failure result into the system log database, generates the corresponding operation feedback prompt through the voice synthesis technology, and returns to the user end for broadcast through the full-duplex communication channel.

[0084] The second aspect embodiment of the application provides an electronic device, including a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the program to realize the method in any embodiment of the first aspect.

[0085] Figure 2 An example of an entity structure diagram of an electronic device is shown in Figure 2 As shown, the electronic device can include a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can invoke the logical instructions in the memory 830 to execute the method in any embodiment of the first aspect, which includes: Building an intelligent production voice recognition application; Building a full-duplex communication channel between the intelligent production voice recognition application and the intelligent production software service; Building a software operation component; Recognizing voice data in a production operation scene in the workshop through the intelligent production voice recognition application to obtain a voice recognition result; Transmitting the voice recognition result to the software operation component through the full-duplex communication channel; Analyzing user intention and performing corresponding software operation through the software operation component based on the voice recognition result.

[0086] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or parts of the present application that essentially contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0087] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the method provided by the above-mentioned methods, and the method comprises: building an intelligent production voice recognition application; building a full-duplex communication channel between the intelligent production voice recognition application and an intelligent production software service; building a software operation component; recognizing voice data in a production operation scene of a workshop through the intelligent production voice recognition application to obtain a voice recognition result; transmitting the voice recognition result to the software operation component through the full-duplex communication channel; analyzing a user intent based on the voice recognition result and performing a corresponding software operation through the software operation component.

[0088] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method provided by the above-mentioned methods, and the method comprises: building an intelligent production voice recognition application; building a full-duplex communication channel between the intelligent production voice recognition application and an intelligent production software service; building a software operation component; recognizing voice data in a production operation scene of a workshop through the intelligent production voice recognition application to obtain a voice recognition result; transmitting the voice recognition result to the software operation component through the full-duplex communication channel; analyzing a user intent based on the voice recognition result and performing a corresponding software operation through the software operation component.

[0089] Example 2 This embodiment describes a specific implementation of voice recognition software interaction in a smart manufacturing workshop. This method realizes a complete closed loop from voice command input to software operation execution by building a dedicated voice recognition application, establishing a full-duplex communication channel, and developing a software operation component. The following describes specific scenarios and technical details.

[0090] In an automobile parts assembly workshop, multiple PC workstations are deployed at each node of the assembly line. When performing assembly tasks, operators need to frequently use production management software for reporting work, quality inspection, and material application operations. Traditional keyboard and mouse operations are inefficient and prone to errors. This embodiment improves operation efficiency through voice interaction. For example, the operator says "submit a report for order A1001, quantity 50 pieces", and the system automatically parses the command and drives the software to complete the corresponding operation.

[0091] The intelligent production voice recognition application is specifically designed for high-noise environments in the workshop, and its construction includes the following steps: Voice data preprocessing module and algorithm recognition public component: After the raw voice signal is collected by the microphone, it is first processed by a low-pass filter. The cutoff frequency of the low-pass filter is set to 8 kHz to filter out common mechanical noise in the workshop (such as conveyor belt operation sound, frequency mostly below 500 Hz). The filtered signal is converted into a digital sequence, and then the Fbank (Filter Bank) method is used to extract acoustic features. Specifically, Fbank processes the voice spectrum through a set of 40 mel-scale triangular filters, calculates the log energy of each filter output, and forms a 40-dimensional mel-frequency cepstral coefficient (MFCC) feature vector. This process simulates human auditory characteristics, highlights key frequency bands in speech, and ignores irrelevant noise.

[0092] BiGRU voice recognition algorithm module component with attention mechanism: The voice recognition model is built based on a bidirectional gated recurrent unit (BiGRU) and integrates an attention mechanism to improve recognition accuracy. BiGRU model structure: BiGRU is composed of a forward GRU layer and a backward GRU layer. The forward GRU processes the input feature sequence in chronological order (where is the Fbank feature at time t), outputting a forward hidden state sequence ; the backward GRU processes in reverse order, outputting a backward hidden state sequence . The complete hidden state at each time step is , capturing the context information of the voice.

[0093] Attention mechanism: In the decoding phase, the attention mechanism assigns weights to the encoder outputs. Specifically, for decoder time step i, the attention energy is computed where is the previous state of the decoder, is the encoder hidden state, , and are trainable parameters. The attention weights are computed by softmax: . The weighted context vector is , which is used to focus on key speech segments (e.g., instruction verbs).

[0094] Output text representation: Finally, the hidden state sequence is mapped to a text sequence, e.g., converting speech to "submit work order A1001 quantity 50", by connecting a time-distributed classification (CTC) output layer.

[0095] Multi-dimensional high-quality corpus: The corpus is divided into two layers: Industry standard kernel corpus: covering typical business processes such as work reporting, quality inspection, material taking, etc., and including standard terms such as "work reporting", "quality inspection passed", "material shortage", etc., a total of 10,000 annotated sentences.

[0096] Enterprise expandable corpus: based on enterprise-specific processes, such as adding terms like "torque calibration value 25Nm" and "welding seam detection standard ISO5817", a total of 5,000 sentences.

[0097] Corpus depth optimization is achieved through multi-modal data collection, including workshop environment recording, video sound accompaniment, etc. The collected data is cleaned (removing duplicates and invalid segments), and annotated with tags and entity parameters by professionals (e.g., "work reporting 50 pieces" is annotated as operation type "work reporting" and entity "quantity = 50").

[0098] Model training and optimization services: Use the annotated corpus to train the teacher model (based on the DeepSpeech2 architecture), and guide the student model (lightweight BiGRU) training through knowledge distillation. The distillation loss function is where is the cross-entropy loss of the student model, is the KL divergence of the teacher and student output distributions. After training, iterative pruning is used to remove the 10% connections with the smallest absolute weight values in BiGRU, and 8-bit parameter quantization is used to compress the model size, reducing the model size by 70% while maintaining an accuracy of over 95%.

[0099] Build a full-duplex communication channel: Full-duplex communication channel based on WebSocket protocol to realize real-time bidirectional data transmission: Deployment and identification: Deploy the speech recognition application on each workstation and obtain the device unique identification code (such as MAC address "00-1B-44-11-3A-B7"). When the application starts, send the identification code to the server as a signal source.

[0100] Verification and connection: The server verifies whether the identification code exists in the registered whitelist. After verification, a persistent WebSocket connection is established to support parallel data transmission (such as speech recognition result upload and operation feedback download). The channel sets a heartbeat packet (every 30 seconds) to detect the connection state, and automatically reconnects in case of abnormality.

[0101] Build software operation components: Software operation components are integrated into production management software, including the following modules: Instruction parsing and intent mapping and instruction confirmation module: This module performs multi-dimensional recognition based on classification algorithms (such as support vector machine SVM). First, extract features (such as bag-of-words model and dependency syntax features) from the speech recognition text to form a feature vector. Then, through multi-dimensional integration algorithm (such as random forest), classify operation type, object and parameter. For example, for "submit report order A1001 quantity 50", the operation item is "report order submission", the operation method is "button click", the mapping interface is "report API", and the confirmation instruction is "please confirm submission report". For high-risk operations (such as "delete record"), the module triggers the voice confirmation process.

[0102] Operation execution module: The module receives the parsed instruction through the full-duplex channel (such as JSON format: {"action": "submit_report", "order": "A1001", "quantity": 50}). According to the intent mapping result, call the corresponding system interface (such as through the UI automation framework to locate and click the "submit" button). After execution, record the log (such as "timestamp | user | operation | status") and generate a voice prompt (such as "report success") to return to the workstation machine for broadcast.

[0103] Complete interaction process example: The operator speaks in front of the workstation machine: "Submit report for order A1001, quantity 50." The speech recognition application collects speech and extracts Fbank features after preprocessing.

[0104] The BiGRU model combined with the attention mechanism outputs the text "submit report order A1001 quantity 50".

[0105] The text is sent to the server through the WebSocket channel.

[0106] The software operation component parses the instruction, maps it to the work order submission function, and calls the API to automatically fill in the order A1001 and the quantity 50, and clicks the submit button.

[0107] The system records the log and voice broadcasts "work report completed".

[0108] This embodiment significantly reduces the operation threshold and improves the efficiency of the workshop through a special voice recognition model, real-time communication, and intelligent operation components. This method can be extended to other industrial scenarios, such as warehouse management and device monitoring.

[0109] Any unmentioned places in this application can be implemented by using or referring to existing technology.

[0110] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0111] The above is only an embodiment of the present application and is not intended to limit the present application. Various modifications and changes can be made to the present application by those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of the claims of the present application.

Claims

1. A method for implementing voice recognition software interaction based on a production operation scenario in a workshop, characterized in that, The method comprises: constructing an intelligent production voice recognition application; constructing a full-duplex communication channel between the intelligent production voice recognition application and an intelligent production software service; constructing a software operation component; recognizing voice data in a production operation scene in the intelligent production voice recognition application to obtain a voice recognition result; transmitting the voice recognition result to the software operation component through the full-duplex communication channel; analyzing a user intent based on the voice recognition result and performing a corresponding software operation through the software operation component.

2. The method of claim 1, wherein, The construction of the intelligent production voice recognition application comprises: constructing a voice data preprocessing module and an algorithm recognition public component; constructing a BiGRU voice recognition algorithm module component that fuses an attention mechanism; constructing a multi-dimensional high-quality corpus; constructing a model training and optimization service.

3. The method of claim 2, wherein, The construction of the voice data preprocessing module and the algorithm recognition public component comprises: using a low-pass filter to remove background noise and other non-target sound interference; using an Fbank method to extract parameters reflecting the essential characteristics of sound from the preprocessed digital signal.

4. The method of claim 2, wherein, The construction of the BiGRU voice recognition algorithm module component that fuses an attention mechanism comprises: using a bidirectional gated recurrent unit as the basis of a voice recognition model; fusing an attention mechanism to strengthen the model's attention to important parts of the input voice sequence; outputting corresponding text representations.

5. The method of claim 2, wherein, The construction of the multi-dimensional high-quality corpus comprises: constructing an intelligent production industry standard kernel corpus, covering key terms and instructions of typical business processes in the intelligent production industry; constructing an intelligent production enterprise extensible corpus, incorporating enterprise-specific process parameters, quality standards, and production terminology based on enterprise individual needs; optimizing the corpus in depth, including multi-channel data collection, data cleaning, and deep labeling.

6. The method of claim 1, wherein, The construction of the full-duplex communication channel between the intelligent production voice recognition application and the intelligent production software service comprises: deploying the intelligent production voice recognition application on each independent PC workstation; obtaining a unique device workstation terminal identification code and sending it as a signal source to the software server; establishing a persistent two-way communication connection after the server verifies the legality of the identification code.

7. The method of claim 1, wherein, The construction of the software operation component comprises: building an instruction parsing and intent mapping and instruction confirmation module; building an operation execution module.

8. The method of claim 7, wherein, The construction of the instruction parsing and intent mapping and instruction confirmation module comprises: performing multi-dimensional automatic identification based on classification algorithms, involving a feature extraction judgment index system and multi-dimensional integrated algorithms; differential analysis of various scene system software operation items, operation methods, system software mapping interfaces, and system software confirmation instructions in specific industries.

9. The method of claim 7, wherein, The construction of the operation execution module comprises: receiving voice recognition results after analyzing user intent through the full-duplex communication channel; mapping specific buttons or operation items to be executed in the system interface based on the instruction parsing and intent mapping component; performing software operations and recording logs and giving voice prompts based on the execution results.

10. An electronic device comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the steps in the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Conversation interface agent for manufacturing operation information

    CN107438095A

  • Intelligent voice interaction system and method for production management

    CN108427746A

  • Chinese civil aviation air traffic control speech recognition method and system

    CN113160798A

  • Aviation maintenance error prevention system based on voice recognition

    CN116051071A

  • Voice interaction system and method, electronic equipment and storage medium

    CN116417003A