Voice whole-process deployment method and control method for first electric power safety area

By automatically generating text corpora for the power industry and training customized speech recognition models, combined with proprietary communication protocols and container encapsulation, the challenges of corpus construction, model customization, and deployment for speech recognition in the power safety zone were solved, achieving efficient adaptation and deployment of intelligent voice control.

CN120877705APending Publication Date: 2025-10-31NR ELECTRIC CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511182178.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies in the power safety zone have problems such as difficulty in obtaining speech data, insufficient model customization capabilities, incompatible communication protocols, and complex deployment processes, making it difficult for speech recognition applications to be efficiently deployed and adapted in power systems.

Method used

By generating text corpora from the power industry, a text-to-speech dataset is constructed, a customized speech recognition model is trained, and a private communication protocol is adapted. A deployment and installation package is generated using container encapsulation, enabling the automatic generation and one-click deployment of the speech recognition model.

Benefits of technology

It has realized the full-process application of speech recognition in the power safety zone, solved the problems of difficult corpus construction, insufficient model customization capabilities, incompatible communication protocols, and complex deployment process, and supports rapid adaptation and efficient deployment, making it suitable for intelligent voice control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877705A_ABST
    Figure CN120877705A_ABST
Patent Text Reader

Abstract

The invention provides a voice whole-process deployment method and control method for a first power safety area, and belongs to the technical field of power systems. The method comprises the following steps: generating a power industry text corpus according to a preset cue word; performing voice synthesis on the power industry text corpus to generate power industry voice data corresponding to the power industry text corpus; according to the power industry text corpus and the corresponding power industry voice data, constructing a text voice pair data set; training a customized speech recognition model for the data set by using the text speech to obtain a trained customized speech recognition model; and carrying out container packaging on the trained customized voice recognition model, a special communication protocol for the power safety first area constructed based on the development plug-in and an operation environment of the special communication protocol, generating a deployment installation package, deploying the deployment installation package to the power safety first area, and obtaining a voice deployment result of the power safety first area. The problems that corpus construction is difficult, the model customization capacity is insufficient, and protocols are not compatible can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system technology, and in particular to a method for deploying and controlling a full-process voice communication system for power safety zone 1. Background Technology

[0002] In recent years, voice deployment, as an important means of human-computer interaction, has been initially applied in various industry scenarios. Typical voice deployment usually includes processes such as corpus collection, model training, model deployment, and service invocation. Its training relies on structured speech and text data, and its deployment often adopts a microservice architecture, accessing inference services through HTTP or HTTPS interfaces.

[0003] While existing technologies have achieved initial industrial applications of speech recognition, speech recognition applications in power systems, especially in highly isolated environments such as Safety Zone 1, still face the following challenges: The power industry is highly specialized, with numerous specific expressions such as equipment codes, geographical terms, and command structures. General-purpose corpora cannot meet training needs, and there is a lack of effective corpus generation and expansion methods. Existing training schemes are typically geared towards general tasks, lacking the ability to quickly customize complex business semantics for power dispatching and station control systems, making it difficult to support "personalized adaptation for each station" in the field. Mainstream speech recognition service frameworks are mostly based on the HTTP protocol, but the use of HTTP / HTTPS is prohibited in Power Safety Zone 1 for network security reasons, allowing only customized private protocols, making speech service deployment difficult. Speech services depend on runtime environment configuration, and models and services are often strongly bound. Field deployment requires high technical skills from engineers, and there is a lack of standardized, low-barrier one-click deployment capabilities. Therefore, existing speech deployment technologies face difficulties such as difficulty in obtaining corpora, limited model customization capabilities, incompatible communication protocols, and complex deployment processes. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a voice deployment method and control method for secure zone 1, which can solve the problems of difficult corpus construction, insufficient model customization capability, incompatible communication protocols, and complex deployment process.

[0005] To achieve the above objectives, the present invention is implemented using the following technical solution: On one hand, this invention provides a method for deploying a full-process voice communication system for power safety zone 1, including: Generate a text corpus of the power industry based on preset prompts; The power industry text corpus is processed by speech synthesis to generate power industry speech data corresponding to the power industry text corpus; Based on the aforementioned power industry text corpus and its corresponding power industry speech data, a text-speech pair dataset is constructed. A customized speech recognition model is trained using the aforementioned text-speech dataset to obtain the trained customized speech recognition model; The trained customized speech recognition model, the dedicated communication protocol for Power Safety Zone 1 built based on the development plugin, and its operating environment are containerized to generate a deployment package. The deployment package is then deployed to Power Safety Zone 1 to obtain the speech deployment results for Power Safety Zone 1.

[0006] Optionally, before performing speech synthesis on the power industry text corpus, the method further includes: Based on the divergent prompt words, the text corpus of the power industry is semantically transformed to generate divergent text corpus of the power industry; Scan the out-of-vocabulary words in the power industry divergent text corpus, and replace the out-of-vocabulary words with speech synthesis-recognizable pronunciation text according to preset mapping rules.

[0007] Optionally, the power industry text corpus is subjected to speech synthesis to generate power industry speech data corresponding to the power industry text corpus, including: The power industry text corpus is subjected to batch speech synthesis, speech rate adjustment, and timbre selection to generate power industry speech data corresponding to the power industry text corpus.

[0008] Optionally, the training steps of the customized speech recognition model include: A basic deep neural network speech recognition model is obtained by training a deep neural network speech recognition model on a dataset of text and speech. A basic deep neural network speech recognition model is trained using a pre-set project-level corpus to obtain a customized speech recognition model.

[0009] Optionally, the acquisition of the dedicated communication protocol for power safety zone 1 built based on the development plugin includes: By using a plugin adapted to the development framework, the HTTP-based interface is converted into a private communication protocol interface that meets the security requirements of Zone 1 of the Power Safety System, thus obtaining the dedicated communication protocol for Zone 1 of the Power Safety System.

[0010] Optional, also includes: The text-to-speech dataset, the customized speech recognition model, and the deployment package are associated and modeled and stored in the project manager, which includes user permission levels, access log records, and audit trail results.

[0011] Optionally, the deployment package includes a model file package and a deployment script; The deployment script includes deployment runtime parameters, deployment service startup commands, and deployment target mapping configurations.

[0012] Secondly, the present invention provides a voice-based full-process control method for power safety zone 1, comprising: Receive voice / audio commands; The voice audio command is input into the trained customized speech recognition model, and the voice text is output. The speech text is semantically classified to obtain the text semantic category; The text semantic category is converted into a structured command, and the structured command is sent to the scheduling system for voice control; The trained customized speech recognition model is obtained using the method described in the first aspect.

[0013] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: This invention addresses the complex scenarios in the power industry by automatically generating text-to-speech datasets to train customized speech recognition models. These models are then applied to the power safety zone, which can invoke the models via a private protocol. By organically integrating corpus construction, model customization, protocol adaptation, and speech deployment, a closed-loop, engineered approach to speech recognition is formed. This method is particularly suitable for intelligent speech systems deployed in the power safety zone, effectively solving key issues in existing technologies regarding industry adaptability, deployment environment, and security compliance. Attached Figure Description

[0014] Figure 1 The diagram shown is a flowchart of one embodiment of the voice deployment method for power safety zone 1 of the present invention; Figure 2 The diagram shown illustrates the construction process of the text-speech pair dataset of the present invention in one embodiment. Detailed Implementation

[0015] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0016] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0017] Example 1

[0018] like Figure 1As shown in the figure, this embodiment introduces a voice full - process deployment method for the first power security zone, including: Step 1: Generate power industry text corpus, specifically: As Figure 2 shown, build a built - in prompt word template library to guide the large - language model to generate text corpus in specific scenarios of the power industry. For example, the prompt word can be "Please give an oral expression of a dispatching operation". The large - language model generates text content that conforms to the semantics of power business according to the prompt word. The prompt word and the model can be loaded in parallel to form the text corpus in the first stage.

[0019] The large - language model outputs structured and recognizable text internal corpus according to the prompt word, forming an initial text corpus dataset. This dataset reflects elements such as power industry professional terms, numbers, equipment names, control commands, etc., and has domain pertinence.

[0020] Automatically inject divergent prompt words, such as "Please describe it in another way" and "Please convert it into the way that dispatchers may use", to guide the large - language model to perform semantic conversion on the initial text corpus dataset, enhance the diversity and coverage of expressions, and generate power industry divergent text corpus.

[0021] The text after divergence processing constitutes the power industry divergent text corpus in the second stage, which has richer expression methods and is suitable for improving the generalization ability of subsequent customized speech recognition models.

[0022] In this embodiment, the efficient construction and automatic expansion of industry corpus are realized. By introducing the large - language model to generate power industry text corpus and combining prompt words for semantic divergence expansion, the problem of lack of professional corpus in the existing technology is effectively solved. Compared with the traditional manual collection or script synthesis method, the corpus construction in this embodiment is more efficient, has a wider semantic coverage, reduces the data preparation cost, and provides a high - quality and structured input basis for customized speech recognition models.

[0023] Step 2: Perform pronunciation correction on the power industry text corpus, specifically: Scan the power industry divergent text corpus to identify the out - of - vocabulary system in it, and identify the words in the corpus that may be misread by the speech synthesis engine, including abbreviations, numbers, proper nouns such as "kV" and "3# main transformer", and replace them with pronunciation texts recognizable by the speech synthesis engine according to the mapping rules. For example, replace "kV" with "kilovolt" and "3#" with "No. 3", so as to improve the clarity and standardization of the synthesized speech, mainly solving the problem that the speech synthesis engine has inaccurate recognition of industry terms.

[0024] This embodiment improves the accuracy and controllability of pronunciation in the speech synthesis stage by introducing pronunciation correction to identify and replace out-of-vocabulary words in the text, avoiding the problem of the speech synthesis engine pronouncing industry-specific terms incorrectly, and improving the quality of speech synthesis.

[0025] Step 3: Construct a text-to-speech pair dataset, specifically as follows: Input the spoken text from step two into the text-to-speech module to synthesize the corresponding speech file. This module supports batch speech synthesis, speech rate adjustment, and timbre selection to ensure clear and natural speech. It performs batch speech synthesis, speech rate adjustment, and timbre selection on power industry text corpora to generate corresponding power industry speech data.

[0026] The spoken text and the audio file are paired one by one to form a complete text-to-speech dataset, which is automatically archived as a standard dataset for use by subsequent customized speech recognition models.

[0027] Step 4: Train a customized speech recognition model, specifically: A deep neural network speech recognition model is trained using a text-to-speech dataset to obtain a basic deep neural network speech recognition model. A basic deep neural network speech recognition model is trained using a pre-set project-level corpus (such as station control system logs, scheduling dialogues, etc.) to obtain a trained customized speech recognition model. The model weights are then optimized for different regions, systems, or sites to achieve customized adaptation.

[0028] This embodiment uses project-level corpus to train the model. The generated model has strong adaptability to equipment numbering, scheduling terminology, and place name expressions in the power industry, which significantly improves the recognition accuracy in actual scheduling and control scenarios and overcomes the problem of insufficient generalization ability of general speech models in industry contexts.

[0029] Step 5: Construct a dedicated communication protocol for the power safety zone 1, specifically as follows: To address the issue of customized speech recognition model services being unable to use HTTP in the Power Safety Zone 1, a plugin adapted to the Power Safety Zone 1's proprietary communication protocol was integrated into the microservice main program. Through this plugin, which adapts to development frameworks, the HTTP interface is registered as a private protocol interface conforming to the Power Safety Zone 1's communication protocol requirements. This results in the Power Safety Zone 1's proprietary communication protocol, capable of handling request header transformation, parameter encapsulation, and return structure adjustments. It is compatible with mainstream development frameworks such as Spring Boot, Django, and Flask, with the adaptation process completed automatically without requiring modification to business logic. The Power Safety Zone 1 can then use this private protocol to call the customized speech recognition model.

[0030] This embodiment breaks through the limitations of communication protocols, adapts to the deployment environment of Power Safety Zone 1, and supports the automatic conversion of HTTP-based voice service interfaces into service components compatible with the dedicated communication protocols of the safety zone. It meets the requirements of Power Safety Zone 1 for network security and protocol closure, and realizes the effective deployment of voice services in the security isolation network. This is an important improvement over the existing voice service deployment mechanism.

[0031] Step Six: Generate the deployment and installation package and apply it to Power Security Zone 1, specifically: The trained customized speech recognition model, the dedicated communication protocol for power safety zone 1 built based on the development plugin, and its operating environment are containerized to generate a deployment and installation package. The deployment and installation package includes a model file package and a deployment script. The deployment script includes deployment and running parameters, deployment service startup commands, deployment directory mapping configurations, etc. The deployment and installation package can be transferred to the Power Safety Zone 1 site via a mobile medium. Power Safety Zone 1 can call the customized speech recognition model using a private protocol, directly execute the script to complete the deployment, and obtain the speech deployment results from Power Safety Zone 1. It also supports the separate output of the model file package and the deployment and installation package, and supports one-click deployment, flexible model replacement and upgrade at the project site.

[0032] This embodiment adopts a containerized deployment approach, which encapsulates the voice model, service logic, and runtime environment in a unified manner. Combined with an automatically generated one-click deployment installation package, it enables engineers without an AI technical background to efficiently deploy the system on-site, avoiding the complex environment configuration and model binding process in traditional methods, and improving project implementation efficiency and stability.

[0033] Step 7: Data association modeling and storage, specifically: The text-to-speech dataset, customized speech recognition model, and deployment package are associated and modeled and stored in the project manager. The project manager includes user permission levels, access log records, and audit trail results to support role-based access permissions, deployment log query permissions, and training record tracking permissions. It also provides functions such as customized speech recognition model version comparison, rollback, and archiving to ensure that the deployment package is secure and controllable.

[0034] This embodiment realizes centralized management and controllable release of speech recognition resources, supports the association modeling and access control between corpora, models, microservices, and deployment packages, solves the problems of scattered speech resources and inconsistent management in traditional systems, and provides enterprises with a traceable, maintainable and scalable speech model management system.

[0035] Example 2

[0036] Based on Example 1, this example introduces a voice-based full-process control method for power safety zone 1, including: Receive voice audio commands; at this time, the pre-trained customized speech recognition model can be invoked through the dedicated communication protocol for the power safety zone 1 built based on the development plugin and its operating environment. Input the voice audio command into the trained customized speech recognition model, and output the voice text; Semantic classification is performed on the speech text to obtain the text semantic category; The text semantic categories are converted into structured commands, and the structured commands are sent to the scheduling system for voice control. The trained customized speech recognition model is obtained using the method described in Example 1, which will not be elaborated here.

[0037] This method constitutes a complete closed loop for voice communication in the field system, achieving seamless linkage from voice understanding to dispatch control. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0038] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0039] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0040] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0041] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for deploying a full-process voice communication system for power safety zone 1, characterized in that, include: Generate a text corpus of the power industry based on preset prompts; The power industry text corpus is processed by speech synthesis to generate power industry speech data corresponding to the power industry text corpus; Based on the aforementioned power industry text corpus and its corresponding power industry speech data, a text-speech pair dataset is constructed. A customized speech recognition model is trained using the aforementioned text-speech dataset to obtain the trained customized speech recognition model; The trained customized speech recognition model, the dedicated communication protocol for Power Safety Zone 1 built based on the development plugin, and its operating environment are containerized to generate a deployment package. The deployment package is then deployed to Power Safety Zone 1 to obtain the speech deployment results for Power Safety Zone 1.

2. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, Before performing speech synthesis on the aforementioned power industry text corpus, the following steps are also included: Based on the divergent prompt words, the text corpus of the power industry is semantically transformed to generate divergent text corpus of the power industry; Scan the out-of-vocabulary words in the power industry divergent text corpus, and replace the out-of-vocabulary words with speech synthesis-recognizable pronunciation text according to preset mapping rules.

3. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, The aforementioned power industry text corpus is subjected to speech synthesis to generate power industry speech data corresponding to the power industry text corpus, including: The power industry text corpus is subjected to batch speech synthesis, speech rate adjustment, and timbre selection to generate power industry speech data corresponding to the power industry text corpus.

4. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, The training steps for the customized speech recognition model include: A basic deep neural network speech recognition model is obtained by training a deep neural network speech recognition model on a dataset of text and speech. A basic deep neural network speech recognition model is trained using a pre-set project-level corpus to obtain a customized speech recognition model.

5. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, The acquisition of the dedicated communication protocol for power safety zone 1 built based on the development plugin includes: By using a plugin adapted to the development framework, the HTTP-based interface is converted into a private communication protocol interface that meets the security requirements of Zone 1 of the Power Safety System, thus obtaining the dedicated communication protocol for Zone 1 of the Power Safety System.

6. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, Also includes: The text-to-speech dataset, the customized speech recognition model, and the deployment package are associated and modeled and stored in the project manager, which includes user permission levels, access log records, and audit trail results.

7. The voice-based full-process deployment method for power safety zone 1 according to claim 1, characterized in that, The deployment package includes a model file package and a deployment script; The deployment script includes deployment runtime parameters, deployment service startup commands, and deployment target mapping configurations.

8. A voice-based full-process control method for power safety zone 1, characterized in that, include: Receive voice / audio commands; The voice audio command is input into the trained customized speech recognition model, and the voice text is output. The speech text is semantically classified to obtain the text semantic category; The text semantic category is converted into a structured command, and the structured command is sent to the scheduling system for voice control; The trained customized speech recognition model is obtained using the method described in any one of claims 1-7.