A method and system for establishing and evaluating an intelligent speech engine capability sharing model

By establishing a shared model for an intelligent voice engine, the problems of low interaction efficiency and underutilization of power dispatch recording data in the unified communication customer service system were solved, achieving efficient intelligent customer service and recording data analysis, and improving customer experience and dispatch command quality.

CN116366768BActive Publication Date: 2025-12-30CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310251669.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-14
Publication Date
2025-12-30
Estimated Expiration
2043-03-14

AI Technical Summary

Technical Problem

Existing technologies are insufficient to meet the high-efficiency interaction requirements of unified communications customer service systems. The low interaction efficiency between human agents and traditional self-service voice response systems leads to long customer waiting times and increases the workload of human agents. Furthermore, the recording data of power dispatching is not fully utilized, affecting the analysis and evaluation of dispatching and command behavior.

Method used

Establish a shared model for an intelligent voice engine. By acquiring dispatcher recording data, perform preprocessing, hierarchical evaluation, and template synthesis, develop an intelligent customer service system and a recording quality inspection system to achieve natural language interaction and instant message response. Combine this with intelligent voice analysis for data mining and evaluation.

Benefits of technology

It improves the interaction efficiency of the customer service system, reduces the workload of manual call center operations, lowers operating costs, and enhances the quality of dispatch and command and the utilization efficiency of recording data through intelligent analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116366768B_ABST
    Figure CN116366768B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent voice engine capability sharing model establishment and evaluation method and system include, obtain the audio data generated by dispatcher in the dispatching process when power system operation, and the data is preprocessed;The data after pre-processing is divided into first data set and second data set, first data set is used to establish voice engine sharing model, and second data set is used to verify voice engine sharing model;According to the verification result, the performance of voice engine sharing model is graded evaluation.The audio data of dispatching command platform is converted into structured index information, through data intelligent analysis, help accident analysis and disposal, combined with the actual scene of dispatching production intelligent learning and optimization are carried out, improve the accuracy of intelligent voice analysis, realize the scientific evaluation of dispatching language specification, improve the quality of dispatching command.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent voice engine capability sharing model establishment and evaluation technology, and in particular to a method and system for establishing and evaluating intelligent voice engine capability sharing models. Background Technology

[0002] With the widespread adoption of unified communications (Southern Network Communications) across the network and the continuous growth of the user base, the call volume of the communication service hotline serving 300,000 users nationwide will increase dramatically. Simultaneously, as communication services continue to develop, the scope of communication services will expand. Limited by existing human customer service personnel, working hours, and knowledge levels, the current unified communications customer service model is insufficient to meet the growing demand for call inquiries. Furthermore, the interaction efficiency between users and the system is greatly limited by the button-based interaction method used by human agents and traditional self-service voice response systems. Excessive customer waiting times lead to a poor customer service experience, severely impacting customer satisfaction. When users cannot quickly obtain the services they need, they will turn to human agents, significantly increasing the call volume pressure on human agents and raising operating costs.

[0003] On the other hand, as a crucial link in the safe and stable operation of the power grid, power dispatchers, communication dispatchers, and automation dispatchers are responsible for controlling and directing the operation of the power system. Dispatch consoles record massive amounts of dispatch audio data daily. Currently, this data is scattered across various levels of the system and is mainly used to help analyze the fault handling process by playing back the recordings after abnormal events occur. Furthermore, because audio files occupy a large amount of space and the audio format is not conducive to data summarization and analysis, recordings are deleted after a certain storage period. This makes it impossible to fully explore the value of this large amount of operational data to help conduct in-depth analysis and scientific evaluation of dispatching and command behaviors and effectiveness. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] Therefore, the present invention provides a method for establishing and evaluating a capability-sharing model for intelligent voice engines, which can solve the problems mentioned in the background art.

[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a method for establishing and evaluating a capability-sharing model for an intelligent voice engine, comprising:

[0008] Acquire audio recordings generated by dispatchers during the dispatching process of the power system and preprocess the data;

[0009] The preprocessed data is divided into a first dataset and a second dataset. The first dataset is used to establish a shared model for the speech engine, and the second dataset is used to verify the shared model for the speech engine.

[0010] Based on the verification results, the performance of the shared model of the speech engine is evaluated in a graded manner.

[0011] A system for establishing and evaluating a capability-sharing model of an intelligent voice engine, characterized by comprising a data acquisition unit, an intelligent prediction unit, a preprocessing unit, an adjustment and analysis unit, a synthesis output unit, a user management unit, and an evaluation unit.

[0012] The data acquisition unit is used to acquire, on a periodic basis, the dispatch recording data generated by power dispatchers, communication dispatch operators and automation dispatch operators when they are responsible for controlling and directing the operation of the power system.

[0013] The intelligent prediction unit is used to predict punctuation characters in sentences by combining speech recognition function, realize the separation of long frequency and short frequency, and predict logical dialogue based on normal calling context sentences and neural network training to realize template synthesis.

[0014] The preprocessing unit is used to perform detection, noise reduction, segmentation, and synthesis processing on the recording data collected by the data acquisition unit;

[0015] The adjustment and analysis unit is used to perform voice adjustment and tone tuning on the data processed by the preprocessing unit, and to perform high-precision text analysis.

[0016] The synthesis output unit is used to combine the analysis results of the adjustment analysis unit into a multi-character set, realize the data format conversion, generate the recording synthesis template required by the user, and establish a recording synthesis template model.

[0017] The user management unit is used to assign different permissions to different users based on their responsibilities.

[0018] The evaluation unit is used to grade and evaluate the performance of the recording synthesis template model established by the synthesis output unit.

[0019] As a preferred embodiment of the intelligent voice engine capability sharing model establishment and evaluation system described in this invention, the preprocessing unit includes an endpoint detection module, a noise cancellation module, a scene segmentation module, and a scene combination module.

[0020] The endpoint detection module is communicatively connected to the data acquisition unit. After receiving the audio data from the data acquisition unit, the endpoint detection module analyzes the audio data stream to obtain the start and end points of the user's speech in this audio segment. The start point is denoted as audio start point N, and the end point is denoted as audio end point N. N is the number of audio segments detected in this period, and the period is one month.

[0021] If the endpoint detection module does not detect the user speaking in the audio or detects that the user has not spoken for a long time, it exits the current recognition process and releases the memory occupied by the detection.

[0022] If the endpoint detection module detects a user speaking in the audio, it transmits the detected audio data to the noise cancellation module. The noise cancellation module and the endpoint detection module have a one-way communication connection, and the scene segmentation module and the noise cancellation module have a two-way communication connection. The noise cancellation module performs background noise reduction processing on the audio data and transmits the processed audio to the scene segmentation module. The scene segmentation module performs segmentation processing on the audio emitted by different characters in the audio.

[0023] When audio overlap occurs during the recognition process, the scene segmentation module transmits this part of the audio to the noise reduction module for secondary noise reduction. The scene combination module and the scene segmentation module have a one-way communication connection. The scene combination module receives audio from different characters transmitted from the scene segmentation module and synthesizes audio from the same character.

[0024] As a preferred embodiment of the intelligent voice engine capability sharing model establishment and evaluation system described in this invention, the adjustment and analysis unit includes a voice adjustment module, a timbre adjustment module, and a high-precision text analysis module.

[0025] The voice adjustment module is communicatively connected to the preprocessing unit. The voice adjustment module receives audio from different characters synthesized by the scene combination module. The voice adjustment module performs voice enhancement processing on the audio of the same character and performs volume, speed, pitch and timbre normalization processing on the enhanced audio. The timbre normalization is processed by the timbre adjustment module.

[0026] After the voice adjustment module and the timbre adjustment module complete the audio processing, the processed audio is transmitted to the high-precision text analysis module for text recognition. The high-precision text analysis module is communicatively connected to the scene combination module.

[0027] If the scene combination module determines that the role is a user, it directly transmits the user's audio to the high-precision text analysis module, and only performs text analysis processing on the user's audio.

[0028] If the scene combination module determines that the role is not a user, the audio will be transmitted to the voice adjustment module for processing.

[0029] The high-precision text analysis module and the timbre adjustment module are connected in full-duplex communication. After the user's audio is text-recognized, the recognized text file is transmitted to the timbre adjustment module. The timbre adjustment module normalizes the content of the text file and generates a unified timbre audio file.

[0030] As a preferred embodiment of the intelligent voice engine capability sharing model establishment and evaluation system described in this invention, the timbre adjustment module includes:

[0031] The timbre adjustment module obtains sample voices of the different characters after they have been processed by the voice adjustment module.

[0032] The timbre adjustment module extracts timbre features from the sample speech using a timbre encoder to obtain candidate timbre features for different roles;

[0033] The timbre adjustment module extracts phoneme-level prosodic information from the sample speech to obtain sample prosodic information.

[0034] The timbre adjustment module upsamples the sample prosody information to the speech frame level to obtain the speech frame features corresponding to the sample speech.

[0035] The timbre adjustment module integrates the candidate timbre features of the different roles into the speech frame features corresponding to the sample speech, and generates the predicted speech corresponding to the sample speech based on the speech frame features after integrating the timbre features.

[0036] The timbre adjustment module adjusts the candidate timbre features of different roles based on the difference between the sample speech and the predicted speech.

[0037] As a preferred embodiment of the intelligent speech engine capability sharing model establishment and evaluation system described in this invention, the synthesis output unit includes:

[0038] The synthesis output unit is communicatively connected to the adjustment and analysis unit. When performing voice synthesis for customer service personnel, the synthesized voice is divided into fixed format and non-fixed format voice. The fixed format voice is predicted by the intelligent prediction unit.

[0039] The synthesis output unit pre-sets a voice template library corresponding to each type of business. The voice template library contains corresponding customer service personnel voice templates. The synthesis output unit also pre-sets the voice style corresponding to each customer service voice template. The voice style includes Chinese male voice, Chinese female voice, foreign language male voice, and foreign language female voice. The synthesis output unit also pre-sets hot words corresponding to each business type and the weight corresponding to each hot word.

[0040] As a preferred embodiment of the intelligent voice engine capability sharing model establishment and evaluation system described in this invention, the evaluation unit includes:

[0041] The data processed by the preprocessing unit is divided into a first dataset and a second dataset. The first dataset is used to establish a shared model for the speech engine, and the second dataset is used to verify the shared model for the speech engine.

[0042] The second dataset is input into the speech engine shared model, and speech recognition is evaluated on different role templates. If the recognition accuracy is less than 80%, the second dataset is re-divided into the third dataset and the fourth dataset in a 1:1 ratio. The third dataset is combined with the first dataset, and the speech engine shared model is re-established by adjusting the analysis unit and the synthesis output unit. The fourth dataset is then used for re-evaluation and verification.

[0043] If the recognition accuracy is greater than 80% but less than 90%, the first dataset is reprocessed through the preprocessing unit to rebuild the voice engine shared model, and the second dataset is used for evaluation and verification.

[0044] If the recognition accuracy is greater than 90%, the audio template of this role is deemed qualified, and other roles are evaluated until all role audio templates are qualified, thus establishing the final optimized voice engine shared model.

[0045] As a preferred embodiment of the intelligent voice engine capability sharing model establishment and evaluation system described in this invention, it further includes:

[0046] The intelligent prediction unit divides the data acquired by the data acquisition unit into a training set and a test set in a 1:1 ratio. The training set is used to train the model to be trained to obtain an intelligent speech prediction model. The intelligent speech prediction model can realize logical prediction of different speech.

[0047] The user management unit includes different roles, including customer service personnel, customer service managers, super system administrators, dispatch personnel, dispatch managers, and user personnel. Each role is held by at least one person.

[0048] The customer service personnel have access to view system-related information but cannot perform any operations on models or reports. The customer service managers have the authority to create, delete, modify, and preview models, as well as manage the modification, deletion, and preview of reports. The super system administrator has full access to the intelligent voice analysis system. The dispatch personnel have access to view intelligent voice analysis system information but cannot perform any operations on models or reports. The dispatch managers have the authority to create, delete, modify, and preview models, as well as manage the modification, deletion, and preview of reports.

[0049] A computer device includes a memory and a processor, the memory storing a computer program, characterized in that the processor executes the computer program to implement the steps of the method described above.

[0050] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described above.

[0051] The beneficial effects of this invention are as follows: This invention proposes a method and system for establishing and evaluating a capability-sharing model for an intelligent voice engine. Utilizing the capabilities provided by an intelligent voice engine open platform, it develops an intelligent customer service system based on natural language interaction, applicable to a unified communication customer service system serving 300,000 employees across the network. Simultaneously, it develops an intelligent customer service system based on instant messaging interaction, allowing users to submit service requests through the unified communication instant messaging interface, with the system automatically responding, significantly reducing manual call center workload and lowering labor costs. Utilizing the capabilities provided by the intelligent voice engine open platform, it develops an intelligent voice analysis system, applied to the quality inspection of unified communication customer service recordings, and more importantly, to the intelligent analysis of recordings on the dispatch and command platform. It transforms the recording data from the dispatch and command platform into structured indexed information, and through intelligent data analysis, aids in accident analysis and handling. Combined with actual dispatch and production scenarios, it performs intelligent learning and optimization, improving the accuracy of intelligent voice analysis, enabling scientific evaluation of the standardization of dispatch terminology, and enhancing the quality of dispatch and command. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0053] Figure 1 A flowchart illustrating a method and system for establishing and evaluating a capability-sharing model of an intelligent voice engine, provided in one embodiment of the present invention;

[0054] Figure 2 A system architecture diagram of a method and system for establishing and evaluating a capability sharing model of an intelligent voice engine, provided in one embodiment of the present invention;

[0055] Figure 3 This is an internal structural diagram of a computer device for a method and system for establishing and evaluating a capability-sharing model of an intelligent voice engine, provided as an embodiment of the present invention. Detailed Implementation

[0056] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0058] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0059] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0060] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0061] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0062] Example 1

[0063] Reference Figure 1-3 This is the first embodiment of the present invention, which provides a method and system for establishing and evaluating a capability-sharing model for an intelligent voice engine, including:

[0064] Step 102: Obtain the audio recordings generated by the dispatcher during the dispatching process while the power system is running, and preprocess the data.

[0065] Step 104: Divide the preprocessed data into a first dataset and a second dataset. The first dataset is used to build a shared model for the speech engine, and the second dataset is used to verify the shared model for the speech engine.

[0066] Step 106: Based on the verification results, the performance of the shared model of the speech engine is evaluated in a graded manner.

[0067] A system for establishing and evaluating a capability-sharing model for an intelligent voice engine, characterized by comprising a data acquisition unit 201, an intelligent prediction unit 202, a preprocessing unit 203, an adjustment and analysis unit 204, a synthesis and output unit 205, a user management unit 206, and an evaluation unit 207.

[0068] The data acquisition unit 201 is used to acquire, on a periodic basis, the dispatch recording data generated by power dispatchers, communication dispatch operators and automation dispatch operators when they are responsible for controlling and directing the operation of the power system.

[0069] The intelligent prediction unit 202 is used to predict punctuation characters in sentences by combining speech recognition function, realize the separation of long frequency and short frequency, and predict logical dialogue by combining neural network training based on normal calling context sentences to realize template synthesis.

[0070] The preprocessing unit 203 is used to perform detection, noise reduction, segmentation and synthesis processing on the recording data collected by the data acquisition unit 201;

[0071] The adjustment and analysis unit 204 is used to perform voice adjustment and tone tuning on the data processed by the preprocessing unit 203, and to perform high-precision text analysis.

[0072] The synthesis output unit 205 is used to combine the analysis results of the adjustment analysis unit 204 into a multi-character set, realize the data format conversion, generate the recording synthesis template required by the user, and establish a recording synthesis template model.

[0073] User management unit 206 is used to assign corresponding permissions to different users based on their responsibilities.

[0074] Evaluation unit 207 is used to grade and evaluate the performance of the recording synthesis template model established by synthesis output unit 205.

[0075] The preprocessing unit 203 includes an endpoint detection module 203a, a noise cancellation module 203b, a scene segmentation module 203c, and a scene combination module 203d.

[0076] Furthermore, the endpoint detection module 203a is communicatively connected to the data acquisition unit 201. After receiving the audio data acquired by the data acquisition unit 201, the endpoint detection module 203a analyzes the audio data stream to obtain the start and end points of the user's speech in this audio segment. The start point is denoted as audio start point N, and the end point is denoted as audio end point N. N is the number of audio segments detected in this period, and the period is one month.

[0077] Furthermore, if the endpoint detection module 203a does not detect the user speaking in the audio or detects that the user has not spoken for a long time, it will exit the current recognition process and release the memory occupied by the detection.

[0078] Furthermore, if the endpoint detection module 203a detects a user speaking in the audio, it transmits the detected audio data to the noise cancellation module 203b. The noise cancellation module 203b and the endpoint detection module 203a are connected in a one-way communication, and the scene separation module 203c and the noise cancellation module 203b are connected in a two-way communication. The noise cancellation module 203b performs background noise reduction processing on the audio data and transmits the processed audio to the scene separation module 203c. The scene separation module 203c separates the audio emitted by different characters in the audio.

[0079] It should be noted that when there is audio overlap during the recognition process, the scene segmentation module 203c transmits this part of the audio to the noise reduction module 203b for secondary noise reduction. The scene combination module 203d is connected to the scene segmentation module 203c in a one-way communication. The scene combination module 203d receives different character audios transmitted from the scene segmentation module 203c and synthesizes the same character audios.

[0080] Furthermore, the adjustment analysis unit 204 includes a voice adjustment module 204a, a timbre adjustment module 204b, and a high-precision text analysis module 204c.

[0081] It should be noted that the voice adjustment module 204a is communicatively connected to the preprocessing unit 203. The voice adjustment module 204a receives audio from different characters synthesized by the scene combination module 203d. The voice adjustment module 204a performs voice enhancement processing on the audio of the same character and performs volume, speed, pitch and timbre normalization processing on the enhanced audio. The timbre normalization is processed by the timbre adjustment module 204b.

[0082] Furthermore, after the voice adjustment module 204a and the timbre adjustment module 204b complete the audio processing, the processed audio is transmitted to the high-precision text analysis module 204c for text recognition. The high-precision text analysis module 204c is communicatively connected to the scene combination module 203d.

[0083] Furthermore, if the scene combination module 203d determines that the role is a user, it directly transmits the user's audio to the high-precision text analysis module 204c, and only performs text analysis processing on the user's audio.

[0084] Furthermore, if the scene combination module 203d determines that the character is not a user, the audio will be transmitted to the voice adjustment module 204a for processing.

[0085] It should be noted that the high-precision text analysis module 204c and the timbre adjustment module 204b are connected in full-duplex communication. After the user's audio is processed for text recognition, the recognized text file is transmitted to the timbre adjustment module 204b. The timbre adjustment module 204b normalizes the content of the text file and generates a unified timbre audio file.

[0086] Furthermore, the timbre adjustment module 204b includes obtaining sample voices of different characters after being processed by the voice adjustment module 204a;

[0087] Furthermore, the timbre adjustment module 204b extracts timbre features from the sample speech using a timbre encoder to obtain candidate timbre features for different roles;

[0088] Furthermore, the timbre adjustment module 204b extracts phoneme-level prosodic information from the sample speech to obtain sample prosodic information;

[0089] Furthermore, the timbre adjustment module 204b upsamples the prosodic information of the sample to the speech frame level to obtain the speech frame features corresponding to the sample speech.

[0090] Furthermore, the timbre adjustment module 204b integrates the candidate timbre features of different roles into the speech frame features corresponding to the sample speech, and generates the predicted speech corresponding to the sample speech based on the speech frame features after integrating the timbre features.

[0091] Furthermore, the timbre adjustment module 204b adjusts the candidate timbre features of different roles based on the difference between the sample speech and the predicted speech.

[0092] It should be noted that the synthesis output unit 205 includes a communication connection with the adjustment and analysis unit 204. When performing customer service personnel voice synthesis, the synthesized voice is divided into fixed format and non-fixed format voice. The fixed format is predicted by the intelligent prediction unit 202.

[0093] Furthermore, the synthesis output unit 205 pre-sets a voice template library corresponding to each type of business. The voice template library contains corresponding customer service personnel voice templates. The synthesis output unit 205 also pre-sets the voice style corresponding to each customer service voice template. The voice styles include Chinese male voice, Chinese female voice, foreign language male voice, and foreign language female voice. The synthesis output unit 205 also pre-sets the hot words corresponding to each business type and the weight corresponding to each hot word.

[0094] Furthermore, the evaluation unit 207 includes dividing the data processed by the preprocessing unit 203 into a first dataset and a second dataset. The first dataset is used to build a shared model of the speech engine, and the second dataset is used to verify the shared model of the speech engine.

[0095] Furthermore, the second dataset is input into the speech engine shared model to evaluate the speech recognition of different role templates. If the recognition accuracy is less than 80%, the second dataset is re-divided into the third and fourth datasets at a 1:1 ratio. The third dataset is then combined with the first dataset. The speech engine shared model is re-established by adjusting the analysis unit 204 and the synthesis output unit 205. The fourth dataset is then used for re-evaluation and verification.

[0096] It should be noted that if the recognition accuracy is greater than 80% but less than 90%, the first dataset will be reprocessed by the preprocessing unit 203 to rebuild the speech engine shared model, and the second dataset will be used for evaluation and verification.

[0097] Furthermore, if the recognition accuracy is greater than 90%, the audio template of this role is deemed qualified, and other roles are evaluated until all role audio templates are qualified, thus establishing the final optimized voice engine shared model.

[0098] It should be noted that the intelligent prediction unit 202 divides the data acquired by the data acquisition unit 201 into a training set and a test set in a 1:1 ratio. The training set is used to train the model to be trained to obtain an intelligent speech prediction model. The intelligent speech prediction model can realize logical prediction of different speech.

[0099] It should be noted that the user management unit 206 includes different roles, including customer service personnel, customer service managers, super system administrators, dispatch personnel, dispatch managers, and user personnel. Each role is held by at least one person.

[0100] Furthermore, customer service personnel have access to view system-related permissions but cannot perform any operations on models or reports. Customer service managers have the authority to create, delete, modify, and preview models, as well as manage the modification, deletion, and preview of reports. The super system administrator has full access to the intelligent voice analysis system. Dispatch personnel have access to view intelligent voice analysis system-related permissions but cannot perform any operations on models or reports. Dispatch managers have the authority to create, delete, modify, and preview models, as well as manage the modification, deletion, and preview of reports.

[0101] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.

[0102] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for establishing and evaluating a shared capability model of an intelligent voice engine. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0103] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0104] Acquire audio recordings generated by dispatchers during the dispatching process of the power system and preprocess the data;

[0105] The preprocessed data is divided into a first dataset and a second dataset. The first dataset is used to build a shared model for the speech engine, and the second dataset is used to validate the shared model for the speech engine.

[0106] Based on the verification results, the performance of the shared model of the speech engine was evaluated in a graded manner.

[0107] Example 2

[0108] Reference Figure 1-3 As an embodiment of the present invention, a method and system for establishing and evaluating a capability sharing model of an intelligent voice engine are provided. In order to verify the beneficial effects of the present invention, a comparative experiment is conducted for scientific demonstration.

[0109] Table 1. Accuracy of templates for different roles during the establishment of a shared model for a certain speech engine.

[0110] Character Template Name Assessment accuracy % handling methods Accuracy % after reprocessing Customer service personnel 91 none / Customer service management personnel 84 Rebuild the model 92 Super System Administrator 78 Rebuild the dataset 83 Dispatch Salesperson 82 Rebuild the model 90 Dispatch and management personnel 79 Rebuild the dataset 91

[0111] This invention proposes a method and system for establishing and evaluating a capability-sharing model for intelligent voice engines. Utilizing the capabilities provided by an intelligent voice engine open platform, it develops an intelligent customer service system based on natural language interaction, applicable to a unified communications customer service system serving 300,000 employees across the network. Simultaneously, it develops an intelligent customer service system based on instant messaging interaction, allowing users to submit service requests through the unified communications instant messaging interface, with the system automatically responding, significantly reducing the workload of manual call center operations and lowering labor costs. Leveraging the capabilities provided by the intelligent voice engine open platform, it develops an intelligent voice analysis system, applied to the quality inspection of unified communications customer service recordings, and more importantly, to the intelligent analysis of recordings in the dispatch and command platform. It transforms the recording data from the dispatch and command platform into structured indexed information, and through intelligent data analysis, aids in accident analysis and handling. Combined with intelligent learning and optimization based on actual dispatch and production scenarios, it improves the accuracy of intelligent voice analysis, enables scientific evaluation of the standardization of dispatch terminology, and enhances the quality of dispatch and command.

[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

[0113] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0114] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0115] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0117] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0118] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for establishing and evaluating a capability-sharing model for intelligent voice engines, characterized in that: The application relates to an intelligent power system operation audio recognition method and device. Acquire audio data generated by a dispatcher during dispatching, and preprocess the data; If no user speech is detected in the audio or the user has not spoken for a long time, the current recognition process is exited to release the memory occupied by the detection; If user speech is detected in the audio, the detected audio data is subjected to background noise reduction, and audio emitted by different roles is separated; When audio overlap occurs during the recognition process, the audio is subjected to secondary noise reduction, and audio of the same role is synthesized; The preprocessed data is divided into a first data set and a second data set, the first data set is used to establish a voice engine shared model, and the second data set is used to verify the voice engine shared model; According to the verification result, the performance of the voice engine shared model is evaluated; The second data set is input into the voice engine shared model, and voice recognition evaluation is performed on different role templates, including customer service personnel, customer service managers, super system administrators, dispatching personnel, dispatching managers and user personnel, and each role is jointly assumed by at least one person; According to the recognition accuracy obtained by the voice recognition evaluation, the following processing is performed: If the recognition accuracy is less than 80%, the second data set is divided into a third data set and a fourth data set in a 1:1 ratio, the third data set is combined with the first data set, the voice engine shared model is re-established, and the fourth data set is used for re-evaluation verification; If the recognition accuracy is greater than 80% and less than 90%, the first data set is reprocessed, the voice engine shared model is re-established, and the second data set is used for evaluation verification; If the recognition accuracy is greater than 90%, it is determined that the role audio template is qualified, and other roles are evaluated until all role audio templates are qualified, and a final optimized voice engine shared model is established.

2. The intelligent speech engine capability sharing model establishment and evaluation system applied to the method of claim 1, characterized in that: The application relates to an intelligent power system operation audio recognition method and device. The data acquisition unit (201) is used for acquiring dispatching audio data generated by power dispatchers, communication dispatching operation personnel and automatic dispatching operation personnel when they are responsible for controlling and commanding power system operation according to a period; The intelligent prediction unit (202) is used for realizing prediction of punctuation characters in a sentence, realizing long-frequency and short-frequency separation, realizing template synthesis according to normal calling scene sentences and combining neural network training prediction logic dialogues; The preprocessing unit (203) is used for detecting, reducing, separating and synthesizing the audio data collected by the data acquisition unit (201); The adjustment analysis unit (204) is used for performing voice adjustment and tone debugging on the data processed by the preprocessing unit (203), and performing high-precision text analysis; The synthesis output unit (205) is configured to perform multi-character synthesis on the analysis result of the adjustment analysis unit (204), realize data format conversion, generate a voice synthesis template required by a user, and establish a voice synthesis template model; The user management unit (206) is configured to assign different users with corresponding permissions according to the responsibilities of different users. The evaluation unit (207) is configured to perform hierarchical evaluation on the performance of the voice synthesis template model established by the synthesis output unit (205).

3. The intelligent speech engine capability sharing model establishment and evaluation system of claim 2, wherein: The preprocessing unit (203) comprises an endpoint detection module (203a), a noise elimination module (203b), a scene separation module (203c), and a scene combination module (203d). The endpoint detection module (203a) is in communication connection with the data acquisition unit (201). After receiving the audio data acquired by the data acquisition unit (201), the endpoint detection module (203a) analyzes the audio data stream, acquires the starting point and the ending point of the user's speech in the audio, records the starting point as audio starting point N, and records the ending point as audio ending point N, where N is the number of audio detected in the current period, and the period is one month. If the endpoint detection module (203a) does not detect the user's speech in the audio or detects that the user has not spoken for a long time, the current recognition process is exited, and the memory occupied by this detection is released. If the endpoint detection module (203a) detects the user's speech in the audio, the detected audio data is transmitted to the noise elimination module (203b). The noise elimination module (203b) is in one-way communication connection with the endpoint detection module (203a). The scene separation module (203c) is in bidirectional communication connection with the noise elimination module (203b). The noise elimination module (203b) performs background noise reduction processing on the audio data, and transmits the processed audio to the scene separation module (203c). The scene separation module (203c) separates the audio emitted by different roles in the audio. When an audio overlapping part appears in the recognition process, the scene separation module (203c) transmits the part of the audio to the noise elimination module (203b) for secondary noise reduction. The scene combination module (203d) is in one-way communication connection with the scene separation module (203c). The scene combination module (203d) receives different role audios transmitted from the scene separation module (203c), and synthesizes audios of the same role.

4. The intelligent speech engine capability sharing model establishment and evaluation system of claim 3, wherein: The adjustment analysis unit (204) comprises a voice adjustment module (204a), a timbre adjustment module (204b), and a high-precision text analysis module (204c). The voice adjustment module (204a) is in communication connection with the preprocessing unit (203), the voice adjustment module (204a) receives audio of different roles synthesized by the scene combination module (203d), the voice adjustment module (204a) carries out voice enhancement processing to the audio of the same role, and carries out volume, sound speed, pitch and tone normalization processing to the enhanced audio, the tone normalization is processed by the tone adjustment module (204b); When the voice adjustment module (204a) and the tone adjustment module (204b) complete the processing of the audio, the processed audio is transmitted to the high-precision text analysis module (204c) for text recognition, and the high-precision text analysis module (204c) is in communication connection with the scene combination module (203d); If the scene combination module (203d) determines that the role is a user personnel, the audio of the user personnel is directly transmitted to the high-precision text analysis module (204c), and only the audio of the user personnel is subjected to text analysis processing; If the scene combination module (203d) determines that the role is a non-user personnel, the audio is transmitted to the voice adjustment module (204a) for processing; The high-precision text analysis module (204c) is in full-duplex communication connection with the tone adjustment module (204b), and after the audio of the user personnel is subjected to text recognition, the recognized text file is transmitted to the tone adjustment module (204b), and the tone adjustment module (204b) normalizes the content of the text file to generate a unified tone audio file.

5. The intelligent speech engine capability sharing model establishment and evaluation system of claim 4, wherein: The tone adjustment module (204b) comprises, The tone adjustment module (204b) obtains sample voice of different roles processed by the voice adjustment module (204a); The tone adjustment module (204b) extracts tone features of the sample voice through a tone encoder to obtain candidate tone features of the different roles; The tone adjustment module (204b) extracts prosody information of the sample voice at a phoneme level to obtain sample prosody information; The tone adjustment module (204b) up-samples the sample prosody information to a voice frame level to obtain voice frame features corresponding to the sample voice; The tone adjustment module (204b) integrates the candidate tone features of the different roles into the voice frame features corresponding to the sample voice, and generates predicted voice corresponding to the sample voice according to the voice frame features after the tone features are integrated; The tone adjustment module (204b) adjusts the candidate tone features of the different roles according to the difference between the sample voice and the predicted voice.

6. The intelligent speech engine capability sharing model establishment and evaluation system of claim 5, wherein: The synthesis output unit (205) comprises, The synthesis output unit (205) is in communication connection with the adjustment analysis unit (204), when synthesizing voice of a customer service personnel, the synthesized voice is divided into fixed format voice and non-fixed format voice, and the fixed format voice is subjected to logical sentence prediction by the intelligent prediction unit (202); The synthesis output unit (205) is provided with a voice template library corresponding to each type of service, and the voice template library includes corresponding customer service personnel voice templates, and the synthesis output unit (205) is provided with a voice style corresponding to each customer service voice template, the voice style includes Chinese male voice, Chinese female voice, foreign male voice and foreign female voice, and the synthesis output unit (205) is also provided with a hot word corresponding to each type of service, and a weight corresponding to each hot word.

7. The intelligent speech engine capability sharing model establishment and evaluation system of claim 6, wherein: The evaluation unit (207) comprises, The data processed by the preprocessing unit (203) is divided into a first data set and a second data set, the first data set is used to establish a voice engine shared model, and the second data set is used to verify the voice engine shared model; The second data set is input into the voice engine shared model, and the voice recognition evaluation of different role templates is performed, if the recognition accuracy is less than 80%, the second data set is re-divided into a third data set and a fourth data set in a 1:1 ratio, and the third data set is combined with the first data set to re-establish the voice engine shared model by adjusting the analysis unit (204) and the synthesis output unit (205), and the fourth data set is used to re-evaluate and verify; If the recognition accuracy is greater than 80% and less than 90%, the first data set is reprocessed by the preprocessing unit (203) for the second time, the voice engine shared model is re-established, and the second data set is used for evaluation and verification; If the recognition accuracy is greater than 90%, it is determined that the role audio template is qualified, and other roles are evaluated until all role audio templates are qualified, and a final optimized voice engine shared model is established.

8. The intelligent speech engine capability sharing model establishment and evaluation system of claim 7, wherein: Further comprising, The intelligent prediction unit (202) divides the data obtained by the data acquisition unit (201) into a training set and a test set in a 1:1 ratio, trains the to-be-trained model through the training set to obtain an intelligent voice prediction model, and the intelligent voice prediction model can realize logical prediction of different voices; The user management unit (206) comprises different roles, and the different roles comprise customer service personnel, customer service managers, super system administrators, dispatching business personnel, dispatching managers and user personnel, and each role is jointly assumed by at least one person; The customer service personnel can perform system-related viewing permissions, and cannot perform any operation on the model and the report, the customer service managers can perform system creation, deletion, model modification and model preview, report modification, deletion, preview management responsibilities, the super system administrators can perform intelligent voice analysis system permission system all permissions, the dispatching business personnel can perform intelligent voice analysis system related viewing permissions, and cannot perform any operation on the model and the report, and the dispatching managers can perform intelligent voice analysis system creation, deletion, model modification, model preview, report modification, deletion, preview management responsibilities. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the method of claim 1.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method of claim 1.

Citation Information

Patent Citations

  • AI customer service model test method and device, electronic equipment and storage medium

    CN111859985A

  • Customer service recording-based intelligent dialogue method, system and device, and storage medium

    CN112562678A