A speech recognition method, device and medium based on a multi-class environment
By using AI-powered customer service to eliminate noise interference in various environments and generate waveform sets for speech recognition, the problems of low workload and low interaction efficiency of human call center staff have been solved. This has enabled efficient speech recognition and data analysis, and reduced operating costs.
Patent Information
- Application Number
- CN202310213411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2043-03-07
AI Technical Summary
Existing AI platforms suffer from high workload and operating costs due to heavy human call volume, and traditional self-service voice response systems have low interaction efficiency, resulting in poor customer experience and an inability to effectively analyze and schedule the value of recorded data.
By acquiring voice information through AI customer service, analyzing and eliminating the influence of environmental noise, and forming a waveform set for voice recognition, the system supports recognition in both single and complex environments. The waveform set is stored and managed in the cloud to optimize the recognition process.
Reduce the workload of human call center staff, improve customer service quality, optimize dispatch and command efficiency, enhance voice recognition accuracy and data analysis capabilities, and reduce storage space usage.
Smart Images

Figure CN116343780B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, in particular to a speech recognition method and device based on multiple environments and a medium. BACKGROUND
[0002] With the in-depth popularization of unified communication application in the whole network, the user scale is growing, and the traffic pressure of the communication service hotline serving 300,000 users in the whole network will increase sharply. At the same time, with the continuous development of communication business, the communication service business scope will be wider and wider. Limited by the factors such as manpower, working hours, knowledge level of existing artificial customer service, the current unified communication customer service platform has been difficult to meet the growing demand of traffic consultation. Artificial agents and traditional self-service voice response systems use key interaction mode, and the interaction efficiency between users and the system is greatly limited. The long waiting time of customers leads to poor customer service experience, which seriously affects customer experience. When users cannot quickly obtain the required service, they will turn to artificial service, which greatly increases the artificial traffic pressure and operating costs.
[0003] As a key link of safe and stable operation of power grid, power dispatchers, communication dispatchers and automation dispatching operators are responsible for controlling and commanding power system operation. Dispatching stations record massive dispatching audio data of dispatchers every day. At present, these data are scattered in various systems and mainly used to help analyze fault handling process by playing back audio after abnormal events occur. Moreover, due to large space occupied by audio files and inconvenient data induction analysis of audio format, audio data will be deleted after file storage exceeds a certain time limit, and there is no way to fully tap the value of these large amounts of production and operation data to help in-depth analysis and scientific evaluation of dispatching behavior and effectiveness.
[0004] It is required to build a safe, stable, robust, maintainable, advanced and open intelligent voice platform to meet the needs of deep integration of business development and become the artificial intelligence strategic planning platform. SUMMARY
[0005] This section aims to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of the specification to avoid obscuring the purpose of this section, abstract and title, and such simplifications or omissions cannot be used to limit the scope of the present application.
[0006] In view of the above problems, the present application is proposed.
[0007] Therefore, the present application solves the technical problems of the prior art: the existing artificial intelligence platform has greatly increased the artificial traffic pressure, the operating cost is increased, and how to improve the quality of customer service, reduce the labor cost, improve the quality of dispatching command and optimize the effect.
[0008] To solve the above technical problems, the present application provides the following technical solutions: a voice recognition method based on multiple environments, comprising:
[0009] Obtain the voice information of the waveform through the artificial intelligence customer service;
[0010] Provide user selection of multiple environments to determine environmental impact factors;
[0011] Analyze the noise information and store and identify the noise information;
[0012] Remove the environmental factors in the voice information to obtain the voice information excluding environmental interference.
[0013] As the voice recognition method based on multiple environments according to the present application, the environmental impact factors include: judging through the environment selected by the user;
[0014] If the selected environment is a single simple noise environment, the normal environmental information in the obtained voice information is identified;
[0015] If the selected environment is a complex composite noise environment, the composite environmental information in the obtained voice information is judged.
[0016] As the voice recognition method based on multiple environments according to the present application, the composite environmental information includes: superimposing the effect waveform affected by a single environment;
[0017] The superimposed waveform information is represented as:
[0018]
[0019] Where x1, x2, …, x p is the waveform of a single environment, w k1 ,w k2 ,…, w kp is the weight of the waveform in the environment, and p is the total number of environments.
[0020] As the voice recognition method based on multiple environments according to the present application, the storage and identification of noise information includes: forming a waveform set of environmental impact effects according to memory experience and stored characteristics.
[0021] As the voice recognition method based on multiple environments according to the present application, the waveform set further includes:
[0022] When judging the waveform information of the composite environment, each element is superimposed with a weight ratio of 1 to 100, and all the obtained waveforms are stored as a waveform set;
[0023] When other users select the same environmental information as the set, the stored waveform set is directly called;
[0024] If the selected environmental information is exactly the same, the called waveform set is directly used for recognition;
[0025] If the selected environmental information is not exactly the same, and the number of selected environments is more than the existing environments, the stored waveform set is superimposed with the new environmental information;
[0026] When the environmental information selected by other users does not belong to the stored case, the waveform set is re-superimposed;
[0027] When the user selects other environments, the slower reaction time is accepted by default, the single environment is superimposed and matched, and after the superimposed environmental matching degree is met, no further superposition is performed, and no waveform set is formed and stored, and the recognition is only valid for this operation.
[0028] The speech recognition method based on multiple types of environments according to the application, characterized in that the noise information is stored and recognized, including: storing the waveform set formed under the environment selected by the user to the cloud, which can be called at any time; at the same time, the set not called for a long time is cleared:
[0029] When the time of not calling reaches or exceeds 1 year, the waveform set is cleared
[0030] When the time of not calling reaches half a year but does not exceed 1 year, the waveform set is reduced, and the waveform set previously frequently called is reduced
[0031] When the time of not calling is less than half a year, all the waveform sets are retained.
[0032] The speech recognition method based on multiple types of environments according to the application, characterized in that the environmental factors in the speech information are removed, including: removing the environmental interference factors with characteristics according to user selection; at the same time, if the user selects a single environment, but the analysis result shows that it is a composite environment, it is considered as selecting other environments; if the user selects a combination of multiple environments, but the actual analysis result is a simpler combination of environments, the user's environmental selection is corrected.
[0033] The speech recognition method based on multiple types of environments according to the application, characterized in that:
[0034] The environment factor elimination in the voice information also includes: separating the recognized environment waveform from the acquired language information; when the separated voice information can achieve effective recognition, performing voice recognition; when the separated voice information cannot achieve effective recognition, judging, if the environment selection of the user is corrected, rechecking the environment factor; if the environment selection of the user is not corrected, re-executing the stripping program, and if all elements in the set do not meet the condition, regarding the voice recognition processing under other environments.
[0035] Until the separated voice information can achieve effective recognition, the voice information is recognized.
[0036] A computing device, comprising: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, realize the steps of the method in any one of the embodiments of the present application.
[0037] A computer readable storage medium, which stores computer executable instructions, and the computer executable instructions, when executed by a processor, realize the steps of the method in any one of the embodiments of the present application.
[0038] The beneficial effects of the present application: the voice recognition method based on multiple types of environments provided by the present application can automatically identify effective content by the system, greatly reduce the artificial traffic pressure, and reduce the artificial cost; through intelligent analysis of data, intelligent learning and optimization are carried out in combination with the actual scene of dispatching production, the intelligent voice analysis accuracy is improved, the conditions of voice recognition are relaxed, and the purpose of using voice recognition in noisy environment is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:
[0040] Figure 1 A whole flow chart of a voice recognition method based on multiple types of environments provided by the first embodiment of the present application;
[0041] Figure 2 A specific running chart of a voice recognition method based on multiple types of environments provided by the second embodiment of the present application;
[0042] Figure 3 An internal structure diagram of a computer device in the two embodiments of the present application. DETAILED DESCRIPTION
[0043] In order to make the above objectives, features and advantages of the present application more clear and easily understood, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all the other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the present application.
[0044] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. The present application, however, can be practiced in a variety of ways beyond the specific details set forth herein without departing from the scope of the present application. It can be appreciated by those skilled in the art that the present application can be practiced without such specific details.
[0045] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. The "in one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments.
[0046] The present application is described in detail with reference to the accompanying drawings. In the detailed description of the embodiments of the present application, the sectional view of the device structure is partially enlarged without the general proportion for the convenience of description, and the schematic view is only an example, which should not limit the scope of protection of the present application herein. In addition, the three-dimensional spatial dimensions of length, width and depth should be included in actual manufacture.
[0047] Meanwhile, in the description of the present application, it should be noted that the terms "upper, lower, inner and outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first, second or third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.
[0048] In the present application, unless otherwise specifically defined and limited, the terms "mounting, connecting, connection" should be understood broadly, for example: it can be fixed connection, detachable connection or integral connection; it can also be mechanical connection, electrical connection or direct connection, it can also be indirectly connected through intermediate medium, or it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0049] Example 1
[0050] Reference Figure 1For an embodiment of the present application, a speech recognition method based on multi-class environment is provided, comprising:
[0051] S1: obtaining the speech information of the waveform through the artificial intelligence customer service;
[0052] Further, the speech information further comprises noise information of the environment and speech information to be recognized.
[0053] It should be understood that the current speech recognition only recognizes the largest information in the collected sound information, but in the case of a very large environmental impact, it will cause the situation of being unable to recognize or recognizing errors.
[0054] It should be noted that before the artificial intelligence customer service, the user will be provided with a variety of environment selection, which can be a single environment or a complex environment superimposed with the environment. The environment selected by the user is judged; if the user provides the environment, the selected environment is preferentially recognized; if there is no environment, the environment selection is skipped and it is defaulted that the current selection is other environment.
[0055] S2: determining the environmental impact factor, and superimposing the waveform of the effect of the single environment to form the waveform information under the complex environment.
[0056] The superimposed waveform information is represented as:
[0057]
[0058] Among them, x1, x2, …, x p is the waveform of a single environment, w k1 , w k2 , …, w kp is the weight of the waveform in the environment, and p is the total number of environments.
[0059] According to the memory experience and storage characteristics, the waveform set of the environmental impact effect is formed.
[0060] When judging the waveform information of the complex environment, each element is superimposed with a weight ratio of 1 to 100, and all the obtained waveforms are stored as a waveform set; when other users select the same environment information as the set, the stored waveform set is directly called; if the selected environment information is completely the same, the called waveform set is directly used for recognition; if the selected environment information is not completely the same, and the environment selection is more than the existing environment, then the stored waveform set is called and superimposed with the new environment information;
[0061] When the environment information selected by other users does not belong to the storage, the waveform set is superimposed again;
[0062] When the user selects other environment, the slower reaction time will be accepted by default, the superimposed matching of all single environments is performed, and the superimposed environment matching degree is not superimposed and does not form a waveform set storage, and the recognition is only valid for this operation.
[0063] It should be noted that the waveform information formed in the recognition process is not stored in the case of default and selection of other environment, saving storage space, and such cases are less.
[0064] It should also be noted that the selection and analysis of the waveform are only for the characteristic repeated paragraphs, saving space and making the recognition process run faster; at the same time, a part of the set is deleted and updated: when the non-call time reaches or exceeds 1 year, the waveform set is cleared; when the non-call time reaches half a year but does not exceed 1 year, the waveform set is reduced, and the reduced waveform set is the waveform set that was frequently called before; when the non-call time is less than half a year, all the waveform sets are retained.
[0065] S3: Eliminate environmental factors in voice information
[0066] According to user selection, the environmental interference factors with characteristics are eliminated; at the same time, if the user selects a single environment, but the analysis result shows that it is a composite environment, it is considered as selecting other environment; if the user selects a combination of multiple environments, but the actual analysis result is a simpler environment combination, the user's environment selection is corrected.
[0067] It should be known that when selecting the environment, multiple selection or wrong selection may occur, even if the user is not in such an environment, but selects such an environment, in which case the user's selection is corrected to ensure the completeness of the voice information that the user needs to recognize.
[0068] The recognized environmental waveform is separated from the obtained language information; when the separated voice information can achieve effective recognition, voice recognition is performed; when the separated voice information cannot achieve effective recognition, a judgment is made: if the user's environment selection is corrected, the environmental factors are rechecked; if the user's environment selection is not corrected, the stripping program is re-executed, and if all elements in the set do not meet the requirements, it is considered as voice recognition processing under other environment;
[0069] Until the separated voice information can achieve effective recognition, the voice information is recognized.
[0070] Finally, the voice information excluding environmental interference is obtained.
[0071] It should be noted that when recognizing the effective voice information, the fluency of the sentence is ensured; and in the process of correcting the user's environment selection, excessive elimination is avoided.
[0072] The embodiment also provides a computing device, comprising a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the method proposed in the above embodiment.
[0073] The embodiment also provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method proposed in the above embodiment.
[0074] The storage medium proposed in the embodiment belongs to the same inventive concept as the method proposed in the above embodiment, and the technical details not described in the embodiment can be referred to the above embodiment, and the embodiment has the same beneficial effects as the above embodiment.
[0075] Those skilled in the art can understand that all or part of the processes in the above embodiment method can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the above method embodiments. Any reference to memory, database or other medium used in each embodiment provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory, magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory, magnetic memory, ferroelectric memory, phase change memory, graphene memory, etc. Volatile memory can include random access memory or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory or dynamic random access memory, etc. The database involved in each embodiment provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., and is not limited thereto. The processor involved in each embodiment provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.
[0076] Embodiment 2
[0077] Reference Figures 2-3 For one embodiment of the present application, a speech recognition method based on multi-class environment is provided. In order to verify the beneficial effects of the present application, economic benefit calculation and simulation experiments are used for scientific demonstration.
[0078] Table 1 is the speech recognition situation under the influence of five single environments:
[0079] Table 1
[0080]
[0081] Table II is the speech recognition situation under five times of complex environment: Table II
[0082]
[0083] The superposition of the integer values of the waveform weight from 1 to 100 forms a relevant complex waveform set, which enhances the adaptability and accuracy to the environment.
[0084] Taking the sea wave and thunder as examples, after the user selects the environment, a waveform set of 100 elements is formed and stored, the environmental factors are stripped, and the speech information after the stripping is judged to be valid, and the language recognition is executed; when the set is not used again after one year, the system automatically clears the set to save storage space.
[0085] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A speech recognition method based on multiple environmental environments, characterized in that, include: Obtain waveform voice information through AI customer service; Provides users with multiple environmental options to determine environmental influencing factors; The noise information is analyzed, stored, and identified. By removing environmental factors from the voice information, we finally obtain voice information that has eliminated environmental interference. The environmental influencing factors include: those determined by the environment selected by the user; If a simple, single-noise environment is selected, then the ordinary environmental information in the acquired sound information will be identified. If a complex compound noise environment is selected, the compound environment information in the acquired sound information will be judged. The composite environmental information includes: superimposed waveforms of effects from single environmental influences; The superimposed waveform information is represented as follows: Where x1, x2, ..., x p For waveforms in a single environment, w k1 ,w k2 ,…,w kp Let p be the weight of the waveform in the environment, and p be the total number of environments. The storage and identification of noise information includes: forming a waveform set of environmental impact effects based on memory experience and storage characteristics; The waveform set also includes: When judging waveform information in a complex environment, each element is superimposed with a weight ratio of 1 to 100, and all the resulting waveforms are stored as a waveform set. When other users select the same environment information as the set, the stored waveform set is called directly; If the selected environmental information is exactly the same, the waveform set called will be used directly for identification; If the selected environment information is not exactly the same, and there are more environment selections than existing environments, then the new environment information is superimposed by calling the stored waveform set; If the environmental information selected by other users is not stored, the waveform set is re-overlaid; When the user selects other environments, a slower response time will be accepted by default. All individual environments will be overlaid for matching. Once the matching degree of the overlaid environments is met, they will no longer be overlaid or stored as waveform sets. The recognition is only valid for this operation.
2. The speech recognition method based on multiple environments as described in claim 1, characterized in that: The storage and identification of noise information includes: storing the waveform set generated under the user-selected environment in the cloud for easy retrieval; and simultaneously clearing sets that have not been accessed for a long time. If the inactivity period reaches or exceeds one year, the waveform set will be cleared. If the time without calling the waveform reaches six months but does not exceed one year, the waveform set will be reduced to the waveform set that was frequently called in the past. If the time since the last call is less than six months, the entire waveform set will be retained.
3. The speech recognition method based on multiple environments as described in claim 2, characterized in that: The removal of environmental factors from the voice information includes: removing characteristic environmental interference factors based on user selection; if the user selects a single environment but the analysis results show a composite environment, then it is considered that other environments have been selected; if the user selects multiple environment combinations but the actual analysis results show a simpler environment combination, then the user's environment selection is corrected.
4. The speech recognition method based on multiple environments as described in claim 3, characterized in that: The removal of environmental factors from the voice information also includes: separating the recognized environmental waveform from the acquired language information; when the separated voice information can be effectively recognized, then performing voice recognition; when the separated voice information cannot be effectively recognized, then making a judgment: if the user's environment selection has been corrected, then re-checking the environmental factors; if the user's environment selection has not been corrected, then re-executing the stripping procedure; if none of the elements in the set are satisfied, then it is considered as voice recognition processing under other environments. The process continues until the separated and processed speech information can be effectively recognized.
5. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 4.
6. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Acoustic modeling method and device, and speech recognition method and device
CN103514878A
Speech input method, device and system
CN103956169A