Speech Recognition Method, Apparatus, Device, and Storage Medium

By obtaining voice scenes and personnel attribute information and performing voice recognition for recognizing speech, the problem of inaccurate recognition in different voice scenes in the prior art is solved, and higher recognition accuracy and automation are achieved.

CN114155841BActive Publication Date: 2025-06-10ANHUI IFLYREC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111347046.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-06-10
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

Existing speech recognition technology is difficult to achieve accurate recognition in different speech scenes and cannot effectively match the characteristics of speech scenes.

Method used

By obtaining the attribute information of the voice scene and the attribute information of the person, combining this information to recognize the recognized voice for speech, ensuring that the recognition results match the voice scene.

Benefits of technology

It improves the accuracy of speech recognition, so that more accurate recognition results can be obtained in any speech scene, and reduces the need for manual correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155841B_ABST
    Figure CN114155841B_ABST
Patent Text Reader

Abstract

The present application provides a voice recognition method, apparatus, device, and storage medium. The method includes: obtaining attribute information of a voice scenario for generating a voice to be recognized, and / or determining attribute information of a person in the voice scenario for generating the voice to be recognized; determining a voice recognition result of the voice to be recognized according to the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario. By using the above voice recognition method, it is possible to ensure that the recognition result matches the voice scenario, so that voice recognition can be competent for recognizing voices in any voice scenario, and more accurate voice recognition results can be obtained in any voice scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a speech recognition method, apparatus, device, and storage medium. Background Art

[0002] With the development of speech recognition technology, the application of speech recognition in speech scenarios such as meetings and speeches has become increasingly widespread. With the help of speech recognition technology, speech audio generated in speech scenarios can be automatically recognized, and corresponding recognition text can be generated, thus providing great convenience for users to record speech content. For example, in a meeting scenario, with the help of speech recognition technology to recognize the speech of the meeting speaker, the purpose of automatically recording the speaker's speech content can be achieved, which can facilitate users to record meeting content, not only improving the recording efficiency but also reducing the dictation pressure on participating users.

[0003] Conventional speech recognition methods are to process speech audio data to determine the text content corresponding to the speech audio, so as to achieve the purpose of speech recognition. However, in actual applications, the speech audio generated in different speech scenarios has its unique characteristics, and conventional speech recognition methods often cannot be competent for the accurate recognition of speech in various speech scenarios. Summary of the Invention

[0004] Based on the above technical status quo, this application proposes a speech recognition method, apparatus, device, and storage medium. This speech recognition method can ensure that the recognized speech recognition result matches the speech scenario, and obtain a more accurate speech recognition result.

[0005] To achieve the above object, this application proposes the following technical solutions:

[0006] A speech recognition method, comprising:

[0007] Obtaining attribute information of the speech scenario in which the speech to be recognized is generated, and / or determining attribute information of the person in the speech scenario in which the speech to be recognized is generated;

[0008] Determining a speech recognition result of the speech to be recognized according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario.

[0009] Optionally, the determining a speech recognition result of the speech to be recognized according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario includes:

[0010] Performing speech recognition on the speech to be recognized to obtain a speech recognition result;

[0011] Correct the speech recognition result according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, so as to obtain the corrected speech recognition result of the speech to be recognized.

[0012] Optionally, obtaining the attribute information of the speech scenario for generating the speech to be recognized includes:

[0013] Obtain the scene image of the speech scenario for generating the speech to be recognized, and / or obtain the speech data in the speech scenario for generating the speech to be recognized;

[0014] Determine the attribute information of the speech scenario by performing image recognition on the scene image and / or by performing speech recognition on the speech data.

[0015] Optionally, determining the attribute information of the speech scenario by performing image recognition on the scene image includes:

[0016] Recognize the text content in the scene image;

[0017] Determine the scene attribute information of the speech scenario according to the text content recognized from the scene image, where the scene attribute information includes at least one of the region, field, and personnel information of the speech scenario.

[0018] Optionally, determining the attribute information of the speech scenario by performing image recognition on the scene image further includes:

[0019] Recognize the scene type of the speech scenario according to the scene image.

[0020] Optionally, determining the attribute information of the speech scenario by performing speech recognition on the speech data includes:

[0021] Perform speech recognition on the speech data to obtain the text content corresponding to the speech data;

[0022] Perform word segmentation on the text content corresponding to the speech data and analyze the word segmentation result to determine the scene attribute information of the speech scenario, where the scene attribute information includes at least one of the region, field, and personnel information of the speech scenario.

[0023] Optionally, determining the attribute information of the person in the speech scenario for generating the speech to be recognized includes:

[0024] Determine the attribute information of the person in the speech scenario according to the speech content of the person in the speech scenario for generating the speech to be recognized and / or the introduction information of the speech scenario, where the attribute information of the person includes at least one of the role, title, and position of the person.

[0025] Optionally, according to the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario, correct the voice recognition result to obtain the corrected voice recognition result of the voice to be recognized, including:

[0026] Generate a structured template of the speech content corresponding to the voice scenario at least according to the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario;

[0027] Identify the text content that does not conform to the structured template of the speech content from the voice recognition result of the voice to be recognized;

[0028] According to the structured template of the speech content, correct and / or delete the text content that does not conform to the structured template of the speech content to obtain the corrected voice recognition result of the voice to be recognized.

[0029] Optionally, the generating a structured template of the speech content corresponding to the voice scenario at least according to the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario includes:

[0030] Extract keywords from the voice recognition result of the voice to be recognized;

[0031] Perform structured processing and combination on the extracted keywords, the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario to obtain a structured template of the speech content corresponding to the voice scenario.

[0032] A voice recognition device, including:

[0033] An information acquisition unit, configured to acquire the attribute information of the voice scenario for generating the voice to be recognized, and / or determine the attribute information of the person in the voice scenario for generating the voice to be recognized;

[0034] A recognition processing unit, configured to determine the voice recognition result of the voice to be recognized according to the attribute information of the voice scenario and / or the attribute information of the person in the voice scenario.

[0035] A voice recognition device, including:

[0036] A memory and a processor;

[0037] Wherein, the memory is connected to the processor and is configured to store programs;

[0038] The processor is configured to implement the above voice recognition method by running the programs in the memory.

[0039] A storage medium stores a computer program thereon. When the computer program is run by a processor, the above-mentioned voice recognition method is implemented.

[0040] In the voice recognition method proposed in this application, when performing voice recognition on the voice to be recognized, the attribute information of the voice scene where the voice to be recognized is generated is first obtained, and / or the attribute information of the person in the voice scene is determined. On this basis, voice recognition is performed on the voice to be recognized according to the attribute information of the voice scene and / or the attribute information of the person in the voice scene, and the voice recognition result of the voice to be recognized is determined.

[0041] Since both the attribute information of the voice scene and the attribute information of the person in the voice scene can reflect the content attributes of the voice generated in the voice scene. Therefore, performing voice recognition on the voice to be recognized according to the attribute information of the voice scene and / or the attribute information of the person in the voice scene can ensure that the recognition result matches the attribute information of the voice scene and / or the attribute information of the person in the voice scene, that is, it can ensure that the recognition result matches the voice scene, so that voice recognition can be competent for the recognition of voices in any voice scene, and more accurate voice recognition results can be obtained in any voice scene. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0043] Figure 1 It is a flowchart showing a voice recognition method provided by an embodiment of the present application;

[0044] Figure 2 It is a flowchart showing another voice recognition method provided by an embodiment of the present application;

[0045] Figure 3 It is a flowchart showing yet another voice recognition method provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram showing the execution process of the voice recognition method provided by an embodiment of the present application;

[0047] Figure 5 It is a schematic diagram showing the structure of a voice recognition device provided by an embodiment of the present application;

[0048] Figure 6 It is a schematic diagram showing the structure of a voice recognition device provided by an embodiment of the present application. Detailed implementation manners

[0049] The technical solution of the embodiment of the present application is applicable to the voice recognition application scenario. By adopting the technical solution of the embodiment of the present application, the recognition accuracy of the voice in the voice scenario can be improved.

[0050] According to the conventional voice recognition technical solution, in order to improve the accuracy of the recognition result of the voice in the voice scenario, after the voice recognition of the voice data is completed, the voice recognition result is manually verified, and the misrecognized content therein is modified. This processing method requires a large amount of manual labor, consumes a large amount of human resources and time, and has a low processing efficiency.

[0051] Based on the above technical status quo, the embodiment of the present application proposes a voice recognition method, device, equipment and storage medium. By adopting this method, a more accurate voice recognition result can be obtained, thereby eliminating the need for manual correction processing content.

[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0053] The embodiment of the present application proposes a voice recognition method. Refer to Figure 1 As shown, this method includes:

[0054] S101. Obtain the attribute information of the voice scenario where the voice to be recognized is generated, and / or determine the attribute information of the person in the voice scenario where the voice to be recognized is generated.

[0055] Specifically, the above-mentioned voice scenario where the voice to be recognized is generated refers to a voice scenario composed of people participating and speaking. When the people in this voice scenario speak, corresponding speech data is generated. For any speech data generated in this voice scenario, when voice recognition needs to be performed on it, this speech data is the voice to be recognized. Exemplarily, the above-mentioned voice scenario can be a meeting scenario, a speech scenario, a conversation scenario, an interview scenario, etc.

[0056] The attribute information of the voice scenario refers to the information used to characterize the attributes of the voice scenario, such as the field, scale, nature, area, etc. For example, assuming a certain meeting is the "Seminar on Flood Control Plan in North China", then the field attribute of this meeting scenario is the flood control field, its regional attribute is North China, and its nature is the nature of a seminar.

[0057] The attribute information of a voice scenario can, to a certain extent, reflect the content attribute information of the voice data generated in that voice scenario. For example, assuming the voice scenario is a conference in the technology field, then the content of the voice generated in this voice scenario is very likely to be related to the technology field; if the voice scenario is a development conference for a certain region, then the content of the voice generated in this voice scenario is very likely to be related to the development of that region.

[0058] Therefore, the attribute information of a voice scenario can be used to define the content attribute of the voice to be recognized generated in that voice scenario.

[0059] The attribute information of the people in a voice scenario refers to the identity, role, title, etc. of the people participating in the voice scenario. For example, assuming the voice scenario is a conference scenario, then the attribute information of the people in the voice scenario includes, but is not limited to, the identity, role, title, position, etc. of the participants.

[0060] The attribute information of the people in a voice scenario can, to a certain extent, reflect the characteristics of the speech content of the speaker, or can, to a certain extent, reflect the attribute information of the people involved in the speech content.

[0061] For example, assuming the voice data is the voice of the conference host introducing the guests, then in this voice, it may contain the titles of each guest, such as XX Chairman, XX General Manager, XX Director, etc. If the attribute information of the people in this voice scenario can be determined in advance, such as the identity, title, etc. of each person participating in the conference are determined in advance, then when the title of the guest in the host's voice cannot be accurately recognized, the correct guest title can still be determined based on the pre-determined people attribute information.

[0062] Another example, assuming that according to the attribute information of the people in the conference scenario, person A is determined to be the host, then when performing speech recognition on the voice data of person A's speech, it can be determined that their pronunciation content is more inclined to hosting language or introduction language.

[0063] Therefore, both the attribute information of a voice scenario and the attribute information of the people in a voice scenario can, to a certain extent, reflect the attributes, characteristics, etc. of the content of the voice generated in that voice scenario.

[0064] As an exemplary implementation, the above-mentioned acquisition of the attribute information of the voice scenario for generating the voice to be recognized can be determined by text analysis of the name and introduction information of the voice scenario, or can be determined by analyzing the content of the voice data in the voice scenario. For specific implementation methods, please refer to the introduction in the following embodiments. In addition, the attribute information of the voice scenario can also be information that has been previously defined and stored. When it is necessary to recognize the voice to be recognized, directly read the attribute information of the voice scenario for generating the voice to be recognized.

[0065] Similarly, the above-mentioned determination of the attribute information of the person in the voice scenario for generating the voice to be recognized can be determined by analyzing the introduction information of the voice scenario. For example, by text analysis of the meeting summary to determine the attribute information of the participants included therein, which is the attribute information of the persons in the meeting. Or, the attribute information of the persons in the voice scenario can also be information that has been previously defined and stored. When it is necessary to recognize the voice to be recognized, directly read the attribute information of the persons in the voice scenario for generating the voice to be recognized.

[0066] As an optional setting method, the above-mentioned attribute information of the voice scenario can be in the form of an attribute label of the voice scenario, and the above-mentioned attribute information of the persons in the voice scenario can be in the form of a person attribute label.

[0067] S102. Determine the speech recognition result of the speech to be recognized according to the attribute information of the voice scenario and / or the attribute information of the persons in the voice scenario.

[0068] Specifically, when the attribute information of the voice scenario for generating the voice to be recognized is obtained, and / or the attribute information of the persons in the voice scenario for generating the voice to be recognized is determined, when performing speech recognition on the voice to be recognized, use the attribute information of the voice scenario and / or the attribute information of the persons in the voice scenario as a reference to determine the text content corresponding to the voice to be recognized, and obtain the speech recognition result of the voice to be recognized.

[0069] Exemplarily, since the attribute information of the voice scenario, or the attribute information of the persons in the voice scenario, can reflect the attributes of the content of the voice data originating from the voice scenario, or the attributes of the persons in the voice data originating from the voice scenario. Therefore, when performing speech recognition on the voice to be recognized originating from the voice scenario, at least the attribute information of the voice scenario or the attribute information of the persons in the voice scenario can be referred to determine the speech content regarding the attributes of the voice scenario in the voice to be recognized, or to determine the speech content regarding the attributes of the persons in the voice scenario in the voice to be recognized.

[0070] For example, assume that a certain meeting is the "Seminar on Flood Control Plan in North China Region". Then, the domain attribute of this meeting scenario can be determined as the flood control domain, its regional attribute as the North China Region, and its nature as a seminar. When recognizing a certain sentence of speech in this meeting scenario, if "Hebei Region" or "Hubei Region" is recognized, then according to the regional attribute of this meeting scenario, it can be determined that the recognized "Hebei Region" or "Hubei Region" may be incorrect recognition results, and it is necessary to re-recognize, or to re-determine the recognition result in combination with other information, or directly determine that the correct result should be "North China Region". Another example is that when recognizing a certain sentence of speech in this meeting scenario, if "earthquake prevention" is recognized from the speech, then according to the domain attribute of this meeting scenario, it should be able to be determined that the recognized "earthquake prevention" may be an incorrect recognition result. At this time, it is necessary to re-recognize, or to re-determine the recognition result in combination with other information, or directly determine that the correct recognition result should be "flood control".

[0071] Another example is that assume the roles of Zhang XX and Li XX in a certain meeting scenario are determined as the hosts. At the same time, Li XX's position is the dean. When performing speech recognition on Zhang XX's speech "I am very glad to serve as the host of this meeting with Dean Li XX today", if the recognition result is "I am very glad to serve as the supporter of this meeting with Dean Li XX today", then according to the pre-determined attribute information of the personnel in this meeting scenario, it can be determined that Li XX's position should be the dean. Therefore, the position attribute of Li XX reflected in the recognition result "Dean Li XX" is incorrect. At this time, it is necessary to re-recognize, or to re-determine the recognition result in combination with other information, or directly correct "Dean Li XX" to "Dean Li XX".

[0072] As can be seen from the above introduction, in the speech recognition method proposed in the embodiment of the present application, when performing speech recognition on the speech to be recognized, first obtain the attribute information of the speech scenario where the speech to be recognized is generated, and / or determine the attribute information of the personnel in this speech scenario. On this basis, according to the attribute information of this speech scenario, and / or the attribute information of the personnel in this speech scenario, perform speech recognition on the speech to be recognized to determine the speech recognition result of the speech to be recognized.

[0073] Since the attribute information of the speech scenario, as well as the attribute information of the personnel in the speech scenario, can both reflect the content attribute of the speech generated in this speech scenario. Therefore, performing speech recognition on the speech to be recognized according to the attribute information of this speech scenario and / or the attribute information of the personnel in this speech scenario can ensure that the recognition result matches the attribute information of the speech scenario and / or the attribute information of the personnel in the speech scenario, that is, it can ensure that the recognition result matches the speech scenario, so that speech recognition can be competent for the recognition of speech in any speech scenario, and in any speech scenario, a more accurate speech recognition result can be obtained.

[0074] As an alternative implementation, when obtaining the attribute information of the voice scenario for generating the voice to be recognized in the embodiments of the present application, the scenario image of the voice scenario for generating the voice to be recognized may be obtained first, and / or the voice data in the voice scenario for generating the voice to be recognized may be obtained; then, by performing image recognition on the scenario image and / or performing voice recognition on the voice data, the attribute information of the scenario image may be determined.

[0075] The above-mentioned scenario image of the voice scenario for generating the voice to be recognized is specifically an image obtained by photographing the on-site space of the voice scenario for generating the voice to be recognized. For example, assuming the voice scenario is a meeting scenario, then photograph the on-site of the meeting scenario, such as photographing the on-site space of the meeting room where the meeting is held, to obtain the scenario image of the meeting scenario.

[0076] The above-mentioned voice data in the voice scenario for generating the voice to be recognized may be voice data generated by any person speaking in the voice scenario, which may be the above-mentioned voice to be recognized or other voices other than the voice to be recognized.

[0077] It can be seen that the attribute information of the voice scenario can be recognized from the scenario image of the voice scenario, or from the voice data of the voice scenario, or comprehensively obtained from the scenario image and voice data of the voice scenario.

[0078] The following introduces each voice scenario attribute recognition solution respectively.

[0079] First, a solution for obtaining the attribute information of the voice scenario from the scenario image is introduced.

[0080] Exemplarily, by performing image recognition on the above-mentioned scenario image to determine the attribute information of the voice scenario, specifically, the text content in the scenario image may be recognized first, and then at least one of the attribute information such as the region, field, and personnel information of the voice scenario may be determined according to the text content recognized from the scenario image.

[0081] Specifically, taking the meeting scenario as an example, in the scenario image of the meeting scenario, there will be text content such as the meeting title, the names of the participants, and the titles of the seated personnel. Among these text contents, there are meeting attribute information or personnel attribute information. For example, the meeting title contains attribute information such as the region and field of the meeting, and through information such as the names of the participants and the titles of the seated personnel, the attributes such as the names, genders, positions, and titles of the participants can also be determined.

[0082] Therefore, by recognizing the text content from the scenario image and then according to the recognized text content, at least one of the attribute information such as the region, field, and personnel information of the voice scenario can be determined.

[0083] As an exemplary implementation, embodiments of the present application pre-train a scene attribute recognition model for recognizing a scene image to determine scene attribute information. Before recognition, operations such as denoising, smoothing, and transformation are first performed on the scene image to enhance the detailed information of the image, and then the scene image is input into the pre-trained scene attribute recognition model to obtain the scene attribute information output by the model.

[0084] The above-mentioned scene attribute recognition model is trained based on the InceptionResNet V2 neural network. During training, first, scene images are obtained, unqualified images are removed, images with clear imaging are retained, and the images are labeled; then, the qualified scene images are input into the neural network model to train the scene attribute recognition model. Among them, the images in the training dataset are input into the building blocks, convolutional calculations are performed through the blocks, and then a series of operations such as Relu activation and feature splicing are performed to further extract features to form a validation dataset, and the recognition model is verified using the validation dataset and a large number of historical qualified images.

[0085] Furthermore, in addition to recognizing the text content from the scene image and then determining attribute information such as the region, field, and personnel information of the voice scene according to the recognized text content, embodiments of the present application also determine the scene type of the voice scene according to the scene image of the voice scene.

[0086] Specifically, the scene type of the voice scene refers to the type of the voice scene itself. In embodiments of the present application, the scene type of the voice scene is represented by a scene major category and a scene subcategory. For example, the scene major category is a meeting scene, and the scene subcategories are training, discussion, etc.

[0087] The scene type of the voice scene can reflect the characteristics of the voice in the voice scene. For example, for a formal government meeting, its voice content is rigorous and conclusive, while for a discussion meeting, its voice content may be that multiple people speak simultaneously or with a dialect. And the characteristics of the voice in the voice scene directly affect the recognition effect of the voice. For example, for the voice in a formal government meeting, due to its rigorous and standardized characteristics, the recognition result with a general confidence level can be used as the final recognition result output, while for the voice in a discussion meeting, since there may be multiple people speaking simultaneously or there may be a dialect, the recognition result with a higher confidence level should be used as the final recognition result output.

[0088] Therefore, determining the scene type of the voice scene has a certain auxiliary effect on determining the correct voice recognition result.

[0089] Embodiments of the present application determine the scene type of the voice scene by recognizing the scene image of the voice scene.

[0090] As an alternative implementation, in the embodiments of the present application, a voice scene recognition model is pre-trained to recognize the scene type of a voice scene. Then, by inputting the scene image of the voice scene into the voice scene recognition model, the recognition result of the scene type of the voice scene output by the model can be obtained.

[0091] Exemplarily, the above-mentioned voice scene recognition model is trained based on the YOLO v3 neural network. During the training process, multiple voice scene images are obtained, and the images are annotated to determine the scene type of each voice scene image; then, a voice scene recognition model is trained according to the annotated voice scene images.

[0092] During the training process, in order to expand the scene type label library of the voice scene recognition model and enable the voice scene recognition model to recognize more scene types, in the embodiments of the present application, after generating a scene tag according to the preset rule of "image scene element + system time + scene number" for the annotated scene type information of the training images, it is inserted into the scene tag record table in the database of the voice scene recognition model. The recognition result of the scene type output by the voice scene recognition model is obtained, and the recognition result of the scene type is compared with the scene tag record table. If the scene tag record table contains the recognition result of the scene type, the comparison of the subsequent recognition results of the scene type is continued until the scene tag record table does not contain the recognition result of the scene type output by the model. Then, the scene tag record table is updated, that is, the recognition result of the scene type output by the model and not included in the scene tag record table is added to the scene tag record table. Through the above operations, it can be ensured that all scene types are included in the scene tag record table of the voice scene recognition model. By continuously improving the scene tag record table, the voice scene recognition model can realize the recognition of any voice scene type.

[0093] Next, a solution for obtaining the attribute information of a voice scene from the voice data of the voice scene will be introduced.

[0094] Exemplarily, for the above-mentioned method of determining the attribute information of a voice scene by performing voice recognition on the voice data in the voice scene of the generated voice to be recognized, specifically, the voice data can be first subjected to voice recognition to obtain the text content corresponding to the voice data; then, the text content corresponding to the voice data is segmented and the segmentation result is analyzed to determine the attribute information such as the region, field, and personnel information of the voice scene.

[0095] Specifically, taking a meeting scene as an example, voice recognition is performed on the speech data of a person's speech in the meeting to obtain the text content corresponding to the speech data.

[0096] Then, the LTP word segmentation tool and TF-IDF algorithm are used to segment the text content. The open-source Ansj_seg tool is used to complete the attribute annotation of each segmented word. Then, the HanLP tool is used to complete text content classification, various mode word segmentations, and industry noun annotation conversions. Its ability to annotate industry nouns is used to analyze and extract each industry noun in the text as a supplement to complete the identification of attributes such as the conference scenario field and region. At the same time, for the information related to personnel in the text, such as attribute information such as personnel names, titles, and positions, it can also be identified together.

[0097] In the above processing process, in order to expand the dictionary of the word segmentation model or word segmentation tool and algorithm, so that the word segmentation model or word segmentation tool and algorithm can recognize more segmented words and the attributes of more segmented words, in the process of applying the word segmentation model or word segmentation tool and algorithm to segment the text and perform attribute annotation, the dictionary of the word segmentation model or word segmentation tool and algorithm is expanded or updated in real time.

[0098] Exemplarily, each time the text is segmented and the word segmentation attributes are annotated, a new word discovery algorithm is used to find new words not included in the dictionary, and then the discovered new words are added to the dictionary, thereby realizing the expansion of the dictionary.

[0099] Specifically, the above new word discovery algorithm is mainly implemented based on the mutual information and left-right entropy algorithms. Its algorithm logic is mainly divided into three steps: First, the corpus text is converted into a string, and then an n_gram dictionary is generated and the word frequency of each word is counted. Then, the pointwise mutual information is used to screen out candidate new words from the previous n_gram dictionary. Finally, the left-right entropy is used to screen out the finally output new words from the candidate new words.

[0100] For the specific content of the above new word discovery algorithm, reference can be made to the specific processing content of the conventional new word discovery algorithm, which will not be elaborated in this embodiment.

[0101] Based on the above dictionary expansion processing, the word segmentation model or word segmentation tool and algorithm can be used more and more effectively, its recognition and extraction performance of word segmentation is getting stronger and stronger, thereby improving the word segmentation efficiency, and further improving the efficiency of determining the attributes of the speech scenario.

[0102] As an alternative implementation, the above-mentioned determination of the attribute information of the personnel in the speech scenario for generating the speech to be recognized can be to determine the attribute information of the personnel in the speech scenario according to the speech content of the personnel in the speech scenario for generating the speech to be recognized and / or the introduction information of the speech scenario. The attribute information of the personnel includes at least one of the role, title, and position of the personnel.

[0103] Specifically, the introduction information of a voice scenario refers to the introduction of the basic information of the voice scenario. For example, in the case of a meeting scenario, the introduction information of the meeting scenario specifically refers to the information that introduces the basic information such as the meeting theme, meeting purpose, and participants of the meeting. It can be seen that in the introduction information of the voice scenario, there may be attribute information such as the role of the person in the voice scenario, the appellation of the person, and the position.

[0104] At the same time, in the speech content of the person in the voice scenario, there will also be the person's attribute information carried. For example, assume that in a certain meeting scenario, Zhang San's speech content is "Dear leaders and colleagues, good morning. I am very glad to serve as the host of this meeting together with Dean Li Si today." Then, according to Zhang San's speech content, it can be determined that the roles of Zhang San and Li Si are both hosts, and at the same time, it can be known that Li Si's position is dean.

[0105] Thus, it can be seen that in the speech content of the person in the voice scenario and the introduction information of the voice scenario, there may be the person's attribute information in this voice scenario. In the embodiment of the present application, in combination with big data technology, the speech content of the person in the voice scenario for generating the voice to be recognized and / or the introduction information of this voice scenario are recognized, and the content related to the person's attribute information is extracted therefrom, so as to determine one or more of the person's role, appellation, position and other attribute information in this voice scenario.

[0106] As an exemplary implementation manner, the embodiment of the present application records the appellation and position of the person in the following manner:

[0107] Adopt the form of name / surname / name + gender, or name / surname / name + position to record the person's appellation or position.

[0108] For example, assume that a person's name is ABC, then "Mr. / Ms. ABC", "Mr. / Ms. A", "Mr. / Ms. BC", "ABC Position 1", "A Position 1", "BC Position 1" can be used to record ABC's appellation and position.

[0109] Assume that a person's name is AB, then "Mr. / Ms. AB", "Mr. / Ms. A", "AB Position 1", "A Position 1" can be used to record AB's appellation and position.

[0110] Assume that a person's name is ABCD (compound surname), then "Mr. / Ms. ABCD", "Mr. / Ms. AB", "Mr. / Ms. CD", "ABCD Position 1", "AB Position 1", "CD Position 1" can be used to record ABCD's appellation and position.

[0111] Among them, the above positions can be abbreviated. For example, "chairman" can be abbreviated as "Dong", and "general manager" can be abbreviated as "Zong".

[0112] As an alternative implementation, refer to Figure 2 As shown, when determining the speech recognition result of the speech to be recognized according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, the following steps S202 to S203 can be implemented:

[0113] S202. Perform speech recognition on the speech to be recognized to obtain a speech recognition result.

[0114] S203. Correct the speech recognition result according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario to obtain the corrected speech recognition result of the speech to be recognized.

[0115] Specifically, in the embodiment of the present application, conventional speech recognition technology is first used to perform speech recognition on the speech to be recognized to obtain a speech recognition result. Then, according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, the content that does not conform to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario in the speech recognition result is corrected, so that the content in the speech recognition result conforms to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, that is, the speech recognition result of the speech to be recognized is matched with the speech scenario, and the accuracy of the speech recognition result is improved.

[0116] Exemplarily, refer to Figure 3 As shown, the above-mentioned correction of the speech recognition result according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario to obtain the corrected speech recognition result of the speech to be recognized can be specifically implemented by performing the following steps S303 to S305:

[0117] S303. Generate a structured template of the speech content corresponding to the speech scenario at least according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario.

[0118] Specifically, keywords are extracted from the speech recognition result obtained by performing speech recognition on the above-mentioned speech to be recognized, and then the extracted keywords, as well as the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, are structurally processed and combined to obtain relevant description text, which is filled into the structured template to obtain a structured template of the speech content corresponding to the speech scenario.

[0119] In the structured template of the speech content obtained through the above processing, the structured information of the speech content that conforms to the above speech scenario is recorded. The structured template of the speech content can be constructed separately according to the role, position, etc. of the speaker.

[0120] For example, for the host in a meeting scenario, a structured template for the speech content corresponding to the host can be constructed. In this structured template, it includes attribute information such as the area and field of the meeting, the participants, as well as attribute information such as the positions, genders, and titles of the participants. Then, theoretically, the speech content of the host should conform to the meeting attributes and participant attributes reflected in the corresponding structured template for the speech content.

[0121] S304. Identify the text content that does not conform to the structured template for the speech content from the speech recognition result of the speech to be recognized.

[0122] Specifically, compare and match the speech recognition result of the speech to be recognized with the structured template for the speech content corresponding to the speaker of the speech to be recognized, and identify from this speech recognition result the content that does not match the speech scenario attribute information and / or the attribute information of the people in the speech scenario in the structured template for the speech content, which is the text content that does not conform to the structured template for the speech content.

[0123] S305. Correct and / or delete the text content that does not conform to the structured template for the speech content according to the structured template for the speech content, and obtain the corrected speech recognition result of the speech to be recognized.

[0124] Specifically, for the text content in the speech recognition result of the speech to be recognized that does not conform to the structured template for the speech content, if this content is different from the content in the structured template for the speech content, then correct it according to the structured template for the speech content to make it the same as the content in the structured template for the speech content; if this content does not exist in the structured template for the speech content, then delete this content.

[0125] After the above processing, the speech recognition result of the speech to be recognized conforms to the corresponding structured template for the speech content, so it conforms to the speech scenario, that is, its accuracy is higher.

[0126] It should be understood that the above embodiments focus on Figure 2 and Figure 3 in the speech recognition method shown, and specifically introduce the processing steps different from those in the Figure 1 shown speech recognition method. For the processing steps that are the same as those in the Figure 1 shown speech recognition method, reference can be made to the corresponding part in the method embodiment shown in Figure 1 and will not be repeated here.

[0127] Based on the detailed introduction of the above embodiments, taking the meeting scenario as an example, when performing speech recognition on the speech in the meeting scenario, the complete processing process can be seen in Figure 4 shown.

[0128] In this speech recognition method, by recognizing the conference image, attribute labels such as the conference scene type label, conference field, and region are determined. At the same time, by performing speech recognition on the conference speech data, the speaker role label, conference field, region and other attribute labels are determined, and the appellation labels of the participants are determined. Then, the obtained scene attribute labels and personnel attribute labels are structurally processed and combined, and the transcriptions in the speech recognition results that do not conform to the scene attribute labels and personnel attribute labels are removed / corrected to obtain the final transcription.

[0129] Corresponding to the above speech recognition method, an embodiment of the present application also proposes a speech recognition device. Refer to Figure 5 As shown, the device includes:

[0130] An information acquisition unit 100, configured to acquire the attribute information of the speech scene for generating the speech to be recognized, and / or determine the attribute information of the personnel in the speech scene for generating the speech to be recognized;

[0131] A recognition processing unit 110, configured to determine the speech recognition result of the speech to be recognized according to the attribute information of the speech scene and / or the attribute information of the personnel in the speech scene.

[0132] As an optional implementation manner, the determining the speech recognition result of the speech to be recognized according to the attribute information of the speech scene and / or the attribute information of the personnel in the speech scene includes:

[0133] Performing speech recognition on the speech to be recognized to obtain a speech recognition result;

[0134] Correcting the speech recognition result according to the attribute information of the speech scene and / or the attribute information of the personnel in the speech scene to obtain the corrected speech recognition result of the speech to be recognized.

[0135] As an optional implementation manner, the acquiring the attribute information of the speech scene for generating the speech to be recognized includes:

[0136] Acquiring the scene image of the speech scene for generating the speech to be recognized, and / or acquiring the speech data in the speech scene for generating the speech to be recognized;

[0137] Determining the attribute information of the speech scene by performing image recognition on the scene image and / or performing speech recognition on the speech data.

[0138] As an optional implementation manner, determining the attribute information of the speech scene by performing image recognition on the scene image includes:

[0139] Identify the text content in the scene image;

[0140] Determine the scene attribute information of the voice scene according to the text content identified from the scene image, where the scene attribute information includes at least one of the area, field, and personnel information of the voice scene.

[0141] As an optional implementation manner, determining the attribute information of the voice scene by performing image recognition on the scene image further includes:

[0142] Identify the scene type of the voice scene according to the scene image.

[0143] As an optional implementation manner, determining the attribute information of the voice scene by performing voice recognition on the voice data includes:

[0144] Perform voice recognition on the voice data to obtain the text content corresponding to the voice data;

[0145] Perform word segmentation on the text content corresponding to the voice data and analyze the word segmentation results to determine the scene attribute information of the voice scene, where the scene attribute information includes at least one of the area, field, and personnel information of the voice scene.

[0146] As an optional implementation manner, determining the attribute information of the personnel in the voice scene where the to-be-recognized voice is generated includes:

[0147] Determine the attribute information of the personnel in the voice scene according to the speech content of the personnel in the voice scene where the to-be-recognized voice is generated and / or the introduction information of the voice scene, where the attribute information of the personnel includes at least one of the role, appellation, and position of the personnel.

[0148] As an optional implementation manner, correcting the voice recognition result according to the attribute information of the voice scene and / or the attribute information of the personnel in the voice scene to obtain the corrected voice recognition result of the to-be-recognized voice includes:

[0149] Generate a speech content structured template corresponding to the voice scene at least according to the attribute information of the voice scene and / or the attribute information of the personnel in the voice scene;

[0150] Identify the text content in the voice recognition result of the to-be-recognized voice that does not conform to the speech content structured template;

[0151] Correct and / or delete the text content that does not conform to the speech content structured template according to the speech content structured template to obtain the corrected voice recognition result of the to-be-recognized voice.

[0152] As an alternative implementation, generating a structured template of the speech content corresponding to the speech scenario according to at least the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario includes:

[0153] Extracting keywords from the speech recognition result of the speech to be recognized;

[0154] Performing structured processing and combination on the extracted keywords, and the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario to obtain a structured template of the speech content corresponding to the speech scenario.

[0155] Specifically, for the specific working content of each unit of the above speech recognition device, please refer to the corresponding processing content in the above embodiments of the speech recognition method, which will not be repeated here.

[0156] Another embodiment of the present application further proposes a speech recognition device, as shown in Figure 6 shown, the device includes:

[0157] A memory 200 and a processor 210;

[0158] Wherein, the memory 200 is connected to the processor 210 for storing programs;

[0159] The processor 210 is configured to implement the speech recognition method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0160] Specifically, the above speech recognition device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0161] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through a bus. Among them:

[0162] The bus may include a path for transmitting information between various components of the computer system.

[0163] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0164] The processor 210 may include a main processor, and may also include a baseband chip, a modem, etc.

[0165] The memory 200 stores a program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.

[0166] The input device 230 may include devices for receiving user input data and information, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0167] The output device 240 may include devices for allowing information to be output to the user, such as a display screen, a printer, a speaker, etc.

[0168] The communication interface 220 may include devices of any transceiver type for communicating with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0169] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement each step of the voice recognition method provided in the above embodiments of the present application.

[0170] Another embodiment of the present application also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, each step of the voice recognition method provided in the above embodiments of the present application is implemented.

[0171] Specifically, for the specific working content of each part of the above voice recognition device, and the specific processing content when the computer program on the above storage medium is run by a processor, reference can be made to the content of each embodiment of the above voice recognition method, which will not be elaborated here.

[0172] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0173] It should be noted that the embodiments in this specification are all described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For device embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0174] The steps in the methods of the embodiments of this application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0175] The modules and sub-modules in the devices and terminals in the embodiments of this application can be combined, divided, and deleted according to actual needs.

[0176] In several embodiments provided by this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are only illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0177] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or they can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] In addition, in each embodiment of the present application, each functional module or sub-module can be integrated into a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated into one module. The above-mentioned integrated module or sub-module can be implemented in the form of hardware, or in the form of a software functional module or sub-module.

[0179] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0180] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0181] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0182] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech recognition method, characterized in that, it includes: obtaining attribute information of the speech scenario where the speech to be recognized is generated, and / or determining attribute information of the person in the speech scenario where the speech to be recognized is generated; wherein, the attribute information of the speech scenario can reflect the content attribute information of the speech data generated in this speech scenario; the attribute information of the person in the speech scenario can reflect the speech content characteristics of the speaker or the attribute information of the person involved in the speech content; performing speech recognition on the speech to be recognized to obtain a speech recognition result; generating a structured template of the speech content corresponding to the speech scenario at least according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, and in the structured template of the speech content, structured information of the speech content conforming to the speech scenario is recorded; identifying content in the speech recognition result of the speech to be recognized that does not match the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario in the structured template of the speech content to obtain text content that does not conform to the structured template of the speech content; correcting and / or deleting the text content that does not conform to the structured template of the speech content according to the structured template of the speech content to obtain the corrected speech recognition result of the speech to be recognized.

2. The method according to claim 1, characterized in that, the obtaining of the attribute information of the speech scenario where the speech to be recognized is generated includes: obtaining a scene image of the speech scenario where the speech to be recognized is generated, and / or obtaining speech data in the speech scenario where the speech to be recognized is generated; determining the attribute information of the speech scenario by performing image recognition on the scene image and / or by performing speech recognition on the speech data.

3. The method according to claim 2, characterized in that, determining the attribute information of the speech scenario by performing image recognition on the scene image includes: recognizing the text content in the scene image; determining the scene attribute information of the speech scenario according to the text content recognized from the scene image, and the scene attribute information includes at least one of the region, field, and personnel information of the speech scenario.

4. The method according to claim 3, characterized in that, determining the attribute information of the speech scenario by performing image recognition on the scene image further includes: recognizing the scene type of the speech scenario according to the scene image.

5. The method according to claim 2, characterized in that, determining the attribute information of the speech scenario by performing speech recognition on the speech data includes: performing speech recognition on the speech data to obtain text content corresponding to the speech data; performing word segmentation on the text content corresponding to the speech data and analyzing the word segmentation result to determine the scene attribute information of the speech scenario, and the scene attribute information includes at least one of the region, field, and personnel information of the speech scenario.

6. The method according to claim 1, characterized in that, determining the attribute information of the person in the speech scenario where the speech to be recognized is generated includes: Determine the attribute information of the person in the speech scenario where the speech to be recognized is generated based on the speech content of the person in the speech scenario and / or the introduction information of the speech scenario, where the attribute information of the person includes at least one of the role, appellation, and position of the person.

7. According to the method described in claim 1, wherein, generating a structured template of the speech content corresponding to the speech scenario at least according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, including: extracting keywords from the speech recognition result of the speech to be recognized; structuring and combining the extracted keywords, the attribute information of the speech scenario, and / or the attribute information of the person in the speech scenario to obtain a structured template of the speech content corresponding to the speech scenario.

8. A speech recognition device, wherein, comprising: an information acquisition unit configured to acquire the attribute information of the speech scenario where the speech to be recognized is generated, and / or determine the attribute information of the person in the speech scenario where the speech to be recognized is generated; wherein, the attribute information of the speech scenario can reflect the content attribute information of the speech data generated in the speech scenario; the attribute information of the person in the speech scenario can reflect the speech content characteristics of the speaker or the attribute information of the person involved in the speech content; a recognition processing unit configured to perform speech recognition on the speech to be recognized to obtain a speech recognition result; generate a structured template of the speech content corresponding to the speech scenario at least according to the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario, where the structured information of the speech content conforming to the speech scenario is recorded in the structured template of the speech content; identify the content in the speech recognition result of the speech to be recognized that does not match the attribute information of the speech scenario and / or the attribute information of the person in the speech scenario in the structured template of the speech content to obtain the text content that does not conform to the structured template of the speech content; and correct and / or delete the text content that does not conform to the structured template of the speech content according to the structured template of the speech content to obtain the corrected speech recognition result of the speech to be recognized.

9. A speech recognition device, wherein, comprising: a memory and a processor; wherein, the memory is connected to the processor and is configured to store a program; the processor is configured to implement the speech recognition method according to any one of claims 1 to 7 by running the program in the memory.

10. A storage medium, wherein, a computer program is stored on the storage medium, and when the computer program is run by a processor, the speech recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice recognition method, device and equipment and storage medium

    CN112102833A

  • Information processing method and device and computer storage medium

    CN112905825A

  • Information processing method, device and system, conference end and server

    CN113517002A