Voice interaction related systems, methods, devices and equipment

By introducing a speech entity recognition model and entity knowledge base in the speech entity recognition system, semantic understanding and entity recognition are directly carried out from the speech signal, the recognition error problem caused by user inaccurate or unclear pronunciation is solved, and a higher recognition accuracy and interaction success rate is achieved.

CN113889117BActive Publication Date: 2025-05-09ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010628897.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-02
Publication Date
2025-05-09
Estimated Expiration
2040-07-02

AI Technical Summary

Technical Problem

The existing voice entity recognition technology has a high probability of converting the voice signal into incorrect text content when the user's pronunciation is inaccurate or unclear, resulting in the inability to correctly identify the entity name in the voice signal.

Method used

By building multimedia program on-demand system, ordering system, communication connection establishment system, voice interaction system, etc., using the voice entity recognition model and entity knowledge base, semantic understanding and entity recognition are directly carried out from the voice signal, avoiding relying on ASR to convert speech into text.

Benefits of technology

It effectively improves the accuracy of speech entity recognition, improves the accuracy and success rate of speech interaction, and is close to the process of human understanding of speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113889117B_ABST
    Figure CN113889117B_ABST
Patent Text Reader

Abstract

The present application discloses systems, methods, devices and equipment related to voice interaction. Among them, the terminal equipment of the voice interaction system collects voice data and sends the voice data to the server; the server builds an entity knowledge base, determines the entity information in the voice data through the voice entity recognition model and the entity knowledge base; and performs voice interaction processing according to the entity information. This processing method allows the introduction of entity knowledge graph information, directly comparing whether there is an entity pronunciation in the knowledge graph in the voice, and realizing semantic understanding and entity recognition from the voice signal, which is closer to the process of human understanding of voice; therefore, it can effectively improve the accuracy of voice entity recognition, thereby improving the accuracy of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and specifically to a multimedia program on-demand system, method and device, a food ordering system, method and device, a communication connection establishment system, method and device, a speech interaction system, method and device, a speech entity recognition model construction method and device, an entity knowledge base construction method and device, a television program on-demand method and device, a meeting recording method and device, a smart speaker, a smart television, a food ordering machine, a user device, and an electronic device. Background Art

[0002] With the continuous development of Automatic Speech Recognition (ASR) technology, intelligent voice assistants have been widely used, such as the intelligent voice assistant services and smart speakers provided by smartphones to users.

[0003] The core function of artificial intelligence is the focus of voice assistants, which enables voice assistants to better understand user commands, especially entities with specific meanings in user commands, mainly including names of people, places, institutions, songs, movies, phone numbers, proper nouns, etc. Taking smart speakers as an example, users can use the song-ordering service provided by the speakers. For example, the user says to the speaker: "I want to listen to Lei Yuxin's "Memorial", where "Lei Yuxin" and "Memorial" are entities with specific meanings and are the processing objects of song-ordering commands. If these two entities cannot be correctly identified, the song ordered by the user cannot be played correctly. At present, the processing process of a typical voice entity recognition system is: first, the input voice signal is converted into text through the speech recognition technology ASR; then, the entity name in the user's command is identified through the semantic understanding of the text.

[0004] However, in the process of implementing the present invention, the inventors found that the technical solution has at least the following problems: the solution is heavily dependent on the text content output by the upstream ASR. If the user's pronunciation is inaccurate (such as an accent or partial pronunciation errors) or the pronunciation is unclear, the ASR is likely to convert the voice signal into incorrect text content, which will result in the inability to correctly recognize the entity name in the voice signal. In summary, how to improve the accuracy of voice entity recognition, thereby improving the accuracy of voice interaction, has become a problem that technicians in this field urgently need to solve. Summary of the invention

[0005] The present application provides a multimedia program on-demand system to solve the problem of low accuracy of voice entity recognition caused by inaccurate or unclear pronunciation or homophonic words in the prior art. The present application also provides a multimedia program on-demand method and device, a food ordering system, method and device, a communication connection establishment system, method and device, a voice interaction system, method and device, a voice entity recognition model construction method and device, an entity knowledge base construction method and device, a TV program on-demand method and device, a conference recording method and device, a smart speaker, a smart TV, a food ordering machine, a user device, and an electronic device.

[0006] The present application provides a multimedia program on-demand system, comprising:

[0007] The smart speaker is used to collect the on-demand voice data of the multimedia program and send the voice data to the server; and play the multimedia program according to the multimedia program playing processing result of the server;

[0008] The server is used to construct a multimedia program knowledge base; determine the multimedia program information in the voice data through a voice entity recognition model and the knowledge base; and perform multimedia program playback processing according to the multimedia program information.

[0009] The present application also provides a meal ordering system, including:

[0010] The ordering device is used to collect the ordering voice data and send the voice data to the server;

[0011] The server is used to build a meal knowledge base; determine the meal information in the voice data through a voice entity recognition model and the knowledge base; and perform meal preparation according to the meal information.

[0012] The present application also provides a communication connection establishment system, comprising:

[0013] The user equipment is used to collect the communication command voice data and send the voice data to the server;

[0014] The server is used to build a communication user knowledge base; determine the communication user information in the voice data through a voice entity recognition model and the knowledge base; and execute communication connection establishment processing according to the communication user information.

[0015] The present application also provides a voice interaction system, comprising:

[0016] The terminal device is used to collect voice data and send the voice data to the server;

[0017] The server is used to build an entity knowledge base; determine the entity information in the voice data through a voice entity recognition model and the entity knowledge base; and perform voice interaction processing according to the entity information.

[0018] The present application also provides a voice interaction method, comprising:

[0019] Build entity knowledge base;

[0020] Determine entity information in the target speech data through the speech entity recognition model and the entity knowledge base;

[0021] According to the entity information, voice interaction processing is performed.

[0022] Optionally, determining entity information in the target speech data by using the speech entity recognition model and the entity knowledge base includes:

[0023] Determining audio feature data of the speech data by using an audio coding model included in the speech entity recognition model;

[0024] The entity information is determined according to the audio feature data through the entity decoding model included in the speech entity recognition model and the entity knowledge base.

[0025] Optionally, determining the entity information according to the audio feature data by using the entity decoding model and the entity knowledge base included in the speech entity recognition model includes:

[0026] Determine, by means of an entity candidate pronunciation determination module included in the entity decoding model, at least one candidate pronunciation of the entity information according to the audio feature data;

[0027] Determine the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base by an entity pronunciation determination module included in the entity decoding model;

[0028] The entity information is determined according to the pronunciation of the entity information.

[0029] Optionally, determining the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base includes:

[0030] Determining the similarity between the pronunciation of the entity in the entity knowledge base and the candidate pronunciation;

[0031] The pronunciation of the entity information is determined according to the similarity.

[0032] Optionally, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand;

[0033] The program entity knowledge base includes: program related entities with the same pronunciation but different characters, user entities, and entity relationships between program related entities and user entities;

[0034] The constructing of the entity knowledge base includes:

[0035] Determine the user entity according to the user's historical playback information, and construct the entity relationship;

[0036] The determining the entity information according to the pronunciation of the entity information includes:

[0037] Determining a candidate entity according to the pronunciation of the entity information;

[0038] The entity information is determined from the candidate entities according to the user information and the entity relationship.

[0039] Optionally, also include:

[0040] Learning the speech entity recognition model from the training data;

[0041] The training data includes: audio data and entity annotation information.

[0042] Optionally, determining entity information in the target speech data by using the speech entity recognition model and the entity knowledge base includes:

[0043] Determining audio feature data of the speech data by using an audio coding model included in the speech entity recognition model;

[0044] Determining pronunciation feature data of entities in the entity knowledge base through the entity encoding model included in the speech entity recognition model;

[0045] The entity information is determined according to the audio feature data and the entity pronunciation feature data through the entity decoding model included in the speech entity recognition model.

[0046] Optionally, the determining the entity information by the entity decoding model included in the speech entity recognition model according to the audio feature data and the entity pronunciation feature data includes:

[0047] Determining pronunciation similarity between entities in the speech data and entities in the entity knowledge base based on the audio feature data and entity pronunciation feature data;

[0048] The entity information is determined according to the pronunciation similarity.

[0049] Optionally, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand;

[0050] The program entity knowledge base includes: program related entities with the same pronunciation but different characters, user entities, and entity relationships between program related entities and user entities;

[0051] The constructing of the entity knowledge base includes:

[0052] Determine the user entity according to the user's historical playback information, and construct the entity relationship;

[0053] The determining the entity information according to the pronunciation similarity includes:

[0054] Determining a candidate entity according to the pronunciation similarity;

[0055] The entity information is determined from the candidate entities according to the user information and the entity relationship.

[0056] Optionally, also include:

[0057] Learning the speech entity recognition model from the training data;

[0058] The training data includes: audio data, entity knowledge base and entity annotation information.

[0059] Optionally, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand;

[0060] The constructing of the entity knowledge base includes:

[0061] Determine multimedia program related entities to form the entity knowledge base.

[0062] The present application also provides a voice interaction method, comprising:

[0063] Collect voice data and send the voice data to a server so that the server can build an entity knowledge base; determine entity information in the voice data through a voice entity recognition model and the entity knowledge base; and perform voice interaction processing based on the entity information.

[0064] The present application also provides a multimedia program on-demand method, comprising:

[0065] Build a knowledge base of multimedia programs;

[0066] Determining multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base;

[0067] The multimedia program playing process is performed according to the multimedia program information.

[0068] The present application also provides a multimedia program on-demand method, comprising:

[0069] Collect multimedia program on-demand voice data, and send the voice data to a server so that the server builds a multimedia program knowledge base; determine the multimedia program information in the voice data through a voice entity recognition model and the knowledge base; and perform multimedia program playback processing according to the multimedia program information.

[0070] This application also provides a method for ordering food, including:

[0071] Build a food knowledge base;

[0072] Determine the food information in the ordering voice data by using the voice entity recognition model and the entity knowledge base;

[0073] According to the meal information, meal preparation is performed.

[0074] This application also provides a method for ordering food, including:

[0075] Collecting ordering voice data, and sending the voice data to a server so that the server can build a meal knowledge base; determining meal information in the voice data through a voice entity recognition model and the knowledge base; and performing meal preparation based on the meal information.

[0076] The present application also provides a method for establishing a communication connection, comprising:

[0077] Build a knowledge base of communication users;

[0078] Determine the communication user information in the communication command voice data through the voice entity recognition model and the knowledge base;

[0079] A communication connection establishment process is performed according to the communication user information.

[0080] The present application also provides a method for establishing a communication connection, comprising:

[0081] Collect communication instruction voice data, and send the voice data to the server so that the server can build a communication user knowledge base; determine the communication user information in the voice data through the voice entity recognition model and the knowledge base; and perform communication connection establishment processing according to the communication user information.

[0082] The present application also provides a method for constructing a speech entity recognition model, comprising:

[0083] Determine a training data set, wherein the training data includes: speech data, entity annotation information, and an entity knowledge base;

[0084] Constructing a network structure of the model;

[0085] The model is learned from a training data set.

[0086] Optionally, the model includes an audio coding model for determining audio feature data of the speech data;

[0087] The model includes an entity decoding model, which determines entity information in the speech data based on the audio feature data and the entity knowledge base.

[0088] Optionally, the model includes an audio coding model for determining audio feature data of the speech data;

[0089] The model includes an entity encoding model for determining pronunciation feature data of entities in the entity knowledge base;

[0090] The model includes an entity decoding model, which is used to determine entity information in the speech data based on the audio feature data and the entity pronunciation feature data.

[0091] The present application also provides a method for constructing an entity knowledge base, comprising:

[0092] Get the entity name of the target domain;

[0093] An entity knowledge base of a target domain is generated according to the entity name, and the entity knowledge base is used to determine entity information in speech data of the target domain through a speech entity recognition model and the entity knowledge base.

[0094] The present application also provides a method for voice entity recognition, comprising:

[0095] Build entity knowledge base and speech entity recognition model;

[0096] determining target voice data;

[0097] The entity information in the target speech data is determined through the speech entity recognition model and the entity knowledge base.

[0098] The present application also provides a voice interaction device, comprising:

[0099] A knowledge base construction unit, used to construct an entity knowledge base;

[0100] An entity determination unit, used to determine entity information in the target speech data through a speech entity recognition model and the entity knowledge base;

[0101] An interaction processing unit is used to perform voice interaction processing according to the entity information.

[0102] The present application also provides an electronic device, comprising:

[0103] Processor; and

[0104] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: building an entity knowledge base; determining entity information in the target voice data through a voice entity recognition model and the entity knowledge base; and performing voice interaction processing based on the entity information.

[0105] The present application also provides a voice interaction device, comprising:

[0106] A voice data collection unit, used for collecting voice data;

[0107] The voice data sending unit is used to send the voice data to the server so that the server builds an entity knowledge base; determine the entity information in the voice data through the voice entity recognition model and the entity knowledge base; and perform voice interaction processing according to the entity information.

[0108] The present application also provides an electronic device, comprising:

[0109] Processor; and

[0110] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: collecting voice data and sending the voice data to a server so that the server builds an entity knowledge base; determining entity information in the voice data through a voice entity recognition model and the entity knowledge base; and performing voice interaction processing according to the entity information.

[0111] The present application also provides a multimedia program on-demand device, comprising:

[0112] A voice data collection unit, used to collect voice data of multimedia program on-demand;

[0113] The voice data sending unit is used to send the voice data to the server so that the server builds a multimedia program knowledge base; determine the multimedia program information in the voice data through the voice entity recognition model and the knowledge base; and perform multimedia program playback processing according to the multimedia program information.

[0114] The present application also provides a smart speaker, comprising:

[0115] Processor; and

[0116] The memory is used to store a program for implementing a multimedia program on-demand method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting multimedia program on-demand voice data, sending the voice data to a server, so that the server builds a multimedia program knowledge base; determining multimedia program information in the voice data through a voice entity recognition model and the knowledge base; and performing multimedia program playback processing according to the multimedia program information.

[0117] The present application also provides a multimedia program on-demand device, comprising:

[0118] A knowledge base construction unit, used for constructing a multimedia program knowledge base;

[0119] An entity recognition unit, used to determine multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base;

[0120] The program playing processing unit is used to perform multimedia program playing processing according to the multimedia program information.

[0121] The present application also provides an electronic device, including:

[0122] Processor; and

[0123] The memory is used to store a program for implementing a multimedia program on-demand method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: building a multimedia program knowledge base; determining multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base; and performing multimedia program playback processing according to the multimedia program information.

[0124] The present application also provides a food ordering device, comprising:

[0125] A knowledge base construction unit, used to construct a food knowledge base;

[0126] An entity recognition unit, used to determine the food information in the ordering voice data through a voice entity recognition model and the entity knowledge base;

[0127] The meal preparation processing unit is used to perform meal preparation processing according to the meal information.

[0128] The present application also provides an electronic device, including:

[0129] Processor; and

[0130] The memory is used to store a program for implementing the ordering method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a meal knowledge base; determining the meal information in the ordering voice data through the voice entity recognition model and the entity knowledge base; and performing meal preparation processing based on the meal information.

[0131] The present application also provides a food ordering device, comprising:

[0132] A voice data collection unit, used to collect ordering voice data;

[0133] A voice data sending unit is used to send the voice data to a server so that the server can build a meal knowledge base; determine the meal information in the voice data through a voice entity recognition model and the knowledge base; and perform meal preparation processing according to the meal information.

[0134] The present application also provides a food ordering machine, comprising:

[0135] Processor; and

[0136] The memory is used to store a program for implementing the method of ordering food. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting voice data for ordering food, and sending the voice data to the server so that the server can build a food knowledge base; determining food information in the voice data through a voice entity recognition model and the knowledge base; and performing food preparation processing according to the food information.

[0137] The present application also provides a communication connection establishment device, comprising:

[0138] A voice data collection unit, used for collecting communication command voice data;

[0139] The voice data sending unit is used to send the voice data to the server so that the server builds a communication user knowledge base; determine the communication user information in the voice data through the voice entity recognition model and the knowledge base; and perform communication connection establishment processing according to the communication user information.

[0140] The present application also provides a user equipment, including:

[0141] Processor; and

[0142] The memory is used to store a program for implementing a method for establishing a communication connection. After the device is powered on and the program of the method is run through the processor, the following steps are performed: communication instruction voice data is collected and the voice data is sent to a server so that the server builds a communication user knowledge base; communication user information in the voice data is determined through a voice entity recognition model and the knowledge base; and communication connection establishment processing is performed according to the communication user information.

[0143] The present application also provides a communication connection establishment device, comprising:

[0144] A knowledge base construction unit, used to construct a communication user knowledge base;

[0145] An entity recognition unit, used to determine the communication user information in the communication instruction voice data through the voice entity recognition model and the knowledge base;

[0146] The communication connection processing unit is used to execute communication connection establishment processing according to the communication user information.

[0147] The present application also provides an electronic device, comprising:

[0148] Processor; and

[0149] The memory is used to store a program for implementing the ordering method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: building a communication user knowledge base; determining the communication user information in the communication instruction voice data through the voice entity recognition model and the knowledge base; and executing the communication connection establishment process according to the communication user information.

[0150] The present application also provides a speech entity recognition model construction device, comprising:

[0151] A training data determination unit, used to determine a training data set, wherein the training data includes: speech data, entity annotation information, and an entity knowledge base;

[0152] A network construction unit, used to construct the network structure of the model;

[0153] The model training unit is used to learn the model from the training data set.

[0154] The present application also provides an electronic device, comprising:

[0155] Processor; and

[0156] The memory is used to store a program for implementing a method for building a voice entity recognition model. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining a training data set, the training data including: voice data, entity annotation information and an entity knowledge base; building a network structure of the model; and learning the model from the training data set.

[0157] The present application also provides an entity knowledge base construction device, comprising:

[0158] An entity determination unit, used to obtain the entity name of the target domain;

[0159] The knowledge base generation unit is used to generate an entity knowledge base in a target domain, and the entity knowledge base is used to determine entity information in speech data in the target domain through a speech entity recognition model and the entity knowledge base.

[0160] The present application also provides an electronic device, comprising:

[0161] Processor; and

[0162] The memory is used to store a program for implementing the method for building an entity knowledge base. After the device is powered on and the program of the method is run through the processor, the following steps are performed: obtaining the entity name of the target domain; generating an entity knowledge base of the target domain according to the entity name, and the entity knowledge base is used to determine the entity information in the speech data of the target domain through the speech entity recognition model and the entity knowledge base.

[0163] The present application also provides a speech entity recognition device, comprising:

[0164] Model building unit, used for speech entity recognition model;

[0165] A knowledge base construction unit, used to construct an entity knowledge base;

[0166] A voice data determination unit, used to determine target voice data;

[0167] The entity determination unit is used to determine the entity information in the target speech data through the speech entity recognition model and the entity knowledge base.

[0168] The present application also provides an electronic device, comprising:

[0169] Processor; and

[0170] The memory is used to store a program for implementing a voice entity recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing an entity knowledge base and a voice entity recognition model; determining target voice data; and determining entity information in the target voice data through the voice entity recognition model and the entity knowledge base.

[0171] The present application also provides a method for playing a television program, comprising:

[0172] Build a TV program knowledge base;

[0173] Determine the target program name corresponding to the target program playback voice instruction data through the voice entity recognition model and the knowledge base;

[0174] According to the target program name, target program object playback processing is executed.

[0175] Optionally, the knowledge base includes: program-related entities with homophones but different characters, user entities, and entity relationships between program-related entities and user entities;

[0176] The step of constructing a TV program knowledge base includes:

[0177] According to the user's historical playback information, the user entity is determined and the entity relationship is constructed.

[0178] Optionally, performing target program object playback processing according to the target program name includes:

[0179] According to the program list, determine the TV channel and broadcast time corresponding to the target program name;

[0180] Determine a target program object according to the broadcast time and the TV channel;

[0181] A process of playing the target program object is executed.

[0182] Optionally, determining the target program object according to the broadcast time and the TV channel includes:

[0183] displaying a plurality of program objects played at a plurality of times by at least one television channel corresponding to the target program name;

[0184] The program object specified by the user is used as the target program object.

[0185] Optionally, also include:

[0186] If the program table does not include the target program name, determining a program name related to the target program name;

[0187] Display related program names;

[0188] If the user specifies to play the related program object, the process of playing the related program object is executed.

[0189] The present application also provides a method for playing a television program, comprising:

[0190] Smart TVs collect user’s program playback voice command data;

[0191] The voice command data is sent to the server so that the server can build a TV program knowledge base; the target program name corresponding to the voice command data is determined through the voice entity recognition model and the knowledge base; and the target program object playback processing is performed according to the target program name;

[0192] Play the target program object.

[0193] The present application also provides a meeting recording method, including:

[0194] Build a language knowledge base in the conference field;

[0195] Determine entity information in the target conference speech data by using a speech entity recognition model and a language knowledge base in the conference field;

[0196] The text sequence corresponding to the conference voice data is determined through the speech recognition model and the entity information to form a conference record.

[0197] Optionally, also include:

[0198] A conference domain corresponding to the conference voice data is determined.

[0199] The present application also provides a meeting recording method, including:

[0200] Collect voice data of the target meeting;

[0201] The voice data is sent to the server so that the server can build a language knowledge base in the conference field; the entity information in the target conference voice data is determined through the voice entity recognition model and the language knowledge base in the conference field; the text sequence corresponding to the conference voice data is determined through the voice recognition model and the entity information to form a meeting record.

[0202] The present application also provides a television program playing device, comprising:

[0203] A voice data collection unit, used to collect the user's program playback voice command data;

[0204] A voice data sending unit is used to send the voice instruction data to the server so that the server can build a TV program knowledge base; determine the target program name corresponding to the voice instruction data through the voice entity recognition model and the knowledge base; and perform target program object playback processing according to the target program name;

[0205] The program playing unit is used to play the target program object.

[0206] The present application also provides a smart TV, comprising:

[0207] Processor; and

[0208] The memory is used to store a program for implementing a method for playing television programs. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting user's program playing voice instruction data; sending the voice instruction data to a server so that the server can build a television program knowledge base; determining a target program name corresponding to the voice instruction data through a voice entity recognition model and the knowledge base; performing target program object playing processing according to the target program name; and playing the target program object.

[0209] The present application also provides a television program playing device, comprising:

[0210] A knowledge base construction unit, used for constructing a TV program knowledge base;

[0211] An entity recognition unit, used to determine the target program name corresponding to the target program playback voice instruction data through the voice entity recognition model and the knowledge base;

[0212] The play processing unit is used to perform target program object play processing according to the target program name.

[0213] The present application also provides an electronic device, comprising:

[0214] Processor; and

[0215] The memory is used to store a program for implementing a method for playing television programs. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a television program knowledge base; determining a target program name corresponding to target program playing voice instruction data through a voice entity recognition model and the knowledge base; and performing target program object playing processing according to the target program name.

[0216] The present application also provides a conference recording device, comprising:

[0217] A voice data collection unit, used to collect voice data of the target conference;

[0218] A voice data sending unit is used to send the voice data to a server so that the server can build a language knowledge base in the conference field; determine the entity information in the target conference voice data through a voice entity recognition model and the language knowledge base in the conference field; determine the text sequence corresponding to the conference voice data through a voice recognition model and the entity information to form a conference record.

[0219] The present application also provides an electronic device, comprising:

[0220] Processor; and

[0221] The memory is used to store a program for implementing a conference recording method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: collecting voice data of the target conference; sending the voice data to the server so that the server can build a language knowledge base in the conference field; determining entity information in the target conference voice data through a voice entity recognition model and the language knowledge base in the conference field; determining a text sequence corresponding to the conference voice data through the voice recognition model and the entity information to form a conference record.

[0222] The present application also provides a conference recording device, comprising:

[0223] A knowledge base construction unit, used to construct a language knowledge base in the conference field;

[0224] An entity recognition unit, used to determine entity information in the target conference speech data through a speech entity recognition model and a language knowledge base in the conference field;

[0225] The conference record determination unit is used to determine the text sequence corresponding to the conference voice data through the speech recognition model and the entity information to form a conference record.

[0226] The present application also provides an electronic device, comprising:

[0227] Processor; and

[0228] The memory is used to store a program for implementing a conference recording method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a language knowledge base in the conference field; determining entity information in target conference voice data through a speech entity recognition model and the language knowledge base in the conference field; determining a text sequence corresponding to the conference voice data through a speech recognition model and the entity information to form a conference record.

[0229] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned various methods.

[0230] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the above-mentioned various methods.

[0231] Compared with the prior art, this application has the following advantages:

[0232] The multimedia program on-demand system provided in the embodiment of the present application uses the multimedia program on-demand voice data of the intelligent speaker to send the voice data to the server; plays the multimedia program according to the multimedia program playback processing result of the server; the server is used to build a multimedia program knowledge base; determines the multimedia program information in the voice data through the voice entity recognition model and the knowledge base; performs multimedia program playback processing according to the multimedia program information; this processing method introduces the multimedia program knowledge graph information, directly compares whether there is the multimedia program entity pronunciation in the knowledge graph in the multimedia program on-demand voice, realizes semantic understanding and entity recognition from the voice signal, no longer relies on ASR to convert the voice into text, and then recognizes the multimedia program entity name therein through the semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of multimedia program name recognition can be effectively improved, thereby improving the success rate and accuracy of multimedia program on-demand.

[0233] The ordering system provided in the embodiment of the present application collects ordering voice data through the ordering device and sends the voice data to the server; the server builds a meal knowledge base; determines the meal information in the voice data through the voice entity recognition model and the knowledge base; and performs meal preparation processing based on the meal information. This processing method introduces meal knowledge graph information and directly compares whether there is the pronunciation of the meal name in the knowledge graph in the ordering voice, thereby realizing semantic understanding and meal name recognition from the ordering voice signal, and no longer relies on ASR to convert the voice into text, and then recognizes the meal name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of meal name recognition can be effectively improved, thereby improving the success rate and accuracy of ordering.

[0234] The communication connection establishment system provided in the embodiment of the present application collects communication command voice data through user equipment and sends the voice data to the server; the server builds a communication user knowledge base; determines the communication user information in the voice data through the voice entity recognition model and the knowledge base; and executes the communication connection establishment process according to the communication user information. This processing method introduces contact graph information and directly compares whether there is a pronunciation of a person's name in the knowledge graph in the call command voice, thereby realizing semantic understanding and meal name recognition from the call command voice signal, and no longer relies on ASR to convert the voice into text, and then recognizes the person's name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of name recognition can be effectively improved, thereby improving the communication success rate and accuracy.

[0235] The voice interaction system provided in the embodiment of the present application collects voice data through a terminal device and sends the voice data to a server; the server builds an entity knowledge base, and determines the entity information in the voice data through a voice entity recognition model and the entity knowledge base; voice interaction processing is performed based on the entity information; this processing method introduces entity knowledge graph information and directly compares whether there is an entity pronunciation in the knowledge graph in the voice, thereby realizing semantic understanding and entity recognition from the voice signal, and no longer relies on ASR to convert the voice into text, and then recognizes the entity name therein through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of voice entity recognition can be effectively improved, thereby improving the accuracy of voice interaction.

[0236] The television program playback system provided in the embodiment of the present application uses a smart TV to collect the user's program playback voice command data, send the voice command data to the server, and play the target program object; the server is used to build a television program knowledge base; through the voice entity recognition model and the knowledge base, determine the target program name corresponding to the voice command data; according to the target program name, perform the target program object playback processing; this processing method introduces program-related entity graph information, directly compares whether there is a program-related entity pronunciation in the knowledge graph in the program playback voice command, and realizes semantic understanding and recognition of entities such as program names from the program playback voice command signal, no longer relying on ASR to convert speech into text, and then recognizes the program name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of program entity recognition can be effectively improved, thereby improving the success rate and accuracy of television program on demand.

[0237] The conference record system provided in the embodiment of the present application collects voice data of the target conference through a terminal device and sends the voice data to a server; the server builds a language knowledge base in the conference field; determines the entity information in the target conference voice data through a voice entity recognition model and the language knowledge base in the conference field; determines the text sequence corresponding to the conference voice data through the voice recognition model and the entity information to form a conference record; this processing method introduces entity graph information such as conference field-related terms, and directly compares whether there are pronunciations of entities such as conference field-related terms in the knowledge graph in the conference voice data, thereby realizing semantic understanding and recognition of entities such as conference field-related terms from the conference voice data signal, and no longer relies on ASR to convert speech into text, and then recognizes the conference field-related terms therein through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of identifying conference field-related terms can be effectively improved, thereby improving the success rate and accuracy of meeting records. BRIEF DESCRIPTION OF THE DRAWINGS

[0238] Figure 1 A schematic diagram of the structure of an embodiment of a multimedia program on-demand system provided by the present application;

[0239] Figure 2 A schematic diagram of a scenario of an embodiment of a multimedia program on-demand system provided by the present application;

[0240] Figure 3 A schematic diagram of device interaction in an embodiment of a multimedia program on-demand system provided by the present application;

[0241] Figure 4 A schematic diagram of the system architecture of an embodiment of a multimedia program on-demand system provided by the present application;

[0242] Figure 5 A schematic diagram of a voice entity recognition model of an embodiment of a multimedia program on-demand system provided by the present application;

[0243] Figure 6 A schematic diagram of a knowledge graph of an embodiment of a multimedia program on-demand system provided by the present application;

[0244] Figure 7 The present application provides another schematic diagram of a voice entity recognition model of an embodiment of a multimedia program on-demand system. DETAILED DESCRIPTION

[0245] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0246] In the present application, a multimedia program on-demand system, method and device, a food ordering system, method and device, a communication connection establishment system, method and device, a voice interaction system, method and device, a voice entity recognition model construction method and device, an entity knowledge base construction method and device, a TV program on-demand method and device, a conference recording method and device, a smart speaker, a smart TV, a food ordering machine, a user device, and an electronic device are provided. Each solution is described in detail in the following embodiments.

[0247] First embodiment

[0248] Please refer to Figure 1 , which is a schematic diagram of an embodiment of a multimedia program on-demand system of the present application. The multimedia program on-demand system provided in this embodiment includes: a server 1 and an intelligent speaker 2.

[0249] The server 1 may be a server deployed on a cloud server, or may be a server dedicated to implementing a multimedia program on-demand system, and may be deployed in a data center.

[0250] Smart Speaker 2 can be a tool for home consumers to surf the Internet using voice, such as ordering songs, shopping online, or getting weather forecasts. It can also control smart home devices, such as opening curtains, setting refrigerator temperature, and heating up water heaters in advance.

[0251] Please refer to Figure 2 , which is a scene diagram of the multimedia program on-demand system of the present application. The server 1 and the smart speaker 2 can be connected through a network, such as the smart speaker 2 can be connected to the network through WIFI, etc. The user interacts with the smart speaker through voice. In this embodiment, the user issues a song-ordering voice command to the smart speaker 2, and the server determines the song name information in the song-ordering voice command through the voice entity recognition model and the pre-built multimedia program knowledge base; and executes the process of playing the song.

[0252] Please refer to Figure 3 , which is a device schematic diagram of the multimedia program on-demand system of the present application. In this embodiment, the smart speaker is used to collect multimedia program on-demand voice data and send the voice data to the server; the server is used to build a multimedia program knowledge base; through the voice entity recognition model and the knowledge base, the multimedia program information in the voice data is determined; according to the multimedia program information, the multimedia program playback process is performed.

[0253] The multimedia program may be a song, a movie, a TV program, a speech video, etc. The multimedia program knowledge base includes program-related entity information. The program-related entity information includes but is not limited to at least one of the following entities: program name (such as song name, movie name, TV program name, speaker name, etc.), program-related person name (such as singer name, movie director name, etc.), etc.

[0254] The server may construct a multimedia program knowledge base by: determining multimedia program related entities to form a multimedia program knowledge base. Table 1 shows the content of the multimedia program knowledge base in this embodiment.

[0255]

[0256]

[0257] Table 1. Entity data

[0258] Depend on Figure 4It can be seen that on the smart speaker side, the analog voice signal is converted into a digital voice signal after passing through the receiving sensor and the corresponding digital signal processing unit of the smart speaker; this part can also include a part of acoustic processing to produce corresponding voice intermediate results, which may include but are not limited to sound spectrum signals, phonemes, characters, character fragments, pinyin, etc. The smart speaker can upload the digital voice signal to the server, and the server models the digital voice signal through the voice entity recognition model, and at the same time combines the entity information in the knowledge graph (i.e., knowledge base) to identify the entity name mentioned in the voice (such as the song name, etc.).

[0259] In specific implementation, the speech entity recognition model can be a machine learning model (such as a deep neural network model, a Bayesian model, etc.) or a non-machine learning model (such as a heuristic model, etc.).

[0260] The voice entity recognition model can directly perform semantic understanding and entity recognition from voice signals, and no longer rely on voice recognition ASR to convert voice into text, and then recognize the entity name in the text through semantic understanding of the text. By introducing the knowledge graph, it is possible to directly compare whether there is an entity pronunciation in the knowledge graph in the voice, which is closer to the process of human understanding of voice, obtain more accurate entity recognition results, and effectively solve the problems of user inaccurate pronunciation, unclear pronunciation, and multiple words with the same pronunciation.

[0261] The following uses several application scenario examples to illustrate specific problems that can be solved by the system provided in the embodiments of the present application:

[0262] Scenario 1:

[0263] User Xiao Zhao often uses smart speakers to order songs while driving. Due to local accent problems, front nasal / back nasal sounds cannot be distinguished, and rolled tongue sounds / curled tongue sounds cannot be distinguished, which often causes the smart speakers in traditional voice entity recognition systems to execute wrong commands. For example, the Chinese pinyin corresponding to Xiao Zhao's voice command is "wo xiang ting bie zi ji", and the smart speaker recognizes it as the text "I want to listen to other myself" through ASR. Finally, "other myself" will be recognized as the name of a song. However, there is no such song "other myself". The song Xiao Zhao really wants to listen to is "other confidant". Using the system provided in the embodiment of the present application, since the knowledge base includes the song name "other confidant", the song can be accurately identified and played correctly.

[0264] Scenario 2:

[0265] The user Doudou is a 5-year-old child who often uses the smart speaker with a screen at home to watch cartoons. Since Doudou is too young to type, he can only search by voice. However, Doudou's pronunciation is not very clear, and he often fails to place orders on demand, and needs his parents to help him place orders. For example, after Doudou issues a voice order on demand, the smart speaker recognizes it as the text "I want to watch that station" through ASR. In fact, what Doudou wants to express is "I want to watch Nezha". Using the system provided in the embodiment of the present application, since the movie name "Nezha" is included in the knowledge base, the movie can be accurately identified and played correctly.

[0266] Scenario 3:

[0267] When Xiao Wang is cleaning the house, he wants to listen to some music to relieve his fatigue. He orders a song to the smart speaker at home through voice. The pinyin corresponding to the voice command is "wo yao ting lei yu xinde ji nian". The smart speaker device recognizes it as the text "I want to listen to Lei Yuxin's Memorial" through ASR. In fact, Xiao Wang wants to play "Memorial" sung by Lei Yuxin instead of "Memorial". Using the system provided by the embodiment of the present application, since the knowledge base includes the song name "Memorial", the song can be accurately identified and played correctly.

[0268] Please refer to Figure 5 , which is a schematic diagram of the voice entity recognition model of the multimedia program on-demand system of the present application. In one example, the server side needs to determine the entity information in the target voice data through the voice entity recognition model and the knowledge base, and can adopt the following processing process: first, determine the audio feature data of the voice data through the audio encoding model included in the voice entity recognition model; then, determine the multimedia program information according to the audio feature data through the entity decoding model included in the voice entity recognition model and the knowledge base.

[0269] The input data of the audio coding model may include an audio frame sequence, and each audio frame may include acoustic feature data of an audio signal, etc. The output data of the entity decoding model may include pinyin, but in practical applications, it may also be any entity representation form, such as text, phonemes, etc.

[0270] In specific implementation, the original audio signal can be processed by an audio signal processing module (Mel filter, etc.) to form an audio frame sequence. The audio coding model can adopt a commonly used coding model, and the model structure includes but is not limited to LSTM, Transformer, etc.; the network structure of the entity decoding model includes but is not limited to LSTM, Transformer, etc.

[0271] In one example, at least one candidate pronunciation of the multimedia program information is determined according to the audio feature data by the entity candidate pronunciation determination module included in the entity decoding model; the pronunciation of the multimedia program information is determined from the at least one candidate pronunciation according to the knowledge base by the entity pronunciation determination module included in the entity decoding model; the multimedia program information is determined according to the pronunciation of the multimedia program information. The input of the model is an audio signal (audio frame sequence), and the output is the entity name contained in the audio (such as a pinyin sequence). In this model, the pronunciation that will cause confusion will generate candidates for all possible pronunciations in the decoding model, and then the more likely pronunciation will be determined through the entity knowledge base. In specific implementation, the decoding part can automatically find the appropriate position of the encoding network to generate candidate pronunciations.

[0272] In a specific implementation, the method of determining the pronunciation of the multimedia program information from the at least one candidate pronunciation according to the knowledge base may include the following sub-steps: 1) determining the similarity between the pronunciation of the entity in the knowledge base and the candidate pronunciation; 2) determining the pronunciation of the multimedia program information according to the similarity, such as taking the candidate pronunciation with a high similarity ranking and a similarity greater than a similarity threshold as the pronunciation of the multimedia program information. The similarity calculation part may be a neural network automatic learning.

[0273] During specific implementation, the server can also be used to learn the voice entity recognition model from training data; wherein the training data includes: audio data and multimedia program annotation information.

[0274] In specific implementation, the server can also be used to determine the entity knowledge base and the training data set to construct the network structure of the model; use the voice data as the input data of the model, and use the multimedia program annotation information as the output data of the model, and train the network parameters of the model according to the knowledge base.

[0275] In one example, if the knowledge base includes program-related entities with homophones but different characters, such as the song titles of two songs, "Memorial" and "Remember", the knowledge base may also include: a user entity, and an entity relationship between the program-related entity and the user entity. Table 2 shows the entity relationship between the program-related entity and the user entity in this embodiment.

[0276]

[0277]

[0278] Table 2. Entity relationship data

[0279] As can be seen from Table 2, the entity relationship can be not only the corresponding relationship between the user entity and the program entity, but also the corresponding relationship between the user entity and the name of the person related to the program. Table 3 shows the user entity information in this embodiment.

[0280] User entity identifier User account (Taobao account) 1 Abdgdf001 2 Hanhao55 …

[0281] Table 3. User entity data

[0282] In one example, the entity relationship between the program-related entity and the user entity can be determined in the following manner: according to the user's historical playback information, the user entity is determined, and the entity relationship is constructed. Table 4 shows the user's historical playback information in this embodiment.

[0283] Playback record mark User Entity Program-related entities time 1 Zhang San Commemorate 20200526 2 Zhang San Looking back 20200315 … 23 Li Si Jay Chou 20190230 24 Li Si Half pot yarn 20200511 …

[0284] Table 4. User historical playback data

[0285] Please refer to Figure 6 , which is a schematic diagram of the knowledge graph of the multimedia program on-demand system of the present application. In one example, the knowledge base also includes user-type entities, and user entities may have entity relationships with some program-related entities, but not with other program-related entities. For example, if user A has requested the song "Memorial", then it has an association relationship with "Memorial" but not with "Commemoration". Figure 6 It can be seen that the knowledge graph describes entities (points) and the relationships (edges) between entities. The user who issues the audio command is an entity in the knowledge graph, and the song is also an entity in the graph. The user's historical listening records can be used to determine songs with the same pronunciation. In addition, the co-occurrence relationship of songs in the playlist can also be used to determine songs with the same pronunciation.

[0286] Accordingly, the server needs to determine the multimedia program information according to the pronunciation of the multimedia program information, which may include the following sub-steps: 1) determine the candidate entities according to the pronunciation of the multimedia program information, such as "ji nian" corresponds to the two candidate entities "ji niang" and "ji niang"; 2) determine the multimedia program information from the candidate entities according to the user information and the entity relationship. For example, it is finally determined that user A requested the song "ji niang" instead of "ji niang".

[0287] Please refer to Figure 7 , which is another schematic diagram of a voice entity recognition model of the multimedia program on-demand system of the present application. In one example, the server is specifically used to determine the pronunciation feature data of the entity in the knowledge base; through the entity decoding model included in the voice entity recognition model, the multimedia program information is determined according to the audio feature data and the entity pronunciation feature data.

[0288] Depend on Figure 7 It can be seen that the audio encoding model included in the speech entity recognition model can be used to calculate the audio feature vector of the input audio signal (audio frame sequence); the entity encoding model included in the speech entity recognition model can be used to calculate the feature vector of the entity name (such as a pinyin sequence); the speech entity recognition model can make the audio feature vector and the feature vector of the entity name it contains closer. Figure 7 In the model shown, the feature vector is not sensitive to confused sounds, that is, although some sounds are mispronounced, the feature vector calculated from the audio is still relatively close.

[0289] In a specific implementation, the server side shall determine the multimedia program information according to the audio feature data and the entity pronunciation feature data through the entity decoding model included in the voice entity recognition model, which may include the following sub-steps: 1) determining the pronunciation similarity between the entity in the voice data and the entity in the knowledge base according to the audio feature data and the entity pronunciation feature data; 2) determining the multimedia program information according to the pronunciation similarity, such as taking the knowledge base entity with a high similarity ranking whose similarity is greater than a similarity threshold as the multimedia program information.

[0290] The audio encoding model and entity decoding model may adopt commonly used encoding models, and the model structure includes but is not limited to LSTM, Transformer, etc. The entity decoding model may include a distance function, and the pronunciation similarity between the entity in the speech data and the entity in the knowledge base may be determined by the distance function. The distance function may be a dot product, Euclidean distance, cosine distance, etc.

[0291] In specific implementation, the server can also be used to learn the voice entity recognition model and the entity pronunciation feature data from the training data; wherein the training data includes: audio data, entity annotation information and knowledge base. Figure 7 The model shown can calculate and store the entity's pronunciation feature vector offline, and then match it with the audio feature vector online.

[0292] In one example, the knowledge base includes: program-related entities with the same pronunciation but different characters, user entities, and entity relationships between program-related entities and user entities; the server side may construct an entity knowledge base by including the following steps: determining the user entity based on the user's historical playback information, and constructing the entity relationship; and specifically determining candidate entities based on the pronunciation similarity; and determining the entity information from the candidate entities based on the user information and the entity relationship.

[0293] It can be seen from the above embodiments that the multimedia program on-demand system provided by the embodiments of the present application, through the intelligent speaker multimedia program on-demand voice data, sends the voice data to the server; plays the multimedia program according to the multimedia program playback processing results of the server; the server is used to build a multimedia program knowledge base; determines the multimedia program information in the voice data through the voice entity recognition model and the knowledge base; performs multimedia program playback processing according to the multimedia program information; this processing method introduces the multimedia program knowledge graph information, directly compares whether there is a multimedia program entity pronunciation in the knowledge graph in the multimedia program on-demand voice, and realizes semantic understanding and entity recognition from the voice signal, no longer relying on ASR to convert the voice into text, and then identifies the multimedia program entity name therein through the semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of multimedia program name recognition can be effectively improved, thereby improving the success rate and accuracy of multimedia program on-demand.

[0294] Second embodiment

[0295] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a voice interaction method, the execution subject of which can be a server, or a smart speaker, a smart TV, a vending machine, a ticket machine, a chat robot, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0296] The voice interaction method provided in the embodiment of the present application may include the following steps:

[0297] Step 1: Build an entity knowledge base;

[0298] Step 2: Determine entity information in the target speech data through the speech entity recognition model and the entity knowledge base;

[0299] Step 3: Perform voice interaction processing based on the entity information.

[0300] In one example, step 2 may include the following sub-steps:

[0301] Step 2-1: determining audio feature data of the speech data by using the audio coding model included in the speech entity recognition model;

[0302] Step 2-2: Determine the entity information according to the audio feature data through the entity decoding model included in the speech entity recognition model and the entity knowledge base.

[0303] In one example, step 2-2 may include the following sub-steps:

[0304] Step 2-2-1: determining at least one candidate pronunciation of the entity information according to the audio feature data through the entity candidate pronunciation determination module included in the entity decoding model;

[0305] Step 2-2-2: determining the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base through the entity pronunciation determination module included in the entity decoding model;

[0306] Step 2-2-3: Determine the entity information according to the pronunciation of the entity information.

[0307] In one example, step 2-2-2 may include the following sub-steps:

[0308] Step 2-2-2-1: Determine the similarity between the pronunciation of the entity in the entity knowledge base and the candidate pronunciation;

[0309] Step 2-2-2-2: Determine the pronunciation of the entity information based on the similarity.

[0310] In one example, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand; the program entity knowledge base includes: program-related entities with homophones but different characters, user entities, and entity relationships between program-related entities and user entities; step 1 may include the following sub-steps: determining the user entity according to the user's historical playback information, and constructing the entity relationship; correspondingly, step 2-2-3 may include the following sub-steps:

[0311] Step 2-2-3-1: Determine a candidate entity according to the pronunciation of the entity information;

[0312] Step 2-2-3-2: Determine the entity information from the candidate entities based on the user information and the entity relationship.

[0313] In one example, the method may further include the following steps: learning the speech entity recognition model from training data; wherein the training data includes: audio data and entity annotation information.

[0314] In another example, step 2-2 may include the following sub-steps:

[0315] Step 2-2-1': Determine the audio feature data of the speech data through the audio coding model included in the speech entity recognition model;

[0316] Step 2-2-2': Determine the pronunciation feature data of the entity in the entity knowledge base through the entity encoding model included in the speech entity recognition model;

[0317] Step 2-2-3': Determine the entity information according to the audio feature data and the entity pronunciation feature data through the entity decoding model included in the speech entity recognition model.

[0318] In one example, step 2-2-3' may include the following sub-steps:

[0319] Step 2-2-3'-1: Determine the pronunciation similarity between the entity in the speech data and the entity in the entity knowledge base based on the audio feature data and the entity pronunciation feature data;

[0320] Step 2-2-3'-2: Determine the entity information based on the pronunciation similarity.

[0321] In one example, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand; the program entity knowledge base includes: program-related entities with homophones but different characters, user entities, and entity relationships between program-related entities and user entities; step 1 may include the following sub-steps: determining the user entity based on the user's historical playback information, and constructing the entity relationship; accordingly, step 2-2-3'-2: may include the following sub-steps:

[0322] Step 2-2-3'-2-1: Determine a candidate entity based on the pronunciation similarity;

[0323] Step 2-2-3'-2-2: Determine the entity information from the candidate entities based on the user information and the entity relationship.

[0324] In one example, the method may further include the following steps: learning the speech entity recognition model from training data; wherein the training data includes: audio data, an entity knowledge base, and entity annotation information.

[0325] In one example, the entity knowledge base includes: a program entity knowledge base in the field of multimedia program on demand; step 1 may include the following sub-steps: determining multimedia program-related entities to form the entity knowledge base.

[0326] Third embodiment

[0327] In the above embodiment, a voice interaction method is provided, and correspondingly, the present application also provides a voice interaction device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0328] A voice interaction device provided by the present application includes:

[0329] A knowledge base construction unit, used to construct an entity knowledge base;

[0330] An entity determination unit, used to determine entity information in the target speech data through a speech entity recognition model and the entity knowledge base;

[0331] An interaction processing unit is used to perform voice interaction processing according to the entity information.

[0332] Fourth embodiment

[0333] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0334] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing an entity knowledge base; determining entity information in target voice data through a voice entity recognition model and the entity knowledge base; and performing voice interaction processing based on the entity information.

[0335] The electronic device may be a smart speaker, a smart TV, a food ordering machine, a vending machine, a ticket machine, a chat robot, and the like.

[0336] Fifth embodiment

[0337] In the above embodiment, a voice interaction system is provided. Correspondingly, the present application also provides a voice interaction method, and the execution subject of the method can be a smart speaker, a smart TV, a vending machine, a ticket machine, a chat robot, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the second embodiment are not repeated here, please refer to the corresponding parts in the second embodiment.

[0338] A voice interaction method provided in the present application may include the following steps: collecting voice data and sending the voice data to a server so that the server builds an entity knowledge base; determining entity information in the voice data through a voice entity recognition model and the entity knowledge base; and performing voice interaction processing based on the entity information.

[0339] Sixth embodiment

[0340] In the above embodiment, a voice interaction method is provided, and correspondingly, the present application also provides a voice interaction device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0341] A voice interaction device provided by the present application includes:

[0342] A voice data collection unit, used for collecting voice data;

[0343] The voice data sending unit is used to send the voice data to the server so that the server builds an entity knowledge base; determine the entity information in the voice data through the voice entity recognition model and the entity knowledge base; and perform voice interaction processing according to the entity information.

[0344] Seventh embodiment

[0345] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0346] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: voice data is collected and sent to a server so that the server builds an entity knowledge base; entity information in the voice data is determined through a voice entity recognition model and the entity knowledge base; and voice interaction processing is performed according to the entity information.

[0347] Eighth embodiment

[0348] In the above embodiment, a multimedia program on-demand system is provided, and correspondingly, the present application also provides a voice interaction system. The system corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0349] The present application provides a voice interaction system including: a terminal device and a server.

[0350] Among them, the terminal device is used to collect voice data and send the voice data to the server; the server is used to build an entity knowledge base; through the voice entity recognition model and the entity knowledge base, the entity information in the voice data is determined; according to the entity information, voice interaction processing is performed.

[0351] Ninth embodiment

[0352] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a multimedia program on-demand method, and the execution subject of the method can be a smart speaker, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0353] A multimedia program on-demand method provided in the present application may include the following steps: collecting multimedia program on-demand voice data, sending the voice data to a server, so that the server builds a multimedia program knowledge base; determining multimedia program information in the voice data through a voice entity recognition model and the knowledge base; and performing multimedia program playback processing according to the multimedia program information.

[0354] Tenth embodiment

[0355] In the above embodiment, a multimedia program on-demand method is provided, and correspondingly, the present application also provides a multimedia program on-demand device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0356] A multimedia program on-demand device provided by the present application includes:

[0357] A voice data collection unit, used to collect voice data of multimedia program on-demand;

[0358] The voice data sending unit is used to send the voice data to the server so that the server builds a multimedia program knowledge base; determine the multimedia program information in the voice data through the voice entity recognition model and the knowledge base; and perform multimedia program playback processing according to the multimedia program information.

[0359] Eleventh Embodiment

[0360] The present application also provides a smart speaker. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0361] A smart speaker in this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a multimedia program on-demand method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting multimedia program on-demand voice data, and sending the voice data to a server so that the server builds a multimedia program knowledge base; determining multimedia program information in the voice data through a voice entity recognition model and the knowledge base; and performing multimedia program playback processing according to the multimedia program information.

[0362] Twelfth Embodiment

[0363] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a multimedia program on-demand method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0364] The present application provides a multimedia program on-demand method, which may include the following steps:

[0365] Step 1: Build a multimedia program knowledge base;

[0366] Step 2: Determine the multimedia program information in the multimedia program on-demand voice data through the voice entity recognition model and the knowledge base;

[0367] Step 3: Execute multimedia program playing processing according to the multimedia program information.

[0368] Thirteenth Embodiment

[0369] In the above embodiment, a multimedia program on-demand method is provided, and correspondingly, the present application also provides a multimedia program on-demand device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0370] A multimedia program on-demand device provided by the present application includes:

[0371] A knowledge base construction unit, used for constructing a multimedia program knowledge base;

[0372] An entity recognition unit, used to determine multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base;

[0373] The program playing processing unit is used to perform multimedia program playing processing according to the multimedia program information.

[0374] Fourteenth Embodiment

[0375] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0376] An electronic device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a multimedia program on-demand method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: constructing a multimedia program knowledge base; determining multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base; and performing multimedia program playback processing according to the multimedia program information.

[0377] Fifteenth Embodiment

[0378] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a method for constructing a voice entity recognition model, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0379] A method for constructing a speech entity recognition model provided in this application may include the following steps:

[0380] Step 1: Determine a training data set, where the training data includes: speech data, entity annotation information, and an entity knowledge base;

[0381] Step 2: Construct the network structure of the model;

[0382] Step 3: Learn the model from the training data set.

[0383] In one example, the model includes an audio encoding model for determining audio feature data of the speech data; the model includes an entity decoding model for determining entity information in the speech data based on the audio feature data and the entity knowledge base.

[0384] In another example, the model includes an audio encoding model for determining audio feature data of the speech data; the model includes an entity encoding model for determining pronunciation feature data of entities in the entity knowledge base; the model includes an entity decoding model for determining entity information in the speech data based on the audio feature data and the entity pronunciation feature data.

[0385] It can be seen from the above embodiments that the method for constructing a speech entity recognition model provided in the embodiments of the present application determines a training data set, wherein the training data includes: speech data, entity annotation information, and an entity knowledge base; constructs a network structure of the model; and learns the model from the training data set. This processing method introduces entity knowledge graph information and directly compares whether there is an entity pronunciation in the knowledge graph in the speech, thereby realizing semantic understanding and entity recognition from the speech signal. It no longer relies on ASR to convert speech into text, and then identifies the entity name therein through semantic understanding of the text, which is closer to the process of human understanding of speech. Therefore, the accuracy of the speech entity recognition model can be effectively improved.

[0386] Sixteenth Embodiment

[0387] In the above embodiment, a multimedia program on-demand method is provided, and correspondingly, the present application also provides a multimedia program on-demand device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0388] A multimedia program on-demand device provided by the present application includes:

[0389] A training data determination unit, used to determine a training data set, wherein the training data includes: speech data, entity annotation information, and an entity knowledge base;

[0390] A network construction unit, used to construct the network structure of the model;

[0391] The model training unit is used to learn the model from the training data set.

[0392] Seventeenth Embodiment

[0393] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0394] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for building a voice entity recognition model. After the device is powered on and runs the program of the method through the processor, the following steps are performed: determining a training data set, the training data including: voice data, entity annotation information and an entity knowledge base; building a network structure of the model; and learning the model from the training data set.

[0395] Eighteenth Embodiment

[0396] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a method for constructing an entity knowledge base, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0397] The present application provides a method for constructing an entity knowledge base, which may include the following steps:

[0398] Step 1: Get the entity name of the target domain;

[0399] Step 2: Generate an entity knowledge base of the target domain according to the entity name, and the entity knowledge base is used to determine the entity information in the speech data of the target domain through the speech entity recognition model and the entity knowledge base.

[0400] The target fields include but are not limited to: multimedia program on demand, food ordering, communications, and the like.

[0401] It can be seen from the above embodiments that the entity knowledge base construction method provided in the embodiments of the present application obtains the entity name of the target field; based on the entity name, an entity knowledge base of the target field is generated, and the entity knowledge base is used to determine the entity information in the speech data of the target field through the speech entity recognition model and the entity knowledge base; this processing method makes it possible to construct entity knowledge graph information, laying a good data foundation for constructing a speech entity recognition model. The model can introduce entity knowledge graph information, directly compare whether there is an entity pronunciation in the knowledge graph in the speech, and realize semantic understanding and entity recognition from the speech signal, no longer relying on ASR to convert speech into text, and then identify the entity name therein through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of the speech entity recognition model can be effectively improved.

[0402] Nineteenth Embodiment

[0403] In the above embodiment, a method for constructing an entity knowledge base is provided. Correspondingly, the present application also provides an entity knowledge base construction device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.

[0404] The present application provides an entity knowledge base construction device comprising:

[0405] An entity determination unit, used to obtain the entity name of the target domain;

[0406] The knowledge base generation unit is used to generate an entity knowledge base in a target domain, and the entity knowledge base is used to determine entity information in speech data in the target domain through a speech entity recognition model and the entity knowledge base.

[0407] Twentieth Embodiment

[0408] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0409] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for building an entity knowledge base. After the device is powered on and the program of the method is run through the processor, the following steps are performed: obtaining an entity name of a target domain; generating an entity knowledge base of the target domain based on the entity name, wherein the entity knowledge base is used to determine entity information in speech data of the target domain through a speech entity recognition model and the entity knowledge base.

[0410] Twenty-first embodiment

[0411] In the above embodiment, a multimedia program on-demand system is provided. Correspondingly, the present application also provides a voice entity recognition method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0412] A method for voice entity recognition provided by the present application may include the following steps:

[0413] Step 1: Build an entity knowledge base and voice entity recognition model;

[0414] Step 2: Determine the target voice data;

[0415] Step 3: Determine entity information in the target speech data through the speech entity recognition model and the entity knowledge base.

[0416] It can be seen from the above embodiments that the speech entity recognition method provided in the embodiments of the present application constructs an entity knowledge base and a speech entity recognition model; determines the target speech data; and determines the entity information in the target speech data through the speech entity recognition model and the entity knowledge base. This processing method makes it possible to construct entity knowledge graph information, laying a good data foundation for constructing a speech entity recognition model. The model can introduce entity knowledge graph information, directly compare whether there is an entity pronunciation in the knowledge graph in the speech, and realize semantic understanding and entity recognition from the speech signal. It no longer relies on ASR to convert speech into text, and then recognizes the entity name therein through semantic understanding of the text. This is closer to the process of human understanding of speech. Therefore, the accuracy of speech entity recognition can be effectively improved.

[0417] Twenty-second embodiment

[0418] In the above embodiment, a method for speech entity recognition is provided, and correspondingly, the present application also provides a device for speech entity recognition. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as those of the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0419] A speech entity recognition device provided by the present application includes:

[0420] Model building unit, used for speech entity recognition model;

[0421] A knowledge base construction unit, used to construct an entity knowledge base;

[0422] A voice data determination unit, used to determine target voice data;

[0423] The entity determination unit is used to determine the entity information in the target speech data through the speech entity recognition model and the entity knowledge base.

[0424] Twenty-third embodiment

[0425] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0426] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice entity recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing an entity knowledge base and a voice entity recognition model; determining target voice data; and determining entity information in the target voice data through the voice entity recognition model and the entity knowledge base.

[0427] Twenty-fourth embodiment

[0428] In the above embodiment, a multimedia program on-demand system is provided, and correspondingly, the present application also provides a meal ordering system. The system corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0429] The present application provides a food ordering system comprising: food ordering equipment and a server.

[0430] Among them, the ordering device is used to collect ordering voice data and send the voice data to the server; the server is used to build a meal knowledge base; through the voice entity recognition model and the knowledge base, the meal information in the voice data is determined; according to the meal information, the meal preparation process is performed.

[0431] The food knowledge base may include various dish names, such as New Orleans hamburgers, French fries, latte coffee, etc.

[0432] The meal preparation process may include sending the order information to a back kitchen, or sending the order information to a front desk salesperson, and so on.

[0433] For example, a user orders food by voice through an ordering device and sends a voice command to the ordering device, "a New Orleans hamburger and a latte". However, due to the user's slurred speech, the accurate text cannot be recognized by traditional voice recognition algorithms, and other food names may be recognized, or it may be impossible to determine what kind of food the user ordered. The system provided in the embodiment of the present application can accurately determine the food information directly based on the user's voice ordering command through a voice entity recognition model and a food knowledge base.

[0434] It can be seen from the above embodiments that the ordering system provided in the embodiments of the present application collects ordering voice data through the ordering device and sends the voice data to the server; the server builds a meal knowledge base; determines the meal information in the voice data through the voice entity recognition model and the knowledge base; and performs meal preparation processing based on the meal information. This processing method introduces meal knowledge graph information and directly compares whether there is the pronunciation of the meal name in the knowledge graph in the ordering voice, so as to realize semantic understanding and meal name recognition from the ordering voice signal, and no longer relies on ASR to convert the voice into text, and then recognizes the meal name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of meal name recognition can be effectively improved, thereby improving the success rate and accuracy of ordering.

[0435] Twenty-fifth embodiment

[0436] In the above embodiment, a meal ordering system is provided. Correspondingly, the present application also provides a meal ordering method, and the execution subject of the method may be a meal ordering device, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment will not be repeated, and please refer to the corresponding parts in the first embodiment.

[0437] A method for ordering food provided in the present application may include the following steps: collecting ordering voice data, and sending the voice data to a server so that the server builds a food knowledge base; determining food information in the voice data through a voice entity recognition model and the knowledge base; and performing food preparation processing based on the food information.

[0438] Twenty-sixth embodiment

[0439] In the above embodiment, a method for ordering food is provided, and correspondingly, the present application also provides a device for ordering food. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0440] A meal ordering device provided in the present application comprises:

[0441] A voice data collection unit, used to collect ordering voice data;

[0442] A voice data sending unit is used to send the voice data to a server so that the server can build a meal knowledge base; determine the meal information in the voice data through a voice entity recognition model and the knowledge base; and perform meal preparation processing according to the meal information.

[0443] Twenty-seventh embodiment

[0444] The present application also provides a food ordering device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0445] A food ordering device according to the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a food ordering method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting food ordering voice data, and sending the voice data to a server so that the server builds a food knowledge base; determining food information in the voice data through a voice entity recognition model and the knowledge base; and performing food preparation processing according to the food information.

[0446] Twenty-eighth Embodiment

[0447] In the above embodiment, a meal ordering system is provided. Correspondingly, the present application also provides a meal ordering method, and the execution subject of the method may be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0448] A method for ordering food provided by the present application may include the following steps:

[0449] Step 1: Build a food knowledge base;

[0450] Step 2: Determine the food information in the ordering voice data through the voice entity recognition model and the entity knowledge base;

[0451] Step 3: Prepare the meal according to the meal information.

[0452] Twenty-ninth embodiment

[0453] In the above embodiment, a method for ordering food is provided, and correspondingly, the present application also provides a device for ordering food. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0454] A meal ordering device provided in the present application comprises:

[0455] A knowledge base construction unit, used to construct a food knowledge base;

[0456] An entity recognition unit, used to determine the food information in the ordering voice data through a voice entity recognition model and the entity knowledge base;

[0457] The meal preparation processing unit is used to perform meal preparation processing according to the meal information.

[0458] Thirtieth Embodiment

[0459] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0460] An electronic device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for ordering food. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a food knowledge base; determining food information in the food ordering voice data through a voice entity recognition model and the entity knowledge base; and performing food preparation processing based on the food information.

[0461] Thirty-first embodiment

[0462] In the above embodiment, a multimedia program on-demand system is provided, and correspondingly, the present application also provides a communication connection establishment system. The system corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0463] The present application provides a communication connection establishment system including: a user device and a server.

[0464] Among them, the user equipment is used to collect communication command voice data and send the voice data to the server; the server is used to build a communication user knowledge base; through the voice entity recognition model and the knowledge base, the communication user information in the voice data is determined; according to the communication user information, the communication connection establishment process is executed.

[0465] The communication user knowledge base may include user entity information such as contact names.

[0466] The user device may be a mobile communication device such as a mobile phone, a smart phone, or a smart speaker.

[0467] The communication connection establishment process may be a process such as dialing a communication user's phone number.

[0468] For example, a user sends a voice command to a smartphone to call someone, "Call Zihao", but because the user speaks unclearly, the accurate name cannot be recognized by the traditional voice recognition algorithm, and it is possible that another name is recognized. However, the system provided by the embodiment of the present application can accurately determine the name information directly based on the user's voice command to make a call through the voice entity recognition model and the communication user knowledge base.

[0469] In one example, the communication user knowledge base may include "Zihao" and "Zihao", and user A wants to call "Zihao". The correspondence between user A and "Zihao" can be stored in the communication user knowledge base, so that it will not be recognized as "Zihao", thereby effectively improving the call accuracy.

[0470] It can be seen from the above embodiments that the communication connection establishment system provided in the embodiments of the present application collects communication command voice data through user equipment and sends the voice data to the server; the server builds a communication user knowledge base; determines the communication user information in the voice data through the voice entity recognition model and the knowledge base; and executes the communication connection establishment process according to the communication user information; this processing method introduces contact graph information and directly compares whether there is a pronunciation of a person's name in the knowledge graph in the call command voice, thereby realizing semantic understanding and meal name recognition from the call command voice signal, and no longer relies on ASR to convert the voice into text, and then recognizes the person's name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of name recognition can be effectively improved, thereby improving the communication success rate and accuracy.

[0471] Thirty-second embodiment

[0472] In the above embodiment, a communication connection establishment system is provided. Correspondingly, the present application also provides a communication connection establishment method, and the execution subject of the method can be a user device, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0473] A method for establishing a communication connection provided in the present application may include the following steps: collecting communication instruction voice data, and sending the voice data to a server so that the server builds a communication user knowledge base; determining the communication user information in the voice data through a voice entity recognition model and the knowledge base; and executing a communication connection establishment process based on the communication user information.

[0474] Thirty-third embodiment

[0475] In the above embodiment, a communication connection establishment method is provided, and correspondingly, the present application also provides a communication connection establishment device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0476] A communication connection establishment device provided by the present application includes:

[0477] A voice data collection unit, used for collecting communication command voice data;

[0478] The voice data sending unit is used to send the voice data to the server so that the server builds a communication user knowledge base; determine the communication user information in the voice data through the voice entity recognition model and the knowledge base; and perform communication connection establishment processing according to the communication user information.

[0479] Thirty-fourth embodiment

[0480] The present application also provides a user equipment. Since the equipment embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The equipment embodiment described below is only illustrative.

[0481] A user device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a communication connection establishment method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: communication instruction voice data is collected and the voice data is sent to a server so that the server builds a communication user knowledge base; communication user information in the voice data is determined through a voice entity recognition model and the knowledge base; and communication connection establishment processing is performed according to the communication user information.

[0482] Thirty-fifth embodiment

[0483] In the above embodiment, a communication connection establishment system is provided. Correspondingly, the present application also provides a communication connection establishment method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0484] A communication connection establishment method provided by the present application may include the following steps:

[0485] Step 1: Build a communication user knowledge base;

[0486] Step 2: Determine the communication user information in the communication command voice data through the voice entity recognition model and the knowledge base;

[0487] Step 3: Execute communication connection establishment processing according to the communication user information.

[0488] Thirty-sixth embodiment

[0489] In the above embodiment, a communication connection establishment method is provided, and correspondingly, the present application also provides a communication connection establishment device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0490] A communication connection establishment device provided by the present application includes:

[0491] A knowledge base construction unit, used to construct a communication user knowledge base;

[0492] An entity recognition unit, used to determine the communication user information in the communication instruction voice data through the voice entity recognition model and the knowledge base;

[0493] The communication connection processing unit is used to execute communication connection establishment processing according to the communication user information.

[0494] Thirty-seventh embodiment

[0495] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0496] An electronic device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for ordering food. After the device is powered on and the program of the method is run through the processor, the following steps are performed: constructing a communication user knowledge base; determining the communication user information in the communication instruction voice data through a voice entity recognition model and the knowledge base; and executing a communication connection establishment process based on the communication user information.

[0497] Thirty-eighth Embodiment

[0498] In the above embodiment, a multimedia program on-demand system is provided, and correspondingly, the present application also provides a television program playing system. The system corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0499] The present application provides a television program playing system including: a smart television and a server.

[0500] Among them, the smart TV is used to collect the user's program playback voice command data, send the voice command data to the server, and play the target program object; the server is used to build a TV program knowledge base; through the voice entity recognition model and the knowledge base, the target program name corresponding to the voice command data is determined; according to the target program name, the target program object playback processing is executed.

[0501] The television program knowledge base includes program-related entities, such as television program names, actor names, channel names, and other entity words related to the broadcasting of television programs.

[0502] In this embodiment, the smart TV can collect the user's program playback voice command data through a remote control or other device. For example, the user issues a voice command "I want to watch Nezha", but the user is unclear and the actual pronunciation is "I want to watch that station"; the server can identify the program name as "Nezha" through the voice entity recognition model and the knowledge base. The specific implementation method of voice entity recognition in this step can be found in the relevant description of Example 1, which will not be repeated here. After the server determines the target program name, it can execute the target program object playback processing, such as sending the video stream of the program to the requesting party's device and playing the target program object through the terminal device.

[0503] In one example, the knowledge base may include program-related entities with homophones but different characters, such as movie names with the same pronunciation, etc.; accordingly, the knowledge base also includes user entities, and entity relationships between program-related entities with homophones but different characters and user entities, that is, the knowledge base includes the relationship between users and TV programs they have watched; accordingly, the server can adopt the relevant processing method in Example 1 to identify the movie name that the user really wants to watch based on the corresponding relationship, rather than other movie names with the same pronunciation, so that the accuracy of the program name can be effectively improved. The specific implementation method of voice entity recognition in this step can be found in the relevant description of Example 1, which will not be repeated here.

[0504] In one example, the server needs to perform target program object playback processing based on the target program name, and the following processing method may be used: first, according to the program schedule of each TV channel (such as the program schedule of the most recent week, which may include program information broadcast in the most recent week and program information currently being broadcast), determine the TV channel and broadcast time related to the target program name; then, according to the broadcast time and the TV channel, determine the target program object; finally, play the target program object.

[0505] Taking cable TV in a certain area as an example, when a user says "I want to watch that station" to the remote control, the remote control first recognizes that the program the user wants to order or replay is named "Nezha", and then can find out which channel and when "Nezha" was played according to the TV program list that can be replayed within a week. If it is found, it can play the replayable program object or the program object currently being played on the relevant channel. The following table shows the program list in this embodiment.

[0506]

[0507]

[0508] As shown in the table above, the program object corresponding to the identified target program name may include multiple program objects, such as the movie version of "Nezha" was broadcast on June 1, the cartoon version of "Nezha" produced by Shanghai People's Fine Arts Studio was broadcast on the 3rd, and the 52-episode version of the cartoon "Nezha" was broadcast on the 5th. In this case, the server can send multiple program objects broadcast at multiple times by at least one TV channel corresponding to the target program name to the smart TV, and the smart TV displays these program objects; the user can specify the target program object through the smart TV; the server sends the video stream of the target program object specified by the user to the smart TV for playback according to the user's request. This processing method allows all relevant program objects to be displayed on the TV screen for the user to select, and then the target program object specified by the user is played.

[0509] In specific implementation, if the server detects that the program list does not include the target program name, it determines the program name related to the target program name; displays the related program name; and plays the related program object if the user specifies to play the related program object. With this processing method, if the program that the user wants to watch first is not found, related TV programs can also be recommended to the user. For example, if the user wants to watch the documentary "Impression of West Lake", but the documentary has not been played in the past week, other programs related to West Lake can be played, such as "Deciphering Lingyin Temple", "Ten Scenes of West Lake", "Gu Jingzhou" and other programs.

[0510] It can be seen from the above embodiments that the television program playback system provided by the embodiments of the present application uses a smart TV to collect the user's program playback voice command data, sends the voice command data to the server, and plays the target program object; the server is used to build a television program knowledge base; through the voice entity recognition model and the knowledge base, the target program name corresponding to the voice command data is determined; according to the target program name, the target program object playback processing is performed; this processing method introduces program-related entity graph information, directly compares whether there is a program-related entity pronunciation in the knowledge graph in the program playback voice command, and realizes semantic understanding and recognition of entities such as program names from the program playback voice command signal, no longer relying on ASR to convert speech into text, and then identifying the program name through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of program entity recognition can be effectively improved, thereby improving the success rate and accuracy of television program on demand.

[0511] Thirty-ninth embodiment

[0512] In the above embodiment, a television program playing system is provided. Correspondingly, the present application also provides a television program playing method, and the execution subject of the method can be a smart TV, a TV remote controller, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0513] The present application provides a method for playing a television program, which may include the following steps:

[0514] Step 1: Collect the user's program playback voice command data;

[0515] Step 2: Send the voice command data to the server so that the server can build a TV program knowledge base; determine the target program name corresponding to the voice command data through the voice entity recognition model and the knowledge base; and perform target program object playback processing according to the target program name;

[0516] Step 3: Play the target program object.

[0517] Fortieth Embodiment

[0518] In the above embodiment, a method for playing a television program is provided. Correspondingly, the present application also provides a device for playing a television program. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as those of the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.

[0519] A television program playing device provided by the present application comprises:

[0520] A voice data collection unit, used to collect the user's program playback voice command data;

[0521] A voice data sending unit is used to send the voice instruction data to the server so that the server can build a TV program knowledge base; determine the target program name corresponding to the voice instruction data through the voice entity recognition model and the knowledge base; and perform target program object playback processing according to the target program name;

[0522] The program playing unit is used to play the target program object.

[0523] Forty-first embodiment

[0524] The present application also provides a smart TV. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0525] A smart TV of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for playing a television program. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting program playing voice instruction data of the user; sending the voice instruction data to a server so that the server can build a television program knowledge base; determining a target program name corresponding to the voice instruction data through a voice entity recognition model and the knowledge base; performing target program object playing processing according to the target program name; and playing the target program object.

[0526] Embodiment 42

[0527] The present application also provides a remote controller. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0528] A remote controller of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for playing a television program. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting the user's program playing voice command data; sending the voice command data to a server so that the server can build a television program knowledge base; determining the target program name corresponding to the voice command data through a voice entity recognition model and the knowledge base; and performing target program object playing processing according to the target program name.

[0529] Forty-third embodiment

[0530] In the above embodiment, a television program playing system is provided. Correspondingly, the present application also provides a television program playing method, and the execution subject of the method can be a server, a smart TV, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0531] The present application provides a method for playing a television program, which may include the following steps:

[0532] Step 1: Build a TV program knowledge base;

[0533] Step 2: Determine the target program name corresponding to the target program playback voice instruction data through the voice entity recognition model and the knowledge base;

[0534] Step 3: Execute the target program object playback process according to the target program name.

[0535] In one example, the knowledge base includes but is not limited to: program-related entities with the same pronunciation but different characters, user entities, and entity relationships between program-related entities and user entities; step 1 can be implemented in the following manner: determining the user entity based on the user's historical playback information, and constructing the entity relationship.

[0536] In one example, step 3 may include the following sub-steps: 3.1) determining the TV channel and broadcast time corresponding to the target program name according to the program list; 3.2) determining the target program object according to the broadcast time and the TV channel; 3.3) executing the processing of broadcasting the target program object.

[0537] In specific implementation, if the execution subject of the method is a server, the video stream of the target program object can be sent to a smart TV for playback; if the execution subject of the method is a smart TV, the target program object can be played.

[0538] In one example, step 3.2 may include the following sub-steps: 3.2.1) displaying multiple program objects played at multiple times by at least one television channel corresponding to the target program name; 3.2.2) using the program object specified by the user as the target program object.

[0539] In specific implementation, if the executor of the method is a server, multiple program objects played at multiple times by at least one TV channel corresponding to the target program name can be sent to a smart TV for display; if the executor of the method is a smart TV, multiple program objects played at multiple times by at least one TV channel corresponding to the target program name can be directly displayed.

[0540] In one example, step 3.2 may include the following sub-steps: 3.2.3) if the program list does not include the target program name, determine the program name related to the target program name; 3.2.4) display the related program name; 3.2.5) if the user specifies to play the related program object, execute the processing of playing the related program object.

[0541] In specific implementation, if the execution subject of the method is a server, the relevant program names can be sent to a smart TV for display; if the execution subject of the method is a smart TV, the relevant program names can be directly displayed.

[0542] Forty-fourth embodiment

[0543] In the above embodiment, a method for playing a television program is provided. Correspondingly, the present application also provides a device for playing a television program. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as those of the first embodiment are not repeated here. Please refer to the corresponding parts in the first embodiment.

[0544] A television program playing device provided by the present application comprises:

[0545] A knowledge base construction unit, used for constructing a TV program knowledge base;

[0546] An entity recognition unit, used to determine the target program name corresponding to the target program playback voice instruction data through the voice entity recognition model and the knowledge base;

[0547] The play processing unit is used to perform target program object play processing according to the target program name.

[0548] Forty-fifth embodiment

[0549] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0550] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for playing a television program. After the device is powered on and the program of the method is run by the processor, the following steps are performed: constructing a television program knowledge base; determining a target program name corresponding to target program playing voice instruction data through a voice entity recognition model and the knowledge base; and performing target program object playing processing according to the target program name.

[0551] Forty-sixth embodiment

[0552] In the above embodiment, a multimedia program on-demand system is provided, and correspondingly, the present application also provides a conference recording system. The system corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.

[0553] The present application provides a conference recording system including: a terminal device and a server.

[0554] Among them, the terminal device is used to collect voice data of the target meeting and send the voice data to the server; the server is used to build a language knowledge base in the conference field; through the voice entity recognition model and the language knowledge base in the conference field, the entity information in the target meeting voice data is determined; through the voice recognition model and the entity information, the text sequence corresponding to the meeting voice data is determined to form a meeting record.

[0555] The language knowledge base in the conference field includes, but is not limited to, entity vocabulary such as professional terms in the corresponding field. That is, the entity information in the target conference voice data can be professional terms in the corresponding field.

[0556] The conference field can be various application fields, such as computer field, medical field, legal field, patent field, etc. In specific implementation, language knowledge bases of multiple conference fields can be constructed, such as language knowledge bases in computer field, medical field, legal field, patent field, etc.

[0557] In specific implementation, for a conference field, the language knowledge in the field can be determined based on various text materials and multimedia materials in the field to form a corresponding language knowledge base.

[0558] Taking an international conference in the computer field as an example, the conference language is English, and the participants include technicians from many countries, but some personnel's professional vocabulary English pronunciation is not clear and accurate. Using the system provided by the embodiment of the present application, the voice data of the conference speaker can be collected through the terminal equipment at the conference site, and the service end can determine the terminology vocabulary (entity information) of the field in the conference voice data collected on site through the voice entity recognition model and the language knowledge base of the conference field (including various terms in the field), and determine the text sequence corresponding to the conference voice data through the voice recognition model and the recognized terminology vocabulary to form a meeting record.

[0559] In one example, the server is further used to determine the conference domain corresponding to the conference voice data. In specific implementation, the conference domain can be specified by the user, such as specifying the conference domain when starting the conference record, or the conference domain can be automatically determined by other means.

[0560] It can be seen from the above embodiments that the conference recording system provided by the embodiments of the present application collects the voice data of the target conference through the terminal device and sends the voice data to the server; the server builds a language knowledge base in the conference field; determines the entity information in the target conference voice data through the voice entity recognition model and the language knowledge base in the conference field; determines the text sequence corresponding to the conference voice data through the voice recognition model and the entity information to form a conference record; this processing method introduces entity graph information such as conference field-related terms, and directly compares whether there are pronunciations of entities such as conference field-related terms in the knowledge graph in the conference voice data, so as to realize semantic understanding and recognition of entities such as conference field-related terms from the conference voice data signal, and no longer relies on ASR to convert speech into text, and then recognizes the conference field-related terms therein through semantic understanding of the text, which is closer to the process of human understanding of speech; therefore, the accuracy of identifying conference field-related terms can be effectively improved, thereby improving the success rate and accuracy of meeting records.

[0561] Forty-seventh embodiment

[0562] In the above embodiment, a conference recording system is provided. Correspondingly, the present application also provides a conference recording method, and the execution subject of the method can be a terminal device such as a court trial integrated machine. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0563] A conference recording method provided by the present application may include the following steps:

[0564] Step 1: Collect the voice data of the target meeting;

[0565] Step 2: Send the voice data to the server so that the server can build a language knowledge base in the conference field; determine the entity information in the target conference voice data through the voice entity recognition model and the language knowledge base in the conference field; determine the text sequence corresponding to the conference voice data through the voice recognition model and the entity information to form a meeting record.

[0566] Forty-eighth embodiment

[0567] In the above embodiment, a conference recording method is provided, and correspondingly, the present application also provides a conference recording device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0568] A conference recording device provided by the present application includes:

[0569] A voice data collection unit, used to collect voice data of the target conference;

[0570] A voice data sending unit is used to send the voice data to a server so that the server can build a language knowledge base in the conference field; determine the entity information in the target conference voice data through a voice entity recognition model and the language knowledge base in the conference field; determine the text sequence corresponding to the conference voice data through a voice recognition model and the entity information to form a conference record.

[0571] Forty-ninth embodiment

[0572] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0573] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a meeting recording method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting voice data of a target meeting; sending the voice data to a server so that the server can build a language knowledge base in the field of meetings; determining entity information in the target meeting voice data through a voice entity recognition model and the language knowledge base in the field of meetings; determining a text sequence corresponding to the meeting voice data through the voice recognition model and the entity information to form a meeting record.

[0574] The 50th embodiment

[0575] In the above embodiment, a conference recording system is provided. Correspondingly, the present application also provides a conference recording method, and the execution subject of the method can be a server, a court trial integrated machine, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0576] A conference recording method provided by the present application may include the following steps:

[0577] Step 1: Build a language knowledge base in the conference field;

[0578] Step 2: Determine entity information in the target conference speech data through a speech entity recognition model and a language knowledge base in the conference field;

[0579] Step 3: Determine the text sequence corresponding to the conference voice data through the speech recognition model and the entity information to form a meeting record.

[0580] In one example, the method may further include the step of determining a conference area corresponding to the conference voice data.

[0581] Fifty-first embodiment

[0582] In the above embodiment, a conference recording method is provided, and correspondingly, the present application also provides a conference recording device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.

[0583] A conference recording device provided by the present application includes:

[0584] A knowledge base construction unit, used to construct a language knowledge base in the conference field;

[0585] An entity recognition unit, used to determine entity information in the target conference speech data through a speech entity recognition model and a language knowledge base in the conference field;

[0586] The conference record determination unit is used to determine the text sequence corresponding to the conference voice data through the speech recognition model and the entity information to form a conference record.

[0587] Fifty-second embodiment

[0588] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.

[0589] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a meeting recording method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: constructing a language knowledge base in the field of meetings; determining entity information in target meeting voice data through a speech entity recognition model and the language knowledge base in the field of meetings; determining a text sequence corresponding to the meeting voice data through the speech recognition model and the entity information to form a meeting record.

[0590] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.

[0591] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0592] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0593] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.

[0594] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A multimedia program on-demand system, characterized in that: include: The smart speaker is used to collect the voice data of multimedia program on-demand and send the voice data to the server; Play the multimedia program according to the multimedia program playing processing result of the server; A server is used to construct a multimedia program knowledge base, the knowledge base including: program-related entities with homophones but different characters, user entities, and entity relationships between program-related entities and user entities; determining audio feature data of the voice data through an audio encoding model included in a voice entity recognition model; determining at least one candidate pronunciation of multimedia program information in the voice data according to the audio feature data through an entity candidate pronunciation determination module in an entity decoding model included in the voice entity recognition model; determining the pronunciation of the multimedia program information from the at least one candidate pronunciation according to the knowledge base through an entity pronunciation determination module included in the entity decoding model; determining candidate entities according to the pronunciation of the multimedia program information; determining the multimedia program information from the candidate entities according to the user information and the entity relationship; and performing multimedia program playback processing according to the multimedia program information.

2. A food ordering system, characterized in that: include: The ordering device is used to collect the ordering voice data and send the voice data to the server; A server is used to construct a food knowledge base, the knowledge base including: food entities with the same pronunciation but different characters, user entities, and entity relationships between food entities and user entities; determining audio feature data of the speech data through an audio encoding model included in a speech entity recognition model; determining at least one candidate pronunciation of food information in the speech data according to the audio feature data through an entity candidate pronunciation determination module in an entity decoding model included in the speech entity recognition model; determining the pronunciation of the food information from the at least one candidate pronunciation according to the knowledge base through an entity pronunciation determination module included in the entity decoding model; determining candidate entities according to the pronunciation of the food information; determining the food information from the candidate entities according to the user information and the entity relationship; and performing meal preparation according to the food information.

3. A communication connection establishment system, characterized in that: include: The user equipment is used to collect the communication command voice data and send the voice data to the server; A server is used to construct a communication user knowledge base, the knowledge base comprising: communication user entities with the same pronunciation but different characters, user entities, and entity relationships between communication user entities and user entities; determining audio feature data of the voice data through an audio encoding model included in a voice entity recognition model; determining at least one candidate pronunciation of communication user information in the voice data according to the audio feature data through an entity candidate pronunciation determination module in an entity decoding model included in the voice entity recognition model; determining the pronunciation of the communication user information from the at least one candidate pronunciation according to the knowledge base through an entity pronunciation determination module included in the entity decoding model; determining candidate entities according to the pronunciation of the communication user information; determining the communication user information from the candidate entities according to the user information and the entity relationship; and executing communication connection establishment processing according to the communication user information.

4. A voice interaction system, characterized in that: include: The terminal device is used to collect voice data and send the voice data to the server; A server is used to construct an entity knowledge base of a target domain, wherein the knowledge base includes: target domain related entities of homophones, user entities, and entity relationships between target domain related entities and user entities; determining audio feature data of target voice data through an audio encoding model included in a voice entity recognition model; determining at least one candidate pronunciation of entity information in the target voice data according to the audio feature data through an entity candidate pronunciation determination module in an entity decoding model included in the voice entity recognition model; determining the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base through an entity pronunciation determination module included in the entity decoding model; determining candidate entities according to the pronunciation of the entity information; determining the entity information from the candidate entities according to user information and the entity relationship; and performing voice interaction processing according to the entity information.

5. A voice interaction method, used in the server of the system described in claim 4, characterized in that: include: Build entity knowledge base; Determine entity information in the target speech data through the speech entity recognition model and the entity knowledge base; According to the entity information, voice interaction processing is performed.

6. The method according to claim 5, characterized in that The step of determining the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base includes: Determining the similarity between the pronunciation of the entity in the entity knowledge base and the candidate pronunciation; The pronunciation of the entity information is determined according to the similarity.

7. The method according to claim 5, characterized in that Constructing the entity knowledge base includes: According to the user's historical playback information, the user entity is determined and the entity relationship is constructed.

8. The method according to claim 5, characterized in that Also includes: Learning the speech entity recognition model from the training data; The training data includes: audio data and entity annotation information.

9. The method according to claim 5, characterized in that Constructing the entity knowledge base includes: Determine multimedia program related entities to form the entity knowledge base.

10. A voice interaction method, used in a terminal device in the system of claim 4, characterized in that: include: Collecting voice data, and sending the voice data to a server, so that the server can build an entity knowledge base; The entity information in the voice data is determined through the voice entity recognition model and the entity knowledge base; and the voice interaction processing is performed according to the entity information.

11. A multimedia program on-demand method, used in the server of the system of claim 1, characterized in that: include: Build a knowledge base of multimedia programs; Determining multimedia program information in multimedia program on-demand voice data through a voice entity recognition model and the knowledge base; The multimedia program playing process is performed according to the multimedia program information.

12. A multimedia program on-demand method, used in the smart speaker in the system of claim 1, characterized in that: include: Collecting multimedia program on-demand voice data, and sending the voice data to a server, so that the server can build a multimedia program knowledge base; The multimedia program information in the voice data is determined through the voice entity recognition model and the knowledge base; and the multimedia program playing process is performed according to the multimedia program information.

13. A method for ordering food, used in the server of the system of claim 2, characterized in that: include: Build a food knowledge base; Determine the food information in the ordering voice data by using the voice entity recognition model and the knowledge base; The meal preparation process is performed according to the meal information.

14. A method for ordering food, used in the ordering device of the system as claimed in claim 2, characterized in that: include: Collecting ordering voice data and sending the voice data to the server so that the server can build a food knowledge base; The meal information in the voice data is determined through the voice entity recognition model and the knowledge base; and meal preparation processing is performed according to the meal information.

15. A method for establishing a communication connection, used in the server of the system of claim 3, characterized in that: include: Build a knowledge base of communication users; Determine the communication user information in the communication command voice data through the voice entity recognition model and the knowledge base; A communication connection establishment process is performed according to the communication user information.

16. A method for establishing a communication connection, used for a user device in the system of claim 3, characterized in that: include: Collecting communication command voice data, and sending the voice data to the server, so that the server can build a communication user knowledge base; The communication user information in the voice data is determined through the voice entity recognition model and the knowledge base; and the communication connection establishment process is performed according to the communication user information.

17. A method for constructing a speech entity recognition model, characterized in that: include: Determine a training data set, wherein the training data includes: speech data, entity annotation information, and an entity knowledge base of a target domain, wherein the knowledge base includes: target domain related entities of homophones, user entities, and entity relationships between target domain related entities and user entities; Constructing the network structure of the model; the model includes: an audio encoding model, an entity decoding model; the entity decoding model includes: an entity candidate pronunciation determination module, an entity pronunciation determination module; The audio coding model is used to determine audio feature data of target speech data; The entity candidate pronunciation determination module is used to determine at least one candidate pronunciation of the entity information in the target voice data according to the audio feature data; The entity pronunciation determination module is used to determine the pronunciation of the entity information from the at least one candidate pronunciation according to the entity knowledge base; determine the candidate entity according to the pronunciation of the entity information; and determine the entity information from the candidate entities according to the user information and the entity relationship; The model is learned from a training data set.

18. A method for speech entity recognition, characterized in that: include: Constructing an entity knowledge base and a speech entity recognition model for the target domain, wherein the knowledge base includes: target domain related entities of homophones, user entities, and entity relationships between target domain related entities and user entities; determining target voice data; The audio feature data of the target speech data is determined by the audio encoding model included in the speech entity recognition model; at least one candidate pronunciation of the entity information in the target speech data is determined according to the audio feature data by the entity candidate pronunciation determination module in the entity decoding model included in the speech entity recognition model; the pronunciation of the entity information is determined from the at least one candidate pronunciation according to an entity knowledge base by the entity pronunciation determination module included in the entity decoding model; the candidate entity is determined according to the pronunciation of the entity information; and the entity information in the target speech data is determined from the candidate entities according to the user information and the entity relationship.

19. A method for playing a television program, characterized in that: include: Constructing a TV program knowledge base, the knowledge base comprising: program-related entities with the same pronunciation but different characters, user entities, and entity relationships between program-related entities and user entities; Determine the audio feature data of the target program playback voice instruction data through the audio encoding model included in the voice entity recognition model; determine at least one candidate pronunciation of the entity information in the target program playback voice instruction data according to the audio feature data through the entity candidate pronunciation determination module in the entity decoding model included in the voice entity recognition model; determine the pronunciation of the entity information from the at least one candidate pronunciation according to the knowledge base through the entity pronunciation determination module included in the entity decoding model; determine the candidate entity according to the pronunciation of the entity information; determine the target program name corresponding to the target program playback voice instruction data from the candidate entity according to the user information and the entity relationship; According to the target program name, target program object playback processing is executed.

20. A conference recording method, characterized in that: include: Collect voice data of the target meeting; The voice data is sent to a server so that the server can build a language knowledge base in the conference field, wherein the knowledge base includes: conference-related entities with the same pronunciation but different characters, user entities, and entity relationships between conference-related entities and user entities; the audio feature data of the voice data is determined by an audio encoding model included in a voice entity recognition model; at least one candidate pronunciation of entity information in the voice data is determined according to the audio feature data by an entity candidate pronunciation determination module in an entity decoding model included in the voice entity recognition model; the pronunciation of the entity information is determined from the at least one candidate pronunciation according to the language knowledge base in the conference field by an entity pronunciation determination module included in the entity decoding model; a candidate entity is determined according to the pronunciation of the entity information; the entity information is determined from the candidate entities according to the user information and the entity relationship; The text sequence corresponding to the voice data is determined through the speech recognition model and the entity information to form a meeting record.

Citation Information

Patent Citations

  • Voice recognition post-processing method and device as well as voice recognition system

    CN105206274A

  • Method and apparatus for speech recognition

    CN107016994A