A recommendation method and computer device based on speech recognition

By acquiring classification information and keywords from voice information and using a natural language processing model to determine the target recommendation file, the problem of inaccurate intent determination in voice interaction is solved, and more accurate recommendation information query is achieved.

CN114639385BActive Publication Date: 2026-01-23SHENZHEN TCL NEW-TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011383831.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-01
Publication Date
2026-01-23
Estimated Expiration
2040-12-01

AI Technical Summary

Technical Problem

Existing voice interaction technologies cannot accurately determine the user's true intent, resulting in poor accuracy of recommended information.

Method used

By acquiring classification information and keywords from voice information, a natural language processing model is used to determine the target recommendation file, and recommendation information is selected from the target recommendation file. Combining classification information and keywords improves the accuracy of the recommendation information.

Benefits of technology

It improves the accuracy of voice interaction in determining recommended information, enabling users to more accurately retrieve the information they actually want.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114639385B_ABST
    Figure CN114639385B_ABST
Patent Text Reader

Abstract

The application provides a recommendation method based on voice recognition and a computer device, the recommendation method based on voice recognition comprises the following steps: obtaining voice information to be processed, and determining classification information and a keyword corresponding to the voice information; determining a target recommendation file corresponding to the voice information according to the classification information; selecting recommendation information in the target recommendation file according to the keyword, and taking the selected recommendation information as response information corresponding to the voice information. The application determines the classification information and the keyword corresponding to the voice information, can determine the target recommendation file meeting the classification information first, and then determines the recommendation information in the target recommendation file based on the keyword; through the combination of the classification information and the keyword, the recommendation information meeting the user's intention can be queried, and the accuracy of determining the recommendation information in voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of voice interaction, and in particular, to a recommendation method based on voice recognition and a computer device. BACKGROUND

[0002] Voice interaction is to issue instructions to a machine through voice and get feedback from the machine. Voice interaction can control devices, including various smart Internet of Things devices, such as smart televisions, smart refrigerators, smart sound boxes, etc.

[0003] At present, recommendation information can be obtained based on voice interaction, but the existing voice interaction usually extracts keywords and queries the keywords to obtain recommendation information, but the query keywords cannot determine the true intention of the user, for example, the recognized keyword is "fish-flavored shredded pork", and the machine cannot determine whether to query the allusion of "fish-flavored shredded pork" or the merchant related to "fish-flavored shredded pork". In this way, the user's real information cannot be accurately queried, and the accuracy of voice interaction to obtain recommendation information is poor.

[0004] Therefore, the prior art needs to be improved. SUMMARY

[0005] The present application provides a recommendation method based on voice recognition and a computer device, which determines a target knowledge graph corresponding to voice information according to a target classification identifier, determines a query result in the target knowledge graph, and can obtain more accurate query results, thereby improving the accuracy of query through voice recognition.

[0006] In a first aspect, the embodiments of the present application provide a recommendation method based on voice recognition, comprising:

[0007] Obtaining voice information to be processed, and determining classification information and keywords corresponding to the voice information;

[0008] Determining a target recommendation file corresponding to the voice information according to the classification information;

[0009] Selecting recommendation information in the target recommendation file according to the keywords, and taking the selected recommendation information as response information corresponding to the voice information.

[0010] In a further improved solution, the determination of the classification information and the keywords corresponding to the voice information specifically comprises:

[0011] Recognizing the voice information to obtain text information corresponding to the voice information;

[0012] Inputting the text information into a natural language processing model, and outputting the classification information and the keywords corresponding to the voice information through the natural language processing model.

[0013] In a further improved solution, the classification information comprises a target classification identifier; and the determining of the target recommendation file corresponding to the voice information according to the classification information specifically comprises:

[0014] querying a target knowledge graph corresponding to the target classification identifier from a plurality of preset knowledge graphs, and taking the target knowledge graph as the target recommendation file corresponding to the voice information; wherein the classification identifiers of the plurality of knowledge graphs are different from each other.

[0015] In a further improved solution, the classification information further comprises an intent identifier, and the target knowledge graph comprises a plurality of sets, each set having a respective set identifier; and after the taking of the target knowledge graph as the target recommendation file corresponding to the voice information, the method further comprises:

[0016] querying a target set with a set identifier consistent with the intent identifier from the plurality of sets;

[0017] replacing the target recommendation file with the queried target set to obtain a replaced target recommendation file.

[0018] In a further improved solution, the plurality of knowledge graphs at least comprises a recipe knowledge graph, a music knowledge graph, and a video knowledge graph; the classification identifier of the recipe knowledge graph is a first classification identifier, the classification identifier of the music knowledge graph is a second classification identifier, and the classification identifier of the video knowledge graph is a third classification identifier.

[0019] Correspondingly, the querying of the target knowledge graph corresponding to the target classification identifier from the plurality of preset knowledge graphs comprises:

[0020] when the target classification identifier is the first classification identifier, querying the target knowledge graph corresponding to the first classification identifier from the plurality of preset knowledge graphs as the recipe knowledge graph;

[0021] when the target classification identifier is the second classification identifier, querying the target knowledge graph corresponding to the second classification identifier from the plurality of preset knowledge graphs as the music knowledge graph;

[0022] when the target classification identifier is the third classification identifier, querying the target knowledge graph corresponding to the third classification identifier from the plurality of preset knowledge graphs as the video knowledge graph.

[0023] In a further improved solution, the selecting of the recommendation information from the target recommendation file according to the keyword specifically comprises:

[0024] querying a plurality of candidate information corresponding to the keyword from the target recommendation file;

[0025] obtaining a weight value corresponding to each of the candidate information;

[0026] determining a preset number of recommended information from the candidate information based on the obtained weight values, wherein the weight value of each recommended information is greater than that of any non-recommended information, and the non-recommended information is the candidate information other than the recommended information.

[0027] In a further improved solution, after the selected recommended information is taken as the response information corresponding to the voice information, the solution further comprises:

[0028] converting the response information into a voice form to obtain voice response information, and playing the voice response information.

[0029] In a second aspect, an embodiment of the present application provides a query device based on voice recognition, comprising:

[0030] a voice information processing module, configured to obtain voice information to be processed, and determine classification information and a keyword corresponding to the voice information;

[0031] a target recommended file determination module, configured to determine a target recommended file corresponding to the voice information according to the classification information;

[0032] a recommendation module, configured to select recommended information in the target recommended file according to the keyword, and take the selected recommended information as response information corresponding to the voice information.

[0033] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0034] obtaining voice information to be processed, and determining classification information and a keyword corresponding to the voice information;

[0035] determining a target recommended file corresponding to the voice information according to the classification information;

[0036] selecting recommended information in the target recommended file according to the keyword, and taking the selected recommended information as response information corresponding to the voice information.

[0037] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the following steps:

[0038] obtaining voice information to be processed, and determining classification information and a keyword corresponding to the voice information;

[0039] determine the target recommendation file corresponding to the voice information according to the classification information;

[0040] select the recommendation information in the target recommendation file according to the keyword, and take the selected recommendation information as the response information corresponding to the voice information.

[0041] Compared with the prior art, the embodiments of the present application have the following advantages:

[0042] In the embodiments of the present application, the voice information to be processed is obtained, and the classification information and the keyword corresponding to the voice information are determined; the target recommendation file corresponding to the voice information is determined according to the classification information; the recommendation information in the target recommendation file is selected according to the keyword, and the selected recommendation information is taken as the response information corresponding to the voice information. The classification information and the keyword corresponding to the voice information are determined, the target recommendation file meeting the classification information is determined first, and then the recommendation information in the target recommendation file is determined based on the keyword; through the combination of the classification information and the keyword, the recommendation information more meeting the user's intention can be queried, and the accuracy of determining the recommendation information in voice interaction is improved. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Figure 1 The schematic diagram of the application scenario of a recommendation method based on voice recognition in the embodiments of the present application;

[0045] Figure 2 The schematic diagram of the Transformer encoding structure in the embodiments of the present application;

[0046] Figure 3 The schematic diagram of the recipe knowledge graph in the embodiments of the present application;

[0047] Figure 4 The structural schematic diagram of a query device based on voice recognition in the embodiments of the present application;

[0048] Figure 5 The structural schematic diagram of the query device based on voice recognition in the embodiments of the present application when implemented;

[0049] Figure 6 The internal structure diagram of the computer device in the embodiments of the present application. DETAILED DESCRIPTION

[0050] In order to make the objects, technical solutions and effects of the present application clearer and more apparent, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and not intended to limit the present application.

[0051] Those skilled in the art can understand that the singular forms "a", "an" and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the use of the term "include" in the specification of the present application means that the features, integers, steps, operations, elements and / or components exist, but does not exclude the presence or addition of one or more

[0052] other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.

[0053] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those skilled in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.

[0054] The inventors have found through research that voice interaction is to issue instructions to machines through voice and get feedback from machines. Voice interaction can control devices, including various smart Internet of Things devices, such as smart televisions, smart refrigerators, smart sound boxes, etc.

[0055] At present, recommended information can be obtained based on voice interaction, but the existing voice interaction usually extracts keywords and queries the keywords to obtain recommended information, but the query keywords cannot determine the true intention of the user, for example, the keyword "fish-flavored shredded pork" is identified, and the machine cannot determine whether to query the allusion of "fish-flavored shredded pork" or to query the related business of "fish-flavored shredded pork". In this way, the user's real information cannot be accurately queried, and the accuracy of the recommended information obtained by voice interaction is poor.

[0056] To solve the above problems, in the embodiment of the present application, the voice information to be processed is acquired, and the classification information and the keyword corresponding to the voice information are determined; the target recommendation file corresponding to the voice information is determined according to the classification information; the recommendation information is selected in the target recommendation file according to the keyword, and the selected recommendation information is taken as the response information corresponding to the voice information. The classification information and the keyword corresponding to the voice information are determined, the target recommendation file meeting the classification information is determined first, and then the recommendation information in the target recommendation file is determined based on the keyword; through the combination of the classification information and the keyword, the recommendation information meeting the user demand can be queried, and the accuracy of determining the recommendation information in the voice interaction is improved.

[0057] The recommendation method based on voice recognition provided by the embodiment of the present application can be applied to an electronic device, which is a device capable of receiving voice information and processing the voice information, such as a computer, a smart terminal, a smart television, a smart sound box, a smart refrigerator, and the like.

[0058] Referring to Figure 1 The embodiment provides a recommendation method based on voice recognition, which comprises the following steps:

[0059] S1, acquiring voice information to be processed, and determining classification information and a keyword corresponding to the voice information.

[0060] In the embodiment of the present application, the voice information to be processed is voice information used for querying recommendation information. The voice information to be processed can be voice information issued by a user. For example, the user says, “how to make sweet and sour spare ribs”, and “how to make sweet and sour spare ribs” is the voice information to be processed.

[0061] In the embodiment of the present application, the classification information is used for reflecting the classification corresponding to the content involved in the voice information. For example, the voice information is “how to make sweet and sour spare ribs”, and the classification information is a recipe. The voice information is “play music black sweater”, and the classification information is music.

[0062] The keyword is a key information required for querying recommendation information, and the recommendation information meeting the user demand can be accurately queried through the keyword. For example, the voice information is “how to make sweet and sour spare ribs”, and the keyword is sweet and sour spare ribs.

[0063] In the embodiment of the present application, the classification information and the keyword corresponding to the voice information can be determined through existing voice recognition technology. In order to improve the accuracy of the classification information and the keyword, the voice information can be first converted into text information, and the text information is processed by using natural language processing to determine the classification information and the keyword corresponding to the voice information.

[0064] Specifically, the step S1 comprises:

[0065] S11, recognizing the voice information to obtain the text information corresponding to the voice information.

[0066] In the embodiment of the present application, the voice information can be converted into text information through automatic speech recognition technology (ASR). The process of voice recognition by ASR comprises: pre-acquiring a plurality of training voices, each of the plurality of training voices having text corresponding to the training voice, determining training parameters corresponding to each training voice, and storing all the determined training parameters in a voice parameter library; after receiving the voice information to be queried, analyzing the voice information to obtain a plurality of voice parameters corresponding to the voice information, for each voice parameter, comparing the voice parameter with all the training parameters in the voice database to determine the training parameter closest to the voice parameter, taking the text corresponding to the training parameter as the text corresponding to the voice information, and determining the text information corresponding to the voice information according to the text corresponding to each of all the voice parameters.

[0067] S12, inputting the text information into a natural language processing model, and outputting the classification information and the keyword corresponding to the voice information through the natural language processing model.

[0068] In the embodiment of the present application, each word in the text information is classified through the natural language processing model to determine the word marked as classification annotation to obtain the classification information, and to determine the word marked as keyword annotation to obtain the keyword. The word belonging to the classification annotation is taken as the classification information corresponding to the text information, and the word belonging to the keyword annotation is taken as the keyword corresponding to the text information.

[0069] The natural language processing model is a natural language processing model that has been trained, and the natural language processing model comprises a Bidirectional Encoder Representations from Transformers (BERT) network of a converter, a Bi-directional Long Short-Term Memory (BiLSTM) network, and a Conditional Random Field (CRF) network.

[0070] The BERT network can learn the relationship between words in the text information to obtain a word vector. A word can be a Chinese character or a word composed of multiple Chinese characters, or an English word. Specifically, the text information is first segmented to obtain a plurality of words, and then an initial word vector corresponding to each word in the plurality of words is obtained. The plurality of initial word vectors are input into the BERT network to obtain an output word vector corresponding to each word.

[0071] The BERT network is constructed using a Transformer encoding structure. Referring to FIG. 3, a schematic diagram of the Transformer encoding structure is shown. Next, the processing flow of the Transformer encoding structure is described by way of example. Figure 2

[0072] Suppose the input is text information, each word in the text information is converted into an initial word vector corresponding to each word, respectively. The initial word vectors are added with position encoding, which represents the position of each word in the text information and the distance between different words. The word vectors added with the position encoding are input into a multi-head attention model. The word vectors processed by the multi-head attention model and the word vectors not processed by the multi-head attention model are added, and then normalized to obtain an intermediate word vector. The intermediate word vector is input into a feedforward neural network. The intermediate word vectors processed by the feedforward neural network and the intermediate word vectors not processed by the feedforward neural network are added, and then normalized to obtain an output word vector.

[0073] The BiLSTM network belongs to a recurrent neural network and includes a forward LSTM network and a backward LSTM network. The BiLSTM network can determine a label corresponding to each word. The BiLSTM network is pre-configured with a plurality of labels, including at least a classification label corresponding to classification information and a keyword label. After determining the label corresponding to each word in the text information, the words belonging to the classification label are taken as the classification information corresponding to the text information, and the words belonging to the keyword label are taken as the keyword corresponding to the text information.

[0074] ​Specifically, the output word vectors corresponding to the text information are input into the feedforward LSTM network in ascending order to obtain the feedforward memory word vectors for each output word vector. The output word vectors are then input into the feedback LSTM network in reverse order to obtain the feedback memory word vectors for each output word vector. For each output word vector, the feedforward and feedback memory word vectors are merged to obtain the memory word vector corresponding to that output word vector. The output matrix of the BiLSTM network is determined based on each memory word vector. Each element in the memory word vector represents the probability value of each label corresponding to the output word vector. That is, for each output word vector, the probability value of each label corresponding to that word vector can be obtained, and the label corresponding to the highest probability value among the probability values ​​of each label corresponding to that word vector is taken as the label of that word vector.

[0075] For example, for the text information "I love China", the words are divided into "I", "love" and "China". The output word vector corresponding to "I" is t1, the output word vector corresponding to "love" is t2, and the output word vector corresponding to "China" is t3. The forward LSTM network includes at least: the first forward LSTM sub-network (LSTM-l1), the second forward LSTM sub-network (LSTM-l2), and the third forward LSTM sub-network (LSTM-l3). The backward LSTM network includes at least: the first backward LSTM sub-network (LSTM-r1), the second backward LSTM sub-network (LSTM-r2), and the third backward LSTM sub-network (LSTM-r3). The forward input includes: inputting t1 into LSTM-l1 to obtain h-l1; inputting h-l1 and t2 into LSTM-l2 to obtain h-l2; and inputting h-l2 and t3 into LSTM-l3 to obtain h-l3. The backward input includes: inputting t3 into LSTM-r1 to obtain h-r1; inputting h-r1 and t2 into LSTM-r2 to obtain h-r2; and inputting h-r2 and t1 into LSTM-r3 to obtain h-r3. h-l1 and h-r3 are combined to obtain the memory word vector f1 corresponding to t1; h-l2 and h-r2 are combined to obtain the memory word vector f2 corresponding to t2; and h-l3 and h-r1 are combined to obtain the memory word vector f3 corresponding to t3. The output matrix is ​​determined based on f1, f2, and f3.

[0076] Assume that f1 is (x1, x2, x3), f1 is a memory word vector corresponding to t1, wherein x1 represents the probability that t1 belongs to the label y1, x2 represents the probability that t1 belongs to the label y2, and x3 represents the probability that t1 belongs to the label y3, if x1 is the largest in (x1, x2, x3), y1 is taken as the label corresponding to t1. Assume that the label y1 is a keyword label, t1 is a keyword, that is, in 'I love China', the label corresponding to 'I' is a keyword label, and the keyword in the text information is 'I'.

[0077] The CRF network is used to adjust the result output by the BiLSTM network. The output result of the BiLSTM network is an output matrix, which is used to reflect the probability that each word corresponds to each label respectively, and the CRF network adds some constraints to ensure that the predicted label is legal. The output matrix obtained by the BiLSTM network is adjusted through the CRF network to obtain the label corresponding to each word respectively, and according to the label corresponding to each word respectively, the classification information and the keyword corresponding to the text information can be determined.

[0078] S2, determining the target recommendation file corresponding to the voice information according to the classification information.

[0079] In the embodiment of the application, the classification information includes a target classification identifier, and the target classification identifier is an identifier used to reflect the classification corresponding to the content involved in the voice information. The target classification identifier can be represented in the form of text, and the natural language processing model can directly output the classification information and the keyword in the form of text. Therefore, the natural language processing model can directly output the target classification identifier in the form of text.

[0080] In the embodiment of the application, the terminal pre-stores data, and the pre-stored data can be divided into a plurality of data sets, each data set has a respective classification identifier, and the classification identifiers of any two data sets are different. The classification identifier of the data set is used to reflect which classification the data set belongs to. Based on the classification information (the classification information includes the target classification identifier), one data set can be determined from the plurality of data sets, and the determined data set is taken as the target recommendation file.

[0081] Specifically, a plurality of data sets are pre-stored, and the classification identifiers of the plurality of data sets are different from each other; the classification information includes a target classification identifier, the target classification identifier is matched with the classification identifier corresponding to each data set, and a data set with a classification identifier consistent with the target classification identifier is selected, and the selected data set is taken as the target recommendation file.

[0082] For example, the plurality of data sets are A1, A2, A3 and A4 respectively, wherein the classification identifier of A1 is s1, the classification identifier of A2 is s2, the classification identifier of A3 is s3, and the classification identifier of A4 is s4. Assuming that the target classification identifier is s1, A1 is taken as the target recommended file.

[0083] In the embodiment of the present application, the pre-stored data set can be a data set stored in the form of a knowledge graph. The knowledge graph is used to describe each entity, the attributes of each entity, and the association between entities, and can more comprehensively describe data. According to the knowledge graph, the recommended information can be more in line with the user's needs. Each knowledge graph has a classification identifier corresponding to the knowledge graph.

[0084] Specifically, step S2 includes:

[0085] S21, obtain a plurality of pre-stored knowledge graphs, wherein the classification identifiers of the plurality of knowledge graphs are different from each other.

[0086] In the embodiment of the present application, each knowledge graph in the plurality of knowledge graphs is pre-established, and the plurality of knowledge graphs at least include a recipe knowledge graph, a music knowledge graph and a video knowledge graph. Each knowledge graph has a corresponding classification identifier, and the classification identifier of the knowledge graph is used to reflect which classification the knowledge graph belongs to, that is, to reflect the category of the knowledge graph. The classification identifier of the recipe knowledge graph is a first classification identifier, the classification identifier of the music knowledge graph is a second classification identifier, and the classification identifier of the video knowledge graph is a third classification identifier. The first classification identifier, the second classification identifier and the third classification identifier can be represented by text. The first classification identifier can be a recipe, the second classification identifier can be music, and the third classification identifier can be a video.

[0087] Next, the detailed process of establishing a knowledge graph is introduced.

[0088] Take the establishment of a recipe knowledge graph as an example. First, recipe data is crawled in the network, the recipe data is cleaned and de-duplicated, and the original unstructured data is converted into a plurality of csv format files, and the plurality of csv format files respectively represent each ontology in the knowledge graph and the attributes of the ontology. The python script of kg_operate.py and the cypher language of neo4j are used to import the plurality of csv format files into the neo4j graph database to establish the recipe knowledge graph.

[0089] S22, query the target knowledge graph corresponding to the target classification identifier in the plurality of knowledge graphs, and take the target knowledge graph as the target recommended file corresponding to the voice information.

[0090] In the embodiment of the present application, after the target classification identifier is determined, the target classification identifier is matched with the classification identifiers respectively corresponding to the plurality of pre-stored knowledge graphs, so as to determine the target knowledge graph corresponding to the target classification identifier in the plurality of knowledge graphs.

[0091] Specifically, when the target classification identifier is a first classification identifier, the target knowledge graph corresponding to the first classification identifier is a recipe knowledge graph in the plurality of preset knowledge graphs; when the target classification identifier is a second classification identifier, the target knowledge graph corresponding to the second classification identifier is a music knowledge graph in the plurality of preset knowledge graphs; when the target classification identifier is a third classification identifier, the target knowledge graph corresponding to the third classification identifier is a video knowledge graph in the plurality of preset knowledge graphs.

[0092] For example, the voice information is "how to make sweet and sour spare ribs", the target classification identifier is "recipe", and the target knowledge graph can be determined as a recipe knowledge graph; the voice information is "song: black sweater", the target classification identifier is "music", and the target knowledge graph can be determined as a music knowledge graph; the voice information is "recommend an Italian movie", the target classification identifier is "video", and the target knowledge graph can be determined as a video knowledge graph. When the target knowledge graph is determined as a recipe knowledge graph, the recipe knowledge graph is the target recommendation file corresponding to the voice information.

[0093] In order to obtain more accurate recommendation information, the data amount of the target recommendation file can be reduced. After step S22, the target knowledge graph is determined as the target recommendation file, the target knowledge graph includes a plurality of sets, the data included in a set is the data corresponding to a classification in the target knowledge graph, and each set has a set identifier respectively corresponding thereto. The set identifier is used to reflect the classification of a set. The classification information further includes an intent identifier, and the intent identifier is used to reflect the user's intention. The intent identifier in the form of text output by the natural language processing module. Based on the intent identifier, a set can be selected as the target recommendation file in the target knowledge graph.

[0094] Specifically, step S22 further includes:

[0095] S23, in the plurality of sets, a target set with a set identifier consistent with the intent identifier is queried.

[0096] In the embodiment of the present application, the target knowledge graph includes a plurality of sets, and the plurality of sets are obtained by classifying data corresponding to the target knowledge graph from different perspectives. The set corresponding to the intent identifier is matched with the set identifiers corresponding to the sets respectively, to determine a set identifier consistent with the intent identifier, and the set corresponding to the set identifier consistent with the intent identifier is taken as a target set.

[0097] Next, the target graph includes a plurality of sets is introduced by examples.

[0098] Referring to Figure 3 , the recipe knowledge graph includes: a total recipe set (cookbook); a set classified by cuisine (cusine), including: a Sichuan cuisine set (set identifier Sichuan), a Cantonese cuisine set (set identifier Cantonese), etc.; a set classified by type (type), including: a quick recipe set (set identifier quick), a low-fat recipe set (set identifier low-fat), etc.; a set classified by each recipe (recipe), and the set identifier is the name of each dish; a set classified by ingredients (ingredient), and the set identifier is the name of the ingredient, for example, chicken, beef, etc. The ingredient includes ingredient1 and ingredient2 in Figure 3 , ingredient1 can include all sets of one kind, and ingredient2 can include all sets of another kind, for example, ingredient1 includes the ingredient set corresponding to the main food material, and ingredient2 represents the ingredient set corresponding to the ingredient. Among them, the BELONG_TO between cusine, type and cookbook represents the belonging relationship, and the HAS_INGREDIENT between recipe and ingredient represents the relationship of containing ingredients.

[0099] For example, for the text information: “the method of low-fat sandwich”, the target classification identifier is: “recipe”, and the intent identifier is: “low-fat”; the target knowledge graph is the recipe knowledge graph, and according to the intent identifier, the target set can be determined as the low-fat set in the recipe knowledge graph.

[0100] The music knowledge graph comprises: a total music set; a set classified by genre, comprising a popular set (set identifier: popular), a rock set (set identifier: rock), a classical set (set identifier: classical), and the like; a set classified by language, comprising a Chinese set (set identifier: Chinese), a Japanese and Korean set (set identifier: Japanese and Korean), an English set (set identifier: English), and the like; a set classified by each song, the set identifier being the name of each song; and a set classified by performer, the set identifier being the name of each performer. The set classified by genre and the set classified by language both belong to the total music set, the set classified by each song belongs to the set classified by language and also belongs to the set classified by genre, and the set classified by performer belongs to the set classified by each song.

[0101] For example, for the text information: “play White Windmill”, the target category is “music”, the intent identifier is “White Windmill”, the target knowledge graph is the music knowledge graph, and according to the intent, it can be determined that the set is the set classified by each song, and the set identifier is “White Windmill”.

[0102] The video knowledge graph comprises: a total video set (program-book), a set classified by type (sub-program), comprising a TV series set (set identifier: TV series), a movie set (set identifier: movie), a variety show set (set identifier: variety show), and the like; a set classified by language, comprising a Chinese set (set identifier: Chinese), a Japanese and Korean set (set identifier: Japanese and Korean), an English set (set identifier: English), and the like; a set classified by feature (type), comprising a leisure set (set identifier: leisure), a comedy set (set identifier: comedy), a science fiction set (set identifier: science fiction), an education set (set identifier: education), and the like; a set classified by each video, the set identifier being the name of each video; and a set classified by performer, the set identifier being the name of each performer.

[0103] For example, for the text information: “watch the variety show Sisterhood, Breaking Waves”, the target category identifier is “video”, the intent identifier is “variety show”; the target knowledge graph is the video knowledge graph, and according to the intent identifier, it can be determined that the target set is the variety show set.

[0104] In this embodiment of the invention, since both classification information and keywords are textual information output by the natural language processing model, they should be converted into a computer-recognizable expression: {"domain":"cookbook", "intent":"recipe_search", "slot":{"ingredient_name":"eggplant"}}, where domain represents the target classification identifier, intent represents the intent identifier, and slot represents the keyword. The target classification identifier is obtained from the expression, and the target knowledge graph is determined using the target classification identifier. The intent identifier is also obtained from the expression, and the target set is determined in the target knowledge graph using the intent identifier to obtain the target recommendation file.

[0105] S3. Select recommended information from the target recommendation file based on the keywords, and use the selected recommended information as the response information corresponding to the voice information.

[0106] In this embodiment of the invention, recommended information is selected from the target recommendation file based on the keywords. In the target recommendation file, there may be multiple pieces of information that match the keywords, and recommended information needs to be selected from multiple pieces of information that match the keywords.

[0107] In this embodiment of the invention, the target category identifier, the intent identifier, and the keyword each correspond to different query priorities. The query priority of the target category identifier is set as the first priority, the query priority of the intent identifier as the second priority, and the query priority of the keyword as the third priority. After determining the target category identifier, the intent identifier, and the keyword corresponding to the voice information, the process of determining the recommended information using the target category identifier, the intent identifier, and the keyword is to perform a layer-by-layer query in descending order of query priority to determine the recommended information.

[0108] In other words, determining recommended information through the target classification identifier, the intent identifier, and the keywords involves a three-level query process: First, a target knowledge graph is determined based on the target classification identifier, which has the highest query priority; second, a target set is determined in the target knowledge graph based on the intent identifier, which has the highest query priority; and finally, recommended information corresponding to the keywords is queried in the target set based on the keywords, which have the highest query priority.

[0109] For example, for the text information "the recipe of low-fat sandwich", the target classification identifier, the intent identifier and the keyword corresponding to the text information have been obtained through the natural language processing model; wherein, the target classification identifier is "recipe", the intent identifier is "low-fat", and the keyword is "sandwich"; then, according to the target classification identifier, the target knowledge graph is determined as the recipe knowledge graph in the plurality of knowledge graphs, and according to the intent identifier, the target set is determined as the low-fat set in the recipe knowledge graph, and the recommended information corresponding to "sandwich" is queried in the low-fat set.

[0110] Specifically, step S3 comprises:

[0111] S31, querying a plurality of candidate information corresponding to the keyword in the target recommendation file.

[0112] In the embodiment of the present application, a plurality of candidate information corresponding to the keyword is queried in the target recommendation file, for example, in the above example, the recommended information corresponding to "sandwich" is queried in the low-fat set, and there may be a plurality of recipes of sandwich in the low-fat set, and the plurality of recipes of "sandwich" queried are taken as the plurality of candidate information corresponding to the keyword.

[0113] S32, obtaining a weight value corresponding to each candidate information in the plurality of candidate information.

[0114] In the embodiment of the present application, the weight value can be the click volume corresponding to each candidate information, the weight value can be the score corresponding to each candidate information, the weight value can also be the degree of favor corresponding to each candidate information, or the weight value can be a comprehensive value of the click volume, the score and the degree of favor.

[0115] S33, determining a preset number of recommended information in the plurality of candidate information based on the obtained weight values, wherein the weight value of each recommended information is greater than that of any non-recommended information, and the non-recommended information is the candidate information in the plurality of candidate information except the preset number of recommended information.

[0116] In the embodiment of the present application, the plurality of candidate information is sorted according to the weight value corresponding to each candidate information respectively to obtain a candidate information queue. A preset number of recommended information is selected in the candidate information queue. The preset number can be set by the user, for example, set to 25. In the plurality of candidate information, the candidate information except the preset number of recommended information is the non-recommended information.

[0117] In the embodiment of the present application, after the target knowledge graph is determined, the cypher language suitable for knowledge graph query is determined according to the keyword and the intent identifier, and the query result is determined in the target knowledge graph according to the cypher language corresponding to the keyword and the intent identifier.

[0118] For example, if the intent identifier is determined to be "recipes", meaning the target set is "recipe set", then the Cypher language corresponding to the keyword and intent identifier could be: "MATCH (ingredient_name:ingredient{name:"eggplant"})-[:HAS_INGREDIENT]<-(recipes) RETURN recipes LIMIT 25", where the keyword is "eggplant", "ingredient" is the target recommendation file (ingredient set), and 25 is a preset value. This means querying the ingredient set: ingredient for 25 recommendations matching the keyword "eggplant".

[0119] In this embodiment of the invention, the number of candidate information items may be less than a preset value. For example, if the voice information is: "Who is the lead singer of White Windmill?", there may only be one candidate information item. When the number of candidate information items is less than the preset value, the candidate information items will be used as the query result.

[0120] In practice, the preset value can also be set to 1, that is, the recommended information with the largest weight value is selected from the candidate information.

[0121] S4. Convert the response information into voice form to obtain voice response information, and play the voice response information.

[0122] In this embodiment of the invention, a dialogue-based query can be implemented. The user emits voice information, the device receives recommended information, converts the response information into voice response information, and plays the voice response information through a sound unit. Specifically, the response information can be converted into voice format using a Text-to-Speech (TTS) method to obtain voice response information, which is then played through a sound unit in the device.

[0123] In this embodiment of the invention, when the device executing the speech recognition-based recommendation method has a display function, the response information can be displayed.

[0124] In the embodiment of the present application, the voice information to be processed is acquired, and the classification information and the keyword corresponding to the voice information are determined; the target recommendation file corresponding to the voice information is determined according to the classification information; the recommendation information is selected in the target recommendation file according to the keyword, and the selected recommendation information is taken as the response information corresponding to the voice information. The classification information and the keyword corresponding to the voice information are determined, the target recommendation file meeting the classification information is determined first, and then the recommendation information in the target recommendation file is determined based on the keyword; through the combination of the classification information and the keyword, the recommendation information meeting the user's intention can be queried, and the accuracy of determining the recommendation information in the voice interaction is improved.

[0125] Based on the above-mentioned recommendation method based on voice recognition, referring to Figure 4 The embodiment of the present application further provides a query device based on voice recognition, comprising:

[0126] A voice information processing module is configured to acquire voice information to be processed, and determine classification information and a keyword corresponding to the voice information.

[0127] A target recommendation file determination module is configured to determine a target recommendation file corresponding to the voice information according to the classification information.

[0128] A recommendation module is configured to select recommendation information in the target recommendation file according to the keyword, and take the selected recommendation information as response information corresponding to the voice information.

[0129] Further, referring to Figure 5 The voice information processing module comprises a voice collection unit, an automatic speech recognition (ASR) unit and a natural language processing (NLP) unit. The voice collection unit is configured to acquire voice information; the ASR unit is configured to convert the voice information into text information, and the NLP unit is configured to identify the text information to obtain the classification information and the keyword corresponding to the voice information. The query device based on voice recognition further comprises a text-to-speech (TTS) unit, a sound production unit and a display unit. The TTS unit is configured to convert the response information into voice response information, the sound production unit is configured to play the voice response information, and the display unit is configured to display the response information.

[0130] In the implementation, the voice acquisition unit acquires voice information, sends the voice information to the automatic voice recognition unit, converts the voice information into text information through the automatic voice recognition unit, processes the text information through the natural language processing unit to obtain classification information and a keyword corresponding to the text information, determines a target recommendation file corresponding to the voice information according to the classification information through the target recommendation file determination module, and selects recommendation information in the target recommendation file according to the keyword through the recommendation module. That is, a target knowledge graph is determined in the target knowledge graph, and the recommendation information is queried in the target knowledge graph through the keyword. The recommendation information is taken as response information corresponding to the voice information, the application information is converted into voice response information through the text-to-voice unit, the voice response information is played through the occurrence unit, and the response information is displayed through the display unit.

[0131] In one embodiment, the present application provides a computer device which can be a terminal, and the internal structure is as shown in the figure. Figure 6 The computer device includes a processor, a memory, a network model interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network model interface of the computer device is used to communicate with external terminals through a network model connection. The computer program is executed by the processor to implement a recommendation method based on voice recognition. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0132] Those skilled in the art can understand that Figure 6 The figure only shows a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0133] The embodiment of the present application provides a computer device including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:

[0134] Acquire voice information to be processed, and determine classification information and a keyword corresponding to the voice information;

[0135] determine a target recommendation file corresponding to the voice information according to the classification information;

[0136] select recommendation information in the target recommendation file according to the keyword, and take the selected recommendation information as response information corresponding to the voice information.

[0137] The embodiment of the application further provides a computer readable storage medium, which has a computer program stored thereon, and the computer program is executed by a processor to realize the following steps:

[0138] acquire voice information to be processed, and determine classification information and a keyword corresponding to the voice information;

[0139] determine a target recommendation file corresponding to the voice information according to the classification information;

[0140] select recommendation information in the target recommendation file according to the keyword, and take the selected recommendation information as response information corresponding to the voice information.

[0141] The technical features of the above embodiments can be combined in any manner, and to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0142] The above embodiments only express several implementation manners of the application, the description is relatively specific and detailed, however, it should not be understood as the limitation of the patent scope of the application. It should be pointed out that, for the ordinary skilled in the art, several modifications and improvements can be made without departing from the concept of the application, and these all belong to the protection scope of the application. Therefore, the patent protection scope of the application should be subject to the appended claims.

Claims

1. A recommendation method based on speech recognition, characterized in that, include: Acquire the voice information to be processed, and determine the classification information and keywords corresponding to the voice information. The classification information includes a target classification identifier and an intent identifier. Determining the target recommendation file corresponding to the voice information based on the classification information includes: The target knowledge graph corresponding to the target classification identifier is queried among several preset knowledge graphs, and the target knowledge graph is used as the target recommendation file corresponding to the voice information; wherein, the classification identifiers of the several knowledge graphs are different from each other; In the target knowledge graph, query the target set whose set identifier matches the intent identifier. The set is obtained by classifying the data corresponding to the target knowledge graph from different perspectives, and each set has its own corresponding set identifier. The target recommendation file is replaced with the target set found in the query to obtain the replaced target recommendation file; Based on the keywords, recommended information is selected from the target recommendation file, and the selected recommended information is used as the response information corresponding to the voice information.

2. The recommendation method based on speech recognition according to claim 1, characterized in that, The determination of the classification information and keywords corresponding to the voice information specifically includes: The voice information is recognized to obtain the corresponding text information; The text information is input into a natural language processing model, which then outputs the classification information and keywords corresponding to the speech information.

3. The recommendation method based on speech recognition according to claim 1, characterized in that, The knowledge graphs include at least: a recipe knowledge graph, a music knowledge graph, and a video knowledge graph; the recipe knowledge graph is categorized by a first category identifier, the music knowledge graph by a second category identifier, and the video knowledge graph by a third category identifier. Accordingly, querying the target knowledge graph corresponding to the target classification identifier in a preset set of knowledge graphs includes: When the target classification identifier is the first classification identifier, the target knowledge graph corresponding to the first classification identifier is queried from a number of preset knowledge graphs and it is the recipe knowledge graph; When the target classification identifier is the second classification identifier, the target knowledge graph corresponding to the second classification identifier is queried in a number of preset knowledge graphs and it is a music knowledge graph; When the target classification identifier is a third classification identifier, the target knowledge graph corresponding to the third classification identifier is queried from a number of preset knowledge graphs and is a video knowledge graph.

4. The recommendation method based on speech recognition according to claim 3, characterized in that, The step of selecting recommendation information from the target recommendation file based on the keywords specifically includes: Query several candidate information corresponding to the keywords in the target recommendation file; Obtain the weight value corresponding to each of the several candidate information; Based on the obtained weight values, a preset number of recommended information items are determined from the plurality of candidate information items. The weight value of each recommended information item is greater than that of any non-recommended information item. The non-recommended information items are the candidate information items other than the preset number of recommended information items from the plurality of candidate information items.

5. The recommendation method based on speech recognition according to claim 1, characterized in that, After selecting the recommended information as the response information corresponding to the voice information, the method further includes: The response information is converted into voice form to obtain voice response information, and the voice response information is played.

6. A query device based on speech recognition, characterized in that, include: The voice information processing module is used to acquire voice information to be processed and determine the classification information and keywords corresponding to the voice information. The classification information includes a target classification identifier and an intent identifier. The target recommendation file determination module is used to determine the target recommendation file corresponding to the voice information based on the classification information, including: querying the target knowledge graph corresponding to the target classification identifier in a preset number of knowledge graphs, and using the target knowledge graph as the target recommendation file corresponding to the voice information; wherein the classification identifiers of the number of knowledge graphs are different from each other; querying the target set whose set identifier is consistent with the intent identifier in a number of sets of the target knowledge graph, wherein the number of sets are obtained by classifying the data corresponding to the target knowledge graph from different perspectives, and each set has its own corresponding set identifier; replacing the target recommendation file with the queried target set to obtain the replaced target recommendation file; The recommendation module is used to select recommended information from the target recommendation file based on the keywords, and use the selected recommended information as the response information corresponding to the voice information.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the recommendation method based on speech recognition as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the recommendation method based on speech recognition as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Recommendation method and device of goods pictures

    CN107833082A

  • A method and system for content recommendation

    CN109408717A

  • Product recommendation method and device, electronic equipment and storage medium

    CN109615425A