Media asset acquisition method and device based on voice interaction
By extracting the text features of user voice commands and image features of media sources, and performing multimodal matching, the problem of inaccurate media acquisition caused by single-modal feature matching in the prior art is solved, and the accuracy and user experience of media acquisition are improved.
Patent Information
- Application Number
- CN202311552827.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
When obtaining media resources through voice interaction, the prior art only uses single-modal features when matching, which makes it difficult to accurately obtain target media resources that meet user needs when the single-modal features of different media resources are similar, and the user experience is poor.
By receiving user voice commands, the user's text features and keywords are extracted, and the media image features and text keywords are obtained from the media database, and multimodal feature matching is performed to determine media with similarity greater than the threshold as the target media.
It improves the accuracy of obtaining target media funds based on user voice, reduces the difficulty of users to select from a large number of media funds, and improves user experience.
Smart Images

Figure CN120020757A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of voice interaction technologies, and particularly to a method and apparatus for obtaining media resources based on voice interaction. Background Art
[0002] With the gradual progress and development of society, more and more intelligent devices can achieve intelligent interaction with users through voice assistant applications. When a user issues a voice command to an intelligent device through voice, the intelligent device can analyze the user's intention through a natural language understanding module to output a result corresponding to the voice command. For example, a user can play various media resources such as popular TV dramas and songs on the intelligent device through voice search.
[0003] Currently, when obtaining target media resources through voice interaction, the user's voice command is often analyzed to match a large number of media resources to lock the corresponding target media resources. In the prior art, during the matching process, the text features in the user's voice command are matched with the corresponding media resource features, that is, the unimodal features of the user's voice command are used to match the unimodal features of the media resources. However, in the above method, when the matching degree between the unimodal features of the user's voice command and the unimodal features of the media resources only meets the requirements, the media resource is determined as the target media resource. Therefore, in the case where the unimodal features corresponding to different media resources are very similar, a large number of target media resources that meet the user's needs will be obtained, and the user still needs to further select from a large number of target media resources, thus making it impossible to accurately obtain the target media resource that meets the user's needs and resulting in a poor user experience.
[0004] Therefore, how to improve the accuracy of obtaining media resources based on voice interaction has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present application provide a method and apparatus for obtaining media resources based on voice interaction, which are used to improve the accuracy of obtaining target media resources based on voice interaction.
[0006] In a first aspect, embodiments of the present application provide a method for obtaining media resources based on voice interaction, including:
[0007] Receiving a user voice command, and extracting the user text features and user text keywords of the user voice command; the user voice command is a command used by the user to obtain target media resources;
[0008] Obtaining the media image features and media text keywords respectively corresponding to multiple media resources from a media resource database;
[0009] For each of multiple media assets, match the media asset image features with the user text features to obtain a first similarity, and match the media asset text keywords with the user text keywords to obtain a second similarity;
[0010] Determine that the media assets for which the first similarity is greater than a first threshold and / or the second similarity is greater than a second threshold are the target media assets, and output the target media assets.
[0011] As an optional implementation manner of an embodiment of this application, after receiving a user voice command, the method further includes:
[0012] Perform semantic recognition on the user voice command, and determine whether the user voice command conforms to the scenario of obtaining media assets according to the semantic recognition result;
[0013] If so, perform the step of extracting the user text features and user text keywords of the user voice command;
[0014] If not, perform semantic rejection on the user voice command.
[0015] As an optional implementation manner of an embodiment of this application, before obtaining the media asset image features and media asset text keywords corresponding to multiple media assets from the media asset database, the method further includes:
[0016] Perform intent recognition on the user voice command to determine the user intent corresponding to the user voice command; determine a corresponding media asset list according to the user intent;
[0017] and / or;
[0018] Obtain the page identifier corresponding to the current media asset display page; determine a corresponding media asset list based on the page identifier corresponding to the media asset display page;
[0019] The obtaining the media asset image features and media asset text keywords corresponding to multiple media assets from the media asset database includes:
[0020] Obtain the media asset image features and media asset text keywords corresponding to each media asset in the media asset list from the media asset database.
[0021] As an optional implementation manner of an embodiment of this application, the obtaining the media asset image features and media asset text keywords corresponding to each media asset in the media asset list includes:
[0022] Obtain the media asset identifiers corresponding to each media asset in the media asset list;
[0023] Based on a preset correspondence relationship, obtain the media resource image features and media resource text keywords corresponding to the media resource identifiers of the respective media resources from the media resource database; the preset correspondence relationship includes the correspondence relationship between the media resource image features and media resource text keywords and the media resource identifiers.
[0024] As an optional implementation manner of an embodiment of the present application, the method for establishing the media resource database includes:
[0025] Obtain the cover images and media resource texts corresponding to the respective media resources;
[0026] Extract features from the cover images corresponding to the respective media resources to obtain the media resource image features, and extract keywords from the media resource texts to obtain the media resource text keywords;
[0027] Correspondingly save the media resource image features and media resource text keywords corresponding to the respective media resources with the media resource identifiers corresponding to the respective media resources, and obtain a media resource database.
[0028] As an optional implementation manner of an embodiment of the present application, after the step of outputting the target media resource, it further includes:
[0029] When the number of the target media resources is multiple, feedback the multiple target media resources to the user;
[0030] Receive a user supplementary voice instruction, and return to execute the step of extracting the user text features and user text keywords of the user voice instruction until the media resources determined that the first similarity is greater than the first threshold and / or the second similarity is greater than the second threshold are the target media resources and output the target media resources, so as to obtain and output a first target media resource from the multiple target media resources.
[0031] In a second aspect, an embodiment of the present application provides a media resource acquisition device based on voice interaction, including:
[0032] A receiving unit, configured to receive a user voice instruction, and extract the user text features and user text keywords of the user voice instruction; the user voice instruction is an instruction for the user to obtain a target media resource;
[0033] An analysis unit, configured to obtain the media resource image features and media resource text keywords respectively corresponding to multiple media resources from the media resource database;
[0034] A matching unit, configured to, for each of the multiple media resources, match the media resource image features with the user text features to obtain a first similarity, and match the media resource text keywords with the user text keywords to obtain a second similarity;
[0035] A determination unit, configured to determine that the media assets with the first similarity greater than a first threshold and / or the second similarity greater than a second threshold are the target media assets, and output the target media assets.
[0036] As an optional implementation manner of an embodiment of the present application, the media asset acquisition device based on voice interaction further includes: an identification unit, specifically configured to perform semantic recognition on the user voice command, and determine whether the user voice command conforms to the scenario of acquiring media assets according to the semantic recognition result; if so, execute the step of extracting the user text features and user text keywords of the user voice command; if not, perform semantic rejection on the user voice command.
[0037] As an optional implementation manner of an embodiment of the present application, the identification unit is further configured to perform intention recognition on the user voice command to determine the user intention corresponding to the user voice command; determine a corresponding media asset list according to the user intention; and / or; obtain a page identifier corresponding to the current media asset display page; determine a corresponding media asset list based on the page identifier corresponding to the media asset display page; the step of obtaining the media image features and media text keywords respectively corresponding to multiple media assets from the media asset database includes: obtaining the media image features and media text keywords corresponding to each media asset in the media asset list from the media asset database.
[0038] As an optional implementation manner of an embodiment of the present application, the media asset acquisition device based on voice interaction further includes: a building unit, specifically configured to obtain the cover image and media text corresponding to each media asset; perform feature extraction on the cover image corresponding to each media asset to obtain the media image features, and perform keyword extraction on the media text to obtain the media text keywords; correspond and save the media image features and media text keywords corresponding to each media asset with the media identifier corresponding to each media asset to obtain a media asset database.
[0039] As an optional implementation manner of an embodiment of the present application, the determination unit is further configured to, when the number of target media assets is multiple, feedback the multiple target media assets to the user; receive a user supplementary voice command, and return to execute the step of extracting the user text features and user text keywords of the user voice command until the media assets with the first similarity greater than the first threshold and / or the second similarity greater than the second threshold are determined as the target media assets and output the target media assets, so as to obtain and output a first target media asset from the multiple target media assets.
[0040] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program; the processor is used to cause the electronic device to implement the media asset acquisition method based on voice interaction described in any one of the above embodiments when executing the computer program.
[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is caused to implement the media asset acquisition method based on voice interaction described in any one of the above embodiments.
[0042] In a fifth aspect, an embodiment of the present application provides a vehicle, including: the media asset acquisition device based on voice interaction described in the second aspect or the electronic device described in the third aspect.
[0043] The media asset acquisition method based on voice interaction provided by the embodiment of the present application is specifically as follows: receiving a user voice instruction, and extracting the user text feature and user text keyword of the user voice instruction; the user voice instruction is an instruction for the user to acquire a target media asset; obtaining the media image feature and media text keyword corresponding to each of multiple media assets from a media asset database; for each of the multiple media assets, matching the media image feature with the user text feature to obtain a first similarity, and matching the media text keyword with the user text keyword to obtain a second similarity; determining that the media asset with the first similarity greater than a first threshold and / or the second similarity greater than a second threshold is the target media asset, and outputting the target media asset. By obtaining the user text feature and user text keyword of the user voice instruction, and then obtaining the media image feature and media text keyword corresponding to each media asset in the media asset list, and matching the user voice instruction and the media asset from two modalities of image and text to obtain the target media asset. Compared with the prior art where only the single-modal features of the user voice instruction and the media asset are matched, the present application starts from the perspective of multi-modal features and matches the multi-modal features of the user voice instruction and the media asset. Even when the single-modal features corresponding to different media assets are very similar, the matching results of other modalities can be combined to obtain the target media asset, which can improve the accuracy of obtaining the target media asset according to the user voice, and thus improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the attached drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 One of the step flowcharts of the media asset acquisition method based on voice interaction provided by the embodiments of the present application;
[0047] Figure 2 Another step flowchart of the media asset acquisition method based on voice interaction provided by the embodiments of the present application;
[0048] Figure 3 Schematic diagram of the media asset list display interface of the media asset acquisition method based on voice interaction provided by the embodiments of the present application;
[0049] Figure 4 Schematic diagram of the system framework of the media asset acquisition method based on voice interaction provided by the embodiments of the present application;
[0050] Figure 5 Schematic diagram of the structure of the media asset acquisition device based on voice interaction provided by the embodiments of the present application;
[0051] Figure 6 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0052] In order to better understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0053] In the following description, many specific details are set forth to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0054] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In addition, in the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more.
[0055] It should be noted that, in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article, or device that includes the element.
[0056] An embodiment of the present application provides a media asset acquisition method based on voice interaction. Referring to Figure 1 as shown, the media asset acquisition method based on voice interaction includes the following steps S101 - S104:
[0057] S101. Receive a user voice command, and extract the user text feature and user text keyword of the user voice command.
[0058] Wherein, the user voice command is a command used by the user to acquire a target media asset.
[0059] In some embodiments, the user voice command is user voice that conforms to the media asset acquisition scenario and is a command used by the user to acquire a target media asset. Specifically, the user voice command can be obtained by constantly listening to the user voice through a sound collection device provided in the vehicle. Exemplarily, when the user is driving a vehicle, the user voice can be listened to through sound collection devices provided at multiple different positions on the vehicle. The user voice command can be determined from the user voice by capturing whether the user voice includes a pre-set keyword. The pre-set keyword can be phrases or word groups such as "watch a movie" or "listen to music". Specifically, the user voice command can be voice texts with pre-set keywords such as "I want to listen to rock music" or "I want to watch an anime movie". Additionally, the user voice command can also be text input by the user in a specified interaction dialog box. For example, the text "I want to watch a movie" is input in the dialog box of the interaction interface.
[0060] Specifically, the extraction of the text feature of the user voice command can be to convert the user voice command into a text feature through a text encoder (TextEncoder). The method for extracting the user text keyword can be to extract the text keyword of the user voice command through a sentence pattern template. Exemplarily, the user text keyword can be any type of instance extracted from the user voice command text. For example, it can be a person's name, an item name, a time, etc., so as to match with the subsequent girl text keyword.
[0061] S102. Obtain the media asset image features and media asset text keywords corresponding to multiple media assets from the media asset database.
[0062] In the embodiments of the present application, the media asset database includes a large number of different types of media assets; for example: there are various different types of media assets in the media asset database, such as movies, TV dramas, music, variety shows, etc.; to meet different user needs. Multiple media asset image features and media asset text keywords corresponding to multiple media assets can be obtained from the media asset database according to the user voice command. Exemplarily, when the user wants to watch a war movie, the media asset image features and media asset text keywords corresponding to all war movies can be obtained from the media asset database.
[0063] It should be noted that the media asset database in the present application can be set in the cloud, which is a database that has stored the media asset image features and media asset text keywords corresponding to all media assets. For each media asset in the media asset database, the corresponding media asset image features and media asset text keywords will be obtained in advance, and a corresponding relationship will be generated and saved between each media asset and its corresponding media asset image features and media asset text keywords. The media asset text keywords can be any type of instance extracted from the media asset information corresponding to the media asset. For example, they can be the names of people, item names, times, etc. in the media asset name, or the names of people, item names, times, etc. in the media asset cover image. So as to quickly obtain the corresponding media asset image features and media asset text keywords after obtaining at least one media asset list corresponding to the user voice, thereby facilitating subsequent matching with the user text features and user text keywords corresponding to the voice command, improving the calculation efficiency, quickly locking the target media asset, and further improving the user experience.
[0064] S103. For each of the multiple media assets, match the media asset image features with the user text features to obtain a first similarity, and match the media asset text keywords with the user text keywords to obtain a second similarity.
[0065] In some embodiments, since both the user text features and the media asset image features are feature vectors, the method of matching the media asset image features with the user text features to obtain the first similarity can be calculated by cosine similarity. For two vectors, the cosine of their included angle can be used to represent their approximate degree. Specifically, it can refer to the following formula:
[0066]
[0067] Where and respectively represent two different feature vectors. In the embodiments of the present application and They are respectively the vector forms of the user text features and the media asset image features.
[0068] Specifically, the cosine value of the angle between two vectors is the cosine similarity of the two vectors. The above formula directly comes from the definition of the inner product. The numerator is the inner product of the two vectors, and the denominator is the product of the norms (lengths) of the two vectors. Through the formula, it can be intuitively considered that to a certain extent, the influence of the vector length is eliminated, and the cosine similarity reflects the difference in direction.
[0069] Exemplarily, when the angle is 0, the two vectors are in the same direction, which is equivalent to the highest similarity, and the cosine value is 1. When the angle is 90°, the two vectors are perpendicular, and the cosine is 0. When the angle is 180°, the two vectors are in the opposite direction, and the cosine is -1. Cosine similarity is widely used for comparison, such as face comparison and voice comparison, to quickly judge the similarity of two pictures or two pieces of voice, and then judge whether they are from the same person. When an image or voice sample has n-dimensional features, we can consider it as an n-dimensional vector. When comparing two samples using cosine similarity, it is to measure the cosine value of the angle between the two n-dimensional vectors and its magnitude.
[0070] In some embodiments, the method of matching the media asset text keywords with the user text keywords to obtain the second similarity may be to use the word vectors trained by a shallow neural network model such as Word2Vec (a language model). By adding the word vectors bit by bit, the vector representations of the media asset text keywords and the user text keywords are obtained, and then the cosine similarity is also used to calculate the similarity of the two vectors.
[0071] Exemplarily, when the user text keyword is an actor's name "Li Si", and the media asset text keyword also includes the actor's name "Li Si", then the current similarity is 1.
[0072] S104. Determine that the media assets with the first similarity greater than the first threshold and / or the second similarity greater than the second threshold are the target media assets, and output the target media assets.
[0073] Specifically, the maximum similarity is 1. After multiple experimental tests, the first threshold and the second threshold can both be set to 0.7. This application does not make specific limitations on this and can be set according to the actual situation. Then, the media assets in the media asset list with the similarity between the user text features and the media asset image features greater than 0.7 and / or the similarity between the user text keywords and the media asset text keywords greater than 0.7 can be used as the target media assets. That is to say, as long as any one of the first similarity and the second similarity corresponding to the media asset is greater than the threshold, or both are greater than the threshold, the corresponding media asset is determined to be the target media asset.
[0074] It should be noted that there may also be multiple target media assets. When there is 1 target media asset that meets the conditions, the target media asset can be directly played. When there are multiple target media assets that meet the conditions, the user can continue to issue a voice command for selection, and then execute the above steps S101 - S104 until the target media asset that the user finally needs to play is obtained.
[0075] In the embodiment of the present application, since the matching is performed from two modalities of text and image to obtain the similarity, through the multi-modal matching, the target media asset obtained in the present application is also more accurate, and thus more in line with the user's needs.
[0076] The method for obtaining media assets based on voice interaction provided by the embodiment of the present application is specifically as follows: receiving a user voice command, and extracting the user text feature and user text keyword of the user voice command; the user voice command is a command used by the user to obtain a target media asset; obtaining the media image feature and media text keyword respectively corresponding to multiple media assets from a media asset database; for each media asset among the multiple media assets, matching the media image feature with the user text feature to obtain a first similarity, and matching the media text keyword with the user text keyword to obtain a second similarity; determining the media asset for which the first similarity is greater than a first threshold and / or the second similarity is greater than a second threshold as the target media asset, and outputting the target media asset. In the embodiment of the present application, by obtaining the user text feature and user text keyword of the user voice command, and then obtaining the media image feature and media text keyword corresponding to each media asset in the media asset list, and matching the user voice command and the media assets from two modalities of image and text to obtain the target media asset. Compared with the prior art where the single-modal features of the user voice command and the media assets are matched, the present application starts from the perspective of multi-modal features and matches the multi-modal features of the user voice command and the media assets. Even when the single-modal features corresponding to different media assets are very similar, the matching results of other modalities can be combined to obtain the target media asset, which can improve the accuracy of obtaining the target media asset according to the user voice, and thus enhance the user experience.
[0077] As an extension and refinement of the above embodiment, the embodiment of the present application provides a method for obtaining media assets based on voice interaction, referring to Figure 2 As shown, the method for obtaining media assets based on voice interaction includes the following steps S201 - S208:
[0078] S201. Perform semantic recognition on the user voice command, and judge whether the user voice command conforms to the scenario of obtaining media assets according to the semantic recognition result.
[0079] In some embodiments, user speech can be collected by a plurality of sound collection devices disposed at different positions on the vehicle, and then semantic recognition is performed on the user speech command to determine whether the current user speech conforms to the scenario of media asset acquisition. Among them, the scenarios of media asset acquisition can be different scenarios such as watching movies, listening to music, watching variety shows, etc. Specifically, the semantic analysis and recognition refers to the process of deeply understanding and analyzing natural language text to identify semantic information such as entities, relationships, and emotions in the text. In this application, natural language text can be analyzed and understood based on a deep neural network model, or other methods can be used to perform semantic recognition on user speech commands. The embodiments of this application do not make any limitations in this regard.
[0080] In the above step S201, if the user speech command conforms to the scenario of obtaining media assets, the following step S202 is executed:
[0081] S202. Extract the user text features and user text keywords of the user speech command.
[0082] In some embodiments, if the user speech conforms to the scenario of obtaining media assets, it is necessary to further extract the user text features and user text keywords of the user speech command to match with the media image features and media text keywords corresponding to multiple media assets to obtain the target media asset.
[0083] In the above step S201, if the user speech command does not conform to the scenario of obtaining media assets, the following step S203 is executed:
[0084] S203. Perform semantic rejection recognition on the user speech command.
[0085] It should be noted that not all of the user's speech can be used to obtain the target media asset. It is also possible to chat with others, etc. Then, when the user speech does not conform to the current scenario of obtaining media assets, there is no need to perform semantic recognition on the user speech, that is, perform semantic rejection recognition.
[0086] S204. Perform intention recognition on the user speech command to determine the user intention corresponding to the user speech command.
[0087] In some embodiments, in response to the wake-up command, it is necessary to analyze the user speech command to obtain the user intention, and then obtain the corresponding at least one media asset list according to the user intention. This is convenient for subsequently obtaining the corresponding media asset list from the media asset database and feeding it back to the user for the user to further select.
[0088] S205. Determine the corresponding media asset list according to the user intention.
[0089] Specifically, the method for performing intent recognition on the user voice instruction may be as follows: Based on the Natural Language Understanding (NLU) method, obtain the intent of the user voice instruction to obtain the corresponding media list from the media database; and obtain the corresponding media list in the media database according to the user intent corresponding to the user voice. The natural language understanding technology imitates the basic way of human language understanding, that is, sentence segmentation, word segmentation, etc., cuts the text into a series of units with semantics and grammar, and will calculate the "key information" in the text, such as entities, triples, intents, events, etc. based on the text representation data. With this information, the machine can understand the user's language and obtain the user intent.
[0090] In some embodiments, in the media database, media corresponding to multiple types are classified and saved, and are identified according to their corresponding types. After obtaining the user intent, the media type currently requested by the user can be known according to the user intent. After determining the media type, the corresponding media list can be obtained from the media database according to the identifier corresponding to the media type.
[0091] After analyzing the user voice instruction, when the text corresponding to the user voice instruction is "I want to listen to music", obtain the corresponding music list, that is, obtain at least one media list composed of all the songs in the current in-vehicle music application. It should be noted that the number of media in a media list is preset. For example, a media list may include 10 media. When the media that meet the wake-up instruction exceed 10, the 11th media and subsequent media are composed into a second media list, and so on, a media list sequence can be obtained. Then, the cover images and names of all the songs in the first media list are generated into a first-page display interface in a preset display manner and displayed on the in-vehicle screen of the vehicle for the user to further select and play. Exemplarily, the display interface may be in units of 10 songs per page and be displayed in multiple pages according to the actual situation. Refer to Figure 3 As shown, it is a schematic diagram of the display interface, including 10 songs that meet the user intent, as well as the cover images and names corresponding to these 10 songs, and also includes a page sequence 31 and a paging control 32. The present application embodiment does not make any limitation on the display manner of the display interface.
[0092] Optionally, in the embodiments of the present application, the method for determining the media list may also refer to the following steps:
[0093] Step A: Obtain the page identifier corresponding to the current media display page.
[0094] In some embodiments, after the user opens the application for obtaining media assets, the user will directly enter the media asset display page preset in the application. For example, it is the home page of the application; a media asset list including multiple media assets will be displayed on this home page. Specifically, the media asset list may include currently popular TV dramas, currently popular movies, etc.; or it will jump to other pages with the user's touch, and then display the media asset list under other pages. Therefore, in the embodiments of the present application, multiple media assets in different display pages of each application will be obtained in advance to generate media asset lists corresponding to different display pages respectively. Then, corresponding page identifiers will be set for each media asset list, and the page identifiers and media asset lists will be stored in the media asset database in a corresponding manner. Since the multiple media assets in different pages may change at any time according to the settings of the application, the present application can also periodically update the media asset lists corresponding to each display page in the media asset database, as well as the media asset image features and media asset text keywords corresponding to each media asset in each media asset list, and store them in the media asset database.
[0095] Step B: Based on the page identifier corresponding to the media asset display page, determine the corresponding media asset list.
[0096] In some embodiments, the corresponding page representation can be obtained according to the display page where the user is currently located to obtain the media asset list corresponding to this display page.
[0097] Exemplarily, when the user sees the media asset list displayed on the display interface, an instruction can be further issued to obtain the target media asset. For example, if a song by Zhang San appears in the current media asset list and the user wants to listen to this song, then only need to issue the user's voice. The corresponding user voice text can be: "I want to listen to that song by Zhang San". After the sound collection device receives it, the semantic analysis of the user voice will be immediately performed to obtain the intention corresponding to the user voice, and then randomly lock the target media asset as the only song by Zhang San in the current media asset list and play it.
[0098] In some embodiments, after obtaining the user text features and user text keywords corresponding to the user voice instruction, only need to obtain the media asset image features and media asset text keywords corresponding to each media asset in the media asset list from the media asset database according to the preset corresponding relationship, so that when the user needs to play the target media asset, the user text features and user text keywords can be more quickly directly matched with the media asset image features and media asset text keywords corresponding to each media asset in the media asset list to obtain the target media asset.
[0099] In some embodiments, after obtaining the user requirements from the user voice instruction, that is, the user text features and user text keywords, then obtain multiple media information in at least one media list from the media database, that is, the media image features and media text keywords, and then match the user information with the multiple media information in the media list, and for each media in the media list, obtain the matching degree between each media and the user information to determine the media in the media list that best matches the user information, and use it as the target media and output it.
[0100] S206. Obtain the media image features and media text keywords corresponding to each media in the media list from the media database.
[0101] In the above step S206, the implementation method for obtaining the media image features and media text keywords corresponding to each media in the media list can refer to the following steps 1 and 2:
[0102] Step 1. Obtain the media identifiers corresponding to each media in the media list.
[0103] In the embodiments of the present application, all media will store the media identifiers of each media and their corresponding media image features and media text keywords in the media database. After obtaining multiple media that meet the user requirements according to the voice instruction, only need to obtain their corresponding identifiers, that is, the media image features and media text keywords corresponding to the media can be obtained according to the identifiers.
[0104] Step 2. Based on the preset correspondence, obtain the media image features and media text keywords corresponding to the media identifiers of each media from the media database.
[0105] Among them, the preset correspondence includes the correspondence between the media image features and media text keywords and the media identifiers.
[0106] In some embodiments, the method for establishing the media database includes the following steps 2.1 to 2.3:
[0107] Step 2.1. Obtain the cover image and media text corresponding to each media.
[0108] In some embodiments, the media usually has a preset cover image and media text. The cover image can be a movie poster, a TV drama poster, a music album cover, etc.; the media text can be the name of the movie, the name of the TV drama, the name of the song, and the labels recognized from the cover image. The labels can be various semantic labels such as the names of the characters on the cover, the names of the items, and the colors.
[0109] In some embodiments, the media resource database can be established in the cloud, and then the cloud is connected to the corresponding smart device, for example, the smart device can be a smart car machine, a smart phone, a tablet computer, etc. After the target media resource is determined according to the voice command, the corresponding media resource information is sent to the smart device, for example, the target media resource is sent from the cloud database to the smart car machine system, so as to decode and play the target media resource on the vehicle's media display screen.
[0110] It should be noted that the media resource acquisition method based on voice interaction provided in the embodiment of the present application can also periodically update the media resource database.
[0111] In some embodiments, new media may be generated at any time, for example, new movies are released, new songs are released, etc., or some media violates regulations and is banned; the corresponding application for obtaining media adjusts the media list of the preset page, so the media list in the media database also needs to be updated. Specifically, the cloud media database is connected to the network, so that new media can be obtained in real time and updated in the media database. At the same time, the media database can also be manually updated by developers. As media in the media database are added or deleted, the corresponding identifiers and media image features and media text keywords of the media also need to be added or deleted to ensure that the information in the current media database is accurate and does not affect user viewing.
[0112] Step 2.2: Extract features from the cover images corresponding to the media assets to obtain the features of the media asset images, and extract keywords from the media asset text to obtain the keywords of the media asset text.
[0113] In some embodiments, the cover image vector features are extracted to obtain the corresponding media asset image features, and the media asset text is subjected to keyword extraction to obtain the media asset text keywords.
[0114] Specifically, the feature extraction of the cover images corresponding to the various media assets can be performed using commonly used image feature extraction methods such as an image encoder ImageEncoder, a neural network feature extraction (Deep Learning Internet) method, and a scale-invariant feature transformation method. The keyword extraction of the media asset text can be performed using an image entity keyword extractor ImageClassier.
[0115] Step 2.3: Save the media asset image features and media asset text keywords corresponding to each media asset and the media asset identifier corresponding to each media asset in correspondence with each other to obtain a media asset database.
[0116] S207. For each of the multiple media assets, match the media asset image features with the user text features to obtain a first similarity, and match the media asset text keywords with the user text keywords to obtain a second similarity.
[0117] S208. Determine that the media assets with the first similarity greater than a first threshold and / or the second similarity greater than a second threshold are the target media assets, and output the target media assets.
[0118] As a supplement and refinement to the above embodiments, the media asset acquisition method based on voice interaction further includes the following Step 1 and Step 2:
[0119] Step 1. When the number of the target media assets is multiple, feed back the multiple target media assets to the user.
[0120] Step 2. Receive the user supplementary voice instruction, and return to execute the extraction of the user text features and user text keywords of the user voice instruction until the step of determining that the media assets with the first similarity greater than a first threshold and / or the second similarity greater than a second threshold are the target media assets and outputting the target media assets, so as to obtain and output a first target media asset from the multiple target media assets.
[0121] Exemplarily, when the user requests target media assets about rock music and the number of the obtained target media assets is 3, that is, 3 song covers of rock music and corresponding names will be displayed on the display interface fed back to the user. At this time, the user can issue the user supplementary voice instruction according to the image features reflected by the cover image. For example, on the current page, only the cover image of Song 2 includes "a photo of singer Li Si". At this time, if the user locks Song 2 that he wants to listen to according to the photo of the singer, the text of the user supplementary voice instruction that can be issued is: "I want to listen to Li Si's song". After receiving the voice instruction with the text of "I want to listen to Li Si's song", the above steps S101 to S104 are also executed to obtain the target media assets that match the user supplementary voice instruction, that is, the song corresponding to the cover image including "a photo of singer Li Si" and play it to meet the user's needs.
[0122] As an extension and refinement to the above embodiments, the embodiment of the present application provides a system framework diagram based on voice interaction. Refer to Figure 4 As shown, the system framework of the media asset acquisition method based on voice interaction includes the following modules:
[0123] Analysis module 41: used for semantic recognition of the user voice instruction, and judging whether the user voice instruction conforms to the scenario of obtaining media assets according to the semantic recognition result.
[0124] Extraction module 42: configured to extract the user text features and user text keywords of the user voice command when the user voice command meets the scenario of obtaining media resources.
[0125] Media resource information acquisition module 43: configured to obtain the media image features and media text keywords respectively corresponding to multiple media resources from the media resource database.
[0126] Matching module 44: configured to, for each of the multiple media resources, match the media image features with the user text features to obtain a first similarity, and match the media text keywords with the user text keywords to obtain a second similarity.
[0127] Output module 45: configured to determine that the media resources for which the first similarity is greater than a first threshold and / or the second similarity is greater than a second threshold are the target media resources, and output the target media resources.
[0128] Update module 46: configured to periodically update the media resource database, and based on the update of the media resource database, update the identifiers, media image features, and media text keywords corresponding to the media resources, and save them to the preset corresponding relationship.
[0129] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application further provides a media resource acquisition device based on voice interaction. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, details of the foregoing method embodiment will not be repeated one by one in this embodiment. However, it should be clear that the media resource acquisition device based on voice interaction in this embodiment can correspondingly implement all the content of the foregoing method embodiment.
[0130] An embodiment of the present application provides a media resource acquisition device based on voice interaction, Figure 5 which is a structural schematic diagram of the media resource acquisition device based on voice interaction. As Figure 5 shown, the media resource acquisition device 500 based on voice interaction includes:
[0131] Receiving unit 501, configured to receive a user voice command, and extract the user text features and user text keywords of the user voice command; the user voice command is a command used by the user to obtain a target media resource;
[0132] Analysis unit 502, configured to obtain the media image features and media text keywords respectively corresponding to multiple media resources from the media resource database.
[0133] A matching unit 503, configured to match the media asset image features with the user text features for each of multiple media assets to obtain a first similarity, and match the media asset text keywords with the user text keywords to obtain a second similarity;
[0134] A determining unit 504, configured to determine that the media assets for which the first similarity is greater than a first threshold and / or the second similarity is greater than a second threshold are the target media assets, and output the target media assets.
[0135] As an optional implementation manner of an embodiment of the present application, the media asset acquisition device based on voice interaction further includes: an identification unit, specifically configured to perform semantic recognition on the user voice command, and determine whether the user voice command conforms to the scenario of acquiring a media asset according to the semantic recognition result; if so, execute the step of extracting the user text features and user text keywords of the user voice command; if not, perform semantic rejection on the user voice command.
[0136] As an optional implementation manner of an embodiment of the present application, the identification unit is further configured to perform intent recognition on the user voice command to determine the user intent corresponding to the user voice command; determine a corresponding media asset list according to the user intent; and / or; obtain a page identifier corresponding to the current media asset display page; determine a corresponding media asset list based on the page identifier corresponding to the media asset display page; the step of obtaining the media asset image features and media asset text keywords respectively corresponding to multiple media assets from the media asset database includes: obtaining the media asset image features and media asset text keywords respectively corresponding to each media asset in the media asset list from the media asset database.
[0137] As an optional implementation manner of an embodiment of the present application, the media asset acquisition device based on voice interaction further includes: a building unit, specifically configured to obtain the cover image and media asset text corresponding to each media asset; perform feature extraction on the cover image corresponding to each media asset to obtain the media asset image features, and perform keyword extraction on the media asset text to obtain the media asset text keywords; and correspondingly save the media asset image features and media asset text keywords respectively corresponding to each media asset with the media asset identifier corresponding to each media asset to obtain a media asset database.
[0138] As an optional implementation manner of an embodiment of the present application, the determining unit 504 is further configured to, when the number of target media assets is multiple, feed back the multiple target media assets to the user; receive a user supplementary voice instruction, and return to execute the steps of extracting the user text feature and user text keywords of the user voice instruction until the media asset determined to have the first similarity greater than the first threshold and / or the second similarity greater than the second threshold is the target media asset and output the target media asset, so as to obtain a first target media asset from the multiple target media assets and output it.
[0139] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device. Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present disclosure is shown in Figure 6 As shown, the electronic device provided in this embodiment includes: a memory 601 and a processor 602. The memory 601 is used to store a computer program; the processor 602 is configured to execute the audio data processing method provided in the above embodiment when executing the computer program.
[0140] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the computing device is enabled to implement the media asset acquisition method based on voice interaction provided in the above embodiment.
[0141] Based on the same inventive concept, an embodiment of the present application further provides a vehicle, and the vehicle includes the media asset acquisition device based on voice interaction provided in the above embodiment or the electronic device provided in the above embodiment.
[0142] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0143] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0144] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0145] Computer-readable media include permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media, such as modulated data signals and carrier waves.
[0146] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for acquiring media resources based on voice interaction, characterized in that: include: Receive a user voice command, and extract user text features and user text keywords of the user voice command; the user voice command is a command used by the user to obtain target media assets; Acquire media image features and media text keywords corresponding to a plurality of media respectively from a media database; For each of the multiple media assets, matching the media asset image feature with the user text feature to obtain a first similarity, and matching the media asset text keyword with the user text keyword to obtain a second similarity; The media asset whose first similarity is greater than a first threshold and / or whose second similarity is greater than a second threshold is determined as the target media asset, and the target media asset is output.
2. The method according to claim 1, characterized in that After receiving the user's voice command, the method further includes: Performing semantic recognition on the user voice command, and judging whether the user voice command meets the scenario of obtaining media resources according to the semantic recognition result; If yes, then executing the step of extracting the user text features and user text keywords of the user voice instruction; If not, the user voice command is semantically rejected.
3. The method according to claim 1, characterized in that Before acquiring the media asset image features and media asset text keywords respectively corresponding to the plurality of media assets from the media asset database, the method further includes: Performing intent recognition on the user voice command to determine the user intent corresponding to the user voice command; and determining a corresponding media resource list according to the user intent; and / or; Obtaining a page identifier corresponding to the current media resource display page; and determining a corresponding media resource list based on the page identifier corresponding to the media resource display page; The acquiring of media image features and media text keywords respectively corresponding to a plurality of media assets from a media asset database includes: The media asset image features and media asset text keywords corresponding to each media asset in the media asset list are obtained from the media asset database.
4. The method according to claim 3, characterized in that The acquiring of media image features and media text keywords corresponding to each media in the media list from the media database includes: Obtaining a media asset identifier corresponding to each media asset in the media asset list; Based on the preset corresponding relationship, the media asset image features and media asset text keywords corresponding to the media asset identifiers of the respective media assets are obtained from the media asset database; the preset corresponding relationship includes the corresponding relationship between the media asset image features and media asset text keywords and the media asset identifiers.
5. The method according to claim 1, characterized in that The method for establishing the media resource database includes: Get the cover image and text corresponding to each media asset; Performing feature extraction on the cover images corresponding to the respective media assets to obtain the features of the media asset images, and performing keyword extraction on the media asset texts to obtain the keywords of the media asset texts; The media asset image features and media asset text keywords corresponding to each media asset are stored in correspondence with the media asset identifiers corresponding to each media asset to obtain a media asset database.
6. The method according to claim 1, characterized in that After the step of outputting the target media asset, the method further includes: When there are multiple target media assets, feeding back the multiple target media assets to the user; Receive a user's supplementary voice instruction, and return to execute the step of extracting the user text features and user text keywords of the user voice instruction, until the step of determining that the first similarity is greater than a first threshold, and / or the media asset with the second similarity greater than a second threshold is the target media asset and outputting the target media asset, so as to obtain the first target media asset from the multiple target media assets and output it.
7. A media resource acquisition device based on voice interaction, characterized in that: include: A receiving unit, used to receive a user voice command and extract user text features and user text keywords of the user voice command; the user voice command is a command for the user to obtain target media assets; An analysis unit, used to obtain media image features and media text keywords corresponding to a plurality of media assets respectively from a media asset database; a matching unit, configured to match, for each of the plurality of media assets, the image feature of the media asset with the feature of the user text to obtain a first similarity, and to match the media asset text keyword with the user text keyword to obtain a second similarity; A determination unit is configured to determine that the media asset whose first similarity is greater than a first threshold and / or whose second similarity is greater than a second threshold is the target media asset, and output the target media asset.
8. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to enable the electronic device to implement the media resource acquisition method based on voice interaction as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computing device, the computing device implements the media resource acquisition method based on voice interaction as described in any one of claims 1 to 6.
10. A vehicle, characterized in that: include: The media resource acquisition device based on voice interaction as described in claim 7 or the electronic device as described in claim 8.