Method and apparatus for determining multimedia content, electronic device, and storage medium
By obtaining target material and text information, and using preset models to determine material characteristics and intention types, the problem of high soundtrack search error rate in the prior art is solved, and efficient matching between soundtrack and material content is achieved.
Patent Information
- Application Number
- PCT/CN2024/127779
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the soundtrack method based on keyword retrieval has a problem that the error rate is high and does not match the video content.
By obtaining target material and text information, using preset models to determine material features, text features and intent types, determining target multimedia content in the target object set based on the intent type, matching text features and material features, optimizing the soundtrack search process.
It improves the accuracy of soundtrack search and improves the matching between the soundtrack and the material content.
Smart Images

Figure CN2024127779_03072025_PF_FP_ABST
Abstract
Description
Method, device, electronic device and storage medium for determining multimedia content
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed on December 29, 2023, with application number 202311869528.8 and invention name “A method, device, electronic device and storage medium for determining multimedia content”. The entire contents of that application are incorporated by reference into this application. Technical Field
[0003] The present disclosure relates to data processing technology, and more particularly to a method, device, electronic device, and storage medium for determining multimedia content. Background Art
[0004] More and more users are sharing their videos on short video platforms. When sharing videos, they usually choose a piece of music as the soundtrack to make the video more interesting.
[0005] Currently, it is possible to search pre-built music libraries based on keywords to obtain music clips containing the corresponding keywords as the soundtrack for the video. However, the soundtracks obtained by keyword search have a high error rate and do not match the video content.
[0006] Summary of the Invention
[0007] The present disclosure provides a method, device, electronic device and storage medium for determining multimedia content, which can improve the accuracy of music retrieval and enhance the matching degree between music and material content.
[0008] In a first aspect, an embodiment of the present disclosure provides a method for determining multimedia content, including:
[0009] Acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0010] Determining material features, text features, and intent type based on the target material and text information using a preset model;
[0011] According to the intent type, target multimedia content corresponding to the target material is determined based on at least one target object in a target object set, wherein the target object set includes keywords in text features, text vectors corresponding to text features, and material vectors corresponding to material features.
[0012] In a second aspect, an embodiment of the present disclosure further provides a device for determining multimedia content, including:
[0013] An information acquisition module, configured to acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0014] A feature determination module, configured to determine material features, text features, and intent type based on the target material and text information using a preset model;
[0015] A content determination module is used to determine the target multimedia content corresponding to the target material based on the intention type and at least one target object in the target object set, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
[0016] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0017] one or more processors;
[0018] a storage device for storing one or more programs,
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining multimedia content as described in any embodiment of the present disclosure.
[0020] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the method for determining multimedia content as described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0022] FIG1 is a schematic flow chart of a method for determining multimedia content provided by an embodiment of the present disclosure;
[0023] FIG2 is a flow chart of another method for determining multimedia content provided by an embodiment of the present disclosure;
[0024] FIG3 is a schematic diagram of a music matching link provided by an embodiment of the present disclosure;
[0025] FIG4 is a schematic structural diagram of a device for determining multimedia content provided by an embodiment of the present disclosure;
[0026] FIG5 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0034] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0037] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0038] Figure 1 is a flow chart of a method for determining multimedia content provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of music retrieval. The method can be executed by a multimedia content determination device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC or a server, etc.
[0039] As shown in FIG1 , the method includes:
[0040] S110: Acquire target material and text information.
[0041] The text information represents the natural language used to describe multimedia content. The multimedia content may be a music clip used for soundtrack, and the music clip database is used to store the music clip. Optionally, the music clip database may store the music clip in the form of a vector. For example, the music information such as the lyrics or song title of the music clip is mapped to the form of a music text vector, and the music text vector is saved. The music information such as the lyrics or song title of the music clip is mapped to the form of a music text vector, and the music video of the music clip is mapped to the form of a music video vector. In the embodiment of the present disclosure, the text information may include information associated with the soundtrack intention input by the user. For example, the text information is determined through the content of the conversation between the user and the intelligent robot. The content of the conversation may include setting the scene, setting the location, setting the time, setting the event, and setting the music preference.
[0042] The target material can be a video or image to be matched with music. For example, the target material can include a video or image with a set content theme. Acquiring the target material can involve acquiring a pre-stored collection of videos or images to be matched with music. Alternatively, the target material can involve acquiring a collection of videos or images currently shot and uploaded by the user.
[0043] For example, a video or image to be matched with music is obtained as the target material. Natural language describing the music requirements for the target material is obtained as text. For example, a user-uploaded video is obtained, and the natural language describing the music requirements for the submitted video is obtained. Alternatively, at least one user-uploaded image is obtained, and the natural language describing the music requirements for the image is obtained.
[0044] S120: Determine material features, text features, and intent type based on the target material and text information using a preset model.
[0045] Among them, the preset model representation is a deep learning model trained using a large amount of text data, which is used to generate natural language text, as well as generate image description information corresponding to the image content.
[0046] Material features can characterize the descriptive information corresponding to the target material. For example, material features may include video features or picture features, etc. Video features may include long text about the video description (such as content description text) and target tags, etc. Picture features may include long text about the picture description (such as content description text) and target tags, etc. Text features may include long text that summarizes or describes text information (such as semantic description text) and target tags. The long text includes semantic description text or content description text, etc. Further, keyword extraction is performed on the long text to obtain keywords. A preset template is used to extract slots from the long text to obtain slots. Among them, the slot is an identifier that characterizes the musical intention in the text information. For example, the slot includes the style of music and the author, etc. The preset template is a template that contains pre-set slots.
[0047] The intent type can be a music intent determined by a preset model based on text information. For example, the preset model performs intent analysis on the text information to determine the intent type. Intent types can include a first intent type and a second intent type. The first intent type can indicate that the text information contains a clear description of multimedia content. The second intent type can indicate that the text information contains a vague description of the multimedia content and / or a description of the attributes of the target material. For example, the first intent type can indicate a clear intent, while the second intent type can indicate a vague intent. A clear intent can be a music intent corresponding to a text message such as "I want a video using XX music." A vague intent can be a music intent corresponding to a text message such as "I want a warm video," where "warmth" is a description of the video's attributes. Alternatively, a vague intent can be a music intent corresponding to a text message such as "I want a travel-themed video using songs by XX singer," where "travel" is a description of the video's attributes and "XX singer" is a vague description of the multimedia content. Alternatively, a vague intent can be a music intent corresponding to a text message such as "I want a cheerful song," where "cheerfulness" is a vague description of the multimedia content.
[0048] Exemplarily, determining material features, text features and intent types based on the target material and text information through a preset model includes: inputting the target material and text information into the preset model; outputting material features based on the target material through the preset model; and outputting text features and intent types based on the text information through the preset model.
[0049] In the embodiment of the present disclosure, the target material and text information are respectively input into the preset model. The preset model can understand the text information and obtain the semantic description text and target label corresponding to the text information as text features. In addition, the preset model can also perform intent analysis on the text information to obtain the intent type. The preset model can also understand the target material and obtain the content description text and target label corresponding to the target material as material features. For example, according to the preset correspondence between the video length and the number of frames extracted, the soundtrack video is frame extracted to obtain a video frame sequence. The video frame sequence is input into the preset model, and the content of each video frame is understood by the preset model, and the content description text and target label about the video content are output.
[0050] In some embodiments, keywords are extracted from the semantic description text to obtain keywords from the text features. Slot extraction is performed on the semantic description text using a preset template to obtain slots corresponding to the text features. Keywords are extracted from the content description text to obtain keywords from the material features. Slot extraction is performed on the content description text using a preset template to obtain slots corresponding to the material features.
[0051] S130: Determine target multimedia content corresponding to the target material based on at least one target object in a target object set according to the intent type.
[0052] The target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features. For example, the target object set includes multiple target objects, and the target objects can be keywords in the text features, text vectors corresponding to the text features, or material vectors corresponding to the material features.
[0053] The target multimedia content may represent multimedia content that matches the content of the target material and conforms to the description corresponding to the text information. For example, the target multimedia content may be a music clip or a set of music clip candidates that matches the content of the video to be accompanied by music and conforms to the user's intention. If the target multimedia content is a set of music clip candidates, the set of music clip candidates is displayed for the user to select. For example, based on the selection operation input by the user for the set of music clip candidates, a music clip is determined from the set of music clip candidates as the soundtrack for the target material. Alternatively, based on the selection operation input by the user for the set of music clip candidates, at least two music clips are determined from the set of music clip candidates, and the at least two music clips are spliced together as the soundtrack for the target material.
[0054] Exemplarily, determining, according to the intent type, target multimedia content corresponding to the target material based on at least one target object in the target object set includes:
[0055] If the intent type represents a single intent, the target multimedia content corresponding to the target material is determined based on the single intent and at least one target object in the target object set.
[0056] If the intent type represents at least two intents, the first candidate multimedia content corresponding to the target material is determined based on at least one target object in the target object set according to each intent, the first candidate multimedia content is sorted according to the text features and material features, and the target multimedia content corresponding to the target material is selected from the first candidate multimedia content according to the sorting result. The first candidate multimedia content includes at least candidate multimedia content in the following three cases: candidate multimedia content obtained by retrieving a preset multimedia content library based on keywords, candidate multimedia content obtained by retrieving a preset multimedia content library based on text vectors, and candidate multimedia content obtained by retrieving a preset multimedia content library based on a combination of text vectors and material vectors. The preset multimedia content library may include a music clip database, etc.
[0057] Furthermore, for the first intent type, a preset multimedia content library is searched according to the keywords in the text features to obtain the target multimedia content corresponding to the target material, wherein the first intent type represents that the text information contains a clear description of the multimedia content.
[0058] Furthermore, for the second intent type, a preset multimedia content library is retrieved according to the text vector corresponding to the text feature to obtain a second candidate multimedia content, wherein the second intent type represents that the text information contains a vague description of the multimedia content and / or a description of the attributes of the target material.
[0059] A preset multimedia content library is searched in combination with the text vector corresponding to the text feature and the material vector corresponding to the material feature to obtain a third candidate multimedia content; and a target multimedia content corresponding to the target material is determined based on the second candidate multimedia content and the third candidate multimedia content.
[0060] For fuzzy intent, we retrieve the text vectors corresponding to the text features and the material vectors corresponding to the material features of the target object set based on the fuzzy intent. We then call different services to search the music clip database, obtaining two recall results. These results are then combined to determine the target multimedia content corresponding to the target material.
[0061] Optionally, determining the target multimedia content corresponding to the target material based on the second candidate multimedia content and the third candidate multimedia content includes: using the text features and / or material features as matching objects, matching the content features of the second candidate multimedia content with the matching objects to obtain a first similarity; matching the content features of the third candidate multimedia content with the matching objects to obtain a second similarity; sorting the second candidate multimedia content and the third candidate multimedia content based on the first and second similarities, and determining the target multimedia content corresponding to the target material based on the sorting results. The content features may include music text vectors and music video vectors corresponding to the second candidate multimedia content or the third candidate multimedia content. Determining the similarity between the music text vector of the second candidate multimedia content and the text vector corresponding to the text features, and determining the similarity between the music video vector of the second candidate multimedia content and the material vector corresponding to the material features, and combining these two similarities to obtain the first similarity. A similar method can be used to determine the second similarity corresponding to the third candidate multimedia content, which will not be further described here. The sorting of the second candidate multimedia content and the third candidate multimedia content may be shuffled.
[0062] Optionally, after determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the intent type, the method further includes: obtaining content information corresponding to the target multimedia content; if the target multimedia content meets the copyright verification conditions, generating soundtrack material based on the target multimedia content, content information and target material, wherein the soundtrack material is a target material with soundtrack. The content information may include music information related to the target multimedia content. For example, the content information includes information such as the music clip identifier, song title, author and music clip duration. The target multimedia content is subjected to copyright verification. If the target multimedia content passes the copyright verification, the target multimedia content is encapsulated according to the content information, and the soundtrack material is obtained by combining the encapsulated target multimedia content and the target material.
[0063] The technical solution of the disclosed embodiment inputs target material and text information into a preset model to obtain material features, text features, and intent types output by the preset model. Then, based on the intent type, the corresponding target multimedia content is determined for the target material based on at least one target object in the target object set. The disclosed embodiment uses a preset model to deeply understand and determine the intent of the target material and text information, determining target multimedia content based on different target objects according to different intents, thereby improving the accuracy of soundtrack retrieval. Since soundtrack retrieval is performed in conjunction with the content of the target material, the matching degree between the soundtrack and the material content can be improved.
[0064] FIG2 is a flow chart of another method for determining multimedia content provided by an embodiment of the present disclosure. Based on the above embodiments, the method determines the target multimedia content corresponding to the target material when the intent type represents at least two intents. As shown in FIG2 , the method includes:
[0065] S210: Acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content.
[0066] S220. Determine material features, text features, and intent types based on the target material and text information through a preset model, wherein the intent type represents at least two intents.
[0067] For example, the intent types include at least clear intent and vague intent.
[0068] In some embodiments, a user-uploaded video to be set to music is obtained, and the user enters text information such as "I heard song C while traveling, so I want to make a travel video with music by singer A, preferably with music clips similar to song C." The text information is then used to identify intent using a preset model to determine the intent type. Intent types include at least two types: explicit intent and ambiguous intent. "Song C" corresponds to an explicit intent, and "travel video with music by singer A" corresponds to an ambiguous intent.
[0069] S230: For a clear intention, search a preset multimedia content library according to the keywords in the text feature to obtain at least one candidate multimedia content.
[0070] Exemplarily, in the case of clear intent, keywords in the text features included in the target object set are obtained according to the clear intent, and a music clip database is searched according to the keywords to obtain a clear intent recall result.
[0071] S240 : For the fuzzy intent, search a preset multimedia content library according to the text vector corresponding to the text feature to obtain at least one candidate multimedia content.
[0072] Exemplarily, a music clip database is searched based on the text vector to obtain target multimedia content corresponding to the target material. The music clip database includes music text vectors corresponding to the music clips. The music text vectors can be determined based on the corresponding content, such as the lyrics or song title. A similarity between the text vector and a music text vector in the music clip database is determined, and a first music clip whose similarity exceeds a preset similarity threshold is identified. The first music clip is then selected as the second candidate multimedia content corresponding to the target material.
[0073] S250 : For the fuzzy intent, search a preset multimedia content library by combining the text vector corresponding to the text feature and the material vector corresponding to the material feature to obtain at least one candidate multimedia content.
[0074] Exemplarily, a music clip database is searched in combination with a text vector and a material vector to obtain target multimedia content corresponding to the target material. The music clip database also includes a multimodal vector corresponding to the music clip, and the multimodal vector includes a music text vector and a music video vector. The music video vector can be determined based on the music video content corresponding to the music clip. The similarity between the text vector and the text vector of the music text vector in the music clip database is determined. Also, the similarity between the material vector and the material vector of the music video vector in the music clip database is determined. A second music clip is determined, wherein both the similarity of the text vector and the similarity of the material vector exceed a preset similarity threshold, and the second music clip is used as the third candidate multimedia content corresponding to the target material. The second candidate multimedia content and the third candidate multimedia content are fused and mixed, and the fuzzy intent recall result is determined based on the mixed result.
[0075] Alternatively, if the target video's duration is less than a set time threshold, the number of video frames obtained through video extraction may be too small to fully capture the target material's material features. In this case, the material features are updated based on the target material's content features to generate new material features. The new material features include more features of the target material, enabling more accurate matching of similar music clips from the music clip database.
[0076] S260: Take the candidate multimedia content corresponding to the explicit intent and the candidate multimedia content corresponding to the vague intent as first candidate multimedia content, and sort the first candidate multimedia content according to the text features and the material features.
[0077] Exemplarily, the text features and the material features are used as matching objects, and the content features of the music clip in the clear intent recall result are matched with the matching object to obtain a third similarity. The content of the music clip in the fuzzy intent recall result is matched with the matching object to obtain a fourth similarity. The clear intent recall result and the fuzzy intent recall result are sorted based on the third and fourth similarities to obtain a sorting result of the first candidate multimedia content. The clear intent recall result includes candidate multimedia content corresponding to the clear intent. The fuzzy intent recall result includes candidate multimedia content corresponding to the fuzzy intent.
[0078] S270: Determine target multimedia content among the first candidate multimedia contents according to the ranking result of the first candidate multimedia contents.
[0079] Figure 3 is a schematic diagram of a music composition link provided by an embodiment of the present disclosure. As shown in Figure 3, the user sends text information 301 and target material 302 to the server. After receiving the text information 301 and target material 302, the server inputs the text information 301 and target material 302 into a preset model 303. The preset model 303 performs intent understanding and inference on the text information 301 to obtain text features and intent type. The preset model 303 also performs content understanding and inference on the target material 302 to obtain material features. Based on the intent type, the corresponding search service 304 is called to perform the following steps: If the intent type includes a clear intent 305, the preset multimedia content library is searched based on the keywords included in the text features to obtain a corresponding search result list 306. The top music clip from the search result list 306 is selected as the clear intent recall result 307. If the intent type includes a fuzzy intent 308, the vector similarity between the text vector corresponding to the text feature and the music text vectors of each music clip in the music clip text vector database 309 is determined. The music clip whose vector similarity exceeds a set threshold is selected as the first fuzzy intent recall result 310. A material vector 311 corresponding to the material feature is determined. The multimodal feature vectors of each music clip in a music clip multimodal vector database 313 are retrieved in combination with the text vector 312 and the material vector 311. Music clips whose vector similarity exceeds a set threshold are determined as second fuzzy intent recall results 314. The multimodal feature vectors include music text vectors and music video vectors.
[0080] Optionally, if the duration of target material 302 is less than a set time threshold, the material features are updated using the content features of target material 302 to obtain new material features. The new material features are then mapped to material vectors 311. The multimodal feature vectors of each music segment in the music segment multimodal vector database 313 are retrieved in combination with the text vector 312 and material vector 311. Music segments whose vector similarity exceeds a set threshold are identified as second fuzzy intent recall results 314.
[0081] If the first fuzzy intent recall result 310 includes M music clips and the second fuzzy intent recall result 314 includes N music clips, then the M+N music clips are sorted according to their similarity to the text features and the material features, and the top 1 music clip is determined as the fuzzy intent recall result 315 based on the sorting result. Alternatively, the top x music clips are determined as the fuzzy intent recall result 315 based on the sorting result.
[0082] The clear intention recall result 307 and the fuzzy intention recall result 315 are mixed according to the text features and material features, and the top 1 music clip is determined as the target multimedia content according to the sorting result. The result output module 316 is used to encapsulate the music information related to the target multimedia content and perform copyright verification and other actions. If the target multimedia content meets the copyright verification conditions, the soundtrack material is generated based on the encapsulated target multimedia content and target material, and the soundtrack material is sent to the user end. Among them, the result output module 316 can be a natural language generation (NLG) module in a preset model, etc. The NLG module is determined based on specific rules and neural network models.
[0083] Optionally, after shuffling the explicit intent recall results 307 and the ambiguous intent recall results 315, a list of top x music clips is determined based on the sorting results. The result output model 316 is used to encapsulate the music information associated with the music clips in the music clip list and perform copyright verification. The music clips in the music clip list that meet the copyright verification criteria are then sent to the user for selection. The user's selected music clip is obtained, and a soundtrack material is generated based on the selected music clip and the target material, which is then sent to the user.
[0084] The technical solution of the disclosed embodiment performs a recall operation based on multiple intents corresponding to the intent type in combination with keyword, text feature and material feature retrieval, and determines the target multimedia content for the target material based on the text similarity between the recall result and the text information, thereby optimizing the music retrieval process, improving the retrieval accuracy, and improving the music matching degree of the target material.
[0085] Figure 4 is a schematic diagram of the structure of a multimedia content determination device provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server, etc.
[0086] As shown in FIG. 4 , the apparatus includes: an information acquisition module 410 , a feature determination module 420 , and a content determination module 430 .
[0087] An information acquisition module 410 is configured to acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0088] A feature determination module 420 is configured to determine material features, text features, and intent type based on the target material and text information using a preset model;
[0089] The content determination module 430 is used to determine the target multimedia content corresponding to the target material based on the intention type and at least one target object in the target object set, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
[0090] Furthermore, the content determination module 430 includes:
[0091] a first content determination unit, configured to determine, if the intent type represents a single intent, target multimedia content corresponding to the target material based on at least one target object in the target object set according to the single intent;
[0092] A second content determination unit is configured to determine, if the intent type represents at least two intents, first candidate multimedia content corresponding to the target material based on at least one target object in the target object set according to each intent, sort the first candidate multimedia content according to the text features and the material features, and select target multimedia content corresponding to the target material from the first candidate multimedia content according to the sorting result.
[0093] Optionally, the first content determination unit is specifically configured to:
[0094] For the first intent type, a preset multimedia content library is searched according to the keywords in the text features to obtain the target multimedia content corresponding to the target material, wherein the first intent type represents that the text information contains a clear description of the multimedia content.
[0095] Optionally, the first content determination unit is specifically configured to:
[0096] For the second intent type, searching a preset multimedia content library according to the text vector corresponding to the text feature to obtain second candidate multimedia content, wherein the second intent type indicates that the text information includes a vague description of the multimedia content and / or a description of the attributes of the target material;
[0097] searching a preset multimedia content library in combination with the text vector corresponding to the text feature and the material vector corresponding to the material feature to obtain third candidate multimedia content;
[0098] The target multimedia content corresponding to the target material is determined according to the second candidate multimedia content and the third candidate multimedia content.
[0099] Furthermore, the determining the target multimedia content corresponding to the target material according to the second candidate multimedia content and the third candidate multimedia content includes:
[0100] using the text feature and / or material feature as a matching object, matching the content feature of the second candidate multimedia content with the matching object to obtain a first similarity;
[0101] matching the content feature of the third candidate multimedia content with the matching object to obtain a second similarity;
[0102] The second candidate multimedia content and the third candidate multimedia content are sorted based on the first similarity and the second similarity, and the target multimedia content corresponding to the target material is determined according to the sorting result.
[0103] Optionally, the feature determination module 420 is specifically configured to:
[0104] Inputting the target material and text information into a preset model;
[0105] Outputting material features based on the target material through the preset model;
[0106] Output text features and intent types based on the text information through the preset model.
[0107] Optionally, the device further comprises:
[0108] A music soundtrack module is used to obtain content information corresponding to the target multimedia content after determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the intent type; if the target multimedia content meets the copyright verification conditions, generate music soundtrack material based on the target multimedia content, content information and target material, wherein the music soundtrack material is a target material with music soundtrack.
[0109] The multimedia content determination device provided in the embodiments of the present disclosure can execute the multimedia content determination method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0110] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0111] FIG5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG5 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG5 ) 500 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG5 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0112] As shown in FIG5 , the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.
[0113] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG5 shows the electronic device 500 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0114] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0115] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0116] The electronic device provided by the embodiment of the present disclosure and the method for determining multimedia content provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0117] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the method for determining multimedia content provided in the above embodiment is implemented.
[0118] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0119] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0120] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0121] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0122] Acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0123] Determining material features, text features, and intent type based on the target material and text information using a preset model;
[0124] According to the intent type, target multimedia content corresponding to the target material is determined based on at least one target object in a target object set, wherein the target object set includes keywords in text features, text vectors corresponding to text features, and material vectors corresponding to material features.
[0125] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0127] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0128] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0129] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0130] According to one or more embodiments of the present disclosure, Example 1 provides a method for determining multimedia content, including:
[0131] Acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0132] Determining material features, text features, and intent type based on the target material and text information using a preset model;
[0133] According to the intent type, target multimedia content corresponding to the target material is determined based on at least one target object in a target object set, wherein the target object set includes keywords in text features, text vectors corresponding to text features, and material vectors corresponding to material features.
[0134] According to one or more embodiments of the present disclosure, Example 2 is the method according to Example 1, wherein determining, based on the intent type, target multimedia content corresponding to the target material based on at least one target object in the target object set includes:
[0135] If the intent type represents a single intent, determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the single intent;
[0136] If the intent type represents at least two intents, the first candidate multimedia content corresponding to the target material is determined based on each intent and at least one target object in the target object set, the first candidate multimedia content is sorted according to the text features and material features, and the target multimedia content corresponding to the target material is selected from the first candidate multimedia content according to the sorting result.
[0137] According to one or more embodiments of the present disclosure, Example 3, according to the method of Example 2, wherein if the intent type represents a single intent, determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the single intent includes:
[0138] For the first intent type, a preset multimedia content library is searched according to the keywords in the text features to obtain the target multimedia content corresponding to the target material, wherein the first intent type represents that the text information contains a clear description of the multimedia content.
[0139] According to one or more embodiments of the present disclosure, Example 4, according to the method of Example 2, wherein if the intent type represents a single intent, determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the single intent includes:
[0140] For the second intent type, searching a preset multimedia content library based on the text vector corresponding to the text feature to obtain second candidate multimedia content, wherein the second intent type indicates that the text information includes a vague description of the multimedia content and / or a description of the attributes of the target material;
[0141] searching a preset multimedia content library in combination with the text vector corresponding to the text feature and the material vector corresponding to the material feature to obtain third candidate multimedia content;
[0142] The target multimedia content corresponding to the target material is determined according to the second candidate multimedia content and the third candidate multimedia content.
[0143] According to one or more embodiments of the present disclosure, Example 5 is the method according to Example 4, wherein determining the target multimedia content corresponding to the target material based on the second candidate multimedia content and the third candidate multimedia content includes:
[0144] using the text feature and / or material feature as a matching object, matching the content feature of the second candidate multimedia content with the matching object to obtain a first similarity;
[0145] matching the content feature of the third candidate multimedia content with the matching object to obtain a second similarity;
[0146] The second candidate multimedia content and the third candidate multimedia content are sorted based on the first similarity and the second similarity, and the target multimedia content corresponding to the target material is determined according to the sorting result.
[0147] According to one or more embodiments of the present disclosure, Example 6, according to the method of Example 1, wherein determining the material features, text features, and intent type based on the target material and text information using a preset model includes:
[0148] Inputting the target material and text information into a preset model;
[0149] Outputting material features based on the target material through the preset model;
[0150] Output text features and intent types based on the text information through the preset model.
[0151] According to one or more embodiments of the present disclosure, Example 7, according to the method of Example 1, further includes, after determining, based on the intent type and at least one target object in the target object set, the target multimedia content corresponding to the target material:
[0152] Acquiring content information corresponding to the target multimedia content;
[0153] If the target multimedia content meets the copyright verification condition, a soundtrack material is generated according to the target multimedia content, the content information and the target material, wherein the soundtrack material is a target material with soundtrack.
[0154] According to one or more embodiments of the present disclosure, Example 8 provides a device for determining multimedia content, including:
[0155] An information acquisition module, configured to acquire target material and text information, wherein the text information represents a natural language used to describe multimedia content;
[0156] A feature determination module, configured to determine material features, text features, and intent type based on the target material and text information using a preset model;
[0157] A content determination module is used to determine the target multimedia content corresponding to the target material based on the intention type and at least one target object in the target object set, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
[0158] According to one or more embodiments of the present disclosure, Example 9 provides an electronic device, the electronic device including:
[0159] one or more processors;
[0160] a storage device for storing one or more programs,
[0161] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining multimedia content as described in any one of Examples 1-7.
[0162] According to one or more embodiments of the present disclosure, Example 10 provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the method for determining multimedia content as described in any one of Examples 1-7.
[0163] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0164] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0165] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for determining multimedia content, comprising: Obtaining a target material and text information, where the text information represents a natural language for describing multimedia content; Determining material features, text features, and intention types based on the target material and the text information through a preset model; Based on the intention type, determining the target multimedia content corresponding to the target material based on at least one target object in a target object set, where the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
2. The method according to claim 1, wherein the determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the intention type includes: If the intention type represents a single intention, determining the target multimedia content corresponding to the target material based on the single intention and at least one target object in the target object set; If the intention type represents at least two intentions, determining a first candidate multimedia content corresponding to the target material based on each intention and at least one target object in the target object set, sorting the first candidate multimedia content according to the text features and material features, and selecting the target multimedia content corresponding to the target material from the first candidate multimedia content according to the sorting result.
3. The method according to claim 2, wherein the determining the target multimedia content corresponding to the target material based on the single intention and at least one target object in the target object set if the intention type represents a single intention includes: For a first intention type, retrieving a preset multimedia content library according to the keywords in the text features to obtain the target multimedia content corresponding to the target material, where the first intention type represents that the text information contains a clear description of the multimedia content.
4. The method according to claim 2, wherein the determining the target multimedia content corresponding to the target material based on the single intention and at least one target object in the target object set if the intention type represents a single intention includes: For a second intention type, retrieving a preset multimedia content library according to the text vector corresponding to the text features to obtain a second candidate multimedia content, where the second intention type represents that the text information contains a fuzzy description of the multimedia content and / or an attribute description of the target material; Retrieving a preset multimedia content library by combining the text vector corresponding to the text features and the material vector corresponding to the material features to obtain a third candidate multimedia content; Determining the target multimedia content corresponding to the target material according to the second candidate multimedia content and the third candidate multimedia content.
5. The method according to claim 4, wherein the determining the target multimedia content corresponding to the target material according to the second candidate multimedia content and the third candidate multimedia content includes: Using the text features and / or material features as matching objects, match the content features of the second candidate multimedia content with the matching objects to obtain a first similarity; Match the content features of the third candidate multimedia content with the matching objects to obtain a second similarity; Combine the first similarity and the second similarity to sort the second candidate multimedia content and the third candidate multimedia content, and determine the target multimedia content corresponding to the target material according to the sorting result.
6. The method according to claim 1, wherein the determining the material features, text features and intention type based on the target material and text information by a preset model includes: Input the target material and text information into the preset model; Output the material features based on the target material by the preset model; Output the text features and intention type based on the text information by the preset model.
7. The method according to claim 1, wherein after determining the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the intention type, further includes: Obtain the content information corresponding to the target multimedia content; If the target multimedia content meets the copyright verification condition, generate a background music material according to the target multimedia content, content information and target material, wherein the background music material is the target material with background music.
8. A device for determining multimedia content, comprising: An information acquisition module, configured to acquire a target material and text information, wherein the text information represents a natural language for describing multimedia content; A feature determination module, configured to determine material features, text features and intention type based on the target material and text information by a preset model; A content determination module, configured to determine the target multimedia content corresponding to the target material based on at least one target object in the target object set according to the intention type, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features and material vectors corresponding to the material features.
9. An electronic device, the electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining multimedia content according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, wherein the computer-executable instructions are used to execute the method for determining multimedia content according to any one of claims 1-7 when executed by a computer processor.
Citation Information
Patent Citations
Search recommendation method and device
CN104063418A
Music searching method and apparatus
CN104991943A
Image searching method and system
CN110083729A
Method and system for intelligently recommending background music based on video multi-dimensional features
CN110704682A
Background music generating method, storage medium and terminal equipment
CN110767201A
Cited By
Vector retrieval method and device for multiple service scenes, storage medium and program product
CN121301445A