Background music recommendation method and apparatus, electronic device and storage medium
By combining the model recommendation method of material characteristics and intention types, the problem of inaccurate soundtrack recommendation in the existing technology is solved, and the accurate matching of soundtrack and material content is achieved, and the effect of video editing is improved.
Patent Information
- Application Number
- PCT/CN2024/127737
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the recommendation of soundtracks in video editing relies on the determination of search terms, which easily returns inappropriate soundtracks, and the soundtrack does not match the video content, resulting in poor results.
The first target soundtrack is determined based on the target material through the first model, and the second model is determined based on the text information based on the text information, and the second target soundtrack is determined using the keywords, text vectors and material vectors in the target object set to replace the soundtrack of the target material.
It improves the accuracy of soundtrack recommendations, enhances the matching between the soundtrack and the material content, and optimizes the soundtrack search process.
Smart Images

Figure CN2024127737_03072025_PF_FP_ABST
Abstract
Description
Music recommendation method, device, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed on December 29, 2023, with application number 202311865321.3 and invention name “A music recommendation method, device, electronic device and storage medium”. The entire contents of that application are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present disclosure relate to data processing technology, and more particularly to a music recommendation method, device, electronic device, and storage medium. Background Art
[0003] With the development of mobile communication technology and Internet technology, it has become possible to instantly edit and produce videos shot by users and quickly share them on social platforms.
[0004] During video editing, choosing the right soundtrack helps optimize the video's presentation and, in turn, expand its reach. Currently, you can retrieve soundtracks by searching pre-built music libraries using keywords. However, this method relies on the accuracy of the search terms. If the search terms are incorrect, inappropriate soundtracks may be returned, or even no soundtracks may be retrieved at all. Furthermore, the soundtracks retrieved based on the search terms may not match the video content, resulting in poor soundtrack quality.
[0005] Summary of the Invention
[0006] The present disclosure provides a soundtrack recommendation method, device, electronic device, and storage medium, which can improve the accuracy of soundtrack recommendation and enhance the matching degree between soundtrack and material content.
[0007] In a first aspect, an embodiment of the present disclosure provides a soundtrack recommendation method, comprising: obtaining target material, determining a first target soundtrack based on the target material through a first model, and adding soundtrack to the target material based on the first target soundtrack, wherein the first model is trained based on target material samples determined based on historical interactive operations and soundtrack samples corresponding to the target material samples; obtaining text information, determining material features, text features and intent types based on the target material and text information through a second model, wherein the text information is a natural language representing the soundtrack intent of the target material; and determining a second target soundtrack corresponding to the target material based on at least one target object in a target object set according to the intent type, and replacing the soundtrack of the target material based on the second target soundtrack, wherein the target object set includes keywords in text features, text vectors corresponding to text features, and material vectors corresponding to material features.
[0008] In a second aspect, an embodiment of the present disclosure also provides a music recommendation device, which includes: a first music determination module, used to obtain target material, determine a first target music based on the target material through a first model, and add music to the target material based on the first target music, wherein the first model is trained based on target material samples determined by historical interactive operations and music samples corresponding to the target material samples; a feature determination module, used to obtain text information, determine material features, text features and intent types based on the target material and text information through a second model, wherein the text information is a natural language that characterizes the music intent of the target material; and a second music determination module, used to determine a second target music corresponding to the target material based on at least one target object in a target object set according to the intent type, and replace the music of the target material based on the second target music, wherein the target object set includes keywords in text features, text vectors corresponding to text features, and material vectors corresponding to material features.
[0009] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the music recommendation method as described in any embodiment of the present disclosure.
[0010] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the music recommendation method as described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0012] FIG1 is a schematic flow chart of a music recommendation method provided by an embodiment of the present disclosure;
[0013] FIG2 is a schematic diagram of a flow chart of a training method for a first model provided in an embodiment of the present disclosure;
[0014] FIG3 is a flow chart of another music recommendation method provided by an embodiment of the present disclosure;
[0015] FIG4 is a schematic diagram of a music matching link provided by an embodiment of the present disclosure;
[0016] FIG5 is a schematic structural diagram of a music recommendation device provided by an embodiment of the present disclosure;
[0017] FIG6 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0020] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0025] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0026] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0028] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0029] Figure 1 is a flow chart of a music recommendation method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of music recommendation. The method can be executed by a music recommendation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server, etc.
[0030] As shown in FIG1 , the method includes:
[0031] S110 , acquiring a target material, determining a first target soundtrack based on the target material using a first model, and adding soundtrack to the target material based on the first target soundtrack.
[0032] The first model is trained based on target material samples determined by historical interaction operations and the corresponding music samples. For example, the first model may be a neural network model trained based on target material samples determined by historical interaction operations and the corresponding music samples. Historical interaction operations may represent interaction information corresponding to historical material within a sampling period. For example, historical interaction operations include searching for, collecting, liking, or sharing the music of historical material.
[0033] The soundtrack represents a musical clip that is presented in conjunction with the target material. The music complements the content of the target material to enhance the artistic effect of the target material. In the disclosed embodiments, a music clip database can be used to store the music clips. Optionally, the music clip database can store the music clips in the form of vectors. For example, musical information such as lyrics or song titles of a music clip can be mapped into a music text vector and stored. A music video of a music clip can be mapped into a music video vector, and both the music text vector and the music video vector can be stored in the music clip database.
[0034] Optionally, determining the target material sample based on historical interaction operations may include sorting candidate material samples based on the frequency of historical interaction operations, and determining the target material sample based on the sorting results. Target material is obtained, and music text features and music video features of the soundtrack in the target material sample are obtained. A first model is trained based on the target material sample, the music text features, and the music video features, so that the first model learns the correlation between the features of the target material sample and the soundtrack features. The music text features may include music text vectors mapped to music information such as lyrics or song titles of a music clip. The music video features may include music video vectors mapped to the music video of the music clip. For example, within a set observation time period, the candidate materials are ranked based on the frequency of occurrence of the same slot feature when users search for soundtracks, with candidate materials corresponding to the most frequently occurring slot features being selected as target materials. Slot features are identification information in textual information that characterizes the intent of the soundtrack. For example, slot features include genre, author, and language. Optionally, a search request input by a user is obtained, and slot features are extracted from the search request using a preset template to obtain slot features. A preset template is a template that includes pre-set slots.
[0035] The target material can be a video or image to be matched with music. For example, the target material can include a video or image with a set content theme. Acquiring the target material can involve acquiring a pre-stored collection of videos or images to be matched with music. Alternatively, the target material can involve acquiring a collection of videos or images currently shot and uploaded by the user.
[0036] For example, a video or picture to be matched with music is obtained as a target material, the target material is input into a first model, a first target soundtrack is determined based on material features of the target material by the first model, and a soundtrack is added to the target material based on the first target soundtrack.
[0037] Optionally, music information corresponding to the first target soundtrack is obtained; if the first target soundtrack meets copyright verification conditions, soundtrack material is generated based on the first target soundtrack, the music information, and the target material, where the soundtrack material is the target material with the soundtrack. The music information may include music information related to the soundtrack. For example, the music information includes information such as the music clip identifier, song title, author, and duration of the music clip. A copyright verification is performed on the first target soundtrack. If the first target soundtrack passes the copyright verification, the first target soundtrack is packaged based on the music information, and the packaged first target soundtrack and the target material are combined to obtain the soundtrack material.
[0038] Due to the limitations of the training samples of the first model, some new features in the target material may not be perceived. Therefore, the first target soundtrack may not match the user's soundtrack intention. In order to better recommend soundtracks, the embodiment of the present disclosure also needs to obtain the user's soundtrack intention.
[0039] S120: Acquire text information, and determine material features, text features, and intent type based on the target material and the text information using a second model.
[0040] The text information is natural language that represents the musical intent of the target material. The musical intent of the target material is obtained by parsing the text information. For example, the text information may include user-entered information related to the musical intent. For example, the text information may be determined based on the content of a conversation between the user and the intelligent robot. The conversation content may include setting a scene, setting a location, setting a time, setting an event, and setting music preferences.
[0041] The second model represents a deep learning model trained using a large amount of text data, which is used to generate natural language text, as well as image description information corresponding to the image content, etc.
[0042] Material features can represent the descriptive information corresponding to the target material. For example, material features can include video features or image features. Video features can include long text describing the video (e.g., content description text) and target tags. Image features can include long text describing the image (e.g., content description text) and target tags. Text features can include long text summarizing or describing text information (e.g., semantic description text) and target tags. The long text includes semantic description text or content description text.
[0043] Optionally, keyword extraction is performed on the long text to obtain keywords. Optionally, slot extraction is performed on the long text using a preset template to obtain slot features.
[0044] The intent type may be the music intent determined by the second model based on the text information. For example, the second model performs intent analysis on the text information to obtain the intent type. The intent type may include a first intent type, a second intent type, and a third intent type, etc. The first intent type indicates that the text information contains a clear description of the music intent for the target material. The second intent type indicates that the text information contains a vague description of the music intent for the target material and / or a description of the attributes of the target material. The attribute description may include the theme attributes of the target material, etc. The third intent type indicates that the text information contains the music requirement for the target material, but does not contain a description of the music intent. For example, the first intent type may indicate a clear intent, the second intent type may indicate a vague intent, and the third intent type may indicate a completely vague intent. A clear intent may be the music intent corresponding to a text such as "I want a video using XX music." A vague intent may be the music intent corresponding to a text such as "I want a warm video," where warmth is an attribute description of the video. Alternatively, a fuzzy intent could be a music intent corresponding to a text such as "I want a travel-themed video with a song by singer xx," where "travel theme" describes the video's attributes and "xx singer" is a vague description of the music. Alternatively, a fuzzy intent could be a music intent corresponding to a text such as "I want a cheerful song," where "cheerful" is a vague description of the music. A completely fuzzy intent could be something like "I want to add music to my video."
[0045] Exemplarily, determining material features, text features and intent types based on the target material and text information through a second model includes: inputting the target material and text information into the second model; outputting material features based on the target material through the second model; and outputting text features and intent types based on the text information through the second model.
[0046] In the disclosed embodiment, the target material and text information are respectively input into the second model, and the second model can understand the text information and obtain the semantic description text and target label corresponding to the text information as text features. In addition, the second model can also perform intent analysis on the text information to obtain the intent type. The second model can also understand the target material and obtain the content description text and target label corresponding to the target material as material features. For example, according to the correspondence between the preset video length and the number of extracted frames, the soundtrack video is frame extracted to obtain a video frame sequence. The video frame sequence is input into the second model, and the content of each video frame is understood by the second model, and the content description text and target label about the video content are output.
[0047] In some embodiments, keywords are extracted from the semantic description text to obtain keywords from the text features. Slot extraction is performed on the semantic description text using a preset template to obtain slot features corresponding to the text features. Keywords are extracted from the content description text to obtain keywords from the material features. Slot extraction is performed on the content description text using a preset template to obtain slot features corresponding to the material features.
[0048] S130: Determine a second target soundtrack corresponding to the target material based on at least one target object in the target object set according to the intention type, and replace the soundtrack of the target material based on the second target soundtrack.
[0049] The target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features. For example, the target object set includes multiple target objects, and the target objects can be keywords in the text features, text vectors corresponding to the text features, or material vectors corresponding to the material features.
[0050] The second target soundtrack can represent a music clip that matches the content of the target material and conforms to the soundtrack intent contained in the text information. Optionally, the second target soundtrack can also be a candidate set of music clips that match the content of the video to be soundtracked and conform to the soundtrack intent. The candidate set of music clips is displayed for the user to select, and the selected soundtrack is then used to replace the soundtrack of the target material.
[0051] Exemplarily, if the intention type represents a single intention, the second target soundtrack corresponding to the target material is determined based on the single intention and at least one target object in the target object set.
[0052] If the intent type represents at least two intents, a candidate soundtrack set corresponding to the target material is determined based on at least one target object in the target object set according to each intent, the candidate soundtracks in the candidate soundtrack set are sorted according to historical interaction operations, and a second target soundtrack corresponding to the target material is determined based on the sorting result.
[0053] Among them, the candidate soundtrack set includes at least two of the soundtrack recall results corresponding to clear intent, the soundtrack recall results corresponding to fuzzy intent, and the soundtrack recall results corresponding to completely fuzzy intent. The soundtrack recall results corresponding to clear intent include candidate soundtracks obtained by searching a preset multimedia content library based on keywords. The soundtrack recall results corresponding to fuzzy intent include candidate soundtracks obtained by searching a preset multimedia content library based on text vectors, and candidate soundtracks determined by a first model based on text vectors and material vectors. The soundtrack recall results corresponding to completely fuzzy intent include candidate soundtracks output by the first model based on material vectors and text vectors. The preset multimedia content library may include a music clip database, etc.
[0054] Furthermore, if the intent type represents a single intent, the second target soundtrack corresponding to the target material is determined based on the single intent and at least one target object in the target object set, including: for the first intent type, searching a preset multimedia content library based on keywords in the text features to obtain the second target soundtrack corresponding to the target material, wherein the first intent type represents that the text information contains a clear description of the soundtrack intention of the target material.
[0055] Furthermore, for the second intent type, a preset multimedia content library is searched based on the text vector corresponding to the text feature to obtain a first candidate soundtrack set, wherein the second intent type indicates that the text information contains a vague description of the soundtrack intent for the target material and / or a description of the attributes of the target material. The material vector and the text vector are input into the first model, and a second candidate soundtrack set is determined based on the material vector and the text vector using the first model. A second target soundtrack corresponding to the target material is determined based on the first and second candidate soundtrack sets.
[0056] Optionally, determining the second target soundtrack corresponding to the target material based on the first candidate soundtrack set and the second candidate soundtrack set includes: sorting the first candidate soundtrack set and the second candidate soundtrack set according to historical interaction operations corresponding to the first candidate soundtrack set and historical interaction operations corresponding to the second candidate soundtrack set, and determining the second target soundtrack corresponding to the target material based on the sorting result.
[0057] Specifically, historical interaction data for each candidate soundtrack in the first candidate soundtrack set is obtained; historical interaction data for each candidate soundtrack in the second candidate soundtrack set is obtained; the candidate soundtracks are sorted according to the historical interaction data, and a second target soundtrack corresponding to the target material is selected from the first candidate soundtrack set and the second candidate soundtrack set based on the sorting results. For example, the top N candidate soundtracks in the sorting can be selected as the second target soundtrack.
[0058] Furthermore, for the third intention type, the material vector is input into the first model, and the first model determines the second target soundtrack corresponding to the target material based on the material vector.
[0059] In the disclosed embodiment, replacing the soundtrack of the target material based on the second target soundtrack includes: if the intent type is clear intent, directly replacing the first soundtrack in the target material with the second target soundtrack as the soundtrack of the target material.
[0060] If the intent type is the second intent type, the music in the second target soundtrack is sorted in descending order according to the historical interaction operations of each music clip in the second target soundtrack. The music ranked TOP1 is selected to replace the first music in the target material. Alternatively, the music ranked TOP M can be returned to the user end for user selection. According to the single music selection operation for the TOP M music input by the user, the music of the target material is replaced by the selected music clip. Alternatively, according to the multiple music selection operation for the TOP M music input by the user, at least two selected music clips are spliced, and the spliced music is used to replace the music of the target material.
[0061] If the intent type is the third intent type, obtain the second target soundtrack output by the first model based on the material vector and the text vector, and arrange the soundtracks in the second target soundtrack in descending order according to the historical interaction operations of each music clip in the second target soundtrack. Determine the soundtrack of the target material based on the sorting result. The specific implementation method is similar to the above embodiment and will not be repeated here.
[0062] The technical solution of the embodiment of the present disclosure is to determine the first target soundtrack corresponding to the target material through a first model, and add soundtrack to the target material based on the first target soundtrack; obtain text information, conduct in-depth understanding and intent judgment of the target material and text information through a second model, and determine the second target soundtrack based on different target objects according to different intents; then, replace the soundtrack of the target material with the second target soundtrack. The technical solution of the embodiment of the present disclosure first uses the first model to determine the first target soundtrack corresponding to the target material. If text information representing the soundtrack intention is obtained, the second target soundtrack corresponding to the target material is determined in combination with the soundtrack intention and the material content, and the second target soundtrack is used to replace the first target soundtrack as the soundtrack of the target material, thereby achieving accurate recommendation of the soundtrack of the target material, improving the matching degree between the soundtrack and the material content, and thus improving the soundtrack effect of the target material.
[0063] Figure 2 is a flow chart of a training method for a first model provided by an embodiment of the present disclosure. Based on the above embodiments, the training method of the first model is additionally limited to enhance the music recommendation capability of the first model through incremental features.
[0064] S210: Acquire target material, determine a first target soundtrack based on the target material using a first model, and add soundtrack to the target material based on the first target soundtrack.
[0065] S220: Acquire text information, and determine material features, text features, and intent type based on the target material and the text information using a second model.
[0066] S230: Determine a second target soundtrack corresponding to the target material based on at least one target object in the target object set according to the intention type, and replace the soundtrack of the target material based on the second target soundtrack.
[0067] S240: Determine incremental features corresponding to the training sample set of the first model according to the text features and the material features, and use the incremental features to update the first model.
[0068] Incremental features represent features that are not present in the material features corresponding to the target material samples in the training sample set of the first model. For example, if text features and material features include: A feature, B feature, C feature, D feature, E feature, F feature, and G feature, and the material features corresponding to the target material samples in the training sample set of the first model include: A feature, B feature, C feature, and D feature, then E feature, F feature, and G feature are incremental features.
[0069] Exemplarily, the text features and material features are compared with the material features corresponding to the target material samples in the training sample set of the first model to obtain incremental features. The incremental features are recorded, and the interactive operations on the incremental features are monitored. The features used to update the first model are determined from the incremental features based on the interactive operations. For example, the candidate material samples containing E features, F features, or G features are monitored, and the F feature is determined to be a feature whose interaction frequency meets the preset conditions. The candidate material sample containing the F feature is used as the target material sample, and the material features and music corresponding to the target material sample are added to the training sample set, and the first model is retrained.
[0070] Furthermore, updating the first model based on the incremental feature includes: S241, obtaining historical interaction operations of the candidate material sample corresponding to the incremental feature.
[0071] Exemplarily, a candidate material sample containing a single incremental feature is determined, and historical interaction operations of the candidate material sample are obtained through logs.
[0072] S242: Determine a target material sample according to historical interaction operations of the candidate material sample, and update the training sample set according to the target material sample and its corresponding background music sample.
[0073] For example, if the historical interaction operations of the candidate material sample meet a preset condition, the candidate material sample is determined to be the target material sample. The preset condition may be determined based on factors such as the frequency, popularity or preference of the interaction operations.
[0074] For example, candidate material samples containing incremental features are sorted in descending order based on the frequency of interactive operations, and the top X candidate material samples in the sorting results are determined as target material samples. The material features corresponding to the target material samples, and the music text features and music video features corresponding to the soundtrack samples are used as new training samples and added to the training sample set.
[0075] S243: Use the updated training sample set to train the first model.
[0076] Exemplarily, each training sample in the updated training sample set is input into the first model so that the first model learns the association between the material features and the music text features and the music video features respectively, thereby enabling the first model to learn the relationship between the incremental features and the soundtrack recommendation, thereby achieving recommendation enhancement.
[0077] For example, the material features corresponding to the target material sample containing incremental features are mapped into material vectors according to the agreed instruction format and input into the first model. The music text features and music video features corresponding to the background music sample are used to supervise the model output results, and the first model is trained through forward propagation and back propagation.
[0078] The technical solution of the embodiment of the present disclosure obtains incremental features by respectively comparing text features and material features with the training sample set of the first model, determines the target material sample according to the historical interaction operations of the candidate material sample corresponding to the incremental features, then updates the training sample set according to the target material sample and its corresponding music sample, and trains the first model with the updated training sample set to automatically perceive the incremental features of the first model, and enables the first model to learn the correlation between the incremental features and the music recommendation, thereby achieving recommendation enhancement.
[0079] FIG3 is a flow chart of another soundtrack recommendation method provided by an embodiment of the present disclosure. Based on the above embodiments, the second target soundtrack corresponding to the target material is determined when the intent type represents at least two intents. As shown in FIG3 , the method includes:
[0080] S310: Acquire target material, determine a first target soundtrack based on the target material using a first model, and add soundtrack to the target material based on the first target soundtrack.
[0081] S320: Acquire text information, and determine material features, text features, and intent types based on the target material and the text information through a second model, wherein the intent type represents at least two intents.
[0082] For example, the intent type may include a combination of at least two of a clear intent, a vague intent, and a completely vague intent.
[0083] In some embodiments, a video uploaded by a user to be accompanied by music is obtained, and a first target music is determined for the video using a first model. Since the first target music is determined based on video features, it may not satisfy the user's music intention. In this case, the user can enter text information through multiple rounds of dialogue with the robot. For example, the text information entered by the user may include "I took a video while traveling and I want to add music to the video. In addition, I heard song C by singer A. I want a travel video with the music using singer A's songs, and preferably a music clip similar to song C." The second model is used to identify the intent of the text information to obtain the intent type. Among them, the intent type includes three types of intent: clear intent, ambiguous intent, and completely ambiguous intent. Among them, "song C" corresponds to a clear intent. And, "travel video with the music using singer A's songs" corresponds to a ambiguous intent. And, "I want to add music to the video" corresponds to a completely ambiguous intent.
[0084] S330: For a clear intention, search a preset multimedia content library according to the keywords in the text feature to obtain at least one candidate soundtrack.
[0085] Exemplarily, in the case of clear intent, keywords in the text features included in the target object set are obtained according to the clear intent, and a music clip database is searched according to the keywords to obtain a clear intent recall result.
[0086] S340 : For the fuzzy intent, search a preset multimedia content library according to the text vector corresponding to the text feature to obtain at least one candidate soundtrack.
[0087] Exemplarily, a music clip database is searched based on the text vector to obtain at least one candidate soundtrack corresponding to the target material. The music clip data includes a music text vector corresponding to the music clip. The music text vector can be determined based on the lyrics or song title of the music clip. The similarity between the text vector and the music text vector in the music clip database is determined. A first music clip is identified whose similarity exceeds a preset similarity threshold. This first music clip is selected as the at least one candidate soundtrack corresponding to the target material, thereby obtaining a first set of candidate soundtracks.
[0088] S350: For fuzzy intent, input the material vector and the text vector into the first model, and determine at least one candidate soundtrack based on the material vector and the text vector through the first model.
[0089] Exemplarily, the material vector and text vector are mapped into input information for the first model according to a predetermined instruction format. The instruction information is used to inform the first model of its specific task. The first model then recalls and rearranges the soundtrack based on the material vector and text vector, obtaining at least one candidate soundtrack, which serves as the second candidate soundtrack set.
[0090] S360: For the fuzzy intent, at least one candidate soundtrack retrieved and at least one candidate soundtrack output by the first model are mixed to obtain a fuzzy intent recall result.
[0091] S370: For a completely ambiguous intent, obtain at least one candidate soundtrack output by the first model based on the material vector and the text vector.
[0092] Exemplarily, for a completely ambiguous intent, at least one candidate soundtrack output by the first model based on the material vector and the text vector is obtained as a completely ambiguous intent recall result.
[0093] S380: Sort the clear intention recall results, the fuzzy intention recall results, and the completely fuzzy intention recall results according to the historical interactive operations, and determine the second target soundtrack corresponding to the target material according to the sorting results.
[0094] Exemplarily, the historical interaction operations corresponding to each soundtrack in the clear intent recall results, fuzzy intent recall results, and completely fuzzy intent recall results are obtained respectively, the soundtracks are sorted in descending order according to the historical interaction operations, and the soundtrack ranked in the TOP X is used as the second target soundtrack corresponding to the target material.
[0095] S390: Replace the soundtrack of the target material based on the second target soundtrack.
[0096] FIG4 is a schematic diagram of a music matching link provided by an embodiment of the present disclosure. As shown in FIG4 , the user terminal sends the target material 401 to the server terminal. After the server terminal obtains the target material 401, it determines the first target music matching 402 based on the target material 401 through the first model 418. The user terminal sends the text information 403 corresponding to the target material 401 to the server terminal. The text information 403 is a natural language that represents the music matching intention of the target material 401. The second model 404 performs intention understanding and reasoning based on the target material 401 and the text information 403 to obtain text features and intention types. The second model 404 performs content understanding and reasoning on the target material 401 to obtain material features. Then, according to the intention type, the corresponding retrieval service 405 is called to perform the following steps: For the case where the intention type includes a clear intention 406, the preset multimedia content library is searched according to the keywords included in the text features to obtain the corresponding search result list 407, and the top 1 music clip is selected from the search result list 407 as the clear intention recall result 408. For the case where the intent type includes fuzzy intent 409, the vector similarity between the text vector corresponding to the text feature and the music text vector of each music segment in the music segment vector database 410 is determined, and the music segments whose vector similarity exceeds the set threshold are used as the first candidate soundtrack set 411. The material vector and text vector corresponding to the material feature are mapped into the instruction format corresponding to the first model 418 and then input into the first model 418. The recommendation result is determined by the first model 418, which is the second candidate soundtrack set 412. The soundtracks in the first candidate soundtrack set 411 and the second candidate soundtrack set 412 are mixed according to the historical interactive operations, and the fuzzy intent recall result 413 is determined based on the mixing result. For the intent type including completely fuzzy intent 414, the second candidate soundtrack set 412 determined by the first model 418 based on the material features and text features is obtained as the completely fuzzy intent recall result 415. The historical interaction operations corresponding to each target music clip in the clear intention recall result 408, the fuzzy intention recall result 413 and the completely fuzzy intention recall result 415 are obtained respectively, the target music clips are sorted in descending order according to the historical interaction operations, and the second target soundtrack 416 is determined according to the sorting results. The result output module 417 is used to implement the encapsulation of the music information related to the second target soundtrack 416 and the copyright verification and other actions. If the second target soundtrack 416 meets the copyright verification conditions, the first target soundtrack 402 in the target material is replaced according to the encapsulated second target soundtrack 416 to obtain a new soundtrack material, and the new soundtrack material is sent to the user end. Among them, the result output module 417 can be the Natural Language Generation (NLG) module in the second model 404, etc. The NLG module is determined based on specific rules and neural network models.
[0097] The technical solution of the disclosed embodiment performs a recall operation based on multiple intentions corresponding to the intent type in combination with keyword, text feature and material feature retrieval, and determines the second target soundtrack based on the historical interactive operations of the recall results. The second target soundtrack is used to replace the first target soundtrack in the target material, thereby optimizing the soundtrack retrieval process, improving the retrieval accuracy, and improving the soundtrack matching degree of the target material.
[0098] FIG5 is a schematic diagram of the structure of a music recommendation device provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC, or a server.
[0099] As shown in FIG. 5 , the apparatus includes: a first music soundtrack determination module 510 , a feature determination module 520 , and a second music soundtrack determination module 530 .
[0100] The first soundtrack determination module 510 is used to obtain target material, determine a first target soundtrack based on the target material through a first model, and add soundtrack to the target material based on the first target soundtrack, wherein the first model is trained based on target material samples determined based on historical interactive operations and soundtrack samples corresponding to the target material samples.
[0101] The feature determination module 520 is used to obtain text information and determine the material features, text features and intention type based on the target material and text information through a second model, wherein the text information is a natural language that represents the music intention of the target material.
[0102] The second soundtrack determination module 530 is used to determine the second target soundtrack corresponding to the target material based on at least one target object in the target object set according to the intention type, and replace the soundtrack of the target material based on the second target soundtrack, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
[0103] Optionally, the device also includes: an incremental feature determination module, which is used to determine the material features, text features and intent type based on the target material and text information through the second model, and then determine the incremental features corresponding to the training sample set of the first model according to the text features and material features, and the incremental features are used to update the first model.
[0104] Furthermore, the device also includes a model training module, which is used to: obtain historical interaction operations of candidate material samples corresponding to the incremental features; determine target material samples based on the historical interaction operations of the candidate material samples, update the training sample set based on the target material samples and their corresponding soundtrack samples; and use the updated training sample set to train the first model.
[0105] Optionally, the second soundtrack determination module 530 includes: a first soundtrack determination unit, which is used to determine the second target soundtrack corresponding to the target material based on at least one target object in the target object set according to the single intention if the intention type represents a single intention; a second soundtrack determination unit, which is used to determine the candidate soundtrack set corresponding to the target material based on at least one target object in the target object set according to each intention if the intention type represents at least two intentions, sort the candidate soundtracks in the candidate soundtrack set according to historical interaction operations, and determine the second target soundtrack corresponding to the target material according to the sorting result.
[0106] Optionally, the first music soundtrack determination unit is specifically used to: for the first intention type, retrieve a preset multimedia content library according to the keywords in the text features to obtain a second target music soundtrack corresponding to the target material, wherein the first intention type represents that the text information contains a clear description of the music soundtrack intention of the target material.
[0107] Optionally, the first soundtrack determination unit is specifically used to: for the second intent type, retrieve a preset multimedia content library according to the text vector corresponding to the text feature to obtain a first candidate soundtrack set, wherein the second intent type represents that the text information contains a vague description of the soundtrack intention of the target material and / or an attribute description of the target material; input the material vector and the text vector into the first model, and determine the second candidate soundtrack set based on the material vector and the text vector through the first model; and determine the second target soundtrack corresponding to the target material according to the first candidate soundtrack set and the second candidate soundtrack set.
[0108] Optionally, determining the second target soundtrack corresponding to the target material based on the first candidate soundtrack set and the second candidate soundtrack set includes: sorting the first candidate soundtrack set and the second candidate soundtrack set according to historical interaction operations corresponding to the first candidate soundtrack set and historical interaction operations corresponding to the second candidate soundtrack set, and determining the second target soundtrack corresponding to the target material based on the sorting result.
[0109] The music recommendation device provided in the embodiments of the present disclosure can execute the music recommendation method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0110] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0111] FIG6 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG6 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG6 ) 600 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG6 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0112] As shown in FIG6 , the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An edit / output (I / O) interface 605 is also connected to the bus 604.
[0113] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0114] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0115] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0116] The electronic device provided by the embodiment of the present disclosure and the music recommendation method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0117] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the music recommendation method provided in the above embodiment is implemented.
[0118] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0119] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0120] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0121] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain target material, determine a first target soundtrack based on the target material through a first model, and add soundtrack to the target material based on the first target soundtrack, wherein the first model is trained based on target material samples determined by historical interactive operations and soundtrack samples corresponding to the target material samples; obtain text information, determine material features, text features and intent types based on the target material and text information through a second model, wherein the text information is a natural language that characterizes the soundtrack intent of the target material; and determine a second target soundtrack corresponding to the target material based on at least one target object in a target object set according to the intent type, and replace the soundtrack of the target material based on the second target soundtrack, wherein the target object set includes keywords in text features, text vectors corresponding to text features and material vectors corresponding to material features.
[0122] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0124] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0125] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0126] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0127] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0128] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0129] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for music score recommendation, comprising: Obtaining target material, determining a first target music score based on the target material through a first model, and adding a music score to the target material based on the first target music score, wherein the first model is trained based on target material samples determined by historical interaction operations and music score samples corresponding to the target material samples; Obtaining text information, and determining material features, text features, and intention types based on the target material and the text information through a second model, wherein the text information is natural language characterizing the music score intention of the target material; Determining a second target music score corresponding to the target material based on at least one target object in the target object set according to the intention type, and replacing the music score of the target material based on the second target music score, wherein the target object set includes keywords in the text features, text vectors corresponding to the text features, and material vectors corresponding to the material features.
2. The method according to claim 1, wherein after determining the material features, text features, and intention types based on the target material and the text information through the second model, it further comprises: Determining incremental features corresponding to the training sample set of the first model according to the text features and material features, and the incremental features are used to update the first model.
3. The method according to claim 2, wherein updating the first model based on the incremental features includes: Obtaining historical interaction operations of candidate material samples corresponding to the incremental features; Determining target material samples according to the historical interaction operations of the candidate material samples, and updating the training sample set according to the target material samples and the music score samples corresponding thereto; Training the first model with the updated training sample set.
4. The method according to claim 1, wherein the determining the second target music score corresponding to the target material based on at least one target object in the target object set according to the intention type includes: If the intention type represents a single intention, determining the second target music score corresponding to the target material based on the single intention and at least one target object in the target object set; If the intention type represents at least two intentions, determining a candidate music score set corresponding to the target material according to each intention and at least one target object in the target object set, sorting the candidate music scores in the candidate music score set according to historical interaction operations, and determining the second target music score corresponding to the target material according to the sorting result.
5. The method according to claim 4, wherein the if the intention type represents a single intention, determining the second target music score corresponding to the target material based on the single intention and at least one target object in the target object set includes: For the first intention type, retrieving a preset multimedia content library according to the keywords in the text features to obtain the second target music score corresponding to the target material, wherein the first intention type represents that the text information contains a clear description of the music score intention of the target material.
6. The method according to claim 4, wherein if the intention type represents a single intention, determining the second target background music corresponding to the target material based on at least one target object in the target object set according to the single intention includes: For the second intention type, retrieving a preset multimedia content library according to the text vector corresponding to the text feature to obtain a first candidate background music set, wherein the second intention type represents that the text information includes a fuzzy description of the background music intention for the target material and / or a description of the attributes of the target material; Inputting the material vector and the text vector into the first model, and determining a second candidate background music set through the first model based on the material vector and the text vector; Determining the second target background music corresponding to the target material according to the first candidate background music set and the second candidate background music set.
7. The method according to claim 6, wherein determining the second target background music corresponding to the target material according to the first candidate background music set and the second candidate background music set includes: Sorting the first candidate background music set and the second candidate background music set according to the historical interaction operations corresponding to the first candidate background music set and the historical interaction operations corresponding to the second candidate background music set, and determining the second target background music corresponding to the target material according to the sorting result.
8. A background music recommendation device, comprising: A first background music determination module, configured to obtain a target material, determine a first target background music based on the target material through a first model, and add background music to the target material based on the first target background music, wherein the first model is trained based on a target material sample determined by a historical interaction operation and a background music sample corresponding to the target material sample; A feature determination module, configured to obtain text information, and determine a material feature, a text feature, and an intention type based on the target material and the text information through a second model, wherein the text information is a natural language representing the background music intention of the target material; A second background music determination module, configured to determine the second target background music corresponding to the target material based on at least one target object in the target object set according to the intention type, and replace the background music of the target material based on the second target background music, wherein the target object set includes keywords in the text feature, the text vector corresponding to the text feature, and the material vector corresponding to the material feature.
9. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the background music recommendation method according to any one of claims 1-7.
10. A storage medium containing computer-executable instructions, wherein the computer-executable instructions are used to execute the background music recommendation method according to any one of claims 1-7 when executed by a computer processor.
Citation Information
Patent Citations
Game acquisition method and device, computer equipment and storage medium
CN111277859A
Method and device for video music matching
CN111753126A
Music recommendation method and device and readable storage medium
CN113569088A
Video polyphonic ringtone score recommendation method, device and equipment and computer storage medium
CN115048546A
Music soundtrack recommendation engine for videos
US8737817B1