Video dubbling method and device and storage medium

By obtaining video and user feature information and using music generation models to generate music audio files, the problems of insufficient scale of music libraries and immature artificial intelligence in the existing technology are solved, and the accuracy and user experience of video music are improved.

CN120302083APending Publication Date: 2025-07-11BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410034272.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the video soundtracking method is unable to accurately combine music and video content characteristics due to insufficient scale of the music library or immature artificial intelligence technology, resulting in poor soundtracking accuracy and poor user experience.

Method used

By obtaining the feature information of the video to be played and the soundtrack feature information input by the user, a pre-trained music generation model is used to generate the soundtrack audio file, and combining the video features and the soundtrack features expected by the user to perform accurate soundtracks.

Benefits of technology

Improve the accuracy and user experience of the video soundtrack, ensuring that the music style matches the video content, and the rhythm trend is consistent with the storyline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302083A_ABST
    Figure CN120302083A_ABST
Patent Text Reader

Abstract

The invention relates to a video music dubbing method and device and a storage medium. The method comprises the following steps: acquiring video feature information of a target video file to be dubbed with music; acquiring dubbed music feature information input by a user; according to the video feature information and the incidental music feature information, acquiring an incidental music audio file through a pre-trained music generation model; and carrying out music dubbing on the target video file according to the music dubbing audio file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, apparatus, and storage medium for video music scoring. Background Art

[0002] With the rise and prosperity of short videos, video intelligent understanding and editing technologies have emerged. For example, to enhance the effects of short videos, music scoring can be included for the produced short videos. In related technologies, adding music scoring to the produced short videos includes at least the following methods: First, video tags can be extracted for the video to be scored, the video to be scored can be classified into a certain video category through the video tags, and then mapped to the corresponding music type through the video category, and a piece of music is randomly selected from the music library matching the music type to score the video, obtaining the final scored video. Second, artificial intelligence technology can also be directly used to generate music scoring, also known as automatic composition, to create a new music scoring.

[0003] However, for the method based on video tags and music scoring tags, the requirements for the music library are relatively high. When the scale or real-time performance of the library is insufficient, accurate music scoring cannot be obtained, and the technology of directly using artificial intelligence technology to generate music scoring is not mature enough to combine music with the content characteristics of the video, and the accuracy of music scoring cannot be guaranteed either. The generated video music scoring effect is average, and the user experience is not good. Summary of the Invention

[0004] To overcome the problems existing in the related technologies, the present disclosure provides a method, apparatus, and storage medium for video music scoring.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for video music scoring is provided. The method includes:

[0006] Obtain video feature information of a target video file to be scored;

[0007] Obtain music scoring feature information input by a user;

[0008] According to the video feature information and the music scoring feature information, obtain a music scoring audio file through a pre-trained music generation model;

[0009] Score the target video file according to the music scoring audio file.

[0010] Optionally, the obtaining a music scoring audio file through a pre-trained music generation model according to the video feature information and the music scoring feature information includes:

[0011] Generate music scoring description data according to the video feature information and the music scoring feature information;

[0012] Input the soundtrack description data into the music generation model to obtain the soundtrack audio file output by the music generation model.

[0013] Optionally, the generating the soundtrack description data according to the video feature information and the soundtrack feature information includes:

[0014] Determine the feature similarity information between the video feature information and the soundtrack feature information;

[0015] According to the feature similarity information, respectively determine the first weight corresponding to the video feature information and the second weight corresponding to the soundtrack feature information;

[0016] Determine the soundtrack description data according to the video feature information, the soundtrack feature information, the first weight, and the second weight.

[0017] Optionally, the soundtrack feature information includes multiple instrument category information; the generating the soundtrack description data according to the video feature information and the soundtrack feature information includes:

[0018] Determine the target instrument category information from the multiple instrument category information;

[0019] According to the target instrument category information, determine the target soundtrack feature information;

[0020] Determine the soundtrack description data according to the video feature information and the target soundtrack feature information.

[0021] Optionally, the soundtrack feature information further includes emotional feature information, and the determining the target instrument category information from the multiple instrument category information includes:

[0022] According to the emotional feature information, determine the target instrument category information from the multiple instrument category information.

[0023] Optionally, the determining the soundtrack description data according to the video feature information and the target soundtrack feature information includes:

[0024] According to the target soundtrack feature information, determine the corresponding multiple instrument performance feature information;

[0025] From the multiple instrument performance feature information, determine the target instrument feature information corresponding to the emotional feature information;

[0026] Determine the soundtrack description data according to the emotional feature information and the target instrument feature information.

[0027] Optionally, the obtaining the soundtrack audio file according to the video feature information and the soundtrack feature information through a pre-trained music generation model includes:

[0028] Send the video feature information and the background music feature information to a server, so that the server generates a background music audio file through the music generation model according to the video feature information and the background music feature information;

[0029] Receive the background music audio file sent by the server.

[0030] Optionally, the obtaining of the video feature information of the target video file to be background-musicated includes:

[0031] Obtain multiple target images from the target video file;

[0032] Input the multiple target images into a pre-trained image recognition model to obtain the video feature information output by the image recognition model.

[0033] According to a second aspect of the embodiments of the present disclosure, there is provided a video background music method, which is applied to a server, and the method includes:

[0034] Receive the video feature information and the background music feature information sent by a terminal; the video feature information is the video feature information of a target video file to be background-musicated, and the background music feature information is the background music feature information input by a user to the terminal;

[0035] Generate a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information;

[0036] Send the background music audio file to the terminal, so that the terminal background-musicates the target video file according to the background music audio file.

[0037] Optionally, the generating of the background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information includes:

[0038] Generate background music description data according to the video feature information and the background music feature information;

[0039] Input the background music description data into the music generation model to obtain the background music audio file output by the music generation model.

[0040] According to a third aspect of the embodiments of the present disclosure, there is provided a video background music device, and the device includes:

[0041] A first obtaining module, configured to obtain the video feature information of a target video file to be background-musicated;

[0042] A second obtaining module, configured to obtain the background music feature information input by a user;

[0043] A first generation module, configured to obtain a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information;

[0044] A first background music module, configured to add background music to the target video file according to the background music audio file.

[0045] Optionally, the first generation module includes:

[0046] A first generation sub-module, configured to generate background music description data according to the video feature information and the background music feature information;

[0047] A second generation sub-module, configured to input the background music description data into the music generation model to obtain the background music audio file output by the music generation model.

[0048] Optionally, the first generation sub-module is configured to determine feature similarity information between the video feature information and the background music feature information; according to the feature similarity information, respectively determine a first weight corresponding to the video feature information and a second weight corresponding to the background music feature information; according to the video feature information, the background music feature information, the first weight, and the second weight, determine the background music description data.

[0049] Optionally, the background music feature information includes multiple instrument category information; the first generation sub-module is configured to determine target instrument category information from the multiple instrument category information; according to the target instrument category information, determine target background music feature information; according to the video feature information and the target background music feature information, determine the background music description data.

[0050] Optionally, the background music feature information further includes emotional feature information, and the first generation sub-module is configured to determine target instrument category information from the multiple instrument category information according to the emotional feature information.

[0051] Optionally, the first generation sub-module is configured to determine corresponding multiple instrument performance feature information according to the target background music feature information; from the multiple instrument performance feature information, determine target instrument feature information corresponding to the emotional feature information; according to the emotional feature information and the target instrument feature information, determine the background music description data.

[0052] Optionally, the first generation module includes:

[0053] A third generation sub-module, configured to send the video feature information and the background music feature information to a server, so that the server generates a background music audio file through the music generation model according to the video feature information and the background music feature information;

[0054] A receiving sub-module, configured to receive the background music audio file sent by the server.

[0055] Optionally, the first obtaining module includes:

[0056] An obtaining sub-module, configured to obtain multiple frames of target images from the target video file;

[0057] A fourth generating sub-module, configured to input the multiple frames of target images into a pre-trained image recognition model, and obtain video feature information output by the image recognition model.

[0058] According to a fourth aspect of the embodiments of the present disclosure, a video background music device is provided, which is applied to a server. The device includes:

[0059] A receiving module, configured to receive video feature information and background music feature information sent by a terminal; the video feature information is the video feature information of a target video file to be background music, and the background music feature information is the background music feature information input by a user to the terminal;

[0060] A second generating module, configured to generate a background music audio file according to the video feature information and the background music feature information through a pre-trained music generating model;

[0061] A second background music module, configured to send the background music audio file to the terminal, so that the terminal performs background music on the target video file according to the background music audio file.

[0062] Optionally, the second generating module includes:

[0063] A fifth generating sub-module, configured to generate background music description data according to the video feature information and the background music feature information;

[0064] A sixth generating sub-module, configured to input the background music description data into the music generating model, and obtain the background music audio file output by the music generating model.

[0065] According to a fifth aspect of the embodiments of the present disclosure, a video background music device is provided. The device includes:

[0066] A processor;

[0067] A memory for storing instructions executable by the processor;

[0068] Wherein, the processor is configured to implement the steps of the video background music method provided in the first aspect of the present disclosure when executed.

[0069] According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the video scoring method provided in the second aspect of the present disclosure are implemented.

[0070] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0071] By obtaining video feature information of a target video file to be scored; obtaining scoring feature information input by a user; according to the video feature information and the scoring feature information, obtaining a scoring audio file through a pre-trained music generation model; and scoring the target video file according to the scoring audio file. In this way, the video feature information of the target video file to be scored can be combined with the scoring feature information expected by the user, improving the accuracy of scoring the target video file and enhancing the user experience.

[0072] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0073] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0074] Figure 1 is a flowchart of a video scoring method shown according to an exemplary embodiment.

[0075] Figure 2 is a schematic diagram of a terminal display interface shown according to an exemplary embodiment.

[0076] Figure 3 is a flowchart of another video scoring method shown according to an exemplary embodiment.

[0077] Figure 4 is a flowchart of another video scoring method shown according to an exemplary embodiment.

[0078] Figure 5 is a schematic diagram of an implementation environment shown according to an exemplary embodiment.

[0079] Figure 6 is a block diagram of a video scoring device shown according to an exemplary embodiment.

[0080] Figure 7 is according to Figure 6 shown in the embodiment, a block diagram of a first generation module.

[0081] Figure 8 is according to Figure 6Another block diagram of the first generation module shown in the illustrated embodiment.

[0082] Figure 9 is according to Figure 6 A block diagram of a first acquisition module shown in the illustrated embodiment.

[0083] Figure 10 A block diagram of another video music scoring device shown according to an exemplary embodiment.

[0084] Figure 11 is according to Figure 10 A block diagram of a second generation module shown in the illustrated embodiment.

[0085] Figure 12 A block diagram of a video music scoring device shown according to an exemplary embodiment.

[0086] Figure 13 A block diagram of a video music scoring device shown according to an exemplary embodiment. Detailed implementation manners

[0087] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0088] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the corresponding device owner.

[0089] Before introducing the detailed implementation manners of the present disclosure in detail, the application scenarios of the present disclosure will be described first. The present disclosure can be applied to the intelligent music scoring scenario of video files. Currently, with the rise and prosperity of short videos, video intelligent understanding and editing technologies have emerged. For example, to enhance the effect of short videos, music scoring can be added to the produced short videos.

[0090] In the related art, adding music scoring to the produced short videos includes at least the following methods: First, video tags can be extracted for the video to be scored, the video to be scored can be classified into a certain video category through the video tags, and mapped to the corresponding music type through the video category, and a piece of music is randomly selected from the music library matching the music type to score the video, obtaining the final scored video. Second, artificial intelligence technology can also be directly used to generate music scoring, also known as automatic composition, to create a new music scoring.

[0091] However, the method based on video tags and music score tags has relatively high requirements for the music score library. Due to reasons such as platform copyright restrictions and limited capacity of the music library, the scale of the music library may be insufficient, or when the timeliness of the music library is insufficient, accurate music scores cannot be obtained; and the method of generating music scores based on artificial intelligence technology will, due to the immaturity of the technology for generating music scores by artificial intelligence, result in the inability to simulate the creative ability of human musicians. The generated music is often relatively single, lacking variation and innovation, and unable to combine the music with the video content features. The accuracy of the music score cannot be guaranteed either. Moreover, it requires a large amount of computing resources and time, and the time cost of generating a piece of music is relatively high, which is not suitable for real-time applications. As a result, the generated video music score effect is average, leading to a poor user experience.

[0092] In order to overcome the technical problems existing in the above related technologies, the present disclosure provides a video music score method, device, and storage medium. By obtaining the video feature information of the target video file to be scored; obtaining the music score feature information input by the user; according to the video feature information and the music score feature information, obtaining a music score audio file through a pre-trained music generation model; and scoring the target video file according to the music score audio file. In this way, the video feature information of the target video file to be scored can be combined with the music score feature information expected by the user, improving the accuracy of scoring the target video file and enhancing the user experience.

[0093] The following will illustrate the present disclosure in conjunction with specific embodiments.

[0094] Figure 1 is a flowchart of a video music score method shown according to an exemplary embodiment. As Figure 1 shown, this method is used in a terminal, and the method includes the following steps.

[0095] In step S11, obtain the video feature information of the target video file to be scored.

[0096] First of all, it should be noted that considering that the existing video music score scheme performs automatic video music scoring based on video category tags and music category tags, it cannot ensure the matching degree between the video to be scored and the selected music at a finer granularity, resulting in the inability of the existing video music score scheme to obtain a video with a relatively high accuracy of music score. Among them, the accuracy of the video with music score can be reflected in aspects such as whether the music style conforms to the plot shown in the video picture, whether the rhythm trend of the music conforms to the development of the plot in the video, etc. This is not restricted here.

[0097] To solve this technical problem, this embodiment proposes a new video background music scheme. This video background music scheme can select target music that matches the video to be background-music based on a finer granularity from a music library, and use the target music to perform background music for the video to be background-music, so that the finally obtained background-music video has higher accuracy.

[0098] Among them, the video feature information may include information such as the objective content, style, brightness, and emotion of the video.

[0099] In this embodiment, first, multiple target images can be obtained from the target video file; then, the multiple target images can be input into a pre-trained image recognition model to obtain video feature information output by the image recognition model.

[0100] Optionally, first, the target video file can be opened locally on the terminal using a preset library function, and basic information of the target video file, such as frame rate, width, and height, can be obtained. Then, the frame extraction frequency can be determined according to the basic information of the target video file. Then, according to the set frame extraction frequency, in the target video file, frame extraction is performed every predetermined frame according to the frame extraction frequency, and the frame data of the extracted frames is converted into image data and saved as a picture file.

[0101] Among them, the picture format can be selected according to actual needs. When decoding the current video, the compressed video data stream will be restored frame by frame to the original YUV format data. When saving the extracted frames as picture files, the YUV format data needs to be converted into picture format data.

[0102] Then, multiple target images can be input into a pre-trained image recognition model to obtain video feature information output by the image recognition model.

[0103] Exemplarily, the image recognition model may include multiple image recognition sub-models, and each image recognition sub-model is used to recognize different image targets. For example, the first image recognition sub-model is used to recognize the objective content of the image, and the first image recognition sub-model is used to recognize the brightness of the image. After inputting the target image into the first image recognition sub-model, the objective content of the image input by the first image recognition sub-model can be obtained. After inputting the target image into the second image recognition sub-model, the brightness of the image input by the second image recognition sub-model can be obtained, and so on, until all the image recognition sub-models have completed recognition. In another possible implementation manner, the image recognition model may include multiple recognition targets and can recognize different image targets in one recognition process.

[0104] In this way, the time-related video features of the target video file can be obtained by extracting frame data in chronological order, so as to determine whether the music style of the video soundtrack is consistent with the story presented in the video and whether the rhythm of the music is consistent with the development of the story in the video.

[0105] In some embodiments, video feature information of the target video file to be matched with music may be obtained on the terminal.

[0106] For example, a user can use the video editing function in the terminal's photo album to extract five pictures from the target video file stored in the terminal according to the video length as the obtained multi-frame target images, and then the multi-frame target images can be input into a pre-trained image recognition model to obtain video feature information output by the image recognition model, wherein the image recognition model can be set on the terminal or on the server. When the image recognition model is set on the terminal, the video feature information of the target video file can be directly obtained through the image recognition model on the terminal. When the image recognition model is set on the server, the multi-frame target images can be sent to the server through the terminal, and the video feature information of the target video file can be obtained through the image recognition model on the server.

[0107] In step S12, the music feature information input by the user is obtained.

[0108] The music feature information may include the music style expected by the user and the music instrument expected by the user.

[0109] In this step, the music feature information input by the user may be obtained on the terminal.

[0110] For example, after obtaining the video feature information of the target video file to be matched with music, Figure 2 As shown, multiple text boxes can be displayed on the display interface of the terminal, wherein each text box can correspond to a different music style and a different music instrument. In this way, when the user clicks on different text boxes, the music style and music instrument input by the user can be determined to obtain the music feature information input by the user.

[0111] In step S13, the soundtrack audio file is obtained through a pre-trained music generation model according to the video feature information and the soundtrack feature information.

[0112] In some embodiments, the music description data may be first generated according to the video feature information and the music feature information, and then the music description data may be input into the music generation model to obtain the music audio file output by the music generation model.

[0113] Among them, the music score description data may include a set of corresponding relationships, for example, [style: lively; beat: half-time; content: rain; instrument: guitar].

[0114] Optionally, considering that there may be inconsistencies between the video feature information and the music score feature information. For example, the style is determined to be lively from the video feature information of the target video file to be scored, but in the music score feature information input by the user, the style is soothing. In this case, since it is for scoring the video, the elements of the video need to be the main consideration. Therefore, the feature similarity information between the video feature information and the music score feature information can be determined first; then, according to the feature similarity information, the first weight corresponding to the video feature information and the second weight corresponding to the music score feature information can be determined respectively; and then, according to the video feature information, the music score feature information, the first weight, and the second weight, the music score description data can be determined.

[0115] Exemplarily, the video feature information and the music score feature information can be presented in the form of text. Therefore, the text similarity degree between the video feature text and the music score feature text can be determined by a trained text recognition model. Specifically, the text similarity degree can be determined by determining the lexical similarity, semantic similarity, and structural similarity, etc. For example, the lexical similarity is a method of measuring the text similarity based on the similarity between words; for example, the co-occurrence frequency of words, the similarity of words, etc. can be used to calculate the lexical similarity. Semantic similarity: Semantic similarity is a method of measuring the text similarity based on the meaning of the text; for example, semantic vectors, semantic models, etc. can be used to calculate the semantic similarity. Structural similarity: Structural similarity is a method of measuring the text similarity based on the structure of the text; for example, dependency trees, syntactic structures, etc. can be used to calculate the structural similarity. Among them, the text recognition model may include a natural language processing (NLP) model.

[0116] For example, when the feature similarity information between the video feature information and the music score feature information is determined to be 1:1, the first weight and the second weight corresponding to the video feature information and the music score feature information are 0.5 and 0.5 respectively; when the feature similarity information between the video feature information and the music score feature information is determined to be (0.4:0.6) to (0.2:0.8), the first weight and the second weight corresponding to the video feature information and the music score feature information are 0.85 and 0.15 respectively; when the feature similarity information between the video feature information and the music score feature information is determined to be (0.6:0.4) to (0.9:0.1), the first weight and the second weight corresponding to the video feature information and the music score feature information are 0.95 and 0.05 respectively.

[0117] In a possible implementation, the background music feature information includes multiple instrument category information; the target instrument category information can be determined from the multiple instrument category information first; then, based on the target instrument category information, the target background music feature information can be determined; and then, based on the video feature information and the target background music feature information, the background music description data can be determined.

[0118] In this way, when the user inputs multiple instruments, the target instrument category information can be determined from the multiple instrument category information, preventing the background music audio from being too chaotic due to excessive instruments in the background music audio, which may cause the process of generating the background music audio to be too long and inefficient.

[0119] In another possible implementation, the background music feature information further includes emotional feature information; first, based on the target background music feature information, the corresponding multiple instrument performance feature information can be determined; second, from the multiple instrument performance feature information, the target instrument feature information corresponding to the emotional feature information can be determined; finally, based on the emotional feature information and the target instrument feature information, the background music description data can be determined.

[0120] Exemplarily, the emotional feature information is the expected background music emotion input by the user when inputting the background music feature information. Since each instrument can correspond to multiple performance feature databases, each database is used to represent the performance features corresponding to the instrument. For example, it can include rhythm, melody, beat, and style. The rhythm database can include multiple rhythms, and each rhythm corresponds to an emotion. Similarly, melody, beat, and style also correspond to different emotions. Then, based on the emotional feature information, the target instrument feature information corresponding to the emotional feature information can be determined. Finally, based on the emotional feature information and the target instrument feature information, the background music description data can be determined.

[0121] In some embodiments, the obtained video feature information and background music feature information can be sent to the server so that the server can generate a background music audio file through the music generation model based on the video feature information and the background music feature information; then, the background music audio file sent by the server can be received.

[0122] Alternatively, the obtained video feature information and background music feature information can be sent to the server, and the server can perform the above steps for generating the background music audio file and then send the background music audio file to the terminal so that the terminal can receive the background music audio file sent by the server.

[0123] In step S14, the target video file is scored with the background music audio file.

[0124] In this step, it can be performed on the video track corresponding to the target video file and the audio track corresponding to the background music audio file, and then the final background music video can be generated.

[0125] Adopting the above technical solution, by obtaining the video feature information of the target video file to be background-music added; obtaining the background-music feature information input by the user; according to the video feature information and the background-music feature information, obtaining the background music audio file through a pre-trained music generation model; and adding background music to the target video file according to the background music audio file. In this way, the video feature information of the target video file to be background-music added can be combined with the background-music feature information expected by the user, improving the accuracy of adding background music to the target video file and enhancing the user experience.

[0126] Figure 3 is a flowchart of another video background-music adding method shown according to an exemplary embodiment, as Figure 3 shown, this method is used in a server, and this method includes the following steps.

[0127] In step S21, receive the video feature information and the background-music feature information sent by the terminal.

[0128] Wherein, the video feature information is the video feature information of the target video file to be background-music added, and the background-music feature information is the background-music feature information input by the user to the terminal.

[0129] In this step, the terminal can obtain multiple target images from the target video file; then can input the multiple target images into a pre-trained image recognition model to obtain the video feature information output by the image recognition model; and the terminal can obtain the background-music feature information input by the user; then can send the obtained video feature information to the server, and the server receives the video feature information and the background-music feature information sent by the terminal.

[0130] In step S22, according to the video feature information and the background-music feature information, generate a background music audio file through a pre-trained music generation model.

[0131] In some embodiments, background-music description data can be generated according to the video feature information and the background-music feature information; then input the background-music description data into the music generation model to obtain the background music audio file output by the music generation model.

[0132] Optionally, the obtained video feature information and the background-music feature information can be sent to the server, so that the server can generate a background music audio file according to the video feature information and the background-music feature information through the music generation model; then can receive the background music audio file sent by the server.

[0133] Alternatively, the obtained video feature information and the background music feature information may be sent to a server. The server may perform the above steps for generating a background music audio file and then send the background music audio file to a terminal, so that the terminal can receive the background music audio file sent by the server.

[0134] In step S23, the background music audio file is sent to the terminal so that the terminal can add background music to the target video file according to the background music audio file.

[0135] By adopting the above technical solution, the video feature information of the target video file to be added with background music and the background music feature information expected by the user can be combined, improving the accuracy of adding background music to the target video file and enhancing the user experience.

[0136] Figure 4 is a flowchart of another video background music adding method shown according to an exemplary embodiment. As Figure 4 shown, this method is used in a server and a terminal, and the method includes the following steps.

[0137] First, the implementation environment of the video background music adding method provided in the embodiments of the present disclosure may be described. Figure 5 is a schematic diagram of an implementation environment shown according to an exemplary embodiment. As Figure 5 shown. This implementation environment includes a server 20 and at least one terminal 10. The terminal 10 and the server 20 communicate with each other through a wired or wireless network. Among them, the terminal 10 is used to upload a video to be added with background music to the server 20. For example, the terminal 10 may call the Web (i.e., network) interface of the server 20 to upload the video feature information of the video to be added with background music and the background music feature information input by the user to the server 20.

[0138] The server 20 also returns the generated background music audio to the user terminal 10. For example, the server 20 returns the generated background music audio to the terminal 10 in the form of a URL (Uniform Resource Locator).

[0139] It should be noted that in Figure 5In the described implementation environment, the terminal 10 may be an electronic device such as a smart phone, a tablet, a laptop, a computer, etc.; the server 20 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, where multiple servers may form a blockchain, and the server is a node on the blockchain; the server 20 may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and no limitation is made here.

[0140] In step S31, the terminal obtains the video feature information of the target video file to be accompanied by music.

[0141] In step S32, the terminal obtains the music accompaniment feature information input by the user.

[0142] In step S33, the terminal sends the video feature information and the music accompaniment feature information to the server.

[0143] In step S34, the server determines the target instrument category information from multiple instrument category information according to the emotional feature information in the received music accompaniment feature information.

[0144] In step S35, the server determines the target music accompaniment feature information according to the target instrument category information.

[0145] In step S36, the server determines the corresponding multiple instrument performance feature information according to the target music accompaniment feature information.

[0146] In step S37, the server determines the target instrument feature information corresponding to the emotional feature information from the multiple instrument performance feature information.

[0147] In step S38, the server determines the music accompaniment description data according to the emotional feature information and the target instrument feature information.

[0148] In step S39, the server determines the feature similarity information between the received video feature information and the music accompaniment feature information.

[0149] In step S310, the server respectively determines the first weight corresponding to the video feature information and the second weight corresponding to the music accompaniment feature information according to the feature similarity information.

[0150] In step S311, the server determines the music accompaniment description data according to the video feature information, the music accompaniment feature information, the first weight, and the second weight.

[0151] In step S312, the server inputs the background music description data into the music generation model to obtain the background music audio file output by the music generation model.

[0152] In step S313, the server sends the background music audio file to the terminal.

[0153] In step S314, the terminal receives the background music audio file and scores the target video file according to the background music audio file.

[0154] By adopting the above technical solution, the video feature information of the target video file to be scored and the background music feature information expected by the user can be combined, improving the accuracy of scoring the target video file and enhancing the user experience.

[0155] Figure 6 It is a block diagram of a video scoring device shown according to an exemplary embodiment. Refer to Figure 6 , the video scoring device 400 includes a first acquisition module 401, a second acquisition module 402, a first generation module 403, and a first scoring module 404.

[0156] The first acquisition module 401 is configured to acquire the video feature information of the target video file to be scored;

[0157] The second acquisition module 402 is configured to acquire the background music feature information input by the user;

[0158] The first generation module 403 is configured to obtain a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information;

[0159] The first scoring module 404 is configured to score the target video file according to the background music audio file.

[0160] Figure 7 It is according to Figure 6 shown in the embodiment, a block diagram of a first generation module is shown. Refer to Figure 7 , the first generation module 403 includes:

[0161] The first generation sub-module 4031 is configured to generate background music description data according to the video feature information and the background music feature information;

[0162] The second generation sub-module 4032 is configured to input the background music description data into the music generation model to obtain the background music audio file output by the music generation model.

[0163] Optionally, the first generation sub-module 4031 is configured to determine the feature similarity information between the video feature information and the background music feature information; respectively determine the first weight corresponding to the video feature information and the second weight corresponding to the background music feature information according to the feature similarity information; and determine the background music description data according to the video feature information, the background music feature information, the first weight, and the second weight.

[0164] Optionally, the background music feature information includes multiple instrument category information; the first generation sub-module 4031 is configured to determine the target instrument category information from the multiple instrument category information; determine the target background music feature information according to the target instrument category information; and determine the background music description data according to the video feature information and the target background music feature information.

[0165] Optionally, the background music feature information further includes emotional feature information, and the first generation sub-module 4031 is configured to determine the target instrument category information from the multiple instrument category information according to the emotional feature information.

[0166] Optionally, the first generation sub-module 4031 is configured to determine the corresponding multiple instrument performance feature information according to the target background music feature information; determine the target instrument feature information corresponding to the emotional feature information from the multiple instrument performance feature information; and determine the background music description data according to the emotional feature information and the target instrument feature information.

[0167] Figure 8 is based on Figure 6 Another block diagram of the first generation module shown in the embodiment. Refer to Figure 8 , the first generation module 403 includes:

[0168] The third generation sub-module 4033 is configured to send the video feature information and the background music feature information to the server, so that the server generates a background music audio file according to the video feature information and the background music feature information through the music generation model;

[0169] The receiving sub-module 4034 is configured to receive the background music audio file sent by the server.

[0170] Figure 9 is based on Figure 6 A block diagram of a first acquisition module shown in the embodiment. Refer to Figure 9 , the first acquisition module 401 includes:

[0171] The acquisition sub-module 4011 is configured to acquire multiple frames of target images from the target video file;

[0172] The fourth generation sub-module 4012 is configured to input the multi-frame target image into a pre-trained image recognition model to obtain video feature information output by the image recognition model.

[0173] Figure 10 is a block diagram of another video music scoring device shown according to an exemplary embodiment. Refer to Figure 10 The video music scoring device 500 includes a receiving module 501, a second generation module 502, and a second music scoring module 503.

[0174] The receiving module 501 is configured to receive video feature information and music scoring feature information sent by a terminal; the video feature information is the video feature information of a target video file to be scored, and the music scoring feature information is the music scoring feature information input by a user to the terminal;

[0175] The second generation module 502 is configured to generate a scored audio file through a pre-trained music generation model according to the video feature information and the music scoring feature information;

[0176] The second music scoring module 503 is configured to send the scored audio file to the terminal so that the terminal scores the target video file according to the scored audio file.

[0177] Figure 11 is according to Figure 10 shown in the embodiment is a block diagram of a second generation module. Refer to Figure 11 The second generation module 502 includes:

[0178] The fifth generation sub-module 5021 is configured to generate scored description data according to the video feature information and the music scoring feature information;

[0179] The sixth generation sub-module 5022 is configured to input the scored description data into the music generation model to obtain the scored audio file output by the music generation model.

[0180] Regarding the devices in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0181] Using the above device, by obtaining the video feature information of a target video file to be scored; obtaining the music scoring feature information input by a user; obtaining a scored audio file through a pre-trained music generation model according to the video feature information and the music scoring feature information; scoring the target video file according to the scored audio file. In this way, the video feature information of the target video file to be scored can be combined with the music scoring feature information expected by the user, improving the accuracy of scoring the target video file and enhancing the user experience.

[0182] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the video scoring method provided by the present disclosure are implemented.

[0183] Figure 12 FIG. 4 is a block diagram of a video scoring apparatus 1200 shown according to an exemplary embodiment. For example, the apparatus 1200 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0184] Referring to Figure 12 , the apparatus 1200 may include one or more of the following components: a processing component 1202, a memory 1204, a power supply component 1206, a multimedia component 1208, an audio component 1210, an input / output interface 1212, a sensor component 1214, and a communication component 1216.

[0185] The processing component 1202 generally controls the overall operation of the apparatus 1200, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the video scoring method. In addition, the processing component 1202 may include one or more modules to facilitate the interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate the interaction between the multimedia component 1208 and the processing component 1202.

[0186] The memory 1204 is configured to store various types of data to support the operation of the apparatus 1200. Examples of these data include instructions for any application or method operating on the apparatus 1200, contact data, phone book data, messages, pictures, videos, etc. The memory 1204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0187] The power supply component 1206 provides power to various components of the apparatus 1200. The power supply component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the apparatus 1200.

[0188] The multimedia component 1208 includes a screen that provides an output interface between the device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0189] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC) that is configured to receive external audio signals when the device 1200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1204 or transmitted via the communication component 1216. In some embodiments, the audio component 1210 further includes a speaker for outputting audio signals.

[0190] The input / output interface 1212 provides an interface between the processing component 1202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0191] The sensor component 1214 includes one or more sensors for providing an assessment of various aspects of the state of the device 1200. For example, the sensor component 1214 can detect the on / off state of the device 1200, the relative positioning of components, such as the display and keypad of the device 1200. The sensor component 1214 can also detect a change in the position of the device 1200 or a component of the device 1200, the presence or absence of user contact with the device 1200, the orientation or acceleration / deceleration of the device 1200, and the temperature change of the device 1200. The sensor component 1214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1214 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0192] The communication component 1216 is configured to facilitate communication between the device 1200 and other devices in a wired or wireless manner. The device 1200 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0193] In an exemplary embodiment, the device 1200 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the video scoring method.

[0194] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, and the above instructions can be executed by a processor 1220 of the device 1200 to complete the video scoring method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0195] Figure 13 is a block diagram of a video scoring device 1300 shown according to an exemplary embodiment. For example, the device 1300 can be provided as a server. Referring to Figure 13 , the device 1300 includes a processing component 1322, which further includes one or more processors, and memory resources represented by a memory 1332 for storing instructions executable by the processing component 1322, such as application programs. The application programs stored in the memory 1332 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1322 is configured to execute instructions to perform the video scoring method.

[0196] The device 1300 can also include a power component 1326 configured to perform power management of the device 1300, a wired or wireless network interface 1350 configured to connect the device 1300 to a network, and an input / output interface 1358. The device 1300 can operate based on an operating system stored in the memory 1332, such as Windows Server TM, Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0197] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0198] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video music scoring method, characterized in that, The method includes: Obtaining video feature information of a target video file to be scored; Obtaining scored music feature information input by a user; Obtaining a scored music audio file through a pre-trained music generation model according to the video feature information and the scored music feature information; Scoring the target video file according to the scored music audio file.

2. The video music scoring method according to claim 1, wherein The obtaining a scored music audio file through a pre-trained music generation model according to the video feature information and the scored music feature information includes: Generating scored music description data according to the video feature information and the scored music feature information; Inputting the scored music description data into the music generation model to obtain the scored music audio file output by the music generation model.

3. The video background music method according to claim 2, wherein The generating scored music description data according to the video feature information and the scored music feature information includes: Determining feature similarity information of the video feature information and the scored music feature information; Respectively determining a first weight corresponding to the video feature information and a second weight corresponding to the scored music feature information according to the feature similarity information; Determining scored music description data according to the video feature information, the scored music feature information, the first weight, and the second weight.

4. The video background music method according to claim 2, characterized in that, The scored music feature information includes multiple musical instrument category information; the generating scored music description data according to the video feature information and the scored music feature information includes: Determining target musical instrument category information from the multiple musical instrument category information; Determining target scored music feature information according to the target musical instrument category information; Determining scored music description data according to the video feature information and the target scored music feature information.

5. The video music scoring method according to claim 4, wherein, The scored music feature information further includes emotional feature information, and the determining target musical instrument category information from the multiple musical instrument category information includes: Determining target musical instrument category information from the multiple musical instrument category information according to the emotional feature information.

6. The video music scoring method according to claim 5, wherein The determining scored music description data according to the video feature information and the target scored music feature information includes: Determining corresponding multiple musical instrument performance feature information according to the target scored music feature information; Determining target musical instrument feature information corresponding to the emotional feature information from the multiple musical instrument performance feature information; Determining scored music description data according to the emotional feature information and the target musical instrument feature information.

7. The video background music method according to claim 1, wherein The obtaining a scored music audio file through a pre-trained music generation model according to the video feature information and the scored music feature information includes: Sending the video feature information and the scored music feature information to a server so that the server generates a scored music audio file through the music generation model according to the video feature information and the scored music feature information; Receiving the scored music audio file sent by the server.

8. The video music scoring method according to any one of claims 1 to 7, characterized in that The obtaining video feature information of a target video file to be scored includes: Obtaining multiple frames of target images from the target video file; Inputting the multiple frames of target images into a pre-trained image recognition model to obtain video feature information output by the image recognition model.

9. A video background music method, characterized in that, When applied to a server, the method includes: Receive the video feature information and the background music feature information sent by the receiving terminal; the video feature information is the video feature information of the target video file to be background-musicked, and the background music feature information is the background music feature information input by the user to the terminal; Generate a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information; Send the background music audio file to the terminal so that the terminal can perform background music on the target video file according to the background music audio file.

10. The video background music method according to claim 9, wherein The generating a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information includes: Generate background music description data according to the video feature information and the background music feature information; Input the background music description data into the music generation model to obtain the background music audio file output by the music generation model.

11. A video background music device, characterized in that, The device includes: A first acquisition module configured to acquire the video feature information of the target video file to be background-musicked; A second acquisition module configured to acquire the background music feature information input by the user; A first generation module configured to obtain a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information; A first background music module configured to perform background music on the target video file according to the background music audio file.

12. A video music accompaniment device, characterized in that, Applied to a server, the device includes: A receiving module configured to receive the video feature information and the background music feature information sent by the terminal; the video feature information is the video feature information of the target video file to be background-musicked, and the background music feature information is the background music feature information input by the user to the terminal; A second generation module configured to generate a background music audio file through a pre-trained music generation model according to the video feature information and the background music feature information; A second background music module configured to send the background music audio file to the terminal so that the terminal can perform background music on the target video file according to the background music audio file.

13. A video music accompaniment device, characterized in that, The device includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the method according to any one of claims 1-8 when executed.

14. A video background music device, characterized in that, The device includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the method according to claim 9 or 10 when executed.

15. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instruction is executed by the processor, it implements the steps of the method according to any one of claims 1-8.

16. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instruction is executed by the processor, it implements the steps of the method according to claim 9 or 10.