Background audio determination method and device, electronic equipment and storage medium

Through a large language model, videos are processed multi-modal, multiple information of the video are obtained, and audio libraries are screened based on this information, solving the problem of low matching between background audio and video, achieving higher matching accuracy and user needs satisfaction.

CN120030187APending Publication Date: 2025-05-23BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510088484.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the matching degree between background audio and video is low, and the specific content of the audio cannot be accurately reflected, resulting in the selected background audio that does not match the video.

Method used

By calling the large language model to perform multimodal processing on the video, obtain the search terms of the video, including video content information, emotional information and scene information, determine the query order, and filter the audio in the audio library based on this information to obtain the matching background audio.

Benefits of technology

Improve the matching degree of background audio and video, making it more accurate and meet user needs, ensuring that the audio and video match in terms of content, emotion and scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030187A_ABST
    Figure CN120030187A_ABST
Patent Text Reader

Abstract

The invention provides a background audio determination method and device, electronic equipment and a storage medium, and belongs to the technical field of multimedia. The method comprises the steps that a large language model is called to carry out multi-mode processing on a video, search terms of the video are obtained, the search terms comprise various kinds of video information, and the various kinds of video information comprise video content information, video emotion information and video scene information of the video; determining a query sequence based on video content information, video emotion information and video scene information in the various video information; according to the query sequence, the audios in the audio library are screened based on the search terms and the tags of the audios in the audio library, at least one background audio is obtained, and the tag of each audio in the audio library comprises audio content information, audio emotion information and audio scene information of the audios. The method ensures that the background audio and the video are matched in multiple dimensions such as content, emotion and scene, the matching degree is higher, the method is more accurate, and user requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technologies, and particularly to a method, apparatus, electronic device, and storage medium for determining background audio. Background Art

[0002] With the development of multimedia technologies, more and more users are keen on making and publishing videos. When making a video, users usually add a background audio to make the video more attractive.

[0003] In related technologies, usually the content of the video to which the background audio needs to be added is matched with the names of each audio in the audio library, and then the audio with the highest matching degree is selected as the background audio of the video.

[0004] Since the name of the audio cannot accurately reflect the specific content of the audio, there are often situations where the selected background audio does not match the video in the above technical solution, that is, the matching degree between the background audio and the video is low. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, electronic device, and storage medium for determining background audio, which ensures a higher matching degree between the background audio and the video, is more accurate, and meets the user's needs. The technical solution of the present disclosure is as follows:

[0006] According to one aspect of the embodiments of the present disclosure, a method for determining background audio is provided, including:

[0007] Invoking a large language model to perform multimodal processing on the video to obtain a search term for the video, the search term including various video information, and the various video information including video content information, video emotion information, and video scene information of the video;

[0008] Based on the video content information, the video emotion information, and the video scene information in the various video information, determining a query order, where the query order is the order for querying background audio;

[0009] According to the query order, based on the search term and the tags of each audio in the audio library, screening the audios in the audio library to obtain at least one background audio, where the tag of each audio in the audio library includes audio content information, audio emotion information, and audio scene information of the audio, and the audio content information, the audio emotion information, and the audio scene information of the background audio respectively match the video content information, the video emotion information, and the video scene information of the video one by one.

[0010] According to another aspect of the embodiments of the present disclosure, a device for determining background audio is provided, including:

[0011] A first processing unit is configured to execute calling a large language model to perform multimodal processing on a video to obtain a search term for the video, wherein the search term includes a plurality of video information, and the plurality of video information includes video content information, video emotion information, and video scene information of the video;

[0012] A determining unit is configured to determine a query order based on the video content information, the video emotion information, and the video scene information in the multiple video information, wherein the query order is an order for querying background audio;

[0013] The screening unit is configured to perform screening of the audios in the audio library in accordance with the query order based on the search terms and the labels of the respective audios in the audio library to obtain at least one background audio, wherein the label of each audio in the audio library includes audio content information, audio emotion information and audio scene information of the audio, and the audio content information, audio emotion information and audio scene information of the background audio are matched one by one with the video content information, video emotion information and video scene information of the video, respectively.

[0014] In some embodiments, the determination unit is configured to execute obtaining the priority among the video content information, the video emotion information and the video scene information in the multiple video information, the priority of each video information being used for the importance of the video information for matching the background audio; and the order of priority from high to low is determined as the query order.

[0015] In some embodiments, the screening unit is configured to perform multiple rounds of screening of the audio in the audio library in accordance with the query order, based on the search terms and the tags of each audio in the audio library, to obtain the at least one background audio, and the audio used in each round of screening is obtained in the previous round of screening.

[0016] In some embodiments, the query order is the video content information, the video emotion information, and the video scene information;

[0017] The screening unit is configured to perform screening of multiple first audios whose content matching scores meet the content matching condition from the audio library based on the video content information and the audio content information of each audio in the audio library; screening of multiple second audios whose emotion matching scores meet the emotion matching condition from the multiple first audios based on the video emotion information and the audio emotion information of each audio in the multiple first audios; and screening of at least one background audio whose scene matching score meets the scene matching condition from the multiple second audios based on the video scene information and the audio scene information of each audio in the multiple second audios.

[0018] In some embodiments, the first processing unit is configured to execute calling the large language model to perform multimodal processing on the video to obtain multiple video description information, and obtain the multiple video information from the multiple video description information, and the angles of describing the videos in the multiple video description information are not completely the same.

[0019] In some embodiments, the first processing unit includes:

[0020] An acquisition subunit, configured to acquire multiple frames of images from the video;

[0021] The processing subunit is configured to input the multiple frames of images into the large language model, perform image recognition and natural language processing on the multiple frames of images through the large language model, and obtain the multiple video description information.

[0022] In some embodiments, the acquisition subunit is configured to perform at least one of the following:

[0023] Randomly acquire multiple frames of images from the video;

[0024] Based on a preset time interval, sampling images from the video to obtain multiple frames of images;

[0025] Based on the duration of the video, multiple frames of images are obtained from the video, and the number of the multiple frames of images is positively correlated with the duration of the video.

[0026] In some embodiments, the first processing unit is configured to execute the steps of obtaining at least one content entity from the multiple video description information; obtaining at least one associated word of the content entity for any one of the at least one content entity, wherein the at least one associated word is used to represent something related to the content entity; and summing up the at least one content entity and the corresponding at least one associated word to obtain the video content information.

[0027] In some embodiments, the first processing unit is configured to execute obtaining at least one emotional keyword from the multiple video description information; for any emotional keyword among the at least one emotional keyword, filtering out a preset emotional word matching the emotional keyword from an emotional vocabulary, the emotional vocabulary including multiple preset emotional words; summing up the at least one emotional keyword and the preset emotional vocabulary corresponding to the at least one emotional keyword to obtain the video emotional information.

[0028] In some embodiments, there are multiple emotional keywords;

[0029] The device also includes:

[0030] The second processing unit is configured to execute if there are contradictory preset emotional words among the preset emotional words corresponding to the multiple emotional keywords, then classify the preset emotional words corresponding to the multiple emotional keywords, and count the number of each category of preset emotional words; from the preset emotional words corresponding to the multiple emotional keywords, filter out the preset emotional words of other categories except the preset emotional words belonging to the target category, and the number of preset emotional words in the target category is the largest.

[0031] In some embodiments, the first processing unit is configured to execute obtaining at least one scene keyword from the multiple video description information; for any scene keyword among the at least one scene keyword, filtering out a preset scene word matching the scene keyword from a scene vocabulary, the scene vocabulary including multiple preset scene words; summing up the at least one scene keyword and the preset scene vocabulary corresponding to the at least one scene keyword to obtain the video scene information.

[0032] In some embodiments, the apparatus further comprises:

[0033] A first acquisition unit is configured to acquire confidence scores of the plurality of video description information, wherein the confidence score of each video description information is used to indicate the accuracy of the video description information;

[0034] The first processing unit is configured to obtain the multiple video information from the video description information whose confidence score reaches a preset standard.

[0035] In some embodiments, the apparatus further comprises:

[0036] A second acquisition unit is configured to acquire at least one additional information of time information, season information, character identification and building type in the video;

[0037] The screening unit is further configured to perform screening of the at least one background audio based on the at least one additional information to obtain at least one third audio, wherein the at least one background audio is audio that has been screened out based on the multiple video information, and the at least one third audio matches the at least one additional information of the video.

[0038] In some embodiments, the apparatus further comprises:

[0039] A third acquisition unit is configured to acquire, for any background audio of the at least one background audio, at least one of style information, quality information, and interaction information of the background audio, wherein the interaction information includes at least one of an interactive behavior and quantity in a video using the background audio and a number of times the background audio is used;

[0040] The screening unit is further configured to perform screening of the at least one background audio based on at least one of the style information, the quality information and the interaction information to obtain at least one fourth audio, wherein the at least one fourth audio meets application conditions in at least one dimension of style, quality and interaction, and the application conditions refer to conditions that the audio must meet when the audio is used in combination with video.

[0041] According to another aspect of an embodiment of the present disclosure, there is provided an electronic device, the electronic device comprising:

[0042] one or more processors;

[0043] A memory for storing program codes executable by the processor;

[0044] The processor is configured to execute the program code to implement the above-mentioned method for determining the background audio.

[0045] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When a program code in the computer-readable storage medium is executed by a processor of an electronic device, the electronic device can execute the above-mentioned method for determining background audio.

[0046] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the above-mentioned method for determining background audio when executed by a processor.

[0047] The solution provided by the disclosed embodiment, in the process of querying background audio for a video, performs multimodal processing on the video by calling a large language model, which can not only obtain a variety of video information in multiple dimensions such as video content information, video emotion information and video scene information, but also make the video information richer and more complete, and ensure the accuracy of the information in each dimension in the search term; then, according to the query order of the multiple video information in the search term, the audio in the audio library is screened, that is, the matching audio is screened in a more fine-grained manner, ensuring that at least one of the screened background audio matches various information in the search term of the video. In summary, this solution performs recall in three dimensions, namely, content, emotion and scene, in the process of querying background audio, so that the queried background audio and video match in the three dimensions of content, emotion and scene, with a higher degree of matching, more accurate and in line with user needs.

[0048] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0050] Figure 1 The figure is a schematic diagram of an implementation environment of a method for determining background audio according to an exemplary embodiment.

[0051] Figure 2 The figure is a flowchart of a method for determining background audio according to an exemplary embodiment.

[0052] Figure 3 The figure is a flowchart of another method for determining background audio according to an exemplary embodiment.

[0053] Figure 4 The figure is a schematic diagram showing a method of querying background audio according to an exemplary embodiment.

[0054] Figure 5 The diagram is a framework diagram of querying background audio according to an exemplary embodiment.

[0055] Figure 6 is a framework diagram of another method of querying background audio according to an exemplary embodiment.

[0056] Figure 7 The invention is a block diagram of a device for determining background audio according to an exemplary embodiment.

[0057] Figure 8 It is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0058] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0059] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0060] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the videos, background audio, content information, emotional information, and scene information involved in this disclosure are all obtained under sufficient authorization.

[0061] Figure 1 It is a schematic diagram of an implementation environment of a method for determining background audio shown according to an exemplary embodiment. Taking the electronic device being provided as a server as an example, refer to Figure 1 , this implementation environment specifically includes: a terminal 101 and a server 102.

[0062] The terminal 101 is at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player, an MP4 player, and a laptop portable computer. An application program that supports background audio matching runs on the terminal 101. This application program can be a clip application program, a multimedia application program (such as a short video application program), a social application program, etc., and the embodiments of this disclosure do not limit this. The user can log in to this application program through the terminal 101 to obtain the services provided by this application program. The terminal 101 can be connected to the server 102 through a wireless network or a wired network, and thus can send the video to which background audio is to be added to the server 102. The server 102 queries the audio library for this video, finds the matching background audio for this video, and returns it to the terminal 101. The terminal 101 can display the background audio queried by the server 102 to the user.

[0063] The terminal 101 generally refers to one of multiple terminals. In this embodiment, the terminal 101 is used as an example for illustration. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals can be several, or the above terminals can be dozens or hundreds, or a larger number. The embodiments of this disclosure do not limit the number and device types of the terminals.

[0064] In some embodiments, server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data and artificial intelligence platforms. Server 102 is used to provide background services for applications that support background audio matching. In some embodiments, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 use a distributed computing architecture for collaborative computing.

[0065] Figure 2 is a flowchart of a method for determining background audio according to an exemplary embodiment. Figure 2 The background audio determination method is applied to a server and includes the following steps:

[0066] In step 201, the server calls a large language model to perform multimodal processing on the video to obtain a search term for the video, where the search term includes a variety of video information, and the variety of video information includes video content information, video emotion information, and video scene information.

[0067] In an embodiment of the present disclosure, upon receiving a background audio query instruction for a video, in response to the background audio query instruction, the server invokes a large language model to perform multimodal processing on the video to obtain a search term for the video.

[0068] The large language model can be a BERT (Bidirectional Encoder Representations from Transformers) model, a GPT-3 (Generative Pretrained Transformer 3) model, or a LaMDA (Language Models for Dialog Applications), etc., which is not limited in the embodiments of the present disclosure. Multimodal processing refers to processing information in multiple modes, such as visual information (such as images in the video), auditory information (such as audio in the video), and text information (such as text in the video) in the video to obtain video information. The embodiments of the present disclosure do not limit the specific processing method.

[0069] The search words of a video include various video information of the video. Various video information may include video content information, video emotion information, and video scene information. That is, the search words of a video may include content search words, emotion search words, and scene search words. Content search words are used to indicate the content of a video, such as "cat in a video". Emotion search words are used to indicate the emotion of a video, such as "positive". Scene search words are used to indicate scenes in a video, such as "field".

[0070] In step 202, the server determines a query order based on video content information, video emotion information, and video scene information in the multiple video information, and the query order is the order of querying the background audio.

[0071] In the disclosed embodiment, the server determines the query order between the various video information in the search term. For example, the server determines the query order between the video content information, the video emotion information, and the video scene information in the search term. The query order refers to the order between querying the audio for content matching, querying the audio for emotion matching, and querying the audio for scene matching. The disclosed embodiment does not limit the query order. For example, the server determines to query the audio for content matching first, then query the audio for emotion matching, and finally query the audio for scene matching.

[0072] In step 203, the server filters the audios in the audio library in accordance with the query order, based on the search terms and the tags of each audio in the audio library, to obtain at least one background audio. The tag of each audio in the audio library includes the audio content information, audio emotion information and audio scene information of the audio. The audio content information, audio emotion information and audio scene information of the background audio are matched one by one with the video content information, video emotion information and video scene information of the video, respectively.

[0073] In the disclosed embodiment, the server compares the content, emotion, scene and other information of the audio in the audio library one by one according to the query order and based on the content, emotion, scene and other information of the video to determine at least one background audio that matches the content, emotion and scene of the video.

[0074] The disclosed embodiment provides a method for determining background audio. In the process of querying background audio for a video, a large language model is called to perform multimodal processing on the video, which can not only obtain a variety of video information in multiple dimensions such as video content information, video emotion information, and video scene information, but also make the video information richer and more complete, and ensure the accuracy of the information in each dimension in the search term; then, according to the query order of the multiple video information in the search term, the audio in the audio library is screened, that is, the matching audio is screened in a more fine-grained manner, ensuring that at least one of the screened background audio matches various information in the search term of the video. In summary, this solution performs recall in three dimensions, namely, content, emotion, and scene, in the process of querying background audio, so that the queried background audio and video match in the three dimensions of content, emotion, and scene, with a higher degree of matching, more accurate, and in line with user needs.

[0075] In some embodiments, the query order is determined based on the video content information, the video emotion information, and the video scene information in the multiple video information, including:

[0076] Obtaining priorities among video content information, video emotion information, and video scene information in multiple video information, wherein the priority of each video information is used for the importance of the video information to matching the background audio;

[0077] The order of priority from high to low is determined as the query order.

[0078] The solution provided by the embodiment of the present disclosure, since the priority of each video information can reflect the importance of the video information to matching the background audio, determines the query order by arranging the priorities of the various video information in the search terms from high to low, so that the audio matching the information with higher priority can be queried subsequently first, that is, it is guaranteed that the background audio and video queried subsequently match in the dimension of the information with higher priority, thereby ensuring the accuracy of the background audio.

[0079] In some embodiments, the audio in the audio library is screened in the order of the query based on the search terms and the tags of each audio in the audio library to obtain at least one background audio, including:

[0080] According to the query order, based on the search terms and the labels of each audio in the audio library, multiple rounds of screening are performed on the audio in the audio library to obtain at least one background audio. The audio used in each round of screening is the audio obtained in the previous round of screening.

[0081] The solution provided by the embodiment of the present disclosure is that the audio used in each round of screening is obtained in the previous round of screening, so that in each round of screening, there is no need to match with all the audio in the audio library, but only needs to match another kind of information with the audio obtained in the previous round of screening, thereby reducing the calculation of the amount of data in the matching process, thereby improving the efficiency of audio query.

[0082] In some embodiments, the query order is video content information, video emotion information, and video scene information;

[0083] According to the query order, based on the search terms and the labels of each audio in the audio library, the audio in the audio library is screened multiple times to obtain at least one background audio, including:

[0084] Based on the video content information and the audio content information of each audio in the audio library, a plurality of first audios whose content matching scores meet the content matching condition are screened out from the audio library;

[0085] Based on the video emotion information and the audio emotion information of each of the multiple first audios, screening out multiple second audios whose emotion matching scores meet the emotion matching condition from the multiple first audios;

[0086] Based on the video scene information and the audio scene information of each of the multiple second audios, at least one background audio whose scene matching score meets the scene matching condition is screened out from the multiple second audios.

[0087] The solution provided by the embodiment of the present disclosure, in the process of querying background audio, first queries the audio that matches the content, then queries the audio that matches both the content and the emotion from the audio that matches the content, and then queries the audio that matches the content, emotion and scene from the audio that matches both the content and emotion, thereby ensuring that the queried audio and video match to a higher degree, are more accurate and meet user needs; and in each round of query, there is no need to match with all the audio in the audio library, which reduces the calculation of the amount of data during the matching process, thereby improving the efficiency of audio query.

[0088] In some embodiments, a large language model is called to perform multimodal processing on a video to obtain search terms for the video, including:

[0089] A large language model is called to perform multimodal processing on the video to obtain multiple video description information. From the multiple video description information, multiple video information is obtained. The angles of describing the videos in the multiple video description information are not completely the same.

[0090] The solution provided by the disclosed embodiment is beneficial to the powerful language understanding and text generation capabilities of the large language model, and performs multimodal processing on the video, so that the obtained multiple video description information can describe the video from different angles, and describe the video information in multiple dimensions such as content, emotion and scene in the video as much as possible, thereby ensuring that the acquired video information is more complete, and then on this basis, extract information in various dimensions such as video content information, video emotion information and video scene information, thereby ensuring the accuracy of information in various dimensions.

[0091] In some embodiments, a large language model is called to perform multimodal processing on a video to obtain multiple video description information, including:

[0092] Get multiple frames of images from a video;

[0093] Multiple frames of images are input into the large language model, and image recognition and natural language processing are performed on the multiple frames of images through the large language model to obtain multiple video description information.

[0094] The solution provided by the disclosed embodiment utilizes the powerful image recognition and natural language processing capabilities of a large language model to process multiple frames of images in a video. This not only ensures that the obtained multiple video description information can more accurately reflect various information in the video, but also reduces the amount of processed data compared to processing the entire video, thereby improving the efficiency of obtaining multiple video description information, and further helping to improve the efficiency of obtaining background audio.

[0095] In some embodiments, acquiring multiple frames of images from a video includes at least one of the following:

[0096] Randomly obtain multiple frames of images from the video;

[0097] Based on a preset time interval, sampling images from the video to obtain multiple frames of images;

[0098] Based on the duration of the video, multiple frame images are obtained from the video, and the number of the multiple frame images is positively correlated with the duration of the video.

[0099] The solution provided by the embodiments of the present disclosure not only enriches the methods of obtaining multiple-frame images from a video, but also randomly obtains multiple-frame images, which is simple to operate and can improve the efficiency of obtaining multiple-frame images; obtaining multiple-frame images based on a preset time interval or video duration ensures that the multiple-frame images can more accurately reflect the overall content of the video, which is conducive to the subsequent acquisition of more accurate video description information.

[0100] In some embodiments, a plurality of video information is obtained from the plurality of video description information, including:

[0101] At least one content entity is obtained from multiple video description information, and video content information is determined based on the at least one content entity.

[0102] The solution provided by the embodiment of the present disclosure determines the video content information based on the content entities in multiple video description information, thus ensuring the accuracy of the video content information, because the content entity intuitively reflects the things appearing in the video.

[0103] In some embodiments, based on at least one content entity, determining video content information includes:

[0104] For any content entity of the at least one content entity, obtaining at least one associated word of the content entity, wherein the at least one associated word is used to represent something related to the content entity;

[0105] At least one content entity and at least one corresponding associated word are summed up to obtain video content information.

[0106] The solution provided by the embodiment of the present disclosure can obtain at least one associated word of each content entity in the video, and obtain video content information by summarizing the content entities in the video and the corresponding associated words. This not only enriches the video content information and facilitates subsequent queries to obtain richer audio, but also the associated words are things related to the content entities, ensuring that the audio after subsequent queries still conforms to the video content.

[0107] In some embodiments, a plurality of video information is obtained from the plurality of video description information, including:

[0108] At least one emotion keyword is obtained from multiple video description information, and based on the at least one emotion keyword, video emotion information is determined.

[0109] The solution provided by the embodiment of the present disclosure determines the video emotion information based on the emotion keywords in the video description information, thus ensuring the accuracy of the video emotion information, because the emotion keywords in the video description information intuitively reflect the emotion in the video.

[0110] In some embodiments, determining video emotion information based on at least one emotion keyword includes:

[0111] For any emotion keyword in the at least one emotion keyword, a preset emotion word matching the emotion keyword is selected from an emotion word library, where the emotion word library includes a plurality of preset emotion words;

[0112] At least one emotional keyword and at least one preset emotional vocabulary corresponding to the emotional keyword are summed up to obtain video emotional information.

[0113] The solution provided by the embodiment of the present disclosure, since the vocabulary for describing emotions is relatively rich, maps the emotion keywords extracted from the video to the preset emotion words in the emotion lexicon. On the basis of ensuring that the determined preset emotion words can accurately reflect the emotions in the video, the video emotion information of the video is normalized (or converged), so that in the subsequent process of emotion matching with the audio, the audio for emotion matching can be accurately queried.

[0114] In some embodiments, there are multiple sentiment keywords;

[0115] The method also includes:

[0116] If there are conflicting preset emotional words among the preset emotional words corresponding to the multiple emotional keywords, classify the preset emotional words corresponding to the multiple emotional keywords, and count the number of each type of preset emotional words;

[0117] From the preset emotional words corresponding to the plurality of emotional keywords, preset emotional words of other categories except the preset emotional words belonging to the target category are filtered out, and the number of preset emotional words in the target category is the largest.

[0118] The solution provided by the disclosed embodiment, when there are contradictory preset emotion words between the preset emotion words corresponding to multiple emotion keywords, classifies the preset emotion words. Since the target category has the most preset emotion words, it means that there are more emotions belonging to the target category in the video. By filtering out the preset emotion words of other categories outside the target category, the preset emotion words in the target category are used as the video emotion information, thereby ensuring the accuracy of the video emotion information.

[0119] In some embodiments, a plurality of video information is obtained from the plurality of video description information, including:

[0120] At least one scene keyword is obtained from multiple video description information, and video scene information is determined based on the at least one scene keyword.

[0121] The solution provided by the embodiment of the present disclosure determines the video scene information based on the scene keywords in the video description information, thereby ensuring the accuracy of the video scene information, because the scene keywords in the video description information intuitively reflect the scenes appearing in the video.

[0122] In some embodiments, based on at least one scene keyword, determining video scene information includes:

[0123] For any scene keyword of the at least one scene keyword, a preset scene word matching the scene keyword is selected from a scene word library, the scene word library including a plurality of preset scene words;

[0124] The at least one scene keyword and the at least one preset scene vocabulary corresponding to the scene keyword are summed up to obtain the video scene information.

[0125] The solution provided by the embodiment of the present disclosure, since the vocabulary for describing the scene may be relatively rich, by mapping the scene keywords extracted from the video to the preset scene words in the scene vocabulary, the video scene information of the video is normalized (or converged) on the basis of ensuring that the determined preset scene words can accurately reflect the scene in the video, so as to facilitate the accurate query of the audio after scene matching in the subsequent scene matching process with the audio.

[0126] In some embodiments, the method further comprises:

[0127] Obtain confidence scores of multiple video description information, where the confidence score of each video description information is used to indicate the accuracy of the video description information;

[0128] Get various video information from multiple video description information, including:

[0129] A variety of video information is obtained from video description information whose confidence scores reach a preset standard.

[0130] The solution provided by the embodiment of the present disclosure, since the confidence score of the video description information can reflect the accuracy of the video description information, the video description information with the confidence score reaching the preset standard is screened out from multiple video description information to obtain multiple videos, thereby ensuring the accuracy of the video information.

[0131] In some embodiments, the method further comprises:

[0132] Acquire at least one additional information of time information, season information, character identification, and building type in the video;

[0133] Based on at least one additional information, at least one background audio is screened to obtain at least one third audio, the at least one background audio is audio that has been screened based on multiple video information, and the at least one third audio matches at least one additional information of the video.

[0134] The solution provided by the embodiment of the present disclosure further filters at least one background audio that has been filtered out through at least one piece of information including time information, season information, character identification, and building type in the video, thereby further ensuring the accuracy of the background audio, that is, making it more closely matched with the video.

[0135] In some embodiments, the method further comprises:

[0136] For any background audio of the at least one background audio, obtaining at least one of style information, quality information, and interaction information of the background audio, wherein the interaction information includes at least one of an interaction behavior and quantity in a video using the background audio and a number of times the background audio is used;

[0137] Based on at least one of the style information, quality information and interaction information, at least one background audio is screened to obtain at least one fourth audio, and the at least one fourth audio meets the application conditions in at least one dimension of style, quality and interaction. The application conditions refer to the conditions that the audio must meet when the audio is used in combination with video.

[0138] The solution provided by the embodiment of the present disclosure further screens at least one background audio through at least one of the style information, quality information and interaction information of the background audio, thereby further ensuring better quality of the background audio.

[0139] Above Figure 2 The following is only a basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 FIG. 1 is a flow chart of another method for determining background audio according to an exemplary embodiment. Figure 3 , the method comprising:

[0140] In step 301, the server calls a large language model to perform multimodal processing on the video to obtain multiple video description information, and obtains multiple video information from the multiple video description information. The multiple video information includes video content information, video emotion information and video scene information. The angles of describing the videos in the multiple video description information are not exactly the same.

[0141] In an embodiment of the present disclosure, in the process of matching background audio for a video, in response to a user triggering a background audio query operation, the terminal sends a background audio query instruction to the server. The background audio query instruction may include a video. After receiving the background audio query instruction for the video, the server calls the large language model to perform multimodal processing on the video in the background audio query instruction to obtain multiple video description information of the video. Then, the server can continue to extract multiple video information such as video content information, video emotion information, and video scene information from the multiple video description information through the large language model.

[0142] Among them, each video description information can describe the information in the video from at least one angle. The at least one angle may include content angle, emotion angle, scene angle, time angle, space angle, etc., which is not limited by the embodiments of the present disclosure. The angles of describing the video in different video description information are not exactly the same. The solution provided by the embodiments of the present disclosure is conducive to the powerful language understanding and text generation capabilities of the large language model, and performs multimodal processing on the video, so that the obtained multiple video description information can describe the video from different angles, and describe the video in multiple dimensions such as content, emotion and scene in the video as much as possible, ensuring that the acquired video information is more complete, and then extracting information in various dimensions such as video content information, video emotion information and video scene information on this basis, ensuring the accuracy of information in various dimensions.

[0143] In the process of obtaining multiple video description information, the server may call the large language model to perform multimodal processing on the entire video to generate multiple video description information. Alternatively, the server may also call the large language model to perform multimodal processing on some images in the video to generate multiple video description information, etc., which is not limited in the embodiments of the present disclosure.

[0144] In some embodiments, the server calls a large language model to perform multimodal processing on some images in the video to generate multiple video description information. Accordingly, the server obtains multiple frames of images from the video. Then, the server inputs multiple frames of images into the large language model, and performs image recognition and natural language processing on the multiple frames of images through the large language model to obtain multiple video description information. The solution provided by the embodiment of the present disclosure utilizes the powerful image recognition and natural language processing capabilities of the large language model to process multiple frames of images in the video, which not only ensures that the obtained multiple video description information can more accurately reflect the various information in the video, but also reduces the amount of data processed compared to processing the entire video, thereby improving the efficiency of obtaining multiple video description information, and further facilitating the improvement of the efficiency of obtaining background audio.

[0145] The method of acquiring multiple frames of images from a video is not limited in the embodiments of the present disclosure. Three methods are described below by way of example, but are by no means limited thereto.

[0146] Method 1: The server randomly obtains multiple frames of images from the video. For example, the server obtains the first frame image and the last frame image from the video.

[0147] In the second method, the server samples images from the video based on a preset time interval to obtain multiple frames of images. The preset time interval can be customized by the user or determined based on the duration of the video, such as preset time interval = duration of the video / preset number of images, which is not limited in the embodiments of the present disclosure.

[0148] Method 3: The server obtains multiple frames of images from the video based on the duration of the video. The number of multiple frames of images is positively correlated with the duration of the video. The above three methods can be freely combined to obtain a new method, which is not limited in the embodiments of the present disclosure.

[0149] The solution provided by the embodiments of the present disclosure not only enriches the methods of obtaining multiple-frame images from a video, but also randomly obtains multiple-frame images, which is simple to operate and can improve the efficiency of obtaining multiple-frame images; obtaining multiple-frame images based on a preset time interval or video duration ensures that the multiple-frame images can more accurately reflect the overall content of the video, which is conducive to the subsequent acquisition of more accurate video description information.

[0150] After obtaining multiple video description information, the server can directly extract video content information, video emotion information and video scene information from the multiple video description information through the large language model. Alternatively, in the process of the large language model generating multiple video description information, the large language model can also provide the confidence of the multiple video description information. Then, the server uses the large language model to filter multiple video description information whose confidences meet the conditions from the multiple video description information, and then extracts video content information, video emotion information and video scene information from the multiple video description information whose confidences meet the conditions. That is, the server obtains the confidence scores (i.e., confidences) of the multiple video description information. The confidence score of each video description information is used to indicate the accuracy of the video description information. Then, the server obtains multiple video information from the video description information whose confidence scores reach the preset standard. The preset standard can be that the confidence score reaches a preset value; or, a preset number of video description information that is at the top in the order from high to low according to the confidence score, etc., which is not limited in the embodiments of the present disclosure. For example, the server extracts video content information, video emotion information, and video scene information from the top 10 video description information (ie, the Hashtag text of the video) in the order of confidence scores from high to low.

[0151] After obtaining multiple video description information, the server can extract multiple video information such as video content information, video emotion information, and video scene information from the multiple video description information through the large language model. Accordingly, the server obtains at least one content entity from the multiple video description information, and determines the video content information based on the at least one content entity. The server obtains at least one emotion keyword from the multiple video description information, and determines the video emotion information based on the at least one emotion keyword. The server obtains at least one scene keyword from the multiple video description information, and determines the video scene information based on the at least one scene keyword.

[0152] Among them, the content entity can be an object, animal or person in the video, etc., which is not limited in the embodiment of the present disclosure. Then, the server can summarize the obtained content entities to obtain video content information. Emotional keywords can be words that can reflect emotions such as positive, negative, sad, etc., which are not limited in the embodiment of the present disclosure. Then, the server can summarize the obtained emotional keywords to obtain video content information. Scene keywords can be words that can reflect scenes such as playground, grassland, classroom, etc., which are not limited in the embodiment of the present disclosure. Then, the server can summarize the obtained scene keywords to obtain video content information.

[0153] The solution provided by the embodiment of the present disclosure determines information of various dimensions such as video content information, video frequency emotion information and video scene information according to content entities, emotion keywords and scene keywords in multiple video description information, so that the video information is richer and more complete, and the accuracy of information of various dimensions is guaranteed.

[0154] In the process of obtaining video content information, in addition to composing the above-obtained content entities into video content information, the server can also associate and expand on the basis of the content entities to enrich the video content (i.e., content search terms). Accordingly, for any content entity in at least one content entity, the server obtains at least one associated word of the content entity. At least one associated word is used to represent things related to the content entity. Then, the server summarizes at least one content entity and at least one corresponding associated word to obtain video content information. For example, the associated word of the content entity "football" can be "World Cup". The solution provided by the embodiment of the present disclosure can obtain at least one associated word of each content entity in the video, and summarize the content entities in the video and the corresponding associated words to obtain video content information, which not only enriches the video content information and facilitates subsequent queries to obtain richer audio, but also the associated words are things related to the content entity, ensuring that the audio after subsequent queries still conforms to the video content.

[0155] In the process of obtaining associated words, the large language model can also output the confidence of each associated word and select associated words that meet the confidence conditions for output. For example, for any content entity, the first three associated words with higher confidence among multiple associated words are used as the associated words of the content entity.

[0156] In the process of obtaining the emotional information of the video, considering that emotions are relatively subjective and the vocabulary for describing emotions is relatively rich, in order to better describe the emotions in the video, the server can normalize the emotional keywords of the video. Accordingly, the process of the server determining the emotional information of the video based on at least one emotional keyword includes: for any emotional keyword in at least one emotional keyword, the server screens out a preset emotional word that matches the emotional keyword from the emotional vocabulary. The emotional vocabulary includes multiple preset emotional words. Then, the server sums up at least one emotional keyword and at least one preset emotional vocabulary corresponding to the emotional keyword to obtain the emotional information of the video. The solution provided by the embodiment of the present disclosure, since the vocabulary for describing emotions is relatively rich, by mapping the emotional keywords extracted from the video to the preset emotional words in the emotional vocabulary, on the basis of ensuring that the determined preset emotional words can accurately reflect the emotions in the video, the video emotional information of the video is normalized (or converged), so that in the subsequent process of emotional matching with the audio, the audio that is matched with the emotion can be accurately queried.

[0157] In some embodiments, the server can normalize any of the above-mentioned emotional keywords through a large language model. For example, the server inputs a prompt message Prompt into the large language model, and the prompt message Prompt is "Please tell me which of the following emotional lexicon the emotional keyword {text} maps to: [sad, inspirational, happy, healing, sweet, lonely, catharsis, missing, fresh, romantic, cute, funny, suspenseful, dynamic, sexy, quiet, atmospheric, excited, lyrical, relaxed, dreamy, nervous, passionate, lazy, grateful]. Then, the server can calculate the similarity between the emotional keyword {text} and each preset emotional word in the emotional lexicon through the large language model, and use the preset emotional word with the greatest similarity as the preset emotional word corresponding to the emotional keyword.

[0158] In some embodiments, there are multiple emotional keywords obtained as described above. If there are contradictory preset emotional words among the preset emotional words corresponding to the multiple emotional keywords, the server classifies the preset emotional words corresponding to the multiple emotional keywords and counts the number of preset emotional words of each category. Then, the server filters out preset emotional words of other categories except the preset emotional words belonging to the target category from the preset emotional words corresponding to the multiple emotional keywords. The number of preset emotional words in the target category is the largest. That is, the server uses the preset emotional words in the target category with the largest number of words as the main emotions in the video. The scheme provided by the embodiment of the present disclosure, in the case where there are contradictory preset emotional words among the preset emotional words corresponding to the multiple emotional keywords, classifies the preset emotional words. Since the target category has the most preset emotional words, it indicates that there are more emotions belonging to the target category in the video. By filtering out preset emotional words of other categories outside the target category, the preset emotional words in the target category are used as the video emotional information, thereby ensuring the accuracy of the video emotional information.

[0159] In the process of obtaining video scene information, considering that the description of the scene is relatively subjective and the vocabulary for describing the scene is relatively rich, in order to better describe the scene in the video, the server can normalize the scene keywords of the video. Accordingly, the process of determining the video scene information based on at least one scene keyword by the server includes: for any scene keyword in at least one scene keyword, the server selects a preset scene word that matches the scene keyword from the scene vocabulary. The scene vocabulary includes multiple preset scene words. Then, the server sums up at least one scene keyword and at least one preset scene vocabulary corresponding to the scene keyword to obtain the video scene information. The solution provided by the embodiment of the present disclosure, since the vocabulary for describing the scene may be relatively rich, by mapping the scene keywords extracted from the video to the preset scene words in the scene vocabulary, on the basis of ensuring that the determined preset scene words can accurately reflect the scene in the video, the video scene information of the video is normalized (or converged), so that in the subsequent process of scene matching with the audio, the audio of the scene matching can be accurately queried.

[0160] The disclosed embodiment does not limit the specific content and quantity of the multiple video information in the search term. Optionally, the multiple video information may also include video style information, video frame rate information and other video information. That is, the server may also obtain video style information, video frame rate information and other video information from the above-mentioned multiple video description information, so that the server may subsequently filter out audio whose style matches the video from the audio library, or filter out audio whose audio rhythm matches the frame rate of the video from the audio library, etc. The disclosed embodiment does not limit the method for obtaining information such as video style information and video frame rate information.

[0161] In step 302, the server determines a query order based on video content information, video emotion information, and video scene information in the multiple video information, and the query order is the order of querying the background audio.

[0162] In the embodiment of the present disclosure, the query order between multiple video information can be customized by the user, or can be determined based on the importance (or contribution) of the information to matching the background audio. The embodiment of the present disclosure does not limit the method for determining the query order.

[0163] In some embodiments, the server obtains the priority among video content information, video emotion information, and video scene information in a variety of video information. The priority of each type of information is used to determine the importance of the information to matching the background audio. Then, the server determines the order of priority from high to low as the query order. In the solution provided by the embodiment of the present disclosure, since the priority of each type of information can reflect the importance of the information to matching the background audio, by determining the order of priority from high to low among the various video information in the search term as the query order, the subsequent search for audio that matches the information with higher priority can be prioritized, that is, it is guaranteed that the background audio and video subsequently queried match in the dimension of higher priority information, thereby ensuring the accuracy of the background audio.

[0164] For example, the priority of video content information in the search term is higher than that of video emotion information, and the priority of video emotion information is higher than that of video scene information. In the subsequent query of background audio, the audio matching the content is first queried based on the video content information, and then the audio matching the emotion is queried from the audio matching the content based on the video emotion information. Finally, the audio matching the scene is queried from the audio matching both the content and emotion based on the video scene information.

[0165] After determining the query order, the server can perform multiple rounds of screening on the audio in the audio library based on the information of multiple videos to search for audio that matches the information of multiple videos. In each round of screening, the server can match the information of each audio in the audio library based on the information of the video. Prior to this, the server needs to obtain the information of each audio in the audio library, see step 303 for details.

[0166] In step 303, the server obtains tags of each audio in the audio library, and the tag of each audio in the audio library includes audio content information, audio emotion information, and audio scene information of the audio.

[0167] In an embodiment of the present disclosure, the audio content information of each audio may include at least one of the name of the audio (such as a song title), the text content of the audio (such as lyrics), the alias of the audio, the singer of the audio, the theme of the audio, the vertical label of the audio, and the user portrait adapted for the audio. Vertical labels are used to represent things associated with the audio, similar to the associative words in the above-mentioned video content information. For any audio, the server can count the usage of the audio in real situations to determine the vertical label of the audio. For example, if the audio is the theme song of a certain "Football World Cup", the vertical label of the audio can be "Football".

[0168] For any audio in the audio library, the server can also use a large language model to obtain the audio content information, audio emotion information and audio scene information of the audio to form a label for the audio.

[0169] For example, the server inputs a prompt message Prompt into the large language model as "Given the lyrics {lyrics}, please use 5 keywords to describe the theme of this song"; then, the large language model outputs 5 theme keywords. The server inputs a prompt message Prompt into the large language model as "Given the lyrics {lyrics}, please use 5 keywords to describe the emotion of this song"; then, the large language model outputs 5 emotion keywords. The server inputs a prompt message Prompt into the large language model as "Given the lyrics {lyrics}, please use 2 keywords to describe the scene of this song"; then, the large language model outputs 2 scene keywords.

[0170] The process of the server obtaining the audio tag in step 303 is similar to the principle of the server obtaining the video content information, video emotion information and video scene information in step 301, and will not be repeated here.

[0171] In some embodiments, the label of each audio in the audio library may also include time, season, gender of the adapted user, style, interactive information, etc., which is not limited in the embodiments of the present disclosure. The interactive information may include at least one of the interactive behaviors (such as likes, shares, etc.) in the video using the audio and the number of corresponding interactive behaviors and the number of times the audio is used.

[0172] For example, the server inputs the prompt information Prompt into the large language model as "Given the lyrics {lyrics}, use 1-2 keywords to describe the season, time, and gender of the people suitable for this song."

[0173] It should be noted that the disclosed embodiment does not limit the timing of obtaining the tags of each audio in the audio library. For example, the server can first obtain the tags of each audio in the audio library, then obtain the search term of the video, and then search (query). That is, the server first executes step 303, and then executes steps 301-302 and step 304.

[0174] In step 304, the server screens the audio in the audio library in accordance with the query order based on the search terms and the tags of each audio in the audio library to obtain at least one background audio.

[0175] In an embodiment of the present disclosure, the server screens the audio in the audio library according to the query order. The present disclosure does not limit the screening method. In some embodiments, the server can screen out the audio set that matches the video content, the audio set that matches the video emotion, and the audio set that matches the video scene from the audio library according to the query order, and then use the audio that exists in each of the above audio sets as the screened background audio. In other embodiments, the server can perform multiple rounds of screening on the audio in the audio library according to the query order based on the search terms and the tags of each audio in the audio library to obtain at least one background audio. The audio used in each round of screening is obtained in the previous round of screening. In the solution provided by the embodiment of the present disclosure, since the audio used in each round of screening is obtained in the previous round of screening, it is not necessary to match all the audio in the audio library in each round of screening, but only to match another type of information with the audio obtained in the previous round of screening, which reduces the calculation of the amount of data in the matching process, thereby improving the efficiency of audio queries.

[0176] In each screening process, the server matches the search terms with the tags of each audio in the audio library to obtain multiple audios in this round of screening. Then, based on the multiple audios that have been screened, the server performs the next round of screening. The audio content information of the background audio matches the video content information of the video; the audio emotion information of the background audio matches the video emotion information of the video; and the audio scene information of the background audio matches the video scene information of the video.

[0177] In some embodiments, the query order is video content information, video emotion information, and video scene information. The server can perform content matching, emotion matching, and scene matching in sequence. Accordingly, based on the video content information and the audio content information of each audio in the audio library, the server screens out multiple first audios whose content matching scores (i.e., content matching degree or content similarity) meet the content matching conditions from the audio library. Based on the video emotion information and the audio emotion information of each audio in the multiple first audios, the server screens out multiple second audios whose emotion matching scores (i.e., emotion matching degree or emotion similarity) meet the emotion matching conditions from the multiple first audios. Based on the video scene information and the audio scene information of each audio in the multiple second audios, the server screens out at least one background audio whose scene matching score (i.e., scene matching degree or scene similarity) meets the scene matching conditions from the multiple second audios.

[0178] The solution provided by the embodiment of the present disclosure, in the process of querying background audio, first queries the audio that matches the content, then queries the audio that matches both the content and the emotion from the audio that matches the content, and then queries the audio that matches the content, emotion and scene from the audio that matches both the content and emotion, thereby ensuring that the queried audio and video match to a higher degree, are more accurate and meet user needs; and in each round of query, there is no need to match with all the audio in the audio library, which reduces the calculation of the amount of data during the matching process, thereby improving the efficiency of audio query.

[0179] If the number of audios screened out by the server after a round of screening is zero, the server can use the audios screened out in the previous round as the background audios for the video. If the number of audios screened out by the server after a round of screening is 1, the server stops the next round of screening and determines the screened audio as the background audio for the video.

[0180] In some embodiments, the server may also obtain at least one additional information of the time information, season information, character identification, and building type in the video. Then, based on the at least one additional information, the server filters at least one background audio to obtain at least one third audio. Among them, the at least one background audio is the audio that has been filtered out based on the above-mentioned multiple video information. The at least one third audio matches at least one additional information of the video and matches the above-mentioned multiple video information. The solution provided by the embodiment of the present disclosure further filters at least one background audio through at least one information of the time information, season information, character identification, and building type in the video, thereby further ensuring the accuracy of the background audio, that is, it is more matched with the video.

[0181] In some other embodiments, for any background audio in at least one background audio, the server may also obtain at least one of the style information, quality information, and interaction information of the background audio. The interaction information includes at least one of the interactive behavior and quantity in the video using the background audio and the number of times the background audio is used. Then, the server screens the at least one background audio based on at least one of the style information, quality information, and interaction information to obtain at least one fourth audio. The at least one fourth audio meets the application condition in at least one dimension of style, quality, and interaction.

[0182] Among them, the application condition refers to the condition that the audio must meet when the audio is used in combination with the video. The application condition can be that the style of the audio conforms to the preset style, or the correlation between the style of the audio and the preset style does not exceed the preset value, etc. Among them, the preset style can match at least one of the user's attribute information and the style of the video. The user's attribute information may include static attribute information such as the user's age, gender, and constellation, as well as dynamic attribute information such as the user's interactive behavior for other videos. Alternatively, the application condition may also be that the quality of the audio reaches a preset quality standard, for example, the signal-to-noise ratio of the audio is higher than a preset value. Alternatively, the application condition may also be that the number of times the audio is used reaches a number threshold, or the number of positive interactive behaviors in the video using the audio reaches a number threshold, etc. The present disclosure embodiment does not impose any restrictions on the application conditions. Positive interactive behaviors may include interactive behaviors that can reflect positive emotions such as likes, shares, positive comments, and collections.

[0183] The solution provided by the embodiment of the present disclosure further screens at least one background audio through at least one of the style information, quality information and interaction information of the background audio, thereby further ensuring better quality of the background audio.

[0184] For example, Figure 4 FIG. 1 is a schematic diagram showing a method of querying background audio according to an exemplary embodiment. Figure 4, the server uses a large language model to perform multimodal processing on the video to which background audio is added, and obtains video content information such as video content entities, associative words, video emotional information, video scene information, and at least one of the time information, season information, character identification, and building type in the video. The server first matches the video content information with the audio content information of each audio in the audio library (such as song name, lyrics summary, song theme, vertical category label) to obtain multiple first audios whose content matching degree meets the conditions. Then, the server matches the video emotional information with the audio emotional information of each audio in the multiple first audios (such as preset emotional words, emotional keywords, emotional summary) to obtain multiple second audios whose emotional matching degree meets the conditions. Then, the server matches the video scene information with the audio scene information of each audio in the multiple second audios (such as scene keywords, preset scene words) to obtain multiple background audios whose emotional matching degree meets the conditions. Finally, the server can further filter multiple background audios based on at least one of the time information, season information, character identification, and building type to obtain at least one background audio. Alternatively, the server may further filter the at least one background audio again based on at least one of the style information, quality information, and interaction information of the background audio to obtain the final at least one background audio. The at least one background audio may be sorted based on any sorting strategy, and the present disclosure embodiment does not limit the sorting strategy.

[0185] During the screening process, the server can perform feature extraction on the search terms of the video to obtain video features (embedding). Video features include video content features, video emotion features, and video emotional features). Then, the server uses the video features to query the audio in the audio library. During the query process, for any audio in the audio library, the server can calculate the similarity between the audio features and the video features of the audio, and use the audio that meets the similarity conditions as the queried audio that matches the video. Among them, the audio features of each audio (including audio content features, audio emotional features, and audio scene features) can be pre-extracted and stored, so that when the background audio is subsequently queried, the audio features of each audio can be directly obtained and used. The condition for satisfying the similarity can be that the similarity reaches a preset value, or it can also be a preset number of audios that are at the top in the order of similarity from high to low, etc., and the embodiments of the present disclosure are not limited to this.

[0186] The server may use a cosine similarity calculation formula to calculate the similarity between the video and the audio. Please refer to the following formula 1:

[0187]

[0188] Among them, A is used to represent the video features of the video, which can be any one of the video content features, video emotion features and video scene features; B is used to represent the audio features of the audio, which can be any one of the audio content features, audio emotion features and audio scene features; n is used to represent the dimension of the feature; similarity is used to represent the similarity between the video and the audio in any dimension of content, emotion and scene.

[0189] For example, Figure 5 FIG. 1 is a framework diagram showing a method of querying background audio according to an exemplary embodiment. Figure 5 , the server uses a large language model to perform multimodal processing on the video to obtain multiple video description information of the video. Then, the server can perform keyword extraction on the multiple video description information to determine the video content information, video emotion information and video scene information. Then, the server performs feature extraction on the video content information, video emotion information and video scene information respectively to obtain video features such as video content features, video emotion features and video emotion features. Then, the server calculates the similarity between the video features and the audio features, and uses the audio that meets the similarity conditions as the background audio of the video.

[0190] In order to more clearly describe the method for determining background audio provided by the embodiment of the present disclosure, the method for determining background audio is further described below in conjunction with the accompanying drawings. Figure 6 FIG. 1 is another framework diagram of querying background audio according to an exemplary embodiment. Figure 6Before querying the background audio of the video, for any audio in the audio, the server can use a large language model to perform multimodal processing on the audio to obtain the audio content information (content + vertical category label), audio emotional information and audio scene information. That is, in order to enrich the audio labels in the audio library, this solution combines the large language model and expands the audio labels by analyzing the lyrics content. In addition to the song title and singer name, the theme, scene, emotion and other text descriptions are also extracted from the lyrics as the song label. At the same time, the time, season, gender and other information described in the lyrics can also be extracted. Then, the server can splice the text of the song title, singer name, and theme keyword together to form the content label of the audio; the vertical category adapted by the song is composed of the vertical category label of the audio; the emotional summary, emotion label and emotional keyword are composed of the emotional label of the audio; the scene label and scene keyword are composed of the scene label of the audio. Then, the server extracts the corresponding audio features from the content label, vertical category label, emotional label and scene label of the audio using the text representation model in the cross-modal retrieval model, and stores each type of feature separately. For example, the server can use an audio content list to store the audio content features of each audio in the audio library, use an audio vertical category list to store the audio vertical category features of each audio in the audio library (or the vertical category label can be directly classified into the audio content, which is not limited in the embodiments of the present disclosure), use an audio emotion list to store the audio emotion features of each audio in the audio library, and use an audio scene list to store the audio scene features of each audio in the audio library. In short, the server can use a large language model to recall the content, emotion, scene, and vertical category dimensions of each audio in the audio library, obtain the audio information in each dimension, and store them separately.

[0191] When searching for background audio of a video, in order to ensure that the search terms of the video are consistent with the labels of the audio, the server also uses a large language model to perform multimodal processing on the video to determine the video content information, video emotion information and video scene information. The video content information, video emotion information and video scene information correspond one-to-one with the audio content information (content + vertical category labels), audio emotion information and audio scene information. For example, the prompt information of the video is "Give the associative words of the content entity "{hashtag}", and give None if there are no associative words; at the same time, use one or two keywords to describe the season, time, scene, building type, gender, character, scene and emotion information in the video description information {hashtag}, and give a confidence score. If it does not exist, use None to indicate it. Give the answer directly". Then, the server will extract video information such as content entities, associative words, emotions and scenes, and use the text representation model in the cross-modal retrieval model to extract video features as the search terms (Query embedding) of the video.

[0192] Then, the server queries the audio in the audio library based on the search terms of the video. That is, the server can query from a list of vertical tags storing audio content information and audio based on video content information such as content entities and associative words in the search terms to determine multiple first audios that match the content. Then, the server queries from multiple first audios based on the video emotion information in the search terms to determine multiple second audios that match the emotion. Then, the server queries from multiple second audios based on the video scene information in the search terms to determine at least one background audio that matches the scene. Then, the server can further filter and sort the at least one background audio obtained based on additional information such as season and time, as well as music quality, interactive information, etc., to obtain a background audio list corresponding to the video.

[0193] The disclosed embodiment provides a method for determining background audio. In the process of querying background audio for a video, a large language model is called to perform multimodal processing on the video, which can not only obtain a variety of video information in multiple dimensions such as video content information, video emotion information, and video scene information, but also make the video information richer and more complete, and ensure the accuracy of the information in each dimension in the search term; then, according to the query order of the multiple video information in the search term, the audio in the audio library is screened, that is, the matching audio is screened in a more fine-grained manner, ensuring that at least one of the screened background audio matches various information in the search term of the video. In summary, this solution performs recall in three dimensions, namely, content, emotion, and scene, in the process of querying background audio, so that the queried background audio and video match in the three dimensions of content, emotion, and scene, with a higher degree of matching, more accurate, and in line with user needs.

[0194] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0195] Figure 7 is a block diagram of a device for determining background audio according to an exemplary embodiment. Figure 7 , the device comprises:

[0196] The first processing unit 701 is configured to execute calling the large language model to perform multimodal processing on the video to obtain a search word of the video, where the search word includes a variety of video information, and the variety of video information includes video content information, video emotion information, and video scene information of the video;

[0197] The determining unit 702 is configured to determine a query order based on video content information, video emotion information, and video scene information in the multiple video information, where the query order is the order of querying the background audio;

[0198] The screening unit 703 is configured to perform screening of the audios in the audio library in accordance with the query order based on the search terms and the labels of each audio in the audio library to obtain at least one background audio, wherein the label of each audio in the audio library includes the audio content information, audio emotion information and audio scene information of the audio, and the audio content information, audio emotion information and audio scene information of the background audio are matched one by one with the video content information, video emotion information and video scene information of the video respectively.

[0199] In some embodiments, the determination unit 702 is configured to execute the acquisition of priorities among video content information, video emotion information, and video scene information in multiple video information, where the priority of each video information is used for the importance of the video information for matching background audio; and the order of priority from high to low is determined as the query order.

[0200] In some embodiments, the screening unit 703 is configured to perform multiple rounds of screening of the audio in the audio library in the order of the query, based on the search terms and the tags of each audio in the audio library, to obtain at least one background audio, and the audio used in each round of screening is obtained in the previous round of screening.

[0201] In some embodiments, the query order is video content information, video emotion information, and video scene information;

[0202] The screening unit 703 is configured to perform, based on the video content information and the audio content information of each audio in the audio library, screening out multiple first audios whose content matching scores meet the content matching conditions from the audio library; based on the video emotion information and the audio emotion information of each audio in the multiple first audios, screening out multiple second audios whose emotion matching scores meet the emotion matching conditions from the multiple first audios; based on the video scene information and the audio scene information of each audio in the multiple second audios, screening out at least one background audio whose scene matching score meets the scene matching conditions from the multiple second audios.

[0203] In some embodiments, the first processing unit 701 is configured to execute a call to a large language model to perform multimodal processing on the video, obtain multiple video description information, and obtain multiple video information from the multiple video description information, and the angles of describing the videos in the multiple video description information are not exactly the same.

[0204] In some embodiments, the first processing unit 701 includes:

[0205] An acquisition subunit is configured to acquire multiple frames of images from a video;

[0206] The processing subunit is configured to input multiple frames of images into the large language model, perform image recognition and natural language processing on the multiple frames of images through the large language model, and obtain multiple video description information.

[0207] In some embodiments, an acquisition subunit is configured to perform at least one of the following:

[0208] Randomly acquire multiple frames of images from a video;

[0209] Perform image sampling from a video based on a preset time interval to obtain multiple frames of images;

[0210] Acquire multiple frames of images from a video based on the duration of the video, and the number of multiple frames of images is positively correlated with the duration of the video.

[0211] In some embodiments, a first processing unit 701 is configured to obtain at least one content entity from multiple video description information, and determine video content information based on the at least one content entity.

[0212] In some embodiments, for any content entity in at least one content entity, the first processing unit 701 is configured to obtain at least one associated word of the content entity, where the at least one associated word is used to represent things related to the content entity; aggregate the at least one content entity and the corresponding at least one associated word to obtain video content information.

[0213] In some embodiments, the first processing unit 701 is configured to obtain at least one emotion keyword from multiple video description information, and determine video emotion information based on the at least one emotion keyword.

[0214] In some embodiments, for any emotion keyword in at least one emotion keyword, the first processing unit 701 is configured to screen out preset emotion words that match the emotion keyword from an emotion word library, where the emotion word library includes multiple preset emotion words; aggregate the at least one emotion keyword and the preset emotion words corresponding to the at least one emotion keyword to obtain video emotion information.

[0215] In some embodiments, there are multiple emotion keywords;

[0216] The apparatus further includes:

[0217] A second processing unit is configured to, if there are conflicting preset emotion words among the preset emotion words corresponding to multiple emotion keywords, classify the preset emotion words corresponding to the multiple emotion keywords, and count the number of preset emotion words in each category; filter out other categories of preset emotion words except the preset emotion words belonging to the target category from the preset emotion words corresponding to the multiple emotion keywords, where the number of preset emotion words in the target category is the largest.

[0218] In some embodiments, the first processing unit 701 is configured to obtain at least one scene keyword from multiple video description information, and determine video scene information based on the at least one scene keyword.

[0219] In some embodiments, the first processing unit 701 is configured to execute, for any scene keyword among at least one scene keyword, filter out a preset scene word that matches the scene keyword from a scene vocabulary, the scene vocabulary including multiple preset scene words; and sum up at least one scene keyword and at least one preset scene vocabulary corresponding to the scene keyword to obtain video scene information.

[0220] In some embodiments, the apparatus further comprises:

[0221] A first acquisition unit is configured to acquire confidence scores of multiple video description information, where the confidence score of each video description information is used to indicate the accuracy of the video description information;

[0222] The first processing unit 701 is configured to obtain a variety of video information from video description information whose confidence score reaches a preset standard.

[0223] In some embodiments, the apparatus further comprises:

[0224] A second acquisition unit is configured to acquire at least one additional information of time information, season information, character identification and building type in the video;

[0225] The screening unit 703 is also configured to perform screening of at least one background audio based on at least one additional information to obtain at least one third audio, wherein the at least one background audio is an audio that has been screened based on multiple video information, and the at least one third audio matches at least one additional information of the video.

[0226] In some embodiments, the apparatus further comprises:

[0227] A third acquisition unit is configured to acquire, for any background audio of at least one background audio, at least one of style information, quality information, and interaction information of the background audio, wherein the interaction information includes at least one of an interaction behavior and quantity in a video using the background audio and a usage count of the background audio;

[0228] The screening unit 703 is further configured to perform screening on at least one background audio based on at least one of the style information, quality information and interaction information to obtain at least one fourth audio, where the at least one fourth audio meets the application condition in at least one dimension of style, quality and interaction, and the application condition refers to the condition that the audio must meet when the audio is used in combination with video.

[0229] The disclosed embodiment provides a device for determining background audio. In the process of querying background audio for a video, by calling a large language model to perform multimodal processing on the video, not only can a variety of video information in multiple dimensions such as video content information, video emotion information, and video scene information be obtained, the video information is more abundant and complete, and the accuracy of the information in each dimension in the search term is guaranteed; then, according to the query order of the various video information in the search term, the audio in the audio library is screened, that is, the matching audio is screened in a more fine-grained manner, ensuring that at least one of the screened background audio matches various information in the search term of the video. In summary, this solution performs recall in three dimensions, namely, content, emotion, and scene, in the process of querying background audio, so that the queried background audio and video match in the three dimensions of content, emotion, and scene, with a higher degree of matching, more accurate, and in line with user needs.

[0230] It should be noted that the background audio determination device provided in the above embodiment only uses the division of the above functional units as an example when querying the background audio for a video. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the background audio determination device provided in the above embodiment and the background audio determination method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0231] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0232] When the electronic device is provided as a server, Figure 8 This is a block diagram of a server 800 according to an exemplary embodiment. The server 800 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 801 and one or more memories 802, wherein the memory 802 stores at least one program code, and the at least one program code is loaded and executed by the processor 801 to implement the background audio determination method provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 800 may also include other components for implementing device functions, which will not be described in detail here.

[0233] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 802 including instructions, and the instructions can be executed by a processor 801 of a server 800 to complete the above-mentioned method for determining background audio. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0234] A computer program product includes a computer program / instruction, which implements the above-mentioned method for determining background audio when executed by a processor.

[0235] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0236] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for determining background audio, characterized in that: The method comprises: Calling a large language model to perform multimodal processing on the video to obtain a search term for the video, wherein the search term includes a variety of video information, and the variety of video information includes video content information, video emotion information, and video scene information of the video; Determine a query order based on the video content information, the video emotion information, and the video scene information in the multiple video information, wherein the query order is an order for querying background audio; According to the query order, based on the search terms and the labels of each audio in the audio library, the audio in the audio library is screened to obtain at least one background audio, the label of each audio in the audio library includes audio content information, audio emotion information and audio scene information of the audio, and the audio content information, audio emotion information and audio scene information of the background audio are matched one by one with the video content information, video emotion information and video scene information of the video, respectively.

2. The method according to claim 1, characterized in that: The determining of the query order based on the video content information, the video emotion information and the video scene information in the multiple video information includes: Obtaining priorities among the video content information, the video emotion information, and the video scene information in the multiple video information, wherein the priority of each video information is used to indicate the importance of the video information to matching the background audio; The order of priorities from high to low is determined as the query order.

3. The method according to claim 1, characterized in that: The step of screening the audio in the audio library according to the query order and based on the search words and the tags of each audio in the audio library to obtain at least one background audio includes: According to the query order, based on the search terms and the tags of each audio in the audio library, multiple rounds of screening are performed on the audio in the audio library to obtain the at least one background audio, and the audio used in each round of screening is the audio obtained in the previous round of screening.

4. The method according to claim 3, characterized in that The query order is the video content information, the video emotion information, and the video scene information; The step of performing multiple rounds of screening on the audio in the audio library in accordance with the query order and based on the search terms and the tags of each audio in the audio library to obtain the at least one background audio includes: Based on the video content information and the audio content information of each audio in the audio library, a plurality of first audios whose content matching scores meet a content matching condition are screened out from the audio library; Based on the video emotion information and the audio emotion information of each of the multiple first audios, screening out multiple second audios whose emotion matching scores meet the emotion matching condition from the multiple first audios; Based on the video scene information and the audio scene information of each audio in the multiple second audios, the at least one background audio whose scene matching score meets the scene matching condition is screened out from the multiple second audios.

5. The method according to claim 1, characterized in that The calling of the large language model to perform multimodal processing on the video to obtain the search term of the video includes: The large language model is called to perform multimodal processing on the video to obtain a plurality of video description information, and the plurality of video information is obtained from the plurality of video description information, wherein the angles of describing the video in the plurality of video description information are not completely the same.

6. The method according to claim 5, characterized in that The calling of the large language model to perform multimodal processing on the video to obtain multiple video description information includes: Acquire multiple frames of images from the video; The multiple frames of images are input into the large language model, and image recognition and natural language processing are performed on the multiple frames of images through the large language model to obtain the multiple video description information.

7. The method according to claim 6, characterized in that The acquiring of multiple frames of images from the video includes at least one of the following: Randomly acquire multiple frames of images from the video; Based on a preset time interval, sampling images from the video to obtain multiple frames of images; Based on the duration of the video, multiple frames of images are obtained from the video, and the number of the multiple frames of images is positively correlated with the duration of the video.

8. The method according to claim 5, characterized in that The acquiring the multiple video information from the multiple video description information includes: Acquire at least one content entity from the plurality of video description information; For any content entity among the at least one content entity, obtaining at least one associated word of the content entity, wherein the at least one associated word is used to represent something related to the content entity; The at least one content entity and the corresponding at least one associated word are summed up to obtain the video content information.

9. The method according to claim 5, characterized in that The acquiring the multiple video information from the multiple video description information includes: Acquire at least one emotion keyword from the plurality of video description information; For any emotion keyword among the at least one emotion keyword, a preset emotion word matching the emotion keyword is selected from an emotion word library, wherein the emotion word library includes a plurality of preset emotion words; The at least one emotion keyword and the preset emotion vocabulary corresponding to the at least one emotion keyword are summed up to obtain the video emotion information.

10. The method according to claim 9, characterized in that There are multiple emotional keywords; The method further comprises: If there are conflicting preset emotional words among the preset emotional words corresponding to the multiple emotional keywords, classify the preset emotional words corresponding to the multiple emotional keywords, and count the number of each type of preset emotional words; From the preset emotional words corresponding to the plurality of emotional keywords, preset emotional words of other categories except the preset emotional words belonging to the target category are filtered out, and the number of preset emotional words in the target category is the largest.

11. The method according to claim 5, characterized in that The acquiring the multiple video information from the multiple video description information includes: Acquire at least one scene keyword from the plurality of video description information; For any scene keyword among the at least one scene keyword, selecting a preset scene word matching the scene keyword from a scene word library, the scene word library including a plurality of preset scene words; The at least one scene keyword and a preset scene vocabulary corresponding to the at least one scene keyword are summed up to obtain the video scene information.

12. The method according to any one of claims 5 to 11, characterized in that: The method further comprises: Obtaining confidence scores of the multiple video description information, where the confidence score of each video description information is used to indicate the accuracy of the video description information; The acquiring the multiple video information from the multiple video description information includes: The multiple video information is obtained from the video description information whose confidence scores reach a preset standard.

13. The method according to claim 1, characterized in that The method further comprises: Acquire at least one additional information of time information, season information, character identification, and building type in the video; Based on the at least one additional information, the at least one background audio is filtered to obtain at least one third audio, wherein the at least one background audio is audio that has been filtered out based on the multiple video information, and the at least one third audio matches the at least one additional information of the video.

14. The method according to claim 1, characterized in that The method further comprises: For any background audio of the at least one background audio, obtaining at least one of style information, quality information, and interaction information of the background audio, wherein the interaction information includes at least one of an interaction behavior and quantity in a video using the background audio and a number of times the background audio is used; Based on at least one of the style information, the quality information and the interaction information, the at least one background audio is screened to obtain at least one fourth audio, and the at least one fourth audio meets the application conditions in at least one dimension of style, quality and interaction, and the application conditions refer to the conditions that must be met when audio and video are used in combination.

15. A device for determining background audio, characterized in that: The device comprises: A first processing unit is configured to execute calling a large language model to perform multimodal processing on a video to obtain a search term for the video, wherein the search term includes a plurality of video information, and the plurality of video information includes video content information, video emotion information, and video scene information of the video; A determining unit is configured to determine a query order based on the video content information, the video emotion information, and the video scene information in the multiple video information, wherein the query order is an order for querying background audio; The screening unit is configured to perform screening of the audios in the audio library in accordance with the query order based on the search terms and the labels of the respective audios in the audio library to obtain at least one background audio, wherein the label of each audio in the audio library includes audio content information, audio emotion information and audio scene information of the audio, and the audio content information, audio emotion information and audio scene information of the background audio are matched one by one with the video content information, video emotion information and video scene information of the video, respectively.

16. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the method for determining background audio as described in any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for determining background audio as described in any one of claims 1 to 14.

18. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for determining background audio according to any one of claims 1 to 14 is implemented.