Background audio determination method and device, equipment and storage medium

By identifying the main objects in the visual screen of the multimedia resource and matching the background audio, the problem of background audio and multimedia resource mismatch is solved, improving the playback effect and reducing user interaction.

CN120104828APending Publication Date: 2025-06-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510161806.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When selecting background audio, the prior art is difficult to ensure that the background audio matches the content of the multimedia resource, resulting in poor playback results.

Method used

By obtaining visual information of multimedia resources, identifying the main objects in the visual picture, and automatically determining the matching background audio based on the object's clothing type, expression type and other characteristics.

Benefits of technology

It realizes the matching of background audio and multimedia resources, improves the playback effect of multimedia resources, and reduces user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104828A_ABST
    Figure CN120104828A_ABST
Patent Text Reader

Abstract

The invention relates to a background audio determination method and device, equipment and a storage medium, and relates to the technical field of multimedia. The method comprises the following steps: acquiring visual information of multimedia resources; identifying a first object appearing in a visual picture of the multimedia resource based on the visual information; and determining a background audio of the multimedia resource based on the first object, the background audio being matched with the first object. According to the method, the background audio is automatically determined for the multimedia resources, manual selection of a user is not needed, and interactive operation of the user is reduced; and the background audio is matched with the first object in the multimedia resource, so that the background audio is matched with the multimedia resource, and the playing effect of the multimedia resource is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of multimedia technology, and in particular to a method, device, equipment and storage medium for determining background audio. Background Art

[0002] With the development of multimedia technology, multimedia resource editing is becoming more and more convenient. In order to enrich the content of multimedia resources, users often add background audio to multimedia resources.

[0003] In the related art, when selecting background audio, one is often selected from multiple background audios as the background audio. Adding background audio in this way will cause the background audio to not match the content of the multimedia resource. Summary of the invention

[0004] The present disclosure provides a method, device, equipment and storage medium for determining background audio, which enables the background audio to match the multimedia resources, thereby improving the playback effect of the multimedia resources. The technical solution of the present disclosure is as follows:

[0005] According to one aspect of an embodiment of the present disclosure, a method for determining background audio is provided, the method comprising:

[0006] Acquire visual information from multimedia resources;

[0007] Based on the visual information, identifying a first object appearing in a visual picture of the multimedia resource;

[0008] Based on the first object, background audio of the multimedia resource is determined, and the background audio matches the first object.

[0009] In some embodiments, determining the background audio of the multimedia resource based on the first object includes at least one of the following:

[0010] identifying a clothing type of the first object, and determining audio matching the clothing type as background audio of the multimedia resource;

[0011] The expression type of the first object is identified, and audio matching the expression type is determined as background audio of the multimedia resource.

[0012] In some embodiments, the step of identifying, based on the visual information, a first object appearing in a visual picture of the multimedia resource includes:

[0013] In the case where a plurality of second objects are identified based on the visual information, a second object indicated by the text information of the multimedia resource among the plurality of second objects is determined as the first object; or

[0014] In the case where a plurality of second objects are identified based on the visual information, an object among the plurality of second objects whose appearance duration in the multimedia resource satisfies a first condition is determined as the first object.

[0015] In some embodiments, the step of identifying, based on the visual information, a first object appearing in a visual picture of the multimedia resource includes:

[0016] A first object appearing in the visual picture of the multimedia resource is identified based on the visual information through a multimodal recognition model, wherein the multimodal recognition model is a neural network model for identifying objects based on visual information.

[0017] In some embodiments, determining the background audio of the multimedia resource based on the first object includes:

[0018] Supplementary background audio is determined, and the background audio matching the first object and the supplementary background audio are determined as background audio recommended for the multimedia resource.

[0019] In some embodiments, the determining of the supplementary background audio comprises at least one of the following:

[0020] Determine the background audio with the first number of usage times ranked first on the first application as the supplementary background audio, wherein the first application is the application to which the multimedia resource belongs;

[0021] Determine the background audio used by the second number of multimedia resources ranked first in the number of times of playback on the first application as the supplementary background audio;

[0022] Determine the background audios ranked top by a third number of usage times on at least one second application as the supplementary background audios, where the second application is a resource playback application different from the first application;

[0023] The audio with the fourth largest number of playback times on at least one third application is determined as the supplementary background audio, and the third application is an audio playback application.

[0024] In some embodiments, the method further comprises:

[0025] Based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, identifying an element in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0026] The determining, based on the first object, the background audio of the multimedia resource comprises:

[0027] Based on the first object, determining a plurality of first background audios matching the first object;

[0028] The first background audio matching the element among the multiple first background audios is determined as the background audio of the multimedia resource.

[0029] In some embodiments, the identifying the element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource comprises at least one of the following:

[0030] In the case where multiple second scenes are identified, determining a scene whose appearance duration in the multimedia resource satisfies a second condition among the multiple second scenes as the first scene;

[0031] In the case where multiple second actions are identified, an action among the multiple second actions whose appearance duration in the multimedia resource satisfies a third condition is determined as the first action.

[0032] In some embodiments, there are multiple first objects, and determining the first background audio matching the element among the multiple first background audios as the background audio of the multimedia resource includes:

[0033] For each first object, a background audio matching the element is selected from a plurality of first background audios matching the first object.

[0034] In some embodiments, the method further comprises:

[0035] The selected multiple background audios are output according to the priorities of the multiple first objects.

[0036] In some embodiments, the method further comprises:

[0037] If the background audio is not determined based on the first object, identifying an element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0038] Based on the element, background audio of the multimedia resource is determined, the background audio matching the element.

[0039] In some embodiments, determining the background audio of the multimedia resource based on the element includes:

[0040] Determine a plurality of second background audios, wherein the plurality of second background audios include at least one of the following: a first number of background audios ranked first in number of times used on a first application, background audios used by a second number of multimedia resources ranked first in number of times played on the first application, a second number of background audios ranked first in number of times used on at least one second application, and a third number of audios ranked first in number of times played on at least one third application, wherein the first application is an application to which the multimedia resources belong, the second application is a resource playing application different from the first application, and the third application is an audio playing application;

[0041] The second background audio matching the element among the multiple second background audios is determined as the background audio of the multimedia resource.

[0042] In some embodiments, the method further comprises at least one of the following:

[0043] Based on the multimedia resource, extract the target segment in the background audio, and synthesize the target segment into the multimedia resource;

[0044] The video screen of the multimedia resource is synchronously synthesized with the background audio so that the screen switching rhythm of the multimedia resource matches the rhythm of the background audio.

[0045] According to another aspect of an embodiment of the present disclosure, a method for determining background audio is provided, the method comprising:

[0046] Display the music interface of multimedia resources;

[0047] In response to a music matching operation on the music matching interface, displaying at least one background audio recommended for the multimedia resource;

[0048] The background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

[0049] In some embodiments, there are multiple first objects, each of the at least one background audio corresponds to a first object, and the displaying of the at least one background audio recommended for the multimedia resource includes:

[0050] The background audios respectively corresponding to the multiple first objects are displayed in sequence according to the priorities of the multiple first objects.

[0051] In some embodiments, the method further comprises:

[0052] In response to a selection operation on any background audio, an adjustment panel of the background audio is displayed, wherein the adjustment panel is used to implement at least one of the following: extracting a target segment of the background audio, synchronizing and synthesizing a video screen of the multimedia resource with the background audio, adjusting the volume of the background audio, and adjusting the volume of the resource audio contained in the multimedia resource;

[0053] In response to an adjustment operation on the background audio on the adjustment panel, the background audio is adjusted.

[0054] According to another aspect of an embodiment of the present disclosure, a device for determining background audio is provided, the device comprising:

[0055] An acquisition unit, configured to acquire visual information of a multimedia resource;

[0056] an identification unit, configured to identify a first object appearing in a visual picture of the multimedia resource based on the visual information;

[0057] The determining unit is configured to determine the background audio of the multimedia resource based on the first object, where the background audio matches the first object.

[0058] In some embodiments, the determining unit is configured to perform at least one of the following:

[0059] identifying a clothing type of the first object, and determining audio matching the clothing type as background audio of the multimedia resource;

[0060] The expression type of the first object is identified, and audio matching the expression type is determined as background audio of the multimedia resource.

[0061] In some embodiments, the identification unit is configured to perform:

[0062] In the case where a plurality of second objects are identified based on the visual information, a second object indicated by the text information of the multimedia resource among the plurality of second objects is determined as the first object; or

[0063] In the case where a plurality of second objects are identified based on the visual information, an object among the plurality of second objects whose appearance duration in the multimedia resource satisfies a first condition is determined as the first object.

[0064] In some embodiments, the identification unit is configured to perform:

[0065] A first object appearing in the visual picture of the multimedia resource is identified based on the visual information through a multimodal recognition model, wherein the multimodal recognition model is a neural network model for identifying objects based on visual information.

[0066] In some embodiments, the determining unit is configured to execute:

[0067] Supplementary background audio is determined, and the background audio matching the first object and the supplementary background audio are determined as background audio recommended for the multimedia resource.

[0068] In some embodiments, the determining unit is configured to perform at least one of the following:

[0069] Determine the background audio with the first number of usage times ranked first on the first application as the supplementary background audio, wherein the first application is the application to which the multimedia resource belongs;

[0070] Determine the background audio used by the second number of multimedia resources ranked first in the number of times of playback on the first application as the supplementary background audio;

[0071] Determine the background audios ranked top by a third number of usage times on at least one second application as the supplementary background audios, where the second application is a resource playback application different from the first application;

[0072] The audio with the fourth largest number of playback times on at least one third application is determined as the supplementary background audio, and the third application is an audio playback application.

[0073] In some embodiments, the identification unit is further configured to perform:

[0074] Based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, identifying an element in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0075] The determining unit is configured to execute:

[0076] Based on the first object, determining a plurality of first background audios matching the first object;

[0077] The first background audio matching the element among the multiple first background audios is determined as the background audio of the multimedia resource.

[0078] In some embodiments, the identification unit is configured to perform at least one of the following:

[0079] In the case where multiple second scenes are identified, determining a scene whose appearance duration in the multimedia resource satisfies a second condition among the multiple second scenes as the first scene;

[0080] In the case where multiple second actions are identified, an action among the multiple second actions whose appearance duration in the multimedia resource satisfies a third condition is determined as the first action.

[0081] In some embodiments, the first object is multiple, and the determining unit is configured to execute:

[0082] For each first object, a background audio matching the element is selected from a plurality of first background audios matching the first object.

[0083] In some embodiments, the apparatus further comprises an output unit configured to execute:

[0084] The selected multiple background audios are output according to the priorities of the multiple first objects.

[0085] In some embodiments, the identification unit is further configured to perform:

[0086] If the background audio is not determined based on the first object, identifying an element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0087] The determining unit is further configured to determine background audio of the multimedia resource based on the element, wherein the background audio matches the element.

[0088] In some embodiments, the determining unit is configured to execute:

[0089] Determine a plurality of second background audios, wherein the plurality of second background audios include at least one of the following: a first number of background audios ranked first in number of times used on a first application, background audios used by a second number of multimedia resources ranked first in number of times played on the first application, a second number of background audios ranked first in number of times used on at least one second application, and a third number of audios ranked first in number of times played on at least one third application, wherein the first application is an application to which the multimedia resources belong, the second application is a resource playing application different from the first application, and the third application is an audio playing application;

[0090] The second background audio matching the element among the multiple second background audios is determined as the background audio of the multimedia resource.

[0091] In some embodiments, the apparatus further comprises a synthesis unit configured to perform at least one of the following:

[0092] Based on the multimedia resource, extract the target segment in the background audio, and synthesize the target segment into the multimedia resource;

[0093] The video screen of the multimedia resource is synchronously synthesized with the background audio so that the screen switching rhythm of the multimedia resource matches the rhythm of the background audio.

[0094] According to another aspect of an embodiment of the present disclosure, a device for determining background audio is provided, the device comprising:

[0095] A display unit, configured to execute and display a music interface of a multimedia resource;

[0096] The display unit is further configured to execute, in response to a music matching operation on the music matching interface, display at least one background audio recommended for the multimedia resource;

[0097] The background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

[0098] In some embodiments, there are multiple first objects, each of the at least one background audio corresponds to a first object, and the display unit is configured to execute:

[0099] The background audios respectively corresponding to the multiple first objects are displayed in sequence according to the priorities of the multiple first objects.

[0100] In some embodiments, the display unit is configured to perform:

[0101] In response to a selection operation on any background audio, an adjustment panel of the background audio is displayed, wherein the adjustment panel is used to implement at least one of the following: extracting a target segment of the background audio, synchronizing and synthesizing a video screen of the multimedia resource with the background audio, adjusting the volume of the background audio, and adjusting the volume of the resource audio contained in the multimedia resource;

[0102] The device further includes an adjustment unit configured to adjust the background audio in response to an adjustment operation on the background audio on the adjustment panel.

[0103] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device comprising:

[0104] processor;

[0105] a memory for storing instructions executable by the processor;

[0106] The processor is configured to execute the instructions to implement the above-mentioned method for determining the background audio.

[0107] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned method for determining background audio.

[0108] According to another aspect of an embodiment of the present disclosure, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the method for determining the background audio is implemented.

[0109] The disclosed embodiments provide a method for determining background audio. The method identifies a first object in a multimedia resource based on visual information of the multimedia resource, and then determines background audio that matches the first object. In this way, background audio is automatically determined for the multimedia resource without the need for manual selection by the user, thereby reducing user interaction operations. Furthermore, since the background audio matches the first object in the multimedia resource, the background audio matches the multimedia resource, which is beneficial to improving the playback effect of the multimedia resource.

[0110] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0112] Figure 1 It is a schematic diagram showing an implementation environment according to an exemplary embodiment.

[0113] Figure 2 The figure is a flowchart of a method for determining background audio according to an exemplary embodiment.

[0114] Figure 3 The figure is a flowchart of another method for determining background audio according to an exemplary embodiment.

[0115] Figure 4 The figure is a flowchart of yet another method for determining background audio according to an exemplary embodiment.

[0116] Figure 5 The present invention is a schematic diagram of a process of determining background audio according to an exemplary embodiment.

[0117] Figure 6 is a flowchart of yet another method for determining background audio according to an exemplary embodiment.

[0118] Figure 7 The invention is a block diagram of a device for determining background audio according to an exemplary embodiment.

[0119] Figure 8 It is a block diagram of another device for determining background audio according to an exemplary embodiment.

[0120] Fig. 9 It is a block diagram of a terminal according to an exemplary embodiment.

[0121] Fig.10 It is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0122] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0123] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0124] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the background audio, expression type, clothing type, etc. involved in this disclosure are all obtained with full authorization.

[0125] The method for determining background audio provided by the embodiment of the present disclosure can be executed by an electronic device, and the electronic device can be provided as at least one of a terminal and a server. Figure 1 This is a schematic diagram of an implementation environment provided by the embodiment of the present disclosure, see Figure 1 , the implementation environment includes: a terminal 101 and a server 102.

[0126] In the disclosed embodiment, a target application is installed on the terminal 101, and the target application is used to edit multimedia resources and play multimedia resources. Editing multimedia resources includes configuring background audio for multimedia resources. The server 102 is a background server of the target application. The server 102 is used to determine background audio for multimedia resources, synthesize the background audio into multimedia resources and send the multimedia resources to the terminal 101. Alternatively, the server sends the background audio to the terminal 101, and the terminal 101 synthesizes the background audio into the multimedia resources. Alternatively, the server 102 determines multiple background audios for the multimedia resources, and the server 102 sends the multiple background audios to the terminal 101, and the terminal 101 displays the multiple background audios to the user and synthesizes the background audio selected by the user into the multimedia resources.

[0127] In some embodiments, the terminal 101 sends the multimedia resource to the server 102, the server 102 extracts the visual information of the multimedia resource, identifies the first object appearing in the visual picture of the multimedia resource based on the visual information, and then determines the background audio of the multimedia resource based on the first object, so that the background audio matches the multimedia resource.

[0128] The terminal 101 may be at least one of a smart phone, a smart watch, a desktop computer, a laptop, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop computer. The terminal 101 has a communication function and can access a wired network or a wireless network. The terminal 101 may generally refer to one of a plurality of terminals, and those skilled in the art may know that the number of the above terminals may be more or less. The server 102 may be an independent physical server, or a server cluster or a distributed file system composed of a plurality of physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server 102 is directly or indirectly connected to the terminal 101 via a wired or wireless communication method, which is not limited in the embodiments of the present disclosure. Optionally, the number of the above servers 102 may be more or less, which is not limited in the embodiments of the present disclosure. Of course, the server 102 may also include other functional servers to provide more comprehensive and diversified services. Among them, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 or the terminal 101 can each independently undertake the computing work, which is not limited in the embodiments of the present disclosure.

[0129] Figure 2 is a flow chart of a method for determining background audio according to an exemplary embodiment. Figure 2 As shown, the method is executed by a server, and the method includes the following steps.

[0130] In step S201, the server obtains visual information of multimedia resources.

[0131] In the embodiments of the present disclosure, multimedia resources may be videos, pictures, graphic information or picture collections, etc. Visual information of multimedia resources is information presented visually by multimedia resources, including various information such as objects, animals, plants, people and texts appearing in the visual images of multimedia resources.

[0132] In step S202, the server identifies a first object appearing in a visual picture of a multimedia resource based on visual information.

[0133] In the disclosed embodiment, the visual picture of the multimedia resource includes multiple objects, and the first object may be an object of a specified type, such as a person, an animal, a plant, or other living thing. Optionally, the server identifies multiple objects appearing in the visual picture, and determines an object of the specified type as the first object.

[0134] In step S203, the server determines background audio of the multimedia resource based on the first object, and the background audio matches the first object.

[0135] In the disclosed embodiment, the background audio matches the first object, that is, the background audio is related to the first object. The background audio can be music, special effect audio, human voice reading audio, etc., and can also be audio including characteristic sounds such as wind, rain, and animal calls.

[0136] In the disclosed embodiment, after the server determines the background audio, it can directly synthesize the background audio into the multimedia resource and then send the multimedia resource to the terminal. Alternatively, after the server determines the background audio, it sends the background audio to the terminal, and the user decides whether to synthesize the recommended background audio into the multimedia resource.

[0137] The disclosed embodiments provide a method for determining background audio. The method identifies a first object in a multimedia resource based on visual information of the multimedia resource, and then determines background audio that matches the first object. In this way, background audio is automatically determined for the multimedia resource without the need for manual selection by the user, thereby reducing user interaction operations. Furthermore, since the background audio matches the first object in the multimedia resource, the background audio matches the multimedia resource, which is beneficial to improving the playback effect of the multimedia resource.

[0138] Above Figure 2The following is only the basic process of the present disclosure. The solution provided by the present disclosure is further described based on a specific implementation method. Figure 3 , Figure 3 The present invention is a flowchart of another method for determining background audio according to an exemplary embodiment. The method is executed by a server and includes the following steps.

[0139] In step S301, the server obtains visual information of multimedia resources.

[0140] In the embodiments of the present disclosure, the terminal may extract the visual information of the multimedia resource and then send the visual information to the server, or the terminal may send the multimedia resource to the server and the server may extract the visual information of the multimedia resource. The visual information is the information presented by the multimedia resource visually. For example, if the multimedia resource is a video, the visual information is the information in the multiple frames of the video.

[0141] In step S302, the server identifies a first object appearing in a visual picture of a multimedia resource based on visual information.

[0142] In the embodiment of the present disclosure, the first object is a person as an example for description. The first object may be one or more.

[0143] In some embodiments, multiple objects may appear in the video screen of the multimedia resource, and there may be interference objects of low importance among the multiple objects, which will make it difficult to determine the background audio, so it is necessary to identify the main object among the multiple objects. Optionally, step S302 includes the following implementations.

[0144] 1. In the case where multiple second objects are identified based on the visual information, the server determines the second object indicated by the text information of the multimedia resource among the multiple second objects as the first object.

[0145] In the embodiment of the present disclosure, the first object is taken as an example to illustrate, and the second object can be any character in the multimedia resources.

[0146] The text information of multimedia resources includes the title, copy and text in the visual picture of the multimedia resources. The text information indicates the object, including the name and attributes of the indicated object. Attributes include clothing, expression, action, etc.

[0147] In this implementation, a second object appears in the multimedia resource, and the text of the multimedia resource indicates the second object. In this way, the multimedia resource indicates the second object through multiple information, indicating that the second object is very likely to be the main object in the multimedia resource. Therefore, it is determined as the first object, thereby improving accuracy.

[0148] 2. When multiple second objects are identified based on the visual information, the server determines an object among the multiple second objects whose appearance duration in the multimedia resource meets the first condition as the first object.

[0149] The first condition may be that the appearance duration is the longest among the plurality of second objects or that the appearance duration reaches the first duration. Optionally, the first duration is positively correlated with the overall duration of the multimedia resource, that is, the longer the overall duration of the multimedia resource, the longer the first duration.

[0150] In this embodiment, if the duration of appearance of any second object in a multimedia resource reaches a certain condition, it indicates that the second object has a high frequency of appearance in the multimedia resource and is more likely to be the main object in the multimedia resource. Therefore, such an object is determined as the first object, thereby improving accuracy.

[0151] In some embodiments, step S302 includes the following implementation method: the server recognizes a first object appearing in a visual picture of a multimedia resource based on visual information through a multimodal recognition model, where the multimodal recognition model is a neural network model for recognizing objects based on visual information.

[0152] The server inputs the visual information into the multimodal recognition model, and the multimodal recognition model outputs the first object. Alternatively, the server inputs the multimedia resource into the multimodal recognition model, the multimodal recognition model extracts the visual information of the multimedia resource, and the multimodal recognition model outputs the first object based on the visual information.

[0153] In some embodiments, if the first object is an object in an object library, the multimodal recognition model is a neural network model trained based on a sample pair of an image of an object in the object library and the object composition, so that the multimodal recognition model can quickly recognize the object based on the image of any object in the object library. In this embodiment, the first object appearing in the visual picture of the multimedia resource is recognized by the neural network model, which improves efficiency and convenience.

[0154] In step S303, the server determines background audio matching the first object.

[0155] In the embodiment of the present disclosure, the background audio matching the first object refers to the background audio matching some personal features of the first object. Optionally, the server determines the background audio matching the first object, including at least one of the following implementations: the server identifies the clothing type of the first object, and determines the audio matching the clothing type as the background audio matching the first object; the server identifies the expression type of the first object, and determines the audio matching the expression type as the background audio matching the first object.

[0156] The expression type may include happy, sad, etc. If the expression type is happy, the background audio that matches the expression type may be cheerful music.

[0157] In this implementation, since clothing type, expression type, etc. are personal features that are highly representative of the first object, these are used as references, and the determined background audio can be matched with the first object, thereby improving the accuracy of the background audio.

[0158] In the embodiments of the present disclosure, the first object is described as a person. In other embodiments, the first object may also be an animal or a plant, and the background audio matching the first object may be an animal song or a plant song.

[0159] In some embodiments, after determining the background audio matching the first object, the server directly determines the background audio matching the first object as the background audio of the multimedia resource.

[0160] Among them, if there are multiple background audios matching the first object, the background audio with the highest priority among the multiple background audios can be determined as the background audio of the multimedia resource. The priorities of the multiple background audios can be set as needed, such as the priorities of the multiple background audios are positively correlated with the number of times each is used, which is not specifically limited here.

[0161] In some other embodiments, in order to expand the background audio recommended for the multimedia resource, the method further includes the following steps S304-S305.

[0162] In step S304, the server determines supplementary background audio.

[0163] In some embodiments, the supplementary background audio is currently popular audio, including at least one of the popular background audio on the resource playback application and the popular audio on the audio playback application, then step S304 includes at least one of the following implementation methods: the server determines the background audio with a first number of usage counts on the first application as the supplementary background audio, and the first application is the application to which the multimedia resource belongs. The server determines the background audio used by the multimedia resources with a second number of playback counts on the first application as the supplementary background audio. The server determines the background audio with a third number of usage counts on at least one second application as the supplementary background audio, and the second application is a resource playback application different from the first application. The server determines the audio with a fourth number of playback counts on at least one third application as the supplementary background audio, and the third application is an audio playback application.

[0164] In the disclosed embodiment, the first application and the second application are both resource playing applications, which are used to play multimedia resources, such as videos, audios, pictures, graphic information or picture collections, etc. The third application is an audio playing application, and the main function of the audio playing application is to play audio, such as the third application is a music application.

[0165] The first number, the second number, the third number and the fourth number are all preset numbers and can be the same or different. Since the popular background audio or audio on each application is different at different stages, the background audio or audio that ranks first in a preset time period on each application can be determined as the supplementary background audio. The preset time period can be, for example, the last month or the last week.

[0166] In the disclosed embodiment, since the multimedia resources are published on the first application, the background audio with many uses is supplemented as the background audio, which conforms to the current usage trend of the background audio on the first application, and can increase the popularity and attractiveness of the multimedia resources. The background audio in the multimedia resources with many plays is supplemented as the background audio. Since the number of plays is large, it means that the audience also prefers the background audio in such multimedia resources. Then, such background audio is supplemented as the background audio, so that after the background audio is synthesized to obtain the multimedia resources, it conforms to the preferences of the audience, which is conducive to improving the playback rate of the multimedia resources. Since the second application is also an application for playing multimedia resources, the background audio with many uses on the second application is determined, that is, the popular background audio on other resource playing applications is determined, and then these background audios are supplemented as the background audio, which can increase the popularity and attractiveness of the multimedia resources. Since the third application is mainly used to play audio, and the audio with many plays represents the current popular trend of audio, then these audios are supplemented as the background audio, which can increase the popularity and attractiveness of the multimedia resources.

[0167] It should be noted that the serial numbers of steps S303-S304 are only for convenience of explanation, and step S303 can be executed before step S304 or after step S304, and step S303 and step S304 can also be executed simultaneously, which is not specifically limited here.

[0168] In step S305, the server determines the background audio and the supplementary background audio matching the first object as the background audio of the multimedia resource.

[0169] In some embodiments, the background audio determined in step S305 is the background audio that is ultimately synthesized into the multimedia resource. In the case where the determined background audio is multiple, the background audio with the highest priority can be synthesized into the multimedia resource based on the priorities of the multiple background audios, or the background audio with the highest priority can be recommended to the user and manually synthesized into the multimedia resource by the user. The priorities of the multiple background audios can be set as needed, such as the priorities of the multiple background audios are positively correlated with the number of times they are used, which is not specifically limited here.

[0170] In other embodiments, the background audio of the multimedia resource determined in step S305 is the background audio recommended to the multimedia resource, that is, it is only recommended, and it is up to the user to decide whether to use the recommended background audio and which background audio to synthesize into the multimedia resource.

[0171] In the above embodiment, the supplementary background audio and the background audio matching the first object are determined as the background audio of the multimedia resource as an example. In other embodiments, the background audio matching the first object can also be directly determined as the background audio of the multimedia resource, that is, the process includes at least one of the following implementation methods: the server identifies the clothing type of the first object and determines the audio matching the clothing type as the background audio of the multimedia resource; the server identifies the expression type of the first object and determines the audio matching the expression type as the background audio of the multimedia resource.

[0172] In an embodiment of the present disclosure, after determining the background audio based on the first object in the multimedia resource, supplementary background audio is also determined based on the popular background audio or audio on the resource playback application and the audio playback application, and then the supplementary background audio and the background audio matching the first object are recommended to the user together. This increases the number of background audios that the user can select, facilitates the user to select satisfactory background audio, and improves the user experience.

[0173] Above Figure 3 The following takes determining the background audio based on other elements in the multimedia resource as an example. Figure 4 , Figure 4 The present invention is a flowchart of another method for determining background audio according to an exemplary embodiment. The method is executed by a server and includes the following steps.

[0174] In step S401, the server obtains visual information of multimedia resources.

[0175] In the embodiment of the present disclosure, step S401 is the same as step S301 and will not be described in detail here.

[0176] In step S402, the server identifies a first object appearing in a visual picture of a multimedia resource based on visual information.

[0177] In the embodiment of the present disclosure, step S402 is the same as step S302 and will not be described in detail here.

[0178] In step S403, the server identifies elements in the multimedia resource based on the visual information of the multimedia resource and at least one of the resource audios contained in the multimedia resource, where the elements include at least one of an emotion type, a first scene, and a first action, and the emotion type is used to represent the emotion expressed by the multimedia resource.

[0179] Wherein, if the element includes an emotion type, the server identifies the emotion type based on at least one of the visual information and the resource audio. If the element includes a first scene, the server identifies the first scene based on at least one of the visual information and the resource audio. If the element includes a first action, the server identifies the first action based on at least one of the visual information and the resource audio.

[0180] In some embodiments, there may be multiple scenes and multiple actions in a multimedia resource, and multiple scenes and multiple actions may cause difficulties in determining the background audio, so it is necessary to identify the main scenes in the multiple scenes and the main actions in the multiple actions. Optionally, the above step S403 includes at least one of the following implementations: when multiple second scenes are identified, the server determines the scene whose appearance duration in the multimedia resource meets the second condition among the multiple second scenes as the first scene; when multiple second actions are identified, the server determines the action whose appearance duration in the multimedia resource meets the third condition among the multiple second actions as the first action.

[0181] In the disclosed embodiment, the second scene is any scene in the multimedia resource, and the second action is any action in the multimedia resource. The second condition may be that the appearance duration is the longest among multiple second scenes or the appearance duration reaches the second duration. The third condition may be that the appearance duration is the longest among multiple second actions or the appearance duration reaches the third duration. In the disclosed embodiment, the first scene may be one or more, and the first action may be one or more, thereby facilitating increasing the probability of selecting background audio.

[0182] It should be noted that the first action refers to a type of action. For example, if the first action is a running action, then any running action such as long-distance running and jogging belongs to the first action.

[0183] The first action may be an action of any object in the multimedia resource. Alternatively, the first action may be an action of the first object. Accordingly, when multiple actions are identified, the action of the first object among the multiple actions may be determined as the first action; or after the first object is identified, only the action of the first object may be identified and determined as the first action.

[0184] In this embodiment, if the appearance duration of any scene in the multimedia resource reaches a certain condition, it means that the scene has a high frequency of appearance in the multimedia resource and is likely to be the core scene in the multimedia resource, so such a scene is determined as the first scene, which improves accuracy. Similarly, if the appearance duration of any action in the multimedia resource reaches a certain condition, it means that the action has a high frequency of appearance in the multimedia resource and is likely to be the core action in the multimedia resource, so such an action is determined as the first action, which improves accuracy.

[0185] In some embodiments, the server further determines the first scene and the first action based on the text information of the multimedia resource. Wherein, when multiple second scenes are identified, the server determines the scene indicated by the text information in the multiple second scenes as the first scene. When multiple second actions are identified, the server determines the action indicated by the text information in the multiple second actions as the first action.

[0186] In some embodiments, the server further identifies elements in the multimedia resource through a multimodal recognition model. That is, the above step S403 includes the following implementation: the server identifies elements in the multimedia resource through a multimodal recognition model based on at least one of visual information and resource audio. That is, the multimodal recognition model is a neural network model that also identifies elements in the multimedia resource based on at least one of visual information and resource audio.

[0187] In some embodiments, the server inputs at least one of the visual information and the resource audio into the multimodal recognition model, and the multimodal recognition model outputs the element. Alternatively, the server inputs the multimedia resource into the multimodal recognition model, extracts at least one of the visual information and the resource audio of the multimedia resource through the multimodal recognition model, and then the multimodal recognition model outputs the element based on at least one of the visual information and the resource audio.

[0188] Accordingly, if the element includes an emotion type, the multimodal recognition model is a neural network model that is also trained based on sample pairs consisting of at least one of visual information and resource audio and emotion types. If the element includes a scene, the multimodal recognition model is a neural network model that is also trained based on sample pairs consisting of at least one of visual information and resource audio and scenes. If the element includes an action, the multimodal recognition model is a neural network model that is also trained based on sample pairs consisting of at least one of visual information and resource audio and actions.

[0189] Optionally, the multimodal recognition model includes multiple recognition modules, which are respectively used to recognize objects, emotion types, scenes, actions, etc. In this way, efficiency can be improved by recognizing various information in multimedia resources based on the multimodal recognition model.

[0190] In step S404, the server determines, based on the first object, a plurality of first background audios matching the first object.

[0191] In the embodiment of the present disclosure, the process of determining multiple first background audios matching the first object is the same as the process of determining the background audio matching the first object in step S303, which is not repeated here.

[0192] In step S405, the server determines the first background audio matching the element among the multiple first background audios as the background audio of the multimedia resource.

[0193] If the element includes an emotion type, the first background audio matching the emotion type among the multiple first background audios is determined as the background audio of the multimedia resource. If the element includes a scene, the first background audio matching the scene among the multiple first background audios is determined as the background audio of the multimedia resource. If the element includes an action, the first background audio matching the action among the multiple first background audios is determined as the background audio of the multimedia resource.

[0194] If the elements include two of the emotion type, scene, and action, the first background audio that matches both of the multiple first background audios is determined as the background audio of the multimedia resource to further improve the matching degree between the background audio and the multimedia resource. Alternatively, the first background audio that matches at least one of the two is determined as the background audio of the multimedia resource to improve the probability of determining the background audio of the multimedia resource.

[0195] If the elements include the three items of emotion type, scene and action, the first background audio among the multiple first background audios that matches all the three items is determined as the background audio of the multimedia resource, so as to further improve the matching degree between the background audio and the multimedia resource. Alternatively, the first background audio among the multiple first background audios that matches at least one of the three items is determined as the background audio of the multimedia resource, so as to improve the probability of determining the background audio of the multimedia resource.

[0196] In some embodiments, the background audio of the multimedia resource can also be determined from the supplementary background audio, that is, the background audio matching the element in the supplementary background audio and multiple first background audios is determined as the background audio of the multimedia resource to increase the probability of determining the background audio of the multimedia resource.

[0197] In some embodiments, if there are multiple first objects, the multiple first background audios include multiple first background audios corresponding to the multiple first objects respectively. Optionally, step S405 includes the following steps: for each first object, the server selects the background audio that matches the element from the multiple first background audios that match the first object. In this implementation, when multiple first objects are identified, the background audio that matches the element is selected from the multiple first background audios that match each object, which is beneficial to increase the probability of determining the background audio of the multimedia resource; and more background audios can be recommended to the user, which increases the number of background audios that the user can select, which is beneficial for the user to select a satisfactory background audio and improve the user experience.

[0198] In some embodiments, recommending the selected multiple background audios to the user may include the following method: the server outputs the selected multiple background audios according to the priorities of the multiple first objects, that is, when the selected multiple background audios are displayed to the user, the higher the priority of the first object, the higher the ranking of the corresponding background audio. For example, the background audio corresponding to the first object with the highest priority is ranked first among the selected multiple background audios.

[0199] The priorities of the multiple first objects can be set based on needs. For example, the priority of each first object is positively correlated with the appearance duration of the first object in the multimedia resource, that is, the longer the appearance duration, the higher the priority.

[0200] In some embodiments, there may be multiple first scenes and multiple first actions. Taking multiple first scenes as an example, for each first scene, the server selects background audio that matches the elements including the first scene from multiple first background audios that match the first object, and then selects background audios corresponding to the multiple first scenes. Further, according to the priorities of the multiple first scenes, the selected multiple background audios are output. The situation of multiple first actions is the same and will not be repeated here. The priority of the first scene and the first action can be set as needed, such as the priority is positively correlated with the duration of its appearance in the multimedia resource.

[0201] In some embodiments, the background audio may not be determined based on the first object, and the background audio can be determined directly based on the elements in the multimedia resource. Accordingly, the following implementation method is also included: if the background audio is not determined based on the first object, the server identifies the elements in the multimedia resource based on the visual information of the multimedia resource and at least one of the resource audios contained in the multimedia resource, where the elements include at least one of an emotion type, a first scene, and a first action, and the emotion type is used to represent the emotion expressed by the multimedia resource; based on the elements, the background audio of the multimedia resource is determined, and the background audio matches the elements.

[0202] In some embodiments, the first object may not be identified in the multimedia resource, and the background audio may also be determined directly based on the elements in the multimedia resource.

[0203] In some embodiments, the process of the server determining the background audio of a multimedia resource based on an element includes the following implementation method: the server determines multiple second background audios, and the multiple second background audios include at least one of the following: a first number of background audios ranked first in terms of usage frequency on a first application, background audios used by a second number of multimedia resources ranked first in terms of play frequency on the first application, a second number of background audios ranked first in terms of usage frequency on at least one second application, and a third number of audios ranked first in terms of play frequency on at least one third application, the first application being the application to which the multimedia resource belongs, the second application being a resource playback application different from the first application, and the third application being an audio playback application; the second background audios among the multiple second background audios that match the element are determined as the background audio of the multimedia resource.

[0204] Among them, the process of determining the second background audio matching the element from multiple second background audios is the same as the process of determining the first background audio matching the element from multiple first background audios, and will not be repeated here.

[0205] In this embodiment, multiple second background audios are first determined. Since the second background audio is the currently more popular audio on the resource playback application or the audio playback application, the background audio is then determined on this basis, so that the determined background audio not only matches the elements in the multimedia resources, but also makes the background audio conform to the current audio popularity trend, and then such background audio is synthesized into the multimedia resources, making the multimedia resources more attractive and helping to improve the playback rate of the multimedia resources.

[0206] In an embodiment of the present disclosure, after the server determines the background audio, it can directly synthesize the background audio into the multimedia resource, and may also include at least one of the following implementation methods: the server extracts the target segment in the background audio based on the multimedia resource, and synthesizes the target segment into the multimedia resource; the server synchronously synthesizes the video screen of the multimedia resource with the background audio, so that the screen switching rhythm of the multimedia resource matches the rhythm of the background audio.

[0207] Among them, the target segment may include the following situations: the target segment is a segment from the beginning of the background audio to the target playback progress, and the duration between the beginning and the target playback progress is the duration of the multimedia resource, that is, the duration of the target segment is the same as the duration of the multimedia resource. Alternatively, the target segment is a climax segment in the background audio. Alternatively, the target segment is the segment that is used most times in the first application. Alternatively, only part of the background audio is related to the first object, then the target segment is the audio related to the first object in the target audio. For example, the background audio is the audio of the first object singing with other singers, then the target segment can be a segment of the background audio sung by the first object.

[0208] Matching the rhythm of switching the screens of multimedia resources with the rhythm of background audio means switching the screens of multimedia resources at the important beats of background audio; or, when the multimedia resources include multiple scenes, switching scenes at the important beats of background audio. Or, when the multimedia resources include multiple actions, switching actions at the important beats of background audio; or, when the multimedia resources include multiple objects, switching objects at the important beats of background audio. That is, adjustments are made in combination with the context of multimedia resources to match the rhythm of multimedia resources and background audio.

[0209] In the disclosed embodiment, the target segment in the background audio is extracted based on the multimedia resource, so that the target segment and the multimedia resource have a higher matching degree, and then the target segment is synthesized into the multimedia resource, which can improve the playback effect of the multimedia resource and help improve the playback rate of the multimedia resource. Synchronizing the multimedia resource with the background audio so that the rhythm of the two matches can improve the playback effect of the multimedia resource and help improve the playback rate of the multimedia resource.

[0210] It should be noted that if the server determines multiple background audios for the multimedia resource, the background audio with the highest priority can be synthesized into the multimedia resource based on the priorities of the multiple background audios. The priorities of the multiple background audios can be set as needed, such as the priorities of the multiple background audios are positively correlated with the number of times each is used, which is not specifically limited here.

[0211] In other embodiments, the server can send the determined background audio to the terminal, which is displayed to the user by the terminal, that is, the user decides whether to synthesize the background audio into the multimedia resource. The background audio determined by the server can be one or more, and correspondingly, the background audio displayed by the terminal is one or more. In the case of multiple background audio, the user can select any background audio to be synthesized into the multimedia resource.

[0212] For example, see Figure 5 , Figure 5 1 is a flow chart of determining background audio according to an exemplary embodiment. The server first identifies the basic content of the multimedia resource, and takes the first object as a character as an example for explanation. The identified basic content includes the characters, scenes, actions, etc. in the multimedia resource. Then the core content of the multimedia resource is identified, that is, the core object (first object) among multiple objects is identified, the core scene (first scene) among multiple scenes is identified, and the core action (first action) among multiple actions is identified. Then the clothing type and expression type of the first object are identified. Wherein, when the clothing type and expression type are identified, the background audio is retrieved based on the clothing type and expression type to obtain multiple first background audios, and then the multiple first background audios are further screened based on the elements. Wherein, if multiple first objects, multiple first scenes or multiple first actions are identified, the screened background audios are sorted according to priority. When the clothing type and expression type are not identified, the emotion type of the multimedia resource is identified, and the background audio is determined based on the emotion type, the first scene, and the first action. After the background audio is determined in any case, the background audio is adjusted based on the multimedia resource and merged into the multimedia resource, and the process ends here.

[0213] In the embodiment of the present disclosure, after determining multiple first background audios based on the first object, background audios matching the elements are selected from the multiple first background audios based on the elements in the multimedia resources as the background audio of the multimedia resources. In this way, the determined background audio matches the various information in the multimedia resources, further improving the matching degree between the background audio and the multimedia resources, which is beneficial to improving the playback effect of the multimedia resources.

[0214] Above Figure 2-4 The process of determining background audio is described with the server as the execution subject. The following describes the process of determining background audio with the terminal as the execution subject. Figure 6 , Figure 6 The present invention is a flowchart of another method for determining background audio according to an exemplary embodiment. The method is executed by a terminal and includes the following steps.

[0215] In step S601, the terminal displays a music matching interface of a multimedia resource.

[0216] In the embodiment of the present disclosure, the music matching interface is used to add background audio to the multimedia resource. Optionally, the music matching interface is an interface within the first application. The first application is the application to which the multimedia resource belongs, and the first application is the resource playing application.

[0217] Among them, the music interface can be used to search for background audio, and can also obtain background audio recommended for multimedia resources with one click.

[0218] In step S602, in response to a music operation on the music interface, the terminal displays at least one background audio recommended for the multimedia resource; wherein the background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

[0219] The music matching operation is used to obtain background audio recommended for the multimedia resource. Accordingly, in response to the music matching operation, the terminal sends a recommendation request to the server, and the server determines the background audio of the multimedia resource based on the recommendation request, sends at least one background audio recommended for the multimedia resource to the terminal, and the terminal displays the at least one background audio. The process of the server determining the background audio of the multimedia resource can be found in Figure 2-Figure 4 The embodiments of the present invention will not be described in detail here.

[0220] In some embodiments, there are multiple first objects, and the terminal displays at least one background audio recommended as a multimedia resource, including the following implementation method: the terminal displays the background audios corresponding to the multiple first objects in sequence according to the priorities of the multiple first objects.

[0221] In this embodiment, the higher the priority of any first object, the higher the importance of the first object in the multimedia resources, that is, the background audio corresponding to the first object has a higher degree of match with the multimedia resources. In this way, the background audio with high priority is displayed in front, which increases the probability of the user selecting such background audio, and thus makes the match between the background audio synthesized into the multimedia resources and the multimedia resources as high as possible, which is conducive to improving the playback rate of the multimedia resources.

[0222] In some embodiments, the user is also able to adjust the recommended background audio, and the method also includes: in response to a selection operation on any background audio, the terminal displays an adjustment panel for the background audio, and the adjustment panel is used to implement at least one of the following: extracting a target segment of the background audio, synchronizing and synthesizing the video screen of the multimedia resource with the background audio, adjusting the volume of the background audio, and adjusting the volume of the resource audio contained in the multimedia resource; in response to an adjustment operation on the background audio on the adjustment panel, adjusting the background audio.

[0223] Optionally, in response to a synthesis operation on the adjusted background audio, the adjusted background audio is synthesized into the multimedia resource.

[0224] In this embodiment, after the terminal displays the background audio recommended for the multimedia resource, the user can also fine-tune and edit the background audio, thereby achieving a more personalized and professional soundtrack effect, thereby improving the flexibility of the soundtrack.

[0225] In the disclosed embodiments, by automatically identifying characters, emotions, scenes, actions, etc. in multimedia resources, appropriate background audio is intelligently matched, reducing the need for manual intervention, thereby improving the efficiency of determining background audio and significantly reducing costs. In addition, by matching more accurate and appropriate background audio to multimedia resources, the emotional expression and viewing experience of multimedia resources can be enhanced, user satisfaction can be improved, and user experience can be enhanced. In addition, through the multimodal recognition model, the visual information, sound information, and text information of multimedia resources can be comprehensively analyzed, providing rich contextual information for multimedia resources to determine background audio, making the background audio more vivid and expressive.

[0226] Furthermore, the disclosed solution can adapt to various types of multimedia content, such as music videos, news reports, social media short videos, etc., and can provide matching background audio for them. In this way, the function of automatically matching background audio can promote users to participate in the creation and sharing of multimedia resources, and promote the production of more multimedia resources.

[0227] An embodiment of the present disclosure provides a method for determining background audio. The method identifies a first object in a multimedia resource based on visual information of the multimedia resource, and then determines background audio that matches the first object. In this way, background audio is automatically recommended for the multimedia resource without the need for manual selection by the user, thereby reducing user interaction operations. Furthermore, since the recommended background audio matches the first object in the multimedia resource, the recommended background audio matches the multimedia resource, which is beneficial to improving the playback effect of the multimedia resource.

[0228] Figure 7 is a block diagram of a device for determining background audio according to an exemplary embodiment. Figure 7 , the device comprises:

[0229] The acquisition unit 701 is configured to acquire visual information of multimedia resources;

[0230] The recognition unit 702 is configured to recognize a first object appearing in a visual picture of a multimedia resource based on the visual information;

[0231] The determining unit 703 is configured to determine the background audio of the multimedia resource based on the first object, where the background audio matches the first object.

[0232] In some embodiments, the determining unit 703 is configured to perform at least one of the following:

[0233] identifying a clothing type of the first object, and determining audio matching the clothing type as background audio of the multimedia resource;

[0234] An expression type of the first object is identified, and audio matching the expression type is determined as background audio of the multimedia resource.

[0235] In some embodiments, the identification unit 702 is configured to perform:

[0236] In the case where multiple second objects are identified based on the visual information, the second object indicated by the text information of the multimedia resource among the multiple second objects is determined as the first object; or,

[0237] In the case where a plurality of second objects are identified based on the visual information, an object whose appearance duration in the multimedia resource satisfies a first condition among the plurality of second objects is determined as a first object.

[0238] In some embodiments, the identification unit 702 is configured to perform:

[0239] A first object appearing in a visual picture of a multimedia resource is identified based on visual information through a multimodal recognition model, where the multimodal recognition model is a neural network model used to identify objects based on visual information.

[0240] In some embodiments, the determining unit 703 is configured to perform:

[0241] Supplementary background audio is determined, and the background audio matching the first object and the supplementary background audio are determined as background audio recommended for the multimedia resource.

[0242] In some embodiments, the determining unit 703 is configured to perform at least one of the following:

[0243] Determine the background audio with the first number of usage times ranked first in the first application as the supplementary background audio, where the first application is the application to which the multimedia resource belongs;

[0244] Determine the background audio used by the multimedia resources ranked second in the number of times played on the first application as the supplementary background audio;

[0245] Determine the background audio with the third largest number of usage times in at least one second application as the supplementary background audio, where the second application is a resource playback application different from the first application;

[0246] The audio with the fourth largest number of playback times on at least one third application is determined as supplementary background audio, and the third application is an audio playback application.

[0247] In some embodiments, the identification unit 702 is further configured to perform:

[0248] Based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, identifying an element in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0249] The determining unit 703 is configured to execute:

[0250] Based on the first object, determining a plurality of first background audios matching the first object;

[0251] A first background audio that matches the element among the multiple first background audios is determined as the background audio of the multimedia resource.

[0252] In some embodiments, the identification unit 702 is configured to perform at least one of the following:

[0253] In the case where multiple second scenes are identified, determining a scene whose appearance duration in the multimedia resource satisfies the second condition among the multiple second scenes as the first scene;

[0254] In the case where multiple second actions are identified, an action among the multiple second actions whose appearance duration in the multimedia resource meets the third condition is determined as the first action.

[0255] In some embodiments, the first object is multiple, and the determining unit 703 is configured to execute:

[0256] For each first object, a background audio matching the element is selected from a plurality of first background audios matching the first object.

[0257] In some embodiments, the apparatus further comprises an output unit configured to execute:

[0258] The selected multiple background audios are output according to the priorities of the multiple first objects.

[0259] In some embodiments, the identification unit 702 is further configured to perform:

[0260] If the background audio is not determined based on the first object, identifying an element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource;

[0261] The determining unit 703 is further configured to determine the background audio of the multimedia resource based on the element, and the background audio matches the element.

[0262] In some embodiments, the determining unit 703 is configured to perform:

[0263] Determine a plurality of second background audios, the plurality of second background audios comprising at least one of the following: a first number of background audios ranked first in frequency of use on the first application, a second number of background audios used by multimedia resources ranked first in frequency of play on the first application, a second number of background audios ranked first in frequency of use on at least one second application, and a third number of audios ranked first in frequency of play on at least one third application, the first application being an application to which the multimedia resources belong, the second application being a resource playing application different from the first application, and the third application being an audio playing application;

[0264] The second background audio that matches the element among the multiple second background audios is determined as the background audio of the multimedia resource.

[0265] In some embodiments, the apparatus further comprises a synthesis unit configured to perform at least one of the following:

[0266] Based on multimedia resources, extract the target segment from the background audio and synthesize the target segment into the multimedia resources;

[0267] Synchronize the video images of multimedia resources with the background audio so that the switching rhythm of the multimedia resources matches the rhythm of the background audio.

[0268] An embodiment of the present disclosure provides a background audio determination device, which identifies a first object in a multimedia resource based on visual information of the multimedia resource, and then determines background audio that matches the first object. In this way, background audio is automatically determined for the multimedia resource without the need for manual selection by the user, thereby reducing user interaction operations. Furthermore, since the background audio matches the first object in the multimedia resource, the background audio matches the multimedia resource, which is beneficial to improving the playback effect of the multimedia resource.

[0269] Regarding the device in the above embodiment, the specific manner in which each unit performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.

[0270] Figure 8 is a block diagram of a device for determining background audio according to an exemplary embodiment. Figure 8 , the device comprises:

[0271] The display unit 801 is configured to execute and display the music interface of the multimedia resource;

[0272] The display unit 801 is further configured to display at least one background audio recommended for the multimedia resource in response to a music matching operation on the music matching interface;

[0273] The background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

[0274] In some embodiments, there are multiple first objects, and at least one background audio corresponds to each first object. The display unit 801 is configured to execute:

[0275] The background audios respectively corresponding to the multiple first objects are displayed in sequence according to the priorities of the multiple first objects.

[0276] In some embodiments, the display unit 801 is configured to perform:

[0277] In response to a selection operation on any background audio, an adjustment panel for the background audio is displayed, and the adjustment panel is used to implement at least one of the following: extracting a target segment of the background audio, synchronizing and synthesizing a video screen of a multimedia resource with the background audio, adjusting the volume of the background audio, and adjusting the volume of the resource audio contained in the multimedia resource;

[0278] The device also includes an adjustment unit configured to adjust the background audio in response to an adjustment operation on the background audio on the adjustment panel.

[0279] An embodiment of the present disclosure provides a method for determining background audio. The method identifies a first object in a multimedia resource based on visual information of the multimedia resource, and then determines background audio that matches the first object. In this way, background audio is automatically recommended for the multimedia resource without the need for manual selection by the user, thereby reducing user interaction operations. Furthermore, since the recommended background audio matches the first object in the multimedia resource, the recommended background audio matches the multimedia resource, which is beneficial to improving the playback effect of the multimedia resource.

[0280] Regarding the device in the above embodiment, the specific manner in which each unit performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.

[0281] In some embodiments, the electronic device is provided as a terminal. Fig. 9 The structure block diagram of a terminal 900 provided by an exemplary embodiment of the present disclosure is shown. The terminal 900 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 900 may also be called a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.

[0282] Typically, the terminal 900 includes a processor 901 and a memory 902 .

[0283] The processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0284] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one program code, which is used to be executed by the processor 901 to implement the method for determining the background audio provided by the method embodiment of the present disclosure.

[0285] In some embodiments, the terminal 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902 and the peripheral device interface 903 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 903 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907 and a power supply 908.

[0286] The peripheral device interface 903 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.

[0287] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may also include circuits related to NFC (Near Field Communication), which is not limited in the present disclosure.

[0288] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on the surface or above the surface of the display screen 905. The touch signal can be input to the processor 901 as a control signal for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 905 can be one, and the front panel of the terminal 900 is set; in other embodiments, the display screen 905 can be at least two, which are respectively set on different surfaces of the terminal 900 or are folded; in some other embodiments, the display screen 905 can be a flexible display screen, which is set on the curved surface or folded surface of the terminal 900. Even, the display screen 905 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 905 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode) and the like.

[0289] The camera assembly 906 is used to capture images or videos. Optionally, the camera assembly 906 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0290] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 901 for processing, or input them into the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.

[0291] The power supply 908 is used to power various components in the terminal 900. The power supply 908 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 908 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0292] Those skilled in the art will understand that Fig. 9 The structure shown in the figure does not constitute a limitation on the terminal 900, and the terminal 900 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.

[0293] In some embodiments, the electronic device is provided as a server. Fig.10 This is a block diagram of a server 1000 according to an exemplary embodiment. The server 1000 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1001 and one or more memories 1002, wherein the memory 1002 stores at least one program code, and the at least one program code is loaded and executed by the processor 1001 to implement the background audio determination method provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 1000 may also include other components for implementing device functions, which will not be described in detail here.

[0294] In an exemplary embodiment, a computer-readable storage medium is also provided, and when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned method for determining background audio. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0295] In an exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the method for determining the background audio is implemented.

[0296] In some embodiments, the computer program product involved in the embodiments of the present disclosure may be deployed and executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network. Multiple electronic devices distributed at multiple locations and interconnected by a communication network may constitute a blockchain system.

[0297] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and embodiments are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, which will not be described one by one here.

[0298] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for determining background audio, characterized in that: The method comprises: Acquire visual information from multimedia resources; Based on the visual information, identifying a first object appearing in a visual picture of the multimedia resource; Based on the first object, background audio of the multimedia resource is determined, and the background audio matches the first object.

2. The method for determining background audio according to claim 1, characterized in that: The determining, based on the first object, the background audio of the multimedia resource comprises at least one of the following: identifying a clothing type of the first object, and determining audio matching the clothing type as background audio of the multimedia resource; The expression type of the first object is identified, and audio matching the expression type is determined as background audio of the multimedia resource.

3. The method for determining background audio according to claim 1, characterized in that: The step of identifying a first object appearing in a visual picture of the multimedia resource based on the visual information includes: In the case where a plurality of second objects are identified based on the visual information, a second object indicated by the text information of the multimedia resource among the plurality of second objects is determined as the first object; or In the case where a plurality of second objects are identified based on the visual information, an object among the plurality of second objects whose appearance duration in the multimedia resource satisfies a first condition is determined as the first object.

4. The method for determining background audio according to claim 1, characterized in that: The step of identifying a first object appearing in a visual picture of the multimedia resource based on the visual information includes: A first object appearing in the visual picture of the multimedia resource is identified based on the visual information through a multimodal recognition model, wherein the multimodal recognition model is a neural network model for identifying objects based on visual information.

5. The method for determining background audio according to claim 1, characterized in that: The determining, based on the first object, the background audio of the multimedia resource comprises: Supplementary background audio is determined, and the background audio matching the first object and the supplementary background audio are determined as background audio recommended for the multimedia resource.

6. The method for determining background audio according to claim 5, characterized in that: The determining of supplementary background audio includes at least one of the following: Determine the background audio with the first number of usage times ranked first on the first application as the supplementary background audio, wherein the first application is the application to which the multimedia resource belongs; Determine the background audio used by the second number of multimedia resources ranked first in the number of times of playback on the first application as the supplementary background audio; Determine the background audios ranked top by a third number of usage times on at least one second application as the supplementary background audios, where the second application is a resource playback application different from the first application; The audio with the fourth largest number of playback times on at least one third application is determined as the supplementary background audio, and the third application is an audio playback application.

7. The method for determining background audio according to claim 1, characterized in that: The method further comprises: Based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, identifying an element in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource; The determining, based on the first object, the background audio of the multimedia resource comprises: Based on the first object, determining a plurality of first background audios matching the first object; The first background audio matching the element among the multiple first background audios is determined as the background audio of the multimedia resource.

8. The method for determining background audio according to claim 7, characterized in that: The identifying the element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource comprises at least one of the following: In the case where multiple second scenes are identified, determining a scene whose appearance duration in the multimedia resource satisfies a second condition among the multiple second scenes as the first scene; In the case where multiple second actions are identified, an action among the multiple second actions whose appearance duration in the multimedia resource satisfies a third condition is determined as the first action.

9. The method for determining background audio according to claim 7, characterized in that: There are multiple first objects, and determining the first background audio matching the element among the multiple first background audios as the background audio of the multimedia resource includes: For each first object, a background audio matching the element is selected from a plurality of first background audios matching the first object.

10. The method for determining background audio according to claim 9, characterized in that: The method further comprises: The selected multiple background audios are output according to the priorities of the multiple first objects.

11. The method for determining background audio according to claim 1, characterized in that: The method further comprises: If the background audio is not determined based on the first object, identifying an element in the multimedia resource based on at least one of the visual information of the multimedia resource and the resource audio contained in the multimedia resource, the element comprising at least one of an emotion type, a first scene, and a first action, the emotion type being used to represent the emotion expressed by the multimedia resource; Based on the element, background audio of the multimedia resource is determined, the background audio matching the element.

12. The method for determining background audio according to claim 11, characterized in that: The determining the background audio of the multimedia resource based on the element includes: Determine a plurality of second background audios, wherein the plurality of second background audios include at least one of the following: a first number of background audios ranked first in number of times used on a first application, background audios used by a second number of multimedia resources ranked first in number of times played on the first application, a second number of background audios ranked first in number of times used on at least one second application, and a third number of audios ranked first in number of times played on at least one third application, wherein the first application is an application to which the multimedia resources belong, the second application is a resource playing application different from the first application, and the third application is an audio playing application; The second background audio matching the element among the multiple second background audios is determined as the background audio of the multimedia resource.

13. The method for determining background audio according to claim 1, characterized in that: The method further comprises at least one of the following: Based on the multimedia resource, extract the target segment in the background audio, and synthesize the target segment into the multimedia resource; The video screen of the multimedia resource is synchronously synthesized with the background audio so that the screen switching rhythm of the multimedia resource matches the rhythm of the background audio.

14. A method for determining background audio, characterized in that: The method comprises: Display the music interface of multimedia resources; In response to a music matching operation on the music matching interface, displaying at least one background audio recommended for the multimedia resource; The background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

15. The method for determining background audio according to claim 14, characterized in that: There are multiple first objects, each of the at least one background audio corresponds to a first object, and the display of the at least one background audio recommended for the multimedia resource includes: The background audios respectively corresponding to the multiple first objects are displayed in sequence according to the priorities of the multiple first objects.

16. The method for determining background audio according to claim 14, characterized in that: The method further comprises: In response to a selection operation on any background audio, an adjustment panel of the background audio is displayed, wherein the adjustment panel is used to implement at least one of the following: extracting a target segment of the background audio, synchronizing and synthesizing a video screen of the multimedia resource with the background audio, adjusting the volume of the background audio, and adjusting the volume of the resource audio contained in the multimedia resource; In response to an adjustment operation on the background audio on the adjustment panel, the background audio is adjusted.

17. A device for determining background audio, characterized in that: The device comprises: An acquisition unit, configured to acquire visual information of a multimedia resource; an identification unit, configured to identify a first object appearing in a visual picture of the multimedia resource based on the visual information; The determining unit is configured to determine the background audio of the multimedia resource based on the first object, where the background audio matches the first object.

18. A device for determining background audio, characterized in that: The device comprises: A display unit, configured to execute and display a music interface of a multimedia resource; The display unit is further configured to execute, in response to a music matching operation on the music matching interface, display at least one background audio recommended for the multimedia resource; The background audio is determined based on a first object appearing in a visual picture of the multimedia resource, and the first object is determined based on visual information of the multimedia resource.

19. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for determining background audio according to any one of claims 1 to 13 or the method for determining background audio according to any one of claims 14 to 16.

20. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for determining background audio as described in any one of claims 1 to 13 or the method for determining background audio as described in any one of claims 14 to 16.

21. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program implements the method for determining background audio according to any one of claims 1 to 13 or the method for determining background audio according to any one of claims 14 to 16.