Audio recommendation method and device, electronic equipment, storage medium and program product
By obtaining user preferences and media file description information, using audio to generate large models and cluster analysis, the problem of traditional dubbing lacks personalization is solved, efficient and accurate audio recommendations are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510399982.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional still photos and video dubbing lacks personalization, making it difficult to achieve efficient and accurate audio recommendations.
By obtaining the target user's preference information and the description information of the media file, a pre-trained audio is used to generate a large model, and a cluster analysis and audio generation model is combined to generate a personalized recommended audio.
It realizes the accuracy and personalization of audio recommendations, and improves user satisfaction and user experience.
Smart Images

Figure CN120234444A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to an audio recommendation method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the popularization of smart phones and the rapid development of social media, users' demand for recording and sharing life is increasing day by day. Traditional still photos have limitations in expressing emotions and telling stories. Therefore, users begin to use album video production tools to more vividly present the beautiful moments in life. In order to improve the quality and viewing experience of the produced videos, voiceovers are often required. How to achieve personalized and efficient voiceovers has gradually become an issue of concern. Summary of the Invention
[0003] To overcome the problems in the related art, the present disclosure provides an audio recommendation method, apparatus, electronic device, storage medium, and program product.
[0004] According to a first aspect of an embodiment of the present disclosure, an audio recommendation method is provided, including: Obtaining preference information of a target user and description information of a media file to be processed; Obtaining a recommended audio according to the preference information and the description information.
[0005] Optionally, the audio recommendation method further includes: Obtaining the audio selected by the target user in different preset scenarios, where different types of audio are set in each preset scenario; Performing clustering analysis on the audio selected by the target user in each preset scenario to generate the preference information.
[0006] Optionally, the performing clustering analysis on the audio selected by the target user in each preset scenario to generate the preference information includes: Determining the feature information of the audio selected by the target user in each preset scenario; Performing clustering analysis on the feature information of the audio selected by the target user in each preset scenario and the feature information of the audio selected by a first reference user in each preset scenario, and determining a first reference user with the highest similarity to the target user as a second reference user; Determining the classification number corresponding to the second reference user as a target classification number for characterizing the preference information of the target user, where the classification number is a serial number for characterizing the preference information of the reference user.
[0007] Optionally, the obtaining the audio selected by the target user in different preset scenarios includes: When the target user uses the dubbing function for the first time, obtain the audio selected by the target user in different preset scenarios.
[0008] Optionally, obtaining the recommended audio according to the preference information and the description information includes: Input the preference information and the description information into a pre-trained audio generation large model to obtain the recommended audio; Among them, the audio generation large model is trained in the following manner: Obtain multiple training samples, each training sample includes a sample classification number, sample description information, and sample audio, where the sample description information is used to describe the sample media file, and the sample classification number is used to represent the preference information of the sample user; Use the training samples to train the audio generation large model until the training end condition is met.
[0009] Optionally, obtaining the description information used to describe the media file to be processed includes: Input the media file to be processed into a pre-trained video understanding large model to obtain the description information.
[0010] Optionally, the audio recommendation method further includes: Use the recommended audio to dub the media file to be processed.
[0011] According to the second aspect of the embodiments of the present disclosure, there is provided an audio recommendation device, including: A first acquisition module, configured to acquire the preference information of the target user and the description information of the media file to be processed; A recommendation module, configured to obtain a recommended audio according to the preference information and the description information.
[0012] According to the third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Among them, the processor is configured to execute the executable instructions in the memory to implement the steps of the audio recommendation method provided in the first aspect of the present disclosure.
[0013] According to the fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the audio recommendation method provided in the first aspect of the present disclosure are implemented.
[0014] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the audio recommendation method provided in the first aspect of the present disclosure.
[0015] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: Obtain the preference information of the target user and the description information of the media file to be processed; obtain the recommended audio according to the preference information and the description information. In this way, while fully considering the user's personal preferences, it is possible to accurately match the audio that highly fits the content of the media file to be processed, ensure the accuracy of audio recommendation, achieve personalized recommendation, and improve the user's satisfaction and usage experience.
[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0018] Figure 1 is a flowchart of an audio recommendation method shown according to an exemplary embodiment.
[0019] Figure 2 is a flowchart of an audio recommendation method shown according to an exemplary embodiment.
[0020] Figure 3 is a flowchart of a clustering process shown according to an exemplary embodiment.
[0021] Figure 4 is a block diagram of an audio recommendation apparatus shown according to an exemplary embodiment.
[0022] Figure 5 is a block diagram of an electronic device shown according to an exemplary embodiment.
[0023] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0025] It should be noted that all actions of obtaining signals, information, or data in this disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and with the authorization given by the owner of the corresponding device.
[0026] Figure 1 It is a flowchart of an audio recommendation method shown according to an exemplary embodiment. As Figure 1 shown, the method may include step S101 and step S102.
[0027] In step S101, obtain the preference information of the target user and the description information of the media file to be processed.
[0028] Exemplarily, the holder of the terminal electronic device can be determined as the target user, or the target user can be set based on actual needs. The media file may include a video or a picture. The preference information of the target user can be determined according to the historical media files browsed by the target user. Or when the target user first uses the dubbing function, different media files under multiple preset scenarios can be provided to the target user, and the target user is guided to select the preferred media file to determine the preference information according to the selection of the target user.
[0029] Exemplarily, the description information of the media file to be processed can be the text information used to describe the media file. The description information of the media file to be processed can be determined in the following way: input the media file to be processed into a pre-trained large video understanding model to obtain the corresponding description information. For example, the large video understanding model can be the Apollo model.
[0030] In step S102, obtain the recommended audio according to the preference information and the description information.
[0031] Exemplarily, the preference information and the description information can be input into a pre-trained large audio generation model to obtain the recommended audio. The recommended audio may include at least one of the recommended music and sound effects. In this way, by using the description information and the preference information, high-quality and personalized audio recommendations can be realized.
[0032] In one embodiment, the audio recommendation method provided by this disclosure may further include: using the recommended audio to dub the media file to be processed.
[0033] Taking the recommended music as an example of the recommended audio, the recommended music can be used as the background music of the media file to synthesize a new media file that conforms to the content of the media file and the user's preference.
[0034] In the above technical solution, the preference information of the target user and the description information of the media file to be processed are obtained; according to the preference information and the description information, a recommended audio is obtained. In this way, while fully considering the user's personal preferences, it is possible to accurately match an audio that highly fits the content of the media file to be processed, ensuring the accuracy of audio recommendation, realizing personalized recommendation, and improving the user's satisfaction and usage experience.
[0035] Figure 2 is a flowchart of an audio recommendation method shown according to an exemplary embodiment. As Figure 2 shown, the method may further include step S103 and step S104.
[0036] In step S103, the audio selected by the target user in different preset scenarios is obtained.
[0037] Among them, different types of audio are set in each preset scenario.
[0038] Exemplarily, the preset scenarios can be set based on actual needs. For example, the preset scenarios may include a sports scenario, a food scenario, a travel scenario, a pet scenario, etc. The types of audio can be set based on at least one of the speed of the audio rhythm, the expressed emotion, and the style. Taking the setting of different types of audio based on the speed of the audio rhythm as an example, fast-paced audio, medium-paced audio, and slow-paced audio can be set in each preset scenario.
[0039] The different types of audio set in each preset scenario can be sequentially presented to the target user, and a prompt message can be generated to guide the target user to select the most favorite audio in each preset scenario, so as to obtain the audio selected by the target user in different preset scenarios.
[0040] According to the audio selected by the target user in different preset scenarios, the preference of the target user for different types of audio in different preset scenarios can be determined, so as to provide more personalized recommended audio that better suits the target user's preferences in the future. For example, if the target user selects lively and relaxed audio in the pet scenario and fast-paced audio in the sports scenario, then the system can recommend audio that conforms to the corresponding style to the target user in different scenarios in the future, thereby improving the target user's usage experience.
[0041] In one embodiment, when the target user first uses the dubbing function, the audio selected by the target user in different preset scenarios can be obtained. In this way, the preferences and needs of the target user can be quickly and accurately understood at the initial stage of the target user using the dubbing function, laying a foundation for providing better quality and personalized services in the future.
[0042] In step S104, cluster analysis is performed on the audio selected by the target user in each preset scenario to generate the preference information of the target user.
[0043] In one embodiment, the generation of the target user preference information can be achieved through steps S301 to S303 as shown in Figure 3 the following: In step S301, determine the feature information of the audio selected by the target user in each preset scenario.
[0044] Exemplarily, a neural network model such as a convolutional neural network (CNN) or a recurrent neural network (RNN) can be used to determine the feature information of the audio. The feature information of the audio can characterize the emotion in the audio, or can also characterize the rhythm and melody of the audio.
[0045] In step S302, perform clustering analysis on the feature information of the audio selected by the target user in each preset scenario and the feature information of the audio selected by the first reference user in each preset scenario, and determine the first reference user with the highest similarity to the target user as the second reference user.
[0046] Exemplarily, K-means clustering can be used to achieve the clustering analysis. The first reference user can be a user who has pre-selected audio in each preset scenario and is different from the target user. For each first reference user, a corresponding classification number can be set based on the analysis result of the selected audio. The classification number is a serial number used to characterize the preference information of the reference user. In this way, the preference information of the user can be structurally represented using the classification number, making the subsequent recommendation process more efficient.
[0047] In step S303, determine the target classification number corresponding to the second reference user as the target classification number used to characterize the preference information of the target user.
[0048] In this way, through clustering analysis, the reference user with the highest similarity to the target user can be simply and accurately determined, and then the target classification number that can accurately characterize the preference information of the target user can be obtained.
[0049] In an alternative embodiment, the recommended audio can be obtained according to the preference information and the description information in the following manner: Input the preference information and the description information of the target user into a pre-trained large audio generation model to obtain the recommended audio.
[0050] Among them, the large audio generation model can be trained in the following manner: Obtain multiple training samples, each training sample including a sample classification number, sample description information, and sample audio, where the sample description information is used to describe the sample media file, and the sample classification number is used to characterize the preference information of the sample user; Use the training samples to train the large audio generation model until the training end condition is met.
[0051] Exemplarily, the sample classification number and the sample description information are the input data for training the audio generation large model, and the sample music is the target output for training the audio generation large model. The training end conditions may include at least one of the following: the output value of the loss function of the model is less than or equal to a preset threshold, and the number of iterations reaches a preset number threshold.
[0052] Exemplarily, the audio generation large model may be a generation model constructed based on a diffusion model. Specifically, it may be a generation model constructed based on the Diffusion Transformer (DiT) framework. By using this audio generation large model, high-quality and personalized audio generation can be achieved by combining the description information and the preference information.
[0053] In the training stage, first, the VAE Encoder (the encoder of the variational autoencoder) can compress the music waveform into the latent space, and train the DiT in the latent space. In the inference stage, based on the trained diffusion model, denoise the random noise to obtain the latent (latent space) of the generated music, and finally reconstruct and generate the music through the VAE Decoder (the decoder of the variational autoencoder). The diffusion model adopts the Transformer architecture and can use the attention mechanism to capture the long-term dependencies and complex patterns in the music data. In the forward noise addition and backward denoising processes of the diffusion model, taking the description information as the text information as an example, the text feature vector generated by the Text Encoder (text encoder) according to the text information and the classification number feature vector generated by the Number Encoder (classification number encoder) according to the preference information can be combined to gradually optimize the generation process. By compressing the original audio signal into a low-dimensional latent space, the computational amount is reduced while the key features of the audio are retained. The generated audio latent representation is decoded back to the original audio signal through the VAE Decoder, which can ensure the quality of the generated music.
[0054] Based on the same inventive concept, the present disclosure also provides an audio recommendation device. Figure 4 It is a block diagram of an audio recommendation device 400 shown according to an exemplary embodiment. Referring to Figure 4 , the audio recommendation device 400 may include: A first acquisition module 401, configured to acquire the preference information of the target user and the description information of the media file to be processed; A recommendation module 402, configured to obtain the recommended audio according to the preference information and the description information.
[0055] In the above technical solution, the preference information of the target user and the description information of the media file to be processed are obtained; according to the preference information and the description information, the recommended audio is obtained. In this way, while fully considering the personal preferences of the user, it is possible to accurately match the audio that highly fits the content of the media file to be processed, ensuring the accuracy of audio recommendation, realizing personalized recommendation, and improving the user's satisfaction and usage experience.
[0056] Optionally, the audio recommendation device 400 may further include: A second acquisition module, configured to acquire the audio selected by the target user in different preset scenarios, where different types of audio are set in each of the preset scenarios; A clustering module, configured to perform clustering analysis on the audio selected by the target user in each of the preset scenarios to generate the preference information.
[0057] Optionally, the clustering module is configured to generate the preference information in the following manner, including: Determine the feature information of the audio selected by the target user in each of the preset scenarios; Perform clustering analysis on the feature information of the audio selected by the target user in each of the preset scenarios and the feature information of the audio selected by the first reference user in each of the preset scenarios, and determine the first reference user with the highest similarity to the target user as the second reference user; Determine the classification number corresponding to the second reference user as the target classification number for characterizing the preference information of the target user, where the classification number is the serial number for characterizing the preference information of the reference user.
[0058] Optionally, the second acquisition module is configured to acquire the audio selected by the target user in different preset scenarios in the following manner: When the target user uses the dubbing function for the first time, acquire the audio selected by the target user in different preset scenarios.
[0059] Optionally, the recommendation module 402 is configured to obtain the recommended audio in the following manner: Input the preference information and the description information into a pre-trained large audio generation model to obtain the recommended audio; Wherein, the large audio generation model is trained in the following manner: Obtain multiple training samples, each training sample including a sample classification number, sample description information, and sample audio, where the sample description information is used to describe the sample media file, and the sample classification number is used to characterize the preference information of the sample user; Use the training samples to train the large audio generation model until the training end condition is met.
[0060] Optionally, the audio generation large model is a generative model constructed based on a diffusion model.
[0061] Optionally, the first acquisition module 401 is configured to acquire description information for describing a media file to be processed in the following manner: Input the media file to be processed into a pre-trained video understanding large model to obtain the description information.
[0062] Optionally, the audio recommendation device 400 may further include: A synthesis module, configured to dub the media file to be processed by using the recommended audio.
[0063] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0064] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the audio recommendation method provided by the present disclosure are implemented.
[0065] Figure 5 FIG. is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0066] Referring to Figure 5 , the electronic device 800 may include one or more of the following components: a first processing component 802, a first memory 804, a first power component 806, a multimedia component 808, an audio component 810, a first input / output interface 812, a sensor component 814, and a communication component 816.
[0067] The first processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The first processing component 802 may include one or more first processors 820 to execute instructions to complete all or part of the steps of the above audio recommendation method. In addition, the first processing component 802 may include one or more modules to facilitate the interaction between the first processing component 802 and other components. For example, the first processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the first processing component 802.
[0068] The first memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The first memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0069] The first power supply component 806 provides power to various components of the electronic device 800. The first power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0070] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0071] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the first memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0072] The first input / output interface 812 provides an interface between the first processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0073] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0074] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0075] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above audio recommendation method.
[0076] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 804 including instructions, and the above instructions can be executed by a first processor 820 of the electronic device 800 to complete the above audio recommendation method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0077] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-described audio recommendation method when executed by the programmable device.
[0078] Figure 6 FIG. 4 is a block diagram of an electronic device 1900 shown in accordance with an exemplary embodiment. For example, the electronic device 1900 may be provided as a server. Referring to Figure 6 FIG. 4, the electronic device 1900 includes a second processing component 1922, which further includes one or more processors, and memory resources represented by a second memory 1932 for storing instructions executable by the second processing component 1922, such as application programs. The application programs stored in the second memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the second processing component 1922 is configured to execute instructions to perform the above-described audio recommendation method.
[0079] The electronic device 1900 may further include a second power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and a second input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0080] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functionality for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present application.
[0081] It should be understood that, unless otherwise specifically stated, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more of them; similarly, "at least one of..." includes any one of the related listed items and any combination of any two or more of them.
[0082] Although terms such as "first", "second", and "third" may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are only used to distinguish one component, part, region, layer, or section from another. Thus, the first component, part, region, layer, or section referred to in the examples described herein may also be referred to as the second component, part, region, layer, or section without departing from the teachings of the respective examples. Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description herein, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0083] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Instead, the use of the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified, or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied in any of the foregoing instances. Additionally, unless otherwise specified or clear from the context that it is referring to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".
[0084] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may, as may be desired and advantageous for any given or particular application, be combined with one or more other features of other implementations. Further, with respect to the use of "comprises," "comprising," "has," "having," "includes," or "including" in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "including."
[0085] Other embodiments of the present disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0086] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An audio recommendation method, characterized in that: include: Obtaining preference information of target users and description information of the media files to be processed; According to the preference information and the description information, recommended audio is obtained.
2. The audio recommendation method according to claim 1, characterized in that The audio recommendation method further includes: Acquire the audio selected by the target user in different preset scenarios, wherein different types of audio are set in each preset scenario; Cluster analysis is performed on the audio selected by the target user in each of the preset scenarios to generate the preference information.
3. The audio recommendation method according to claim 2, characterized in that: The performing cluster analysis on the audio selected by the target user in each of the preset scenarios to generate the preference information includes: Determining feature information of the audio selected by the target user in each of the preset scenarios; Performing cluster analysis on feature information of the audio selected by the target user in each of the preset scenarios and feature information of the audio selected by the first reference user in each of the preset scenarios, and determining the first reference user with the highest similarity to the target user as the second reference user; The classification number corresponding to the second reference user is determined as a target classification number for characterizing the preference information of the target user, wherein the classification number is a serial number for characterizing the preference information of the reference user.
4. The audio recommendation method according to claim 2, characterized in that: The obtaining of the audio selected by the target user in different preset scenarios includes: When the target user uses the dubbing function for the first time, the audio selected by the target user in different preset scenarios is obtained.
5. The audio recommendation method according to claim 1, characterized in that: The obtaining of the recommended audio according to the preference information and the description information includes: Inputting the preference information and the description information into a pre-trained audio generation model to obtain the recommended audio; The audio generation model is trained in the following way: Acquire multiple training samples, each of which includes a sample classification number, sample description information, and sample audio, wherein the sample description information is used to describe the sample media file, and the sample classification number is used to represent the preference information of the sample user; The audio generation model is trained using the training samples until a training end condition is met.
6. The audio recommendation method according to claim 1, characterized in that: Get the description information of the media file to be processed, including: The media file to be processed is input into a pre-trained video understanding model to obtain the description information.
7. The audio recommendation method according to claim 1, characterized in that: The audio recommendation method further includes: The recommended audio is used to dub the media file to be processed.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.