Video processing method, mobile terminal and readable storage medium
By automatically selecting and synthesizing background music from the image content of the original video, the problem of low efficiency and poor effect of manual selection is solved, and efficient video background music matching is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI TRANSSION CO LTD
- Filing Date
- 2020-08-07
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, manually selecting background music for videos is inefficient and produces poor results.
By acquiring image content from the original video, matching background music is automatically selected and synthesized with the original video, reducing manual intervention, improving matching efficiency, and enhancing the effect.
It enables automated matching of background music for videos, improving efficiency and enhancing matching results.
Smart Images

Figure CN111917999B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video processing technology, and in particular to a video processing method, a mobile terminal, and a readable storage medium. Background Technology
[0002] With the development of technology, video has become an important way for people to record and share their lives. To make videos more interesting, people often add background music. However, when choosing background music for videos, people usually have to manually select from existing templates, which is inefficient and produces poor results.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a video processing method, a mobile terminal, and a readable storage medium, which aims to solve the technical problems of low efficiency and poor effect when manually matching background music in videos.
[0005] To achieve the above objectives, this application provides a video processing method, which includes the following steps:
[0006] Obtain the original video to be processed, and obtain the image content in the original video;
[0007] Based on the image content, select matching background music;
[0008] The background music is combined with the original video to output a target video with background music.
[0009] Optionally, the step of obtaining image content from the original video includes:
[0010] Obtain motion model information and / or environmental information of the subject in the original video.
[0011] Optionally, the motion model information includes at least one of the following: motion subject information, motion rhythm information, motion amplitude information, and motion type information.
[0012] Optionally, the step of filtering out matching background music based on the image content includes:
[0013] Obtain the music model information of the background music to be matched;
[0014] Based on the image content and the music model information, matching background music is selected.
[0015] Optionally, the music mode information includes at least one of the following: music rhythm information, music timing information, music volume information, or music style information.
[0016] Optionally, the background music to be matched is background music obtained from a video library, characterized in that the step of filtering out matching background music based on the image content includes:
[0017] Obtain at least one of the following video contents from the videos in the video library: subject information in the video, environmental information in the video, and action type information in the video;
[0018] Based on the image content and the video content, select matching background music.
[0019] Optionally, the background music to be matched is background music obtained from an audio library, characterized in that the step of filtering out matching background music based on the image content includes:
[0020] Obtain the audio model from the audio in the audio library;
[0021] Based on the image content and the audio model, matching background music is selected.
[0022] Optionally, the step of filtering out matching background music based on the image content includes:
[0023] The comparison results of the motion model information of the subject in the original video and the music model information of the background music to be matched are obtained;
[0024] If the comparison result meets the preset conditions, then the background music is determined to be the selected background music.
[0025] Optionally, the step of filtering out matching background music based on the image content includes:
[0026] The original video is segmented based on the motion model information of the subject in the image content;
[0027] For the segmented original video, matching background music is selected based on the image content of the segmented original video.
[0028] Optionally, the step of determining the background music as the selected background music if the comparison result meets the preset conditions includes at least one of the following:
[0029] If the similarity between the action rhythm information in the action model information and the music rhythm information in the music model information of the background music to be matched meets the first preset condition, then the background music to be matched is determined to be the selected background music.
[0030] If the similarity between the motion amplitude information in the motion model information and the music volume information in the music model information of the background music to be matched meets the second preset condition, then the background music to be matched is determined to be the selected background music.
[0031] If the similarity between the action subject information in the action model information and the video subject information in the video of the video library meets the third preset condition, then the background music to be matched is determined to be the selected background music.
[0032] If the similarity between the action type information in the action model information and the action type information in the videos of the video library meets the fourth preset condition, then the background music to be matched is determined to be the selected background music.
[0033] If the environmental information in the original video and the environmental information of the videos in the video library have a similarity to the fifth preset condition, then the background music to be matched is determined to be the selected background music.
[0034] Optionally, the step of combining the background music with the original video includes the following steps before:
[0035] Adjust the motion model of the subject in the original video and / or the music model of the background music to adjust the similarity between the two models.
[0036] This application also provides a mobile terminal, the mobile terminal including: a memory, a processor, and a video processing program stored in the memory and executable on the processor, wherein when the video processing program is executed by the processor, it implements the steps of the video processing method described above.
[0037] This application also provides a readable storage medium, characterized in that the readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the video processing method described above.
[0038] This application embodiment obtains the original video to be processed and extracts the image content from the original video. Then, based on the image content, it filters out matching background music. Finally, it synthesizes the matching background music with the original video to obtain the final target video. In this application, after the user provides the original video to be processed, this method automatically extracts the image content from the video and performs background music matching and filtering based on the extracted image content. By automating the execution of each step of the method, manual intervention is reduced, and the overall efficiency of the matching process is improved. At the same time, matching background music based on image content results in a better matching effect between the video content and background music in the synthesized target video. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A schematic diagram of the hardware structure of a mobile terminal to implement the various embodiments of this application;
[0042] Figure 2 A communication network system architecture diagram provided for an embodiment of this application;
[0043] Figure 3 This is a flowchart illustrating the first embodiment of the video processing method of this application;
[0044] Figure 4 In the third embodiment of the video processing method of this application, Figure 3 Detailed flowchart of step S20;
[0045] Figure 5 This is a schematic diagram showing the correspondence between action rhythm feature points and music rhythm feature points in the third embodiment of the video processing method of this application.
[0046] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0047] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0048] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustration and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0049] Terminals can be implemented in various forms. For example, the terminals described in this application may include mobile terminals such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0050] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.
[0051] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0052] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal:
[0053] The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), and TDD-LTE (Time Division Duplexing-Long Term Evolution).
[0054] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0055] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0056] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage medium) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0057] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0058] Display unit 106 is used to display information input by the user or information provided to the user. Display unit 106 may include display panel 1061, which may be a liquid crystal display (LCD) or an organic light-emitting diode (OLED).
[0059] The display panel 1061 can be configured in various ways.
[0060] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Specifically, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands from processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Specifically, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being limited here.
[0061] Furthermore, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0062] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0063] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0064] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0065] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0066] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0067] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.
[0068] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0069] Specifically, UE201 can be the aforementioned terminal 100, which will not be elaborated here.
[0070] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Among them, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface), and eNodeB2021 connects to EPC203. eNodeB2021 can provide UE201 with access to EPC203.
[0071] EPC203 may include MME (Mobility Management Entity) 2031, HSS (Home Subscriber Server) 2032, other MMEs 2033, SGW (Serving Gateway) 2034, PGW (Packet Data Network Gateway) 2035, and PCRF (Policy and Charging Rules Function) 2036, etc. Among them, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 provides registers to manage functions such as the Home Location Register (not shown in the diagram) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0072] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0073] Although the above description uses the LTE system as an example, those skilled in the art should understand that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, and future new network systems, etc., which are not limited here.
[0074] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.
[0075] This application provides a video processing method.
[0076] In a first embodiment of the video processing method, the method includes:
[0077] Step S10: Obtain the original video to be processed, and obtain the image content in the original video;
[0078] The original video can be a video stored locally by the user or a video downloaded from the network. This application does not restrict the source of the video.
[0079] The acquired image content mainly includes motion model information of the subject in the original video and environmental information, as well as other information that can reflect the subject's movements. Motion model information includes subject information, rhythm information, amplitude information, and type information. Subject information includes people, cats, dogs, etc.; rhythm information includes angular velocity and acceleration at the joints of the subject; amplitude information includes displacement at the joints; and type information includes walking, dancing, etc. Environmental information includes objects in the environment, weather conditions, and changes in objects in the background.
[0080] Step S20: Based on the image content, select matching background music;
[0081] Using the previously obtained motion model information and / or environmental information of the subject, a candidate music library is obtained. Based on the motion attributes of the subject, motion feature points and motion rhythm feature points are obtained. The motion feature curve of the original video is obtained based on the motion rhythm feature points. For each candidate music in the candidate music library, the music rhythm feature points are obtained using the time series of the motion rhythm feature points. The music feature curve of each candidate music can be obtained using the music rhythm feature points. Finally, a suitable background music is matched based on the two feature curves.
[0082] Step S30: Combine the background music with the original video to output a target video with background music;
[0083] For the matching background music, this application can combine them or synthesize them individually with the original video. At the same time, before synthesizing the background music with the original video, the rhythm of the background music and / or the original video can be adjusted, and finally the synthesized video is output.
[0084] In this embodiment, the original video to be processed is acquired, and the image content within the original video is obtained. Then, matching background music is selected based on the image content. Finally, the matching background music is synthesized with the original video to obtain the final target video. In this application, after the user provides the original video to be processed, this method automatically extracts the image content from the video and performs background music matching and selection based on the extracted image content. By automating the execution of each step of the method, manual intervention is reduced, improving the overall efficiency of the matching process. Furthermore, matching background music based on image content results in a better match between the video content and background music in the synthesized target video.
[0085] Furthermore, based on the first embodiment of the video processing method of this application, a second embodiment of the video processing method is provided. In the second embodiment, step S10 includes:
[0086] Step A: Obtain motion model information and / or environmental information of the subject in the original video;
[0087] The motion model information includes at least one of the following: motion subject information, motion rhythm information, motion amplitude information, and motion type information.
[0088] The subject of the action can be a person, an animal, or a robot; any objective entity capable of movement can serve as the subject of this application. Furthermore, if a person is chosen as the subject, this application can further analyze information such as the subject's gender, age group, and body type. For example, this application ultimately analyzes the subject to be a graceful young woman or a middle-aged man with a slightly overweight physique. The rhythm and amplitude information of the action can be obtained through inference from action attributes, including angular velocity, acceleration, displacement, and the body part in motion. These attributes can be used to summarize the subject's action. Action attributes are not limited to those listed in this application but also include other information related to the action that can be obtained through image recognition analysis. For instance, the rhythm information of the action can be inferred from the angular velocity and acceleration of each moving part, and the amplitude information can be inferred from the displacement of each moving part. Based on the rhythm and amplitude information, the closest action type can be determined. Environmental information refers to the external environment in which the subject is located, such as a classroom, library, park, or shopping mall. The subject and its environment provide usable conditions for this application to subsequently obtain a candidate music library.
[0089] Specifically, the acquisition of motion model information can be as follows: For a subject capable of generating motion, the subject contains different joints, and the motion process is accompanied by and closely related to the movement of these joints. Commonly used joints include: wrist, elbow, and shoulder joints of the upper limbs; knee and ankle joints of the lower limbs; and in addition, this application will also use hip joints, cervical spine, and lumbar spine. A large angular velocity of a joint indicates a fast motion frequency, and a large displacement of a joint indicates a large motion amplitude. By comparing the angular velocities and displacements of different joints, the part of the body generating the motion can be determined. At the same time, through big data search, the acquired angular velocities and displacements of the joints are compared with the angular velocities and displacements of different types of motion to determine the type of motion. For example, if the joints of the upper and lower limbs of the subject have large angular velocities and large displacements, the subject can be considered to be dancing. If the angular velocities and displacements of the upper limb joints are small, and the angular velocities and displacements of the lower limb joints are relatively large and periodic, then the subject's motion may be walking. After identifying the movement parts of the action subject, we will filter the joints, select all the joints of the movement parts, and use the movement attributes of these joints as the movement attributes of the action subject.
[0090] In this embodiment, motion model information and / or environmental information of the main subject in the original video are obtained. This image content provides a usable reference for subsequent background music matching.
[0091] Furthermore, referring to Figure 3 , Figure 4 and Figure 5 Based on the above embodiments of the video processing method of this application, a third embodiment of the video processing method is provided. In the third embodiment,
[0092] Step S20 includes:
[0093] Step S21: Obtain the music model information of the background music to be matched;
[0094] Based on the action subject information in the original video, preliminary screening is performed on online and / or local video and audio libraries to identify background music that meets certain criteria for matching, forming a candidate music library. This candidate music library helps to initially identify potentially matching background music and appropriately reduce the workload in subsequent steps. For the background music to be matched, music model information is then constructed.
[0095] The music model information includes at least one of the following: music rhythm information, music duration information, music volume information, or music style information. Music duration information refers to the length of the music, and music style includes genres such as pop, classical, and blues. Music rhythm can be obtained in the following way: Music also has rhythm. This application utilizes music attributes that can represent music rhythm. By quantifying these music attributes, a numerical value reflecting the music rhythm is obtained. Simultaneously, to link the action attributes in the video with the music attributes of the candidate music, this application uses action rhythm feature points to obtain music rhythm feature points. Similarly, through the music rhythm feature points and their corresponding values, a music feature curve for each candidate music is obtained. This music feature curve can be continuous or discontinuous. The time series of action rhythm feature points reflects the temporal relationship attributes of the action rhythm feature points in the original video. This application links the original video with the candidate music through the time series. The correspondence between action rhythm feature points and music rhythm feature points established through the time series is referenced... Figure 5 To ensure rhythmic matching between the original video and the candidate music, the relative positions of the rhythmic feature points of the original video and the candidate music on their respective timelines should be roughly consistent. For each musical rhythmic feature point, a musical value is obtained by quantifying its representative musical characteristics. For music, the duration and pitch of its notes can reflect its musical characteristics to a certain extent; therefore, the musical value is positively correlated with the note duration and pitch of the musical rhythmic feature point. The horizontal axis of the musical feature curve represents the time axis of the music, and the vertical axis represents the musical value. Each musical rhythmic feature point has corresponding horizontal and vertical coordinates, resulting in discrete points that constitute the musical feature curve. Continuous musical feature curves can be obtained through interpolation. Furthermore, if the distribution of musical rhythmic feature points is too scattered, we obtain the musical feature curves for relatively densely distributed areas of musical rhythmic feature points, ultimately obtaining the overall segmented and discontinuous musical feature curve.
[0096] Step S22: Based on the image content and the music model information, select matching background music;
[0097] After obtaining the background music to be matched, this application will further filter the background music. The motion model information of the subject will be the filtering standard. In this way, the background music matched after filtering will correspond to the image content of the original video, ensuring the accuracy of the matched background music.
[0098] First, establish the motion feature curves for the original video. The process is as follows:
[0099] The presence of action attributes at specific time points indicates the existence of a subject performing an action. Therefore, this application uses these points as action feature points. The extreme points of angular velocity among these action feature points may represent changes in the action attributes of the subject; thus, this application uses these points as action rhythm feature points. Action rhythm feature points are themselves action feature points, possessing action attributes. This application can quantify these action attributes to obtain a numerical value that reflects them. Therefore, each action rhythm feature point will have a corresponding numerical value. Through these points and their corresponding values, we can obtain the action feature curve corresponding to the original video. This action feature curve can be continuous or discontinuous. The action attributes of action rhythm feature points include some abstract concepts that cannot be compared by the program. Therefore, we need to quantify the information contained in the action attributes of the action rhythm feature points to obtain a value that the program can use for comparison. This value is obtained by quantifying and weighting each piece of information contained in the action attributes. The action value generally reflects the action attributes of the action rhythm feature points. The action value is positively correlated with the angular velocity and displacement of the action rhythm feature points. For the motion feature curve, the horizontal axis is the time axis of the original video and the vertical axis is the motion value. Each motion rhythm feature point has a corresponding horizontal and vertical coordinate, resulting in some discrete points that constitute the motion feature curve. Continuous motion feature curves can be obtained by interpolation. At the same time, if the distribution of motion rhythm feature points is too scattered, we obtain the motion feature curves of relatively densely distributed regions of motion rhythm feature points, and finally obtain the overall segmented and discontinuous motion feature curves.
[0100] Next, based on the motion feature curve of the original video and the music feature curve of the background music to be matched, after obtaining these two feature curves, it is necessary to compare the similarity between the two curves to determine whether the original video and the candidate music are consistent in terms of rhythm, that is, to determine whether the candidate music can be used as the matching background music. There are many methods to compare the similarity between the two curves. For example, you can compare the overall trend changes of the two curves, or you can select enough corresponding points on the two curves and compare the difference between their corresponding values and the trend of the difference.
[0101] The matching process is as follows:
[0102] In this application, a key factor in linking the original video with the candidate music is the time series of motion rhythm feature points and music rhythm feature points. Since the time series of motion rhythm feature points and music rhythm feature points are consistent, a correspondence exists between them. Using the absolute value of the difference more accurately reflects the gap between motion and music values, and avoids bias in the subsequent summation due to the mixing of positive and negative numbers, better reflecting the difference in matching degree between the two curves. The difference between each pair of feature points and the total difference must each be less than a certain value to ensure that the original video and candidate music match rhythmically overall and at key points, thus ensuring that the resulting background music conforms to our standards.
[0103] In this embodiment, all the content constitutes an action-music matching model, which mainly utilizes the rhythmic features in the video and music. In this model, the original video is linked to candidate music through rhythm, and by quantifying the action attributes of the action rhythmic feature points in the original video and the musical attributes of the musical rhythmic feature points in the candidate music, action feature curves and musical feature curves are obtained. Comparing the matching degree between the original video and the candidate music is transformed into comparing the matching degree between the action feature curves and the musical feature curves. This model ensures that the original video matches the final background music, improving the professionalism of the music matching process in this application.
[0104] Furthermore, based on the above embodiments of the video processing method of this application, a fourth embodiment of the video processing method is provided, in which...
[0105] If the background music to be matched is obtained from the video library, then step S20 includes:
[0106] Step B1: Obtain at least one of the following video contents from the videos in the video library: subject information in the video, environmental information in the video, and action type information in the video;
[0107] Step B2: Based on the image content and the video content, filter out matching background music;
[0108] When the music to be matched is obtained from the video library, the process is similar to that of the original video. Based on the video content in the video library, the main body information, environmental information, and action type information in the video are obtained. Then, the image content in the original video is matched with the video content in the video library. For example, whether the main body information and action type information are the same, and whether the environmental information is similar. If the main body information and action type information are the same and the environmental information is similar, it can be determined to be the background music to be matched.
[0109] In this embodiment, for background music obtained from videos in the video library, the content of the videos in the video library, such as subject information, environmental information, and action type information, is obtained, and background music is matched according to the video content, making the matching process faster.
[0110] Furthermore, based on the above embodiments of the video processing method of this application, a fourth embodiment of the video processing method is provided, in which...
[0111] The background music to be matched is obtained from the audio library. Step S20 includes:
[0112] Step C1: Obtain the audio model from the audio in the audio library;
[0113] Step C2: Based on the image content and the audio model, select matching background music;
[0114] For audio in the audio library, the same method as in the third embodiment can be used to obtain the music feature curves of the audio in the audio library. The music feature curves, along with the audio's volume information, style information, etc., form an audio model. After obtaining the music feature curves, the motion feature curves of the original video are matched with the music feature curves of the music in the audio library according to the same process as in the third embodiment. At the same time, the style information in the audio model and the environmental information in the original video are used as auxiliary matching information. When the matching degree between the motion feature curve and the music feature curve is greater than a preset value and the style information and environmental information are similar (e.g., the style is classical music and the environment is a quiet environment), it is determined that the audio in the audio library matches the video content of the original video and can be used as matching background music. Among them, the user can pre-set the corresponding relationship regarding whether the style information in the music model is similar to the environmental information of the video content. For example, in a quiet environment, similar style information is classical music, lyrical music, etc. This application does not limit the specific setting method and corresponding relationship.
[0115] In this embodiment, an audio model is built from the audio in the audio library. The background music is matched with the image content of the original video based on the accuracy of the matching result.
[0116] Furthermore, based on the above embodiments of the video processing method of this application, a fifth embodiment of the video processing method is provided. In the fifth embodiment,
[0117] Step S20 includes:
[0118] Step D1: Obtain the comparison result between the motion model information of the subject in the original video and the music model information of the background music to be matched;
[0119] Step D2: If the comparison result meets the preset conditions, then the background music is determined to be the selected background music;
[0120] The motion model information includes at least one of motion subject information, motion rhythm information, motion amplitude information, and motion type information. The music model information includes at least one of music rhythm information, music timing information, music volume information, or music style information. In this embodiment, it is not necessary to convert the motion model information and music model information into motion feature curves and music feature curves; instead, the comparison results of the motion model information and music model information themselves can be used.
[0121] Step D2 includes at least one of the following steps:
[0122] Step D21: If the similarity between the action rhythm information in the action model information and the music rhythm information in the music model information of the background music to be matched meets the first preset condition, then the background music to be matched is determined to be the selected background music.
[0123] Here, the motion model information and music model information need to be converted into motion feature curves and music feature curves. Then, the similarity between the motion feature curves and music feature curves is compared using the method in the third embodiment of the case. When the similarity is greater than a preset value and meets the first preset condition, the background music to be matched is determined to be the selected background music.
[0124] Step D22: If the similarity between the motion amplitude information in the motion model information and the music volume information in the music model information of the background music to be matched meets the second preset condition, then the background music to be matched is determined to be the selected background music.
[0125] Step D23: If the similarity between the action subject information in the action model information and the video subject information in the video of the video in the video library meets the third preset condition, then the background music to be matched is determined to be the selected background music.
[0126] Step D24: If the similarity between the action type information in the action model information and the action type information in the videos of the video library meets the fourth preset condition, then the background music to be matched is determined to be the selected background music.
[0127] Step D25: If the environmental information in the original video and the environmental information of the video in the video library have the same similarity as the fifth preset condition, then the background music to be matched is determined to be the selected background music.
[0128] For steps D22 to D25, only one feature from the action model information and the music model information needs to be used for judgment. Therefore, as long as the similarity of the corresponding attributes is greater than the corresponding preset value, the corresponding preset condition is met, and thus matching background music can be found. For example, the background music in videos of elderly people taking a walk, or the background music in videos of people in a library, etc.
[0129] In this embodiment, the criteria for finding matching background music are diverse, thereby providing users with more background music options and improving the user experience.
[0130] Furthermore, based on the above embodiments of the video processing method of this application, a sixth embodiment of the video processing method is provided. In the sixth embodiment,
[0131] Step S20 includes:
[0132] Step E1: Segment the original video according to the motion model information of the subject in the image content;
[0133] Step E2: For the segmented original video, select matching background music based on the image content of the segmented original video;
[0134] Based on motion model information, the original video is divided into multiple clearly defined segments. For example, the main subject of the original video may change from a human to a cat or dog, or the original video with fast-paced motion may be distinguished from one with slow-paced motion based on motion rhythm information. This segmentation allows for the accurate identification of the most suitable background music for each segment. The method for matching the segmented original video with the background music can be one included in the aforementioned embodiments.
[0135] In this embodiment, the original video is segmented to obtain the background music that best matches each segment, thereby improving the accuracy of the matching results.
[0136] Furthermore, based on the above embodiments of the video processing method of this application, a seventh embodiment of the video processing method is provided, in which...
[0137] Step S30 includes:
[0138] Step F: Adjust the motion model of the subject in the original video and / or the music model of the background music to adjust the similarity between the two models;
[0139] To make the final merged video with background music more natural and smooth, adjustments need to be made to the motion model or music model.
[0140] Step F includes:
[0141] Step F1, adjust at least one of the following in the motion model of the subject in the original video: motion timing information, motion rhythm information, or motion amplitude information; and / or
[0142] Step F2: Adjust at least one of the following in the background music music model: music rhythm information, music timing information, or music volume information;
[0143] For example, simultaneously adjusting the motion rhythm of the motion model and the music rhythm of the applied model can create a more unified rhythm between the two. Alternatively, increasing the volume of a particular music model can emphasize the corresponding content in the original video. Furthermore, adjustments to the original video and / or background music are not limited to those listed above.
[0144] In this embodiment, the motion model and / or application model are adjusted before synthesis, increasing the appeal of the application and also making the output target video more perfect. Providing different target videos for users to choose from enhances user engagement while respecting their choices.
[0145] In addition, this application also provides a video processing apparatus, which includes:
[0146] The acquisition module is used to acquire the original video to be processed and to acquire the image content in the original video;
[0147] The filtering module is used to filter out matching background music based on the image content;
[0148] The generation module is used to synthesize the background music with the original video and output a target video with background music.
[0149] Optionally, the acquisition module is further configured to:
[0150] Obtain motion model information and / or environmental information of the subject in the original video.
[0151] Optionally, the filtering module further includes:
[0152] The first acquisition unit is used to acquire the music model information of the background music to be matched;
[0153] The first filtering unit is used to filter out matching background music based on the image content and the music model information.
[0154] Optionally, the filtering module further includes:
[0155] The second acquisition unit is used to acquire at least one of the following video contents from the videos in the video library: subject information in the video, environmental information in the video, and action type information in the video;
[0156] The second filtering unit is used to filter out matching background music based on the image content and the video content.
[0157] Optionally, the filtering module further includes:
[0158] The third acquisition unit is used to acquire audio models from the audio in the audio library;
[0159] The third filtering unit is used to filter out matching background music based on the image content and the audio model.
[0160] Optionally, the filtering module further includes:
[0161] The fourth acquisition unit is used to acquire the comparison result between the motion model information of the subject in the original video and the music model information of the background music to be matched;
[0162] The fourth filtering unit is used to determine that the background music is the filtered background music if the comparison result meets the preset conditions.
[0163] Optionally, the filtering module further includes:
[0164] The segmentation unit is used to segment the original video according to the motion model information of the subject in the image content;
[0165] The fifth filtering unit is used to filter out matching background music based on the image content of the segmented original video.
[0166] Optionally, the video processing apparatus further includes:
[0167] An adjustment module is used to adjust the motion model of the subject in the original video and / or the music model of the background music to adjust the similarity between the two models.
[0168] This application also provides an apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described above.
[0169] This application also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described above.
[0170] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the methods described in the various possible implementations above.
[0171] This application also provides a chip including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device equipped with the chip performs the methods described in the various possible embodiments described above.
[0172] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0173] It should be understood that although the terms first, second, third, etc., may be used herein to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if," as used herein, can be interpreted as "when," "when," or "in response to determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising," "including," indicate the presence of the stated feature, step, operation, element, component, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, or mean any one or any combination thereof. Therefore, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B, and C". Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0174] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0175] It should be noted that step designations such as S10 and S20 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the protection scope of this application.
[0176] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0177] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0178] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.
Claims
1. A video processing method, characterized in that, The video processing method includes the following steps: Obtain the original video to be processed, and obtain the image content from the original video, including: Obtain motion model information of the subject in the original video, wherein the motion model information includes at least one of the following: subject information, rhythm information, amplitude information, and type information; Based on the image content, select matching background music; Using the motion model information of the subject of the action as a filtering criterion, the matching background music is filtered, including: The time points in the image content that have action attributes are used as action feature points; The extreme points of angular velocity among the action feature points are taken as action rhythm feature points; Based on the motion rhythm feature points and their corresponding values, obtain the motion feature curve corresponding to the original video; For discrete points in the motion feature curve, continuous motion feature curves can be obtained by interpolation, or motion feature curves of relatively densely distributed regions of motion rhythm feature points can be obtained to obtain overall segmented discontinuous motion feature curves. The matching background music is filtered based on the continuous motion feature curve or the overall segmented discontinuous motion feature curve to determine the filtered background music. The filtered background music is combined with the original video to output the target video with background music.
2. The video processing method as described in claim 1, characterized in that, The step of filtering out matching background music based on the image content includes: Obtain the music model information of the background music to be matched; Based on the image content and the music model information, matching background music is selected.
3. The video processing method as described in claim 2, characterized in that, The music model information includes at least one of the following: music rhythm information, music timing information, music volume information, or music style information.
4. The video processing method as described in claim 2, wherein the background music to be matched is background music obtained from a video library, characterized in that, The step of filtering out matching background music based on the image content includes: Obtain at least one of the following video contents from the videos in the video library: subject information in the video, environmental information in the video, and action type information in the video; Based on the image content and the video content, select matching background music.
5. The video processing method as described in claim 2, wherein the background music to be matched is background music obtained from an audio library, characterized in that, The step of filtering out matching background music based on the image content includes: Obtain the audio model from the audio in the audio library; Based on the image content and the audio model, matching background music is selected.
6. The video processing method as described in claim 2, characterized in that, The step of filtering out matching background music based on the image content includes: The comparison results of the motion model information of the subject in the original video and the music model information of the background music to be matched are obtained; If the comparison result meets the preset conditions, then the background music is determined to be the selected background music.
7. The video processing method as described in claim 1, characterized in that, The step of filtering out matching background music based on the image content includes: The original video is segmented based on the motion model information of the subject in the image content; For the segmented original video, matching background music is selected based on the image content of the segmented original video.
8. The video processing method as described in claim 6, characterized in that, The step of determining that the background music is the selected background music if the comparison result meets the preset conditions includes at least one of the following: If the similarity between the action rhythm information in the action model information and the music rhythm information in the music model information of the background music to be matched meets the first preset condition, then the background music to be matched is determined to be the selected background music. If the similarity between the motion amplitude information in the motion model information and the music volume information in the music model information of the background music to be matched meets the second preset condition, then the background music to be matched is determined to be the selected background music. If the similarity between the action subject information in the action model information and the video subject information in the video in the video library meets the third preset condition, then the background music to be matched is determined to be the selected background music. If the similarity between the action type information in the action model information and the action type information in the videos in the video library meets the fourth preset condition, then the background music to be matched is determined to be the selected background music. If the environmental information in the original video and the environmental information in the video library meet the fifth preset condition, then the background music to be matched is determined to be the selected background music.
9. The video processing method as described in claim 8, characterized in that, Before the step of combining the filtered background music with the original video, the following steps are included: Adjust the motion model of the subject in the original video and / or the music model of the filtered background music to adjust the similarity between the two models.
10. The video processing method as described in claim 9, characterized in that, The steps of adjusting the motion model of the subject in the original video and / or the music model of the filtered background music include: Adjust at least one of the following in the motion model of the subject in the original video: motion timing information, motion rhythm information, or motion amplitude information; and / or Adjust at least one of the following music models of the filtered background music: music rhythm information, music timing information, or music volume information.
11. A mobile terminal, characterized in that, The mobile terminal includes: a memory, a processor, and a video processing program stored in the memory and executable on the processor, wherein the video processing program, when executed by the processor, implements the steps of the video processing method as described in any one of claims 1 to 10.
12. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the video processing method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Method and device for automatically generating background music
CN103795897A
Method and device for dubbing music for animations
CN106503034A
Video processing method and device, terminal and computer readable storage medium
CN110099300A
Background music construction method and device, electronic equipment and storage medium
CN111324773A