Music generation method, device and equipment

By generating and adjusting pitch curves, users can create music directly based on image features, solving the problem of limited selection range in existing technologies and realizing the convenience of customized music generation.

CN120612905APending Publication Date: 2025-09-09ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510954630.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the prior art, users have limited choices when selecting music to match pictures, and are unable to select music that matches the pictures from the existing music database, making it difficult to meet user needs.

Method used

By obtaining the subject outline in the target image, an initial pitch curve is generated with the vertical axis as the pitch data axis and the horizontal axis as the time data axis. The curve is displayed on the terminal interface and the user is allowed to adjust it. The target pitch curve is generated based on the user's adjustment operation, and finally an audio file is generated.

Benefits of technology

Users can directly adjust the pitch trajectory and time distribution based on visual features to generate music melodies that meet their personal needs, which improves the convenience of music generation and eliminates the process of relying on text descriptions or professional music symbols.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612905A_ABST
    Figure CN120612905A_ABST
Patent Text Reader

Abstract

The invention provides a music generation method, apparatus and device. The music generation method comprises the steps of obtaining a main body contour in a target image; generating an initial pitch curve graph based on the main body contour; the longitudinal axis of the initial pitch curve graph is a pitch data axis, and the horizontal axis is a time data axis; displaying the initial pitch curve graph on a terminal interface; obtaining an adjustment operation of a user for the initial pitch curve graph, and obtaining an adjusted target pitch curve graph; the adjustment operation comprises an adjustment operation for pitch data or an adjustment operation for time data; and generating an audio file based on the target pitch curve graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a music generation method, apparatus, and device. Background Art

[0002] When users post picture notes on the application platform, they can match the corresponding music to the picture. Usually, users can select music that suits the image from an existing music database or music recommended by the terminal application; or use the terminal device to analyze the image features, determine the image style, and then obtain music that suits the image style from an existing music database. However, this method of relying on the existing music database to search for matching music is prone to problems such as limited user selection range and inability to select music that suits the picture from the existing music data, which makes it difficult to meet user needs. Summary of the Invention

[0003] In view of this, the embodiments of the present application provide a music generation method, apparatus and device to solve the problem that users cannot select the music they need from the existing music database and cannot meet user needs.

[0004] According to a first aspect of an embodiment of the present application, a music generation method is provided, comprising: Get the subject outline in the target image; Generate an initial pitch curve based on the subject outline; the vertical axis of the initial pitch curve is a pitch data axis, and the horizontal axis is a time data axis; Displaying the initial pitch curve on the terminal interface; Acquiring a user's adjustment operation on the initial pitch curve graph to obtain an adjusted target pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data; An audio file is generated based on the target pitch graph.

[0005] According to a second aspect of an embodiment of the present application, there is provided a music generating apparatus, comprising: An acquisition module, used to acquire the subject outline in the target image; An initial pitch curve generating module, configured to generate an initial pitch curve based on the subject outline; wherein the vertical axis of the initial pitch curve is a pitch data axis, and the horizontal axis is a time data axis; A display module, configured to display the initial pitch curve on a terminal interface; a target pitch curve generating module, configured to obtain an adjustment operation performed by a user on the initial pitch curve to obtain an adjusted target pitch curve; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data; A music melody generation module is used to generate an audio file based on the target pitch curve graph.

[0006] According to a third aspect of an embodiment of the present application, a computing device is provided, comprising a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor implements the steps of the music generation method when executing the computer instructions.

[0007] One embodiment of the present specification can achieve at least the following beneficial effects: by obtaining the outline of the subject in the target image, an initial pitch curve graph is generated based on the subject outline, with the vertical axis being the pitch data axis and the horizontal axis being the time data axis; the initial pitch curve graph is displayed on the terminal interface so that the user can adjust the initial pitch curve graph, and based on the user's adjustment operation on the initial pitch curve graph, an adjusted target pitch curve graph is obtained, and a musical melody is generated based on the target pitch curve graph. In the embodiment of the present specification, after the initial pitch curve graph is generated based on the outline of the subject of the target image, the initial pitch curve graph is displayed to the user, allowing the user to directly visually adjust the pitch trajectory and time distribution according to their own needs to obtain a target pitch curve, and then, based on the customized target pitch curve generated by the user's adjustment operation, a musical melody that meets the user's needs can be generated.

[0008] In addition, the method of generating music melodies directly based on visual features can save users from relying on text descriptions or professional music symbols (such as musical scores) to create music, thereby effectively improving the convenience of music generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0010] Figure 1 This is a schematic diagram of an application scenario of a music generation method provided in an embodiment of this specification; Figure 2 This is a flowchart of a music generation method provided by an embodiment of this specification; Figure 3 A schematic diagram of an initial pitch curve provided in an embodiment of this specification; Figure 4 is a schematic diagram including multiple pitch curves provided in an embodiment of this specification; Figure 5This is a swim lane diagram of a music generation method provided in an embodiment of this specification; Figure 6 This is a structural diagram of a music generating device provided in an embodiment of this specification; Figure 7 This is a structural diagram of a music generating device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0011] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.

[0012] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.

[0013] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0014] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0015] In the prior art, when users post picture notes on application platforms, they can match the picture with corresponding music. Typically, users can select music that matches the image from an existing music database or music recommended by the terminal application; alternatively, the terminal device can analyze the image features, determine the image style, and then retrieve music that matches the image style from an existing music database. However, this method of relying on existing music databases to search for matching music is prone to problems such as limited user selection range and inability to select satisfactory music that matches the picture from the existing music database.

[0016] In order to solve the defects in the prior art, this solution provides the following embodiments.

[0017] Figure 1 A schematic diagram of an application scenario of a music generation method provided in an embodiment of this specification.

[0018] like Figure 1 As shown, the solution may include a user terminal 101 and a server 102. The user can provide a target image that needs to be matched with music through the user terminal 101. The user terminal 101 can send the target image to the server 102. The server 102 can obtain the outline of the subject in the target image, and generate an initial pitch curve graph based on the outline of the subject, with the vertical axis being the pitch data axis and the horizontal axis being the time data axis. The generated initial pitch curve graph can also be sent to the client terminal 101 so that the user terminal 101 can display the initial pitch curve graph on the terminal interface, thereby facilitating the user to adjust the initial pitch curve graph as needed. The server 102 can obtain the user's adjustment operation on the initial pitch curve graph, and generate an adjusted target pitch curve graph based on the user's adjustment operation; it can also generate an audio file based on the target pitch curve graph, and send the generated audio file to the user terminal for user use.

[0019] although Figure 1 As shown in FIG, after receiving the target image, the user terminal 101 can send the target image to the server 102. In actual application, if the computing resources of the user terminal 101 meet the conditions for running the music generation method, the solution of generating an audio file based on the target image can also be executed on the user terminal 101.

[0020] In such Figure 1 In the application scenario shown, the server may be connected to one or more terminal devices via a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Figure 1 The server 102 may include but is not limited to any device, equipment, platform, equipment cluster, etc. with computing and processing capabilities. Figure 1The user terminal 101 may include but is not limited to a smart phone, a tablet computer, a laptop computer, a PDA, a personal computer, a smart home device, a vehicle-mounted device, etc.

[0021] In this application, a music generation method is provided. This application also relates to a music generation device and a computing device, which are described in detail one by one in the following embodiments.

[0022] Figure 2 A flowchart of a music generation method provided in an embodiment of this specification.

[0023] From the program perspective, the execution body of the process can be a program installed on an application server or application terminal. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.

[0024] like Figure 2 As shown, the process may include the following steps: Step 202: Acquire the subject outline in the target image.

[0025] In the embodiments of this specification, the target image may be an image containing curved lines. In practical applications, the target image can be obtained through a variety of methods or channels. Specifically, the target image may be an image captured by a user through a camera or downloaded from a network channel. In terms of type, the target image may include a landscape image (e.g., an image of mountains), an image of a building, an image of a person, or an image of a motion trajectory.

[0026] The subject outline may be a key visual feature used to define the shape of the subject and distinguish the subject from the background or other objects. In the embodiments of this specification, the subject outline may refer to the edge contour line of the subject in the target image. The subject may refer to the core object of visual attention. The subject may be a specific object, such as a person, scenery, animal, building, or object. It may also be an abstract form, such as a dynamic trajectory. In a figure image, the subject may be a person, and the subject outline may refer to the continuous lines that outline the head, body, and limbs of the person. In an image of mountains, the subject may be a mountain range, and the subject outline may refer to the contour lines formed by extracting the boundary lines between the mountain range and background areas such as the sky and vegetation, which can clearly outline the overall shape of the mountain range (such as the undulating trend of the peak, the continuous outline of the mountain body, the concave boundary of the valley, etc.).

[0027] In practical applications, a target image may include multiple subjects, such as a person and a mountain. If the target image includes multiple subjects, the multiple subjects included in the target image can be marked, and the user can also select one of the subjects as the target subject to obtain the subject outline of the target subject.

[0028] Step 204: Generate an initial pitch curve graph based on the subject outline; the vertical axis of the initial pitch curve graph is the pitch data axis, and the horizontal axis is the time data axis.

[0029] In the embodiments of this specification, a pitch curve graph is a graph that visually presents the pattern of pitch changes over time. Specifically, the pitch curve in the pitch curve graph can be connected by a series of continuous points, and each point corresponds to the pitch value at a moment. The ups and downs of the pitch curve can directly reflect the change in pitch. For example, a rising pitch curve can indicate a rise in pitch, such as the ascending "do-re-mi" in the scale; a falling pitch curve can indicate a decrease in pitch, such as the descending "mi-re-do"; a horizontal pitch curve can indicate that the pitch remains stable, such as a continuous single tone. In addition, the steepness of the pitch curve can represent the rate of pitch change; for example, a steep pitch curve indicates a fast pitch change; a flat pitch curve indicates a slow pitch change.

[0030] In the embodiments of this specification, the vertical axis of the initial pitch graph can be a pitch data axis, representing pitch and used to reflect the highness of the tone. The unit of pitch can be Hertz (Hz), semitone, or MIDI pitch value. The horizontal axis of the initial pitch graph can be a time data axis, representing time and used to reflect the dynamic change trend of pitch over time. The time unit can be seconds (s) or milliseconds (ms).

[0031] In practical applications, based on the mapping relationship between the visual contour morphology and the pitch variation law, the spatial characteristics of the subject contour can be converted into an initial pitch curve of the pitch variation over time.

[0032] In the embodiments of this specification, the pitch curve diagram is used to convert the abstract pitch change into an intuitive and visible curve form, and the user can clearly observe the fluctuation trend of the sound tone through the visual curve trend.

[0033] Step 206: Display the initial pitch curve graph on the terminal interface.

[0034] The terminal interface can be a visual interface for users to interact with terminal devices, which can be smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, in-vehicle devices, etc.

[0035] Figure 3 A schematic diagram of an initial pitch curve diagram provided in an embodiment of this specification is shown as follows: Figure 3 shown.

[0036] In the embodiment of this specification, the initial pitch curve is displayed on the terminal interface, which enables the user to intuitively perceive the dynamic change trend of the pitch over time and clearly grasp the overall shape of the pitch contour.

[0037] Step 208: Acquire the user's adjustment operation on the initial pitch curve graph to obtain an adjusted target pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data.

[0038] In the embodiments of this specification, the initial pitch curve graph is editable while displayed on the terminal interface. A user can adjust the pitch curve in the initial pitch curve graph. In practical applications, the user's adjustment operations on the initial pitch curve graph may include adjusting pitch data at a specified time point or within a specified time period, or adjusting time data within a specified time period.

[0039] Furthermore, when the initial pitch curve graph is displayed on the terminal interface, corresponding audio can be synchronously generated based on the initial pitch curve. The terminal device can provide a playback function for the generated audio to play the audio, so that the user can perceive the acoustic effect corresponding to the pitch curve in real time through hearing. Specifically, during the audio playback process, the terminal interface can synchronously display the playback progress in the pitch curve graph, such as indicating the current playback position through dynamic marking points or highlighted segments, so that the user can intuitively compare the auditory effect of the audio with the visual form of the pitch curve, and accurately locate the pitch segment that needs to be adjusted.

[0040] In the embodiments of this specification, the target pitch curve can be a pitch curve that meets the user's needs after the user makes one or more adjustments to the initial pitch curve. Of course, if the audio generated synchronously based on the initial pitch curve meets the user's needs—that is, the audio generated synchronously based on the initial pitch curve is what the user wants—the user can choose not to make adjustments to the initial pitch curve. In this case, the initial pitch curve serves as the target pitch curve.

[0041] During the process of adjusting the initial pitch curve graph, the initial pitch curve graph and each pitch curve graph obtained by performing multiple adjustments on the initial pitch curve graph may be copied. Figure 4 is a schematic diagram including multiple pitch curves provided in an embodiment of this specification, such as Figure 4 As shown, the multiple pitch curve graphs may include an initial pitch curve and multiple pitch curve graphs obtained through adjustments. In practical applications, users can directly adjust the initial pitch curve graph; they can also copy the initial pitch curve graph and make adjustments to the copied initial pitch curve graph. While users are adjusting the initial pitch curve graph or the copy of the initial pitch curve graph, they can also generate audio in real time based on the adjusted pitch curve graph and play the generated audio. In practical applications, the WebAudio API can be used to achieve real-time mapping of adjustment operations to audio changes.

[0042] In an embodiment of the present specification, the initial pitch curve graph is copied, and adjustments are made to the copied initial pitch curve graph. The audio generated in real time based on the real-time adjusted pitch curve graph and the audio generated based on the initial pitch curve graph can be compared to determine the part that needs to be adjusted next.

[0043] As an embodiment, multiple adjustments made by the user to the initial pitch graph may refer to the user adjusting the initial pitch graph once and then adjusting the resulting pitch graph again, or even continuing to adjust the resulting pitch graph after the second adjustment. For example, after the user makes an initial adjustment to the initial pitch graph and generates a first audio based on the initially adjusted pitch graph, if the user still needs to perform a second optimization on the first audio, the user may save and copy the initially adjusted pitch graph, and then perform a second adjustment on the copied pitch graph.

[0044] As another implementation method, the user can make multiple adjustments to the initial pitch graph by copying multiple copies of the initial pitch curve, that is, adjusting the multiple copies of the pitch curve graph separately. The adjustment of multiple copies of the pitch curve graph can be performed in parallel. For example, five copies of the initial pitch curve graph can be copied simultaneously and adjusted in parallel. After the adjustment is completed, corresponding audio is generated based on each copy of the initial pitch curve graph, and then the target audio is filtered out from the five generated audios. Alternatively, the basic audio is first filtered out, and then the pitch curve graph corresponding to the basic audio is further adjusted to finally obtain the target pitch curve graph. In addition, the adjustment of multiple copies can also be performed sequentially. For example, one copy of the initial pitch curve graph can be copied first, and then the one copy of the initial pitch curve graph can be adjusted and audio generated. If the user is not satisfied with the audio, the initial pitch curve graph can be copied again, and the copied copy of the initial pitch curve graph can be adjusted to generate audio. This process is repeated until satisfactory audio is obtained.

[0045] Step 210: Generate an audio file based on the target pitch curve graph.

[0046] It should be understood that in the methods described in one or more embodiments of this specification, the order of some steps can be adjusted according to actual needs, or some steps can be omitted.

[0047] Figure 2The method, by obtaining a subject outline in a target image, generates an initial pitch curve graph based on the subject outline, with the vertical axis being the pitch data axis and the horizontal axis being the time data axis; displays the initial pitch curve graph on a terminal interface so that the user can adjust the initial pitch curve graph, obtains an adjusted target pitch curve graph based on the user's adjustment operation on the initial pitch curve graph, and generates a musical melody based on the target pitch curve graph. In an embodiment of the present specification, after the initial pitch curve graph is generated based on the subject outline of the target image, the initial pitch curve graph is displayed to the user, allowing the user to directly visually adjust the pitch trajectory and time distribution according to their needs, thereby enabling the generation of a musical melody that meets the user's needs based on the customized target pitch curve generated by the user's adjustment operation.

[0048] In addition, the method of generating music melodies directly based on visual features can save users from relying on text descriptions or professional music symbols (such as musical scores) to create music, thereby effectively improving the convenience of music generation.

[0049] based on Figure 2 The present specification also provides some improved implementation methods of the method, which are described below.

[0050] In the embodiments of this specification, for ease of understanding, specific content of obtaining the subject outline of the target image is also provided.

[0051] In one or more embodiments of this specification, optionally, obtaining the subject contour in the target image may specifically include: Depth information is extracted from the target image to obtain a depth image of the target image.

[0052] Threshold segmentation is performed on the depth image to obtain a binary image.

[0053] An edge detection algorithm is used to extract edges from the binary image to generate a subject outline containing subject edge information.

[0054] In the embodiments of this specification, a depth image is an image that can represent the distance relationship of objects in a target image and assist in identifying the subject's outline. In practical applications, one or more monocular image analysis algorithms or deep learning algorithms can be used to extract depth information from the target image, thereby obtaining a depth image of the target image. Monocular image analysis algorithms may include MiDaS, ZoeDepth, and other algorithms. Deep learning algorithms may include neural network models, etc.

[0055] In practical applications, a threshold segmentation algorithm can be used to perform threshold segmentation on the depth image to obtain a binary image. Specifically, the UnityEngine.Rendering.PostProcessing tool can be used to convert the depth image into a binary image.

[0056] In practical applications, one or more edge detection algorithms can be used to extract edges from binary images. For example, gradient-based edge detection algorithms, such as the Sobel operator and the Scharr operator, can be used to extract edges from binary images and generate the subject's outline. Alternatively, edge detection algorithms based on second-order derivatives, such as the Laplacian operator and the Canny edge detector, can be used to extract edges from binary images and generate the subject's outline. Neural network-based edge detection algorithms can also be used, such as the Holistically-Nested Edge Detection (HED) end-to-end edge detection method based on deep learning, and the Canny-EdgeCNN edge detection method, which combines the principles of the Canny algorithm with deep learning.

[0057] As an implementation method, the ControlNet controller and Depth algorithm of the ComfyUI tool can be used to extract the subject contour in the target image.

[0058] For ease of understanding, the embodiments of this specification also provide specific content for generating an initial pitch curve graph.

[0059] Optionally, generating an initial pitch curve based on the subject profile may specifically include: Based on the subject profile map, a frequency spectrum map corresponding to the subject profile is generated.

[0060] The initial pitch curve graph is generated based on the spectrogram.

[0061] In the embodiments of this specification, the spectrogram may be a mel-spectrogram. A mel-spectrogram is an audio feature representation method designed based on the human auditory characteristics. In a mel-spectrogram, the horizontal axis represents time, the vertical axis represents mel-frequency, and color represents energy.

[0062] In practical applications, a correlation between the geometric features of the subject's outline and audio parameters can be established. Based on this correlation, the subject outline image can be converted into a spectrogram. Geometric features of the subject's outline can include shape, curvature, length, etc. Audio parameters can include frequency, time, energy, etc.

[0063] As an implementation method, the geometric features of the subject's outline can be first extracted. Specifically, the image processing tool can be used to perform edge detection on the outline image and extract the coordinates of all contour points on the subject's outline. , and arranged in order, such as arranging the coordinates of the contour points from left to right. The geometric features of the subject's contour can then be mapped to an audio signal. Specifically, the x-coordinate of each contour point can be used as a timestamp; based on the predefined correspondence between the y-coordinate range and the Mel frequency range, the y-coordinate of each contour point can be mapped to a frequency. For example, assuming the x-coordinate range is 0 to 10,000 pixels and the mapping value time is 0 to 50 seconds (i.e., 1 pixel corresponds to 0.005 seconds), then the time t corresponding to a pixel with an x ​​value of 200 is 1 second. For another example, if the y-coordinate range is 0 to 500 pixels and corresponds to a Mel frequency of 0 to 8,000 Mel, then a pixel with a y value of 250 corresponds to 4,000 Mel. The curvature of each contour point can then be calculated, and based on the curvature of each contour point, the energy of the position corresponding to each contour point can be determined. For example, a larger curvature indicates a higher energy. This results in an audio waveform.

[0064] Furthermore, the generated audio waveform can be converted into a Mel-spectrogram. Specifically, the audio waveform can be framed and windowed. For example, the audio waveform can be cut at fixed time intervals (such as 20ms per frame), and a window function (such as a Hanning window) can be applied to each frame to reduce spectral leakage. Then, a short-time Fourier transform can be performed on each frame signal to convert each frame signal from the time domain to the frequency domain to obtain a "frequency-energy" distribution. Then, the result obtained by the short-time Fourier transform can be filtered using a triangular filter Mel filter bank that conforms to human ear perception to convert the linear frequency into a Mel frequency. Finally, the logarithm of the filtered energy value can be taken to obtain a "time-Mel frequency-logarithmic energy" matrix. The above matrix can be visualized to obtain a Mel-spectrogram.

[0065] Converting a Mel-spectrogram to a pitch contour graph is the process of extracting a one-dimensional "time-pitch" sequence from the two-dimensional "time-frequency-energy" spectrum information and visualizing it. As an implementation, the initial pitch contour graph can be generated based on the spectrogram using the following steps.

[0066] S1. Perform normalization preprocessing on the Mel-spectrogram.

[0067] Specifically, energy normalization can be performed on the Mel spectrum graph to map the energy value of each time frame-frequency point to the range of 0-1; then Gaussian filtering is used to remove high-frequency noise, retaining the main energy distribution characteristics, and obtaining the denoised Mel spectrum matrix (dimension is time step × Mel frequency bins).

[0068] S2. Extract the peak frequency.

[0069] Specifically, for the preprocessed Mel spectrum matrix, all Mel frequency bins are traversed in each time frame, and the frequency point with the largest energy value is selected as the candidate peak frequency of the time frame; if there are multiple peaks with close energy (difference ≤ preset threshold), the lowest frequency point is taken as the effective peak frequency.

[0070] S3. Convert frequency to pitch.

[0071] Specifically, the Mel-to-Hertz conversion formula is as shown in Formula 1, which converts the effective peak frequency (Mel frequency) of each time frame into a linear frequency (Hz); Formula 1 Then use Equation 2 to map the linear frequency to a pitch value (such as a MIDI note number). Formula 2 Among them, f is the linear frequency, m is the Mel frequency, and p is the pitch value.

[0072] S4, smoothing the pitch contour; Specifically, a sliding window can be used to perform mean filtering on the pitch values ​​of continuous time frames to eliminate instantaneous jump noise; the window size can be set according to actual needs, such as the window size can be 5-10 time frames; if the difference between the pitch value of a time frame and the adjacent frames exceeds a preset threshold, it can be replaced with the mean of the adjacent frames to the time frame to obtain a smoothed pitch sequence.

[0073] S5. Generate an initial pitch contour map.

[0074] Specifically, the smoothed pitch sequence can be plotted as a continuous curve in chronological order with the time frame as the horizontal axis (in seconds) and the pitch value (MIDI number or Hz) as the vertical axis. The amplitude of the curve corresponds to the pitch value, forming an initial pitch contour diagram.

[0075] In practical applications, after obtaining the initial pitch contour graph, the initial pitch curve graph can be displayed on the terminal interface so that the user can intuitively perceive the dynamic change trend of the pitch over time and clearly grasp the overall shape of the pitch contour.

[0076] The initial pitch curve graph is editable while it is displayed on the terminal interface. Users can edit and adjust the pitch curve in the initial pitch curve graph.

[0077] In actual applications, users can adjust the pitch data at a time point or within a time period so that the audio file obtained based on the pitch curve meets user needs.

[0078] Based on this, the step of obtaining the user's adjustment operation on the initial pitch curve graph to obtain the adjusted target pitch curve graph may specifically include: Acquire the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

[0079] In the embodiments of this specification, the specified time interval can be any time interval or any time point in the pitch curve. The pitch data can refer to the pitch value on the pitch curve. The adjustment operation on the pitch data can refer to adjusting the pitch value on the pitch curve.

[0080] The target pitch curve is a pitch curve obtained after the user makes one or more adjustments to the pitch data.

[0081] For ease of understanding, in the embodiments of this specification, a method is also provided for obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the specific content of the adjusted target pitch curve.

[0082] Optionally, obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the adjusted target pitch curve may specifically include: Get the user's vertical drag operation on the pitch curve within a specified time interval; The position of the pitch curve within the specified interval on the pitch data axis is adjusted according to the dragging operation to obtain an adjusted target pitch curve.

[0083] The user's vertical dragging operation on the pitch curve within a specified time interval can specifically refer to the user selecting the pitch curve within the specified interval and directly dragging the pitch curve up or down to increase or decrease the pitch value; or the user can also increase or decrease the pitch value by performing an up-swipe or down-swipe operation on the terminal interface instead of directly dragging the pitch curve. It should be noted that dragging the pitch curve up or down here does not cause the selected pitch curve to be translated up or down as a whole; rather, it is to gradually adjust the pitch data between the two endpoints while keeping the pitch data of the two endpoints of the pitch curve within the specified time interval unchanged.

[0084] In actual applications, even a drag operation in the vertical direction may have a certain operation angle, but the overall operation is still a vertical drag operation.

[0085] Based on the above adjustment operation, the execution entity can identify the specified time interval (start time , end time ), as well as the direction and distance of the drag; the drag distance can be converted into an offset on the pitch data axis As you can understand, dragging the pitch curve upward or sliding up in the terminal interface will If the value is positive, drag the pitch curve downward or perform a sliding operation in the terminal interface. Then, you can extract the specified time interval All pitch data points within ,in < < , is the original pitch value in the initial pitch curve. Then, the pitch value corresponding to each time point in the specified time interval can be offset to obtain the new pitch value , keeping the time coordinate unchanged; finally, the adjusted pitch point Smoothly connect with the unadjusted pitch points outside the interval to form a complete target pitch curve.

[0086] In addition, in the embodiments of this specification, another method is provided for obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the specific content of the adjusted target pitch curve.

[0087] Optionally, obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the adjusted target pitch curve may specifically include: The pitch data input by the user for the pitch curve within the specified time interval is obtained.

[0088] The position of the pitch curve within the specified interval is adjusted according to the pitch data to obtain an adjusted target pitch curve.

[0089] In the embodiment of this specification, the pitch data input by the user for the pitch curve within the specified time interval may refer to the user selecting the pitch curve within the specified interval and entering precise pitch parameters, thereby replacing the original pitch data within the selected specified time interval with the input pitch parameters.

[0090] In the embodiment of the present specification, the input pitch data can be smoothly connected with the unadjusted pitch points outside the interval to form a complete target pitch curve.

[0091] As an implementation method, the time data in the time data axis of the initial pitch graph may also be adjusted.

[0092] Optionally, obtaining the user's adjustment operation on the initial pitch curve graph to obtain the adjusted target pitch curve graph may specifically include: Acquire the user's adjustment operation on the time data within the specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

[0093] The time data may refer to the duration of the pitch value on the pitch curve. The adjustment operation on the time data may refer to adjusting the duration of the pitch on the pitch curve.

[0094] For ease of understanding, in the embodiments of this specification, a method for obtaining the specific content of the adjusted target pitch curve graph by adjusting the time data within a specified time interval on the initial pitch curve graph is also provided.

[0095] Optionally, obtaining the user's adjustment operation on the time data within a specified time interval on the initial pitch curve to obtain the adjusted target pitch curve may specifically include: Gets the user's horizontal zoom operation on time data within a specified time interval.

[0096] According to the scaling operation, scaling adjustment is performed on the time data axis for the pitch curve within the specified interval to obtain an adjusted target pitch curve.

[0097] In the embodiments of this specification, the user's horizontal scaling operation on the time data within a specified time interval may specifically include the user compressing the time data within the specified time interval, or the user stretching the time data within the specified time interval. It is understood that if the user compresses the time data within the specified time interval, the duration of the pitch data within the specified time interval will be shortened; if the user stretches the time data within the specified time interval, the duration of the pitch data within the specified time interval will be lengthened.

[0098] Furthermore, as an implementation method, the user can compress the time data within a specified time interval in the following manner: the user selects a pitch curve within the specified time interval, simultaneously presses two fingers on both ends of the selected specified time interval, keeps the two fingers within the selected interval, and slides them inward to reduce the distance between the two fingers. The terminal interface can display the compressed interval length in real time until the fingers are released after the expected compression effect is achieved. It can be understood that stretching the time data within the specified time interval is to simultaneously press two fingers on both ends of the selected specified time interval, keep the two fingers within the selected interval, and slide them outward to increase the distance between the two fingers.

[0099] As another embodiment, the user can compress the time data within a specified time interval in the following manner: the user selects the pitch curve within the specified time interval, then selects one endpoint (such as the left endpoint) of the pitch curve within the specified time interval by single-clicking or double-clicking, drags the other endpoint (such as the right endpoint) with one finger, and slides it closer to the selected endpoint (left endpoint). The finger can be released after the desired compression effect is achieved. It can be understood that stretching the time data within the specified time interval is to select one endpoint (such as the left endpoint) of the pitch curve within the specified time interval by single-clicking or double-clicking, and then drags the other endpoint (such as the right endpoint) with one finger and slides it away from the selected endpoint (left endpoint).

[0100] After the adjustment is completed, the system can compress or stretch the time data within the specified time interval according to the current compression ratio, and shift the time data outside the specified time interval, so that the time data outside the specified time interval can be connected with the time data within the specified time interval.

[0101] As another real-time method, users can scale the time data within a specified time interval. This can also be done in the following way: the user selects the pitch curve within the specified time interval and enters the compression coefficient k (a compression coefficient less than 1 is compression, such as compressing to 50% of the original duration; a compression coefficient greater than 1 is stretching, such as compressing to 150% of the original duration). The system can determine the start time based on the specified time interval selected by the user. and end time , and calculate the duration T of the specified time interval, ; According to the compression coefficient, the target duration T' can be calculated. , and all time points t in the interval can also be calculated based on the compression ratio ( ≤ ≤ ) for compression conversion, the new time coordinate for: ; After compression, the original interval ends New end position replace, If the compression coefficient is less than 1, the subsequent time data outside the interval will be shifted forward by the distance T-T'; if the compression coefficient is greater than 1, the subsequent time data outside the interval will be shifted backward by the distance T'-T to avoid overlap or blank spaces on the time axis.

[0102] In actual applications, the user's adjustment operation on the time data within a specified time interval on the initial pitch curve graph may optionally include: cutting the time data within the specified time interval; cutting the time data within the specified time interval may refer to cutting the pitch curve within the specified time interval. The pitch curve within the cut specified time interval may be deleted or copied elsewhere.

[0103] In actual applications, optionally, the user's adjustment operation on the time data within the specified time interval on the initial pitch curve graph can also include: inserting a blank time segment at the specified time point, and then copying or filling the pitch curve in the inserted blank time segment.

[0104] For ease of understanding, the embodiments of this specification also provide specific content of generating an audio file based on a target pitch curve.

[0105] Optionally, generating an audio file based on the target pitch curve graph may specifically include: The target pitch curve is sampled segment by segment along the time data axis to obtain discrete note units.

[0106] A pitch parameter and a time parameter are determined for each note unit.

[0107] The audio file is generated based on the pitch parameter and the time parameter.

[0108] In audio or music processing, a note unit is the smallest independent pitch segment formed by discretizing continuous pitch changes according to certain rules. In the embodiments of this specification, sampling can be performed along the time data axis to obtain the pitch value (Hz or MIDI number) at each time point. Consecutive sampling points with pitch values ​​less than a preset value are merged into a single note, and the start time, duration, and base pitch of the note are determined to obtain discrete note units.

[0109] In the embodiment of this specification, the pitch parameter may refer to the reference pitch of the note unit, and the time parameter may refer to the duration of the note unit.

[0110] In practical applications, generating the audio file based on the pitch parameter and the time parameter may specifically include: converting the pitch parameter and the time parameter into parameters required for audio synthesis, and generating the audio file based on the parameters required for audio synthesis.

[0111] Specifically, the pitch parameter can be mapped to the oscillator frequency. If the pitch parameter is a MIDI number, the MIDI to frequency formula can be used to convert the MIDI value to Hz. If the pitch parameter is in Hz, the Hz value of the pitch parameter can be used directly without conversion. Furthermore, the time parameter can be mapped to the playback duration, and the start time of the note unit can be mapped to the scheduling time (startTime) based on the audio context. Next, the melody can be generated based on the Web Audio API. Generating a melody based on the Web Audio API specifically involves initializing the audio context, creating an oscillator (OscillatorNode) and a gain controller (GainNode) for each note, and constructing the audio signal generation and output chain (oscillator → gain controller → speaker). Playback is then scheduled according to the mapped parameters, and the volume dynamics are controlled using the ADSR envelope to optimize the naturalness of the timbre. Volume dynamics can include attack, decay, sustain, and release. Different timbres can be simulated by adjusting the oscillator type, and frequency gradients can be used for continuous portamento to enhance melodic coherence.

[0112] In practical applications, the generated audio file can be saved. Specifically, the audio file can be named based on the subject of the recognized target image. In addition, the target image can be used as the cover image during audio playback.

[0113] In the embodiments of this specification, the process of synchronously generating the corresponding audio based on the initial pitch curve in the previous text, and generating the audio based on the adjusted pitch curve after the user adjusts the initial pitch curve is similar to the process of generating an audio file based on the target pitch curve, and will not be repeated here.

[0114] The various technical features in the above embodiments can be arbitrarily combined as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of the various technical features in the above embodiments also falls within the scope of disclosure of this specification.

[0115] Figure 5 This is a swim lane diagram of a music generation method provided in an embodiment of this specification. Figure 5 As shown, the process of the music generation method may involve execution entities such as terminal devices and servers, and the process may include a pitch curve adjustment stage and an audio file generation stage.

[0116] During the pitch curve adjustment phase, the execution entities may include a server and a terminal device, and specifically may include the following steps: Step 502: Acquire the subject outline in the target image.

[0117] In the embodiments of this specification, the target image can be an image captured by a user through a camera or downloaded from a network channel. The target image can be a landscape image (such as a mountain image), a building image, a person image, or a motion trajectory image.

[0118] Step 504: Generate an initial pitch curve graph based on the subject outline; the vertical axis of the initial pitch curve graph is the pitch data axis, and the horizontal axis is the time data axis.

[0119] A pitch curve is a graph that visually displays how pitch changes over time. A pitch curve can be composed of a series of connected points, each corresponding to a specific pitch value at a specific moment. The fluctuations in the pitch curve directly reflect changes in pitch.

[0120] Based on the mapping relationship between the visual contour morphology and the pitch variation law, the spatial characteristics of the subject contour can be converted into an initial pitch curve of the pitch variation over time.

[0121] Step 506: The terminal interface displays the initial pitch curve graph.

[0122] Step 508: Acquire the user's adjustment operation on the initial pitch curve graph to obtain an adjusted pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data.

[0123] The initial pitch curve graph is editable while it is displayed on the terminal interface. Users can edit and adjust the pitch curve in the initial pitch curve graph on the terminal interface.

[0124] As an implementation manner, the user may adjust the pitch data within a specified time interval on the initial pitch curve graph.

[0125] Specifically, the user can select a pitch curve within a specified interval and directly drag the pitch curve up or down to increase or decrease the pitch value; or the user can also increase or decrease the pitch value by sliding up or down in the terminal interface instead of directly dragging the pitch curve. At this time, the server can identify the specified time interval (start time) selected by the user. , end time ), as well as the direction and distance the pitch curve is dragged; converting the drag distance into an offset on the pitch data axis Then, extract the specified time interval All pitch data points within ,in, is the original pitch value in the initial pitch curve graph, < < Then, the pitch value corresponding to each time point in the specified time interval is offset to obtain the new pitch value. , keeping the time coordinate unchanged; finally, the adjusted pitch point Smoothly connect with the unadjusted pitch points outside the interval to form a complete target pitch curve.

[0126] Alternatively, the user may select a pitch curve within a specified interval and then input accurate pitch parameters for the selected pitch curve. In this case, the server may replace the original pitch data within the specified time interval with the pitch parameters input by the user.

[0127] As an implementation method, the user may also adjust the time data in the time data axis of the initial pitch graph.

[0128] In actual applications, users can perform horizontal zoom operations on time data within a specified time interval. Specifically, users can select a pitch curve within a specified time interval, press and hold two fingers at both ends of the selected specified time interval, keep the two fingers within the selected interval, and slide them inward to reduce the distance between the two fingers, thereby achieving compression of the time data. It is understandable that users can press and hold two fingers at both ends of the selected specified time interval, keep the two fingers within the selected interval, and slide them outward to increase the distance between the two fingers, thereby achieving stretching of the time data.

[0129] In addition, the user can also select a pitch curve within a specified time interval, then select an endpoint (such as the left endpoint) of the pitch curve within the specified time interval by single-clicking or double-clicking, and drag the other endpoint (such as the right endpoint) with one finger, sliding it closer to the selected endpoint (left endpoint), thereby achieving compression of the time data. It is understandable that the user can also select an endpoint (such as the left endpoint) of the pitch curve within the specified time interval by single-clicking or double-clicking, and then drag the other endpoint (such as the right endpoint) with one finger, sliding it away from the selected endpoint (left endpoint), thereby achieving stretching of the time data.

[0130] After the adjustment is completed, the server can compress or stretch the time data within the specified time interval according to the current compression ratio, and shift the time data outside the specified time interval, so that the time data outside the specified time interval can be connected with the time data within the specified time interval.

[0131] As another real-time method, users can also select the pitch curve within a specified time interval and enter the compression coefficient k (a compression coefficient less than 1 is compression, such as compressing to 50% of the original duration; a compression coefficient greater than 1 is stretching, such as compressing to 150% of the original duration) to achieve the scaling of the time data within the specified time interval. In this case, the server can determine the start time based on the specified time interval selected by the user. and end time , and calculate the duration T of the specified time interval, ; According to the compression coefficient, the target duration T' can be calculated. , and all time points t in the interval can also be calculated based on the compression ratio ( ≤ ≤ ) for compression conversion, the new time coordinate for: ; After compression, the original interval ends New end position replace, If the compression coefficient is less than 1, the subsequent time data outside the interval will be shifted forward by the distance T-T'; if the compression coefficient is greater than 1, the subsequent time data outside the interval will be shifted backward by the distance T'-T to avoid overlap or blank spaces on the time axis.

[0132] In actual applications, the user's adjustment operation on the time data within a specified time interval on the initial pitch curve graph may optionally include: cutting the time data within the specified time interval. Here, cutting the time data within the specified time interval may refer to cutting the pitch curve within the specified time interval. In actual applications, the pitch curve within the specified time interval obtained by cutting may be deleted, or the pitch curve within the specified time interval obtained by cutting may be copied elsewhere.

[0133] In actual applications, optionally, the user's adjustment operation on the time data within the specified time interval on the initial pitch curve graph can also include: inserting a blank time segment at the specified time point, and then copying or filling the pitch curve in the inserted blank time segment.

[0134] Step 510: The terminal interface displays the adjusted pitch curve graph.

[0135] Step 512: Generate intermediate audio based on the adjusted pitch curve graph.

[0136] Specifically, the adjusted pitch curve is sampled segment by segment along the time data axis to obtain discrete initial note units.

[0137] An initial pitch parameter and an initial time parameter are determined for each initial note unit.

[0138] The intermediate audio is generated based on the initial pitch parameter and the initial time parameter.

[0139] Furthermore, sampling can be performed along the time data axis to obtain the pitch value (Hz or MIDI number) at each time point, and continuous sampling points with pitch values ​​less than a preset value can be merged into a single note to determine the start time, duration and reference pitch of the note; thereby obtaining a discrete initial note unit.

[0140] In the embodiment of this specification, the pitch parameter may refer to the reference pitch of the note unit, and the time parameter may refer to the duration of the note unit.

[0141] In practical applications, generating the audio file based on the initial pitch parameters and the initial time parameters may specifically include: converting the initial pitch parameters and the initial time parameters into parameters required for intermediate audio synthesis, and generating intermediate audio based on the parameters required for intermediate audio synthesis.

[0142] Specifically, the initial pitch parameter can be mapped to an oscillator frequency. If the initial pitch parameter is a MIDI number, the MIDI to frequency formula can be used to convert the MIDI value to Hz. If the initial pitch parameter is in Hz, the Hz value of the pitch parameter can be used directly without conversion. Furthermore, the initial time parameter can be mapped to the playback duration, and the start time of the initial note unit can be mapped to the scheduling time (startTime) based on the audio context. Next, a melody can be generated using the Web Audio API. Generating a melody using the Web Audio API specifically involves initializing the audio context, creating an oscillator (OscillatorNode) and a gain controller (GainNode) for each initial note, and constructing the audio signal generation and output chain (oscillator → gain controller → speaker). Playback is then scheduled according to the mapped parameters, and volume dynamics are controlled using an ADSR envelope to optimize timbre naturalness. Volume dynamics can include attack, decay, sustain, and release. Different timbres can be simulated by adjusting the oscillator type, and frequency gradients can be used for continuous portamento to enhance melodic coherence.

[0143] Step 514: Play the intermediate audio.

[0144] Step 516: Acquire the user's adjustment operation on the adjusted pitch curve graph based on the intermediate audio, and obtain the adjusted target pitch curve graph.

[0145] In the embodiment of this specification, the adjustment operation for the adjusted pitch curve graph is similar to the adjustment operation for the initial pitch image described above, and is not described in detail here.

[0146] In the embodiment of this specification, the method for obtaining the adjusted target pitch curve diagram is similar to the method for obtaining the adjusted pitch curve diagram in the previous text, and will not be repeated here.

[0147] During the audio file generation phase, the execution entity may include a server, and the steps may specifically include: Step 518: Generate an audio file based on the target pitch curve.

[0148] In the embodiments of this specification, the specific method of generating an audio file based on the target pitch curve graph has been described in detail above and will not be repeated here.

[0149] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.

[0150] Figure 6 This is a structural diagram of a music generation device provided in an embodiment of this specification.

[0151] like Figure 6 As shown, the device may include: The acquisition module 602 is configured to acquire the subject outline in the target image.

[0152] An initial pitch curve generating module 604 is configured to generate an initial pitch curve based on the subject profile; the vertical axis of the initial pitch curve is a pitch data axis, and the horizontal axis is a time data axis; Display module 606, used for displaying the initial pitch curve graph on the terminal interface; The target pitch curve generating module 608 is configured to obtain an adjustment operation performed by the user on the initial pitch curve to obtain an adjusted target pitch curve; the adjustment operation may include an adjustment operation on pitch data or an adjustment operation on time data; The music melody generation module 610 is configured to generate an audio file based on the target pitch curve.

[0153] based on Figure 6 The present specification also provides some specific implementation plans of the method, which are described below.

[0154] Optionally, the acquisition module 602 may be specifically configured to: Extracting depth information from the target image to obtain a depth image of the target image; Performing threshold segmentation processing on the depth image to obtain a binary image; An edge detection algorithm is used to extract edges from the binary image to generate a subject outline containing subject edge information.

[0155] Optionally, the initial pitch curve graph generating module 604 may be specifically configured to: Based on the subject outline, generating a frequency spectrum corresponding to the subject outline; The initial pitch curve graph is generated based on the spectrogram.

[0156] Optionally, the target pitch curve generating module 608 may specifically include: The first target pitch curve generating unit is configured to obtain a user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

[0157] Optionally, the first target pitch curve graph generating unit may be specifically configured to: Get the user's vertical drag operation on the pitch curve within a specified time interval; The position of the pitch curve within the specified interval on the pitch data axis is adjusted according to the dragging operation to obtain an adjusted target pitch curve.

[0158] Optionally, the first target pitch curve graph generating unit may be specifically configured to: Obtaining pitch data input by a user for the pitch curve within the specified time interval; The position of the pitch curve within the specified interval is adjusted according to the pitch data to obtain an adjusted target pitch curve.

[0159] Optionally, the target pitch curve generating module 608 may specifically include: The second target pitch curve generating unit is configured to obtain a user's adjustment operation on time data within a specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

[0160] Optionally, the second target pitch curve generating unit may be specifically configured to: Get the user's horizontal zoom operation on the time data within the specified time interval; According to the scaling operation, scaling adjustment is performed on the time data axis for the pitch curve within the specified interval to obtain an adjusted target pitch curve.

[0161] Optionally, the music melody generation module 610 may be specifically configured to: Sampling the target pitch curve segment by segment along the time data axis to obtain discrete note units; Determine the pitch parameter and the time parameter for each note unit; The audio file is generated based on the pitch parameter and the time parameter.

[0162] It is understood that the above modules refer to computer programs or program segments for performing one or more specific functions. In addition, the distinction between the above modules does not mean that the actual program codes must also be separated.

[0163] The above is a schematic scheme of a music generation device of this embodiment. It should be noted that the technical scheme of the music generation device and the technical scheme of the music generation method described above are based on the same concept. For details not described in detail in the technical scheme of the music generation device, please refer to the description of the technical scheme of the music generation method described above.

[0164] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.

[0165] Figure 7 This is a structural diagram of a music generating device provided in an embodiment of this specification. Figure 7 As shown, the device 700 may include: at least one processor 710; and, A memory 730 in communication with the at least one processor; wherein, The memory 730 stores instructions 720 that can be executed by the at least one processor 710. The instructions are executed by the at least one processor 710 to enable the at least one processor 710 to: Get the subject outline in the target image.

[0166] An initial pitch curve is generated based on the subject outline; the vertical axis of the initial pitch curve is a pitch data axis, and the horizontal axis is a time data axis.

[0167] The initial pitch curve is displayed on the terminal interface.

[0168] Acquire the user's adjustment operation on the initial pitch curve graph to obtain an adjusted target pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data.

[0169] An audio file is generated based on the target pitch graph.

[0170] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for devices, equipment, and embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The devices, equipment, and methods provided in the embodiments of this specification correspond to each other, so the devices and equipment also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding devices and equipment will not be repeated here.

[0171] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0172] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system onto a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0173] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.

[0174] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0175] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0176] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0177] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0178] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0180] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0181] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0182] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0183] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0184] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0185] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A music generation method, comprising: Get the subject outline in the target image; generating an initial pitch graph based on the subject profile; The vertical axis of the initial pitch curve is the pitch data axis, and the horizontal axis is the time data axis; Displaying the initial pitch curve on the terminal interface; Acquiring a user's adjustment operation on the initial pitch curve graph to obtain an adjusted target pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data; An audio file is generated based on the target pitch graph.

2. The music generation method according to claim 1, wherein obtaining the subject outline in the target image comprises: Extracting depth information from the target image to obtain a depth image of the target image; Performing threshold segmentation processing on the depth image to obtain a binary image; An edge detection algorithm is used to extract edges from the binary image to generate a subject outline containing subject edge information.

3. The music generation method according to claim 1 , wherein generating an initial pitch curve based on the subject profile comprises: Based on the subject outline, generating a frequency spectrum corresponding to the subject outline; The initial pitch curve graph is generated based on the spectrogram.

4. The music generation method according to claim 1, wherein obtaining the user's adjustment operation on the initial pitch curve to obtain the adjusted target pitch curve comprises: Acquire the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

5. The music generation method according to claim 4, wherein obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the adjusted target pitch curve specifically comprises: Get the user's vertical drag operation on the pitch curve within a specified time interval; The position of the pitch curve within the specified interval on the pitch data axis is adjusted according to the dragging operation to obtain an adjusted target pitch curve.

6. The music generation method according to claim 4, wherein obtaining the user's adjustment operation on the pitch data within a specified time interval on the initial pitch curve graph to obtain the adjusted target pitch curve specifically comprises: Obtaining pitch data input by a user for the pitch curve within the specified time interval; The position of the pitch curve within the specified interval is adjusted according to the pitch data to obtain an adjusted target pitch curve.

7. The music generation method according to claim 1 , wherein obtaining the user's adjustment operation on the initial pitch curve to obtain the adjusted target pitch curve comprises: Acquire the user's adjustment operation on the time data within the specified time interval on the initial pitch curve to obtain an adjusted target pitch curve.

8. The music generation method according to claim 7, wherein obtaining the user's adjustment operation on the time data within a specified time interval on the initial pitch curve to obtain the adjusted target pitch curve specifically comprises: Get the user's horizontal zoom operation on the time data within the specified time interval; According to the scaling operation, scaling adjustment is performed on the time data axis for the pitch curve within the specified interval to obtain an adjusted target pitch curve.

9. The music generation method according to claim 1, wherein generating an audio file based on the target pitch graph comprises: Sampling the target pitch curve segment by segment along the time data axis to obtain discrete note units; Determine the pitch parameter and the time parameter for each note unit; The audio file is generated based on the pitch parameter and the time parameter.

10. A device for generating music, comprising: An acquisition module, used to acquire the subject outline in the target image; An initial pitch curve generating module, which generates an initial pitch curve based on the subject profile; The vertical axis of the initial pitch curve is the pitch data axis, and the horizontal axis is the time data axis; A display module, configured to display the initial pitch curve on a terminal interface; a target pitch curve generating module, configured to obtain an adjustment operation performed by a user on the initial pitch curve to obtain an adjusted target pitch curve; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data; A music melody generation module is used to generate an audio file based on the target pitch curve graph.

11. A device for generating music, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Get the subject outline in the target image; Generate an initial pitch curve based on the subject outline; the vertical axis of the initial pitch curve is a pitch data axis, and the horizontal axis is a time data axis; Displaying the initial pitch curve on the terminal interface; Acquiring a user's adjustment operation on the initial pitch curve graph to obtain an adjusted target pitch curve graph; the adjustment operation includes an adjustment operation on pitch data or an adjustment operation on time data; An audio file is generated based on the target pitch graph.