Method, device and equipment for generating music based on image

By generating sound effect segments and splicing music through multiple user operations on images, this solves the problem that ordinary users find it difficult to create personalized music, and realizes the generation of personalized music with simple interaction.

CN121640967APending Publication Date: 2026-03-10ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies have high barriers to entry for music creation methods, making it difficult for ordinary users to create personalized musical works.

Method used

By acquiring multiple operations performed by the user on the image displayed on the terminal device, the element features of the target elements in the image are determined, sound effect segments that match the element features are generated, and then spliced ​​into music according to the operation sequence.

Benefits of technology

Users can generate personalized music through simple interaction without any professional knowledge, which lowers the threshold for music creation and improves the universality and artistry of music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640967A_ABST
    Figure CN121640967A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, device and equipment for generating music based on an image. According to the scheme, the method comprises the steps of obtaining multiple operations executed by a user for an image displayed by the terminal equipment, and for any operation in the multiple operations, determining element features of a target element corresponding to any operation in the image, the element features can comprise at least one of materials, categories, positions in the image and scene information of the image, generating sound effect fragments matched with the element features according to the element features of the target elements corresponding to any operation in the image, and obtaining a plurality of sound effect fragments; and splicing the plurality of sound effect fragments according to a sequence of executing multiple operations by the user to obtain music matched with the image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of computer technology, in particular to a method for generating music based on an image. One or more embodiments of the present specification also relate to an apparatus for generating music based on an image, a computing device, a computer-readable storage medium, and a computer program product. BACKGROUND

[0002] With the continuous development of society, music is applied to different fields, such as the education field to improve children's sound perception through interesting interaction; the psychological healing field to relieve anxiety through natural sound effect combination; the entertainment field, users want to create personalized music works. However, music creation requires professional tools and professional knowledge reserves, and ordinary users are difficult to create personalized music works.

[0003] Therefore, how to provide a method capable of generating music based on simple interaction is a technical problem to be solved. SUMMARY

[0004] Therefore, one or more embodiments of the present specification provide a method, apparatus and device for generating music based on an image to solve the technical problem that the existing music creation method has a high threshold and ordinary users are difficult to create.

[0005] According to a first aspect of one or more embodiments of the present specification, a method for generating music based on an image is provided, including: obtaining a plurality of operations performed by a user on an image displayed by a terminal device; determining an element feature of a target element corresponding to any one of the plurality of operations in the image; the element feature includes at least one of material, category, position in the image, and scene information of the image; generating a sound effect segment matched with the element feature according to the element feature of the target element corresponding to the any one of the plurality of operations in the image, to obtain a plurality of sound effect segments; and splicing the plurality of sound effect segments in the order in which the user performs the plurality of operations to obtain music matched with the image.

[0006] According to a second aspect of one or more embodiments of this specification, an apparatus for generating music based on an image is provided, comprising: an operation acquisition module for acquiring multiple operations performed by a user on an image displayed on a terminal device; an element feature determination module for determining, for any one of the multiple operations, an element feature of a target element corresponding to that operation in the image; the element feature including at least one of material, category, position in the image, and scene information of the image; a sound effect segment determination module for generating a sound effect segment matching the element feature based on the element feature of the target element corresponding to the any one operation in the image, thereby obtaining multiple sound effect segments; and a music generation module for splicing the multiple sound effect segments in the order in which the user performs the multiple operations to obtain music matching the image.

[0007] According to a third aspect of one or more embodiments of this specification, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein the processor, when executing the computer instructions, implements the steps of the image-based music generation method.

[0008] According to a fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions, which, when executed by a processor, implement the steps of the image-based music generation method.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for generating music based on images.

[0010] At least one embodiment of this specification achieves the following beneficial effects: By acquiring multiple operations performed by a user on an image displayed on a terminal device, determining the element features of the target element corresponding to any one of the multiple operations in the image, generating sound effect segments matching the element features based on the element features of the target element corresponding to any one operation, obtaining multiple sound effect segments, and splicing the multiple sound effect segments according to the order in which the user performs the multiple operations to obtain music matching the image. By allowing the user to operate on the image displayed on the terminal device, such as clicking or double-clicking, the server can generate music based on the user's operations on the image. This eliminates the need for the user to possess high professional skills or specialized tools, enabling users to generate corresponding music through simple image interactions. It lowers the professional requirements for users, reduces the barrier to music generation, and has high universality. Furthermore, the server can generate multiple music segments based on multiple user operations and splice them according to the user's operation order to generate music matching the image. This allows for the generation of personalized music matching the user's operation order, thus meeting user needs. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the overall architecture of a method for generating music based on images, provided in one embodiment of this specification.

[0013] Figure 2 This is a schematic flowchart of an embodiment of a method for generating music based on an image provided in this specification;

[0014] Figure 3 This is a swimlane diagram of an image-based music generation method provided in one embodiment of this specification;

[0015] Figure 4 This is a schematic diagram illustrating an embodiment of displaying an image in a terminal device provided in this specification;

[0016] Figure 5 This is a schematic diagram illustrating multiple operations performed by a user on an image, provided in one embodiment of this specification.

[0017] Figure 6 This is a schematic diagram of a music playback page displaying visual sound wave particle effects, provided in one embodiment of this specification.

[0018] Figure 7 This specification provides an embodiment corresponding to... Figure 2 A schematic diagram of a device for generating music from images;

[0019] Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0021] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0022] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this specification. The singular forms “a,” “an,” “an,” “the,” and “the” used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” used in one or more embodiments of this specification includes any or all possible combinations of one or more associated listed items. The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0023] Although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second, and similarly, second may also be referred to as first, without departing from the scope of one or more embodiments of this specification. Ordinal numbers such as first and second do not necessarily indicate order; often they are used to distinguish objects. For example, first server and second server usually refer to two servers. To distinguish these two servers, they are described as first server and second server. Of course, sometimes these two servers may be the same server. Depending on the context, the word "if" as used herein can be interpreted as "when," "when," or "in response to a determination."

[0024] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving data sent by B, or it can be understood as A indirectly receiving data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending data directly to A, or it can be understood as B indirectly sending data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities. In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B", unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is above B", unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation entry points shall be provided for users to choose to authorize or refuse.

[0026] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0027] Figure 1This is a schematic diagram illustrating the overall architecture of a method for generating music based on images, as provided in an embodiment of this specification. Figure 1 As shown, the solution may include a terminal device 1 and a server 2. The terminal device 1 may have a display component for displaying images and may also sense user actions on the displayed images. The user can perform multiple operations on the images displayed on the terminal device 1. The terminal device 1 can send information about these multiple operations to the server 2. The server 2 can determine the element characteristics of the image elements in each operation, such as material and category, and generate sound effect segments that match these element characteristics. The server 2 can then splice the sound effect segments according to the order in which the user performs the operations to obtain music that matches the images. The server 2 can then send the obtained music to the terminal device 1. The terminal device 1 can also play the music. The terminal device 1 can interact with the user through a graphical user interface to invoke the server, thereby implementing the method provided in the embodiments of this specification.

[0028] In such Figure 1 In the application scenarios shown, the server can connect to one or more terminal devices via a local area network (LAN), a wide area network (WAN), an internet connection, or other types of data networks. Figure 1 The servers mentioned can include, but are not limited to, any device, equipment, platform, or equipment cluster with computing and processing capabilities. Figure 1 The terminal devices in this context may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices.

[0029] If the terminal device's operating resources can meet the execution conditions for processing the image-based music generation method, it can also be executed by the terminal device.

[0030] Figure 2 This is a flowchart illustrating a method for generating music based on an image, as provided in an embodiment of this specification.

[0031] From a programming perspective, the executor of the process can be a program hosted on an application server or application terminal. From a hardware perspective, the executor of the process can be a server or terminal. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities.

[0032] like Figure 2 As shown, the process may include the following steps:

[0033] Step 202: Obtain the multiple operations performed by the user on the image displayed on the terminal device.

[0034] In the embodiments of this specification, the image may be acquired by the user using the camera device of the terminal device. Specifically, it may be an image captured by the user using the terminal device, such as by taking a picture or video; or an image scanned by the camera after it is activated, without the user clicking a shooting control, such as a photo button or video recording button. The image may also be selected by the user from the local storage of the terminal device; or, the image may be selected by the user from multiple images provided by the default of the currently displayed application on the terminal device; or, the image may be randomly displayed by the currently displayed application based on multiple default images; or, the image may be generated based on the user's operation on the image control, and so on.

[0035] In practical applications, users can open application pages, H5 pages, or mini-program pages on their terminal devices that can generate music based on images. The page will display a default system image. If the user doesn't like the default image, they can use the image selection controls on the page to access a menu bar. This menu bar can contain controls for local images, system images, and generated images, allowing the user to select the corresponding image to change the default system image. This allows users to perform multiple operations on the new image. The default system image displayed each time the page is opened can be different, or it can be the same each time.

[0036] The operation can be a single or multiple clicks by the user on a certain location in the image; or it can be a long press operation by the user on a certain location in the image, and so on.

[0037] Multiple operations can refer to user actions on different locations within an image, with each operation targeting a different element displayed in the image, and each operation corresponding to one element. Alternatively, multiple operations can also refer to multiple actions performed by the user on the same location within the image. Or, multiple operations can refer to multiple operations performed on different locations of the same element.

[0038] In one or more embodiments of this specification, the operations on the image can be interface operations commonly used by general users, which can reduce the user's learning cost and improve the user experience. Optionally, the operations can include at least one of the following: single-click operation, double-click operation, long-press operation, and selection operation.

[0039] In the embodiments of this specification, a single click operation can be a single click on a location in an image. A double click operation can be a double click on a location in an image. A long press operation can be a press and hold operation on a location in an image for a preset duration. A selection operation can be a circle operation on a region in an image, or a selection operation on a selection box corresponding to an element displayed in the image; or a selection operation on a selection box corresponding to the natural language description of an element displayed in the image, and so on. In practical applications, when operating on elements in an image, the outline of the element in the image can be displayed, allowing the user to determine whether the desired element has been operated on, thus improving the user experience. In practical applications, users can also operate on the background information in the image, and the server can generate corresponding sound effect clips based on the background information of the operation; for example, if there is a blank area in the image without any elements, and the user operates on the blank area, a relatively soft sound effect clip can be generated.

[0040] In one or more embodiments of this specification, the image displayed by the terminal device may include at least one of the following: an image scanned by the terminal device's camera, an image captured by the terminal device, and an image locally stored by the terminal device.

[0041] In the embodiments of this specification, the image can also be an image generated using a text-based graph model based on natural language description information input by the user. Specifically, the terminal device can display an image selection control in the image display interface. After the user clicks on it, a text-based graph control can be provided to the user. The user can click on the text-based graph control to display or jump to a dialog box for inputting natural language description information. The user can input natural language description information for the desired image in the dialog box. The server can call the text-based graph model to process the natural language description information and generate an image corresponding to the text. If the image does not meet the user's expectations, the user can change the natural language description and generate a new image; or, the user can click the refresh control to generate a new image based on the original natural language description information.

[0042] Alternatively, images can be included by default in applications on the terminal device. This provides users with multiple ways to select or confirm images, improving the user experience.

[0043] The elements contained in the image can include objects that can produce sound on their own, such as music boxes, birds, cows, cats, and streams; or objects that cannot produce sound on their own, such as tables, chairs, blackboards, and keyboards.

[0044] In one or more embodiments of this specification, the multiple operations may include at least one of the following: multiple operations by the user on a single frame of an image displayed by the terminal device, or multiple operations by the user on multiple frames of an image displayed by the terminal device.

[0045] In the embodiments of this specification, the image displayed by the terminal device can be a frame of image, such as a photograph, or a frame of image captured by the terminal device. Multiple operations on a frame of image can indicate that the user has performed multiple operations on various elements in the image, and the elements operated on all belong to the same frame of image. For example, if the user performs an operation on element 1, element 2, and element 3 in the image once each, it can indicate that the user has performed multiple operations.

[0046] Alternatively, the image displayed on the terminal device can also be a multi-frame image, such as a video clip displayed on the terminal device, or a multi-frame image scanned or captured by the terminal device. Multiple operations on multi-frame images can indicate that the user has performed operations on elements in different frames, and the operated elements can belong to different frames. For example, if the user operates on elements A and B in the first frame (frame 1), on element C in the second frame (frame 2), and on elements D and E in the third frame (frame 3), it can also indicate that the user has performed multiple operations. The aforementioned first, second, and third frames can represent images of different frames, and can be images of consecutive frames or images of non-consecutive frames; this is not limited here. Multi-frame images can be multiple frames extracted from a single video; or multiple frames extracted from multiple videos; or multiple frames scanned by the terminal device; or multiple frames captured by the terminal device; or multiple frames selected from the terminal device's local storage, etc. Among them, the multi-frame images scanned by the terminal device can represent that the terminal device scans with a camera device but does not take pictures. The server can obtain the user's operation and then store the images scanned by the terminal device when the user operates. The terminal device does not perform local storage processing, thereby reducing the consumption of local storage resources of the terminal device.

[0047] The image-based music generation method described in this specification can be applied to different types of images or videos, enabling the server to perceive user actions on different types of images or videos and generate corresponding music based on those actions, thereby improving the user experience.

[0048] In one or more embodiments of this specification, optionally, acquiring the multiple operations performed by the user on the image displayed on the terminal device may include: acquiring the multiple operations performed by the user on the image displayed on the display interface of the terminal device; or, acquiring video information containing the user's selection operations on several physical objects; and parsing the multiple operations performed by the user on each physical object from the video information.

[0049] In the embodiments of this specification, multiple operations performed by the user on the image displayed on the terminal device's display interface are obtained. Specifically, this could involve displaying a frame of image or video information captured by the terminal device's camera module, with the user performing operations within the display area of ​​the image or video information, resulting in multiple operations. For example, in a children's teaching scenario, the teacher may have pre-captured a video or a frame of image using the terminal device, allowing the child or teacher to perform operations on the captured video or frame displayed on the terminal device, resulting in multiple operations. Alternatively, the terminal device could display an image after scanning has commenced, allowing the user to move the terminal device to change the displayed image, and the user to perform operations on the scanned image, resulting in multiple operations. The server can store the images of user operations, eliminating the need for local storage on the terminal device.

[0050] In the embodiments of this specification, acquiring video information containing a user's selection operations on several physical objects can specifically involve using a terminal device to capture video or image information of the user's operations on multiple physical objects. For example, in a children's teaching scenario, a child picks up or puts down an item on a table in the classroom to select it. The teacher uses a terminal device to record video or images of the child selecting the item, so that the server can determine the element being operated on based on the state of the child holding the item in the video or image. Alternatively, a child uses a specific tool to point at an item on a table in the classroom. The teacher uses a terminal device to record video or images of the child pointing at the item, so that the server can determine the element being operated on based on the pointing of the specified tool in the video or image.

[0051] Alternatively, in a psychological healing scenario, the person seeking healing can take different items in a certain order according to their own wishes, while another user uses a terminal device to record the order in which the person takes the items. During the process of determining the order, if the server detects that the user might prefer a particular item and perform multiple actions on the same item, the server can include each action on the same item in the total number of actions.

[0052] In the embodiments described in this specification, the server can parse the elemental features of the physical object operated by the user from the video information, as well as the order in which the user operates on the object. Based on the elemental features, it can generate matching sound effect segments, and then splice these sound effect segments together based on the order of operations to generate corresponding music for playback. In practical applications, the server can also generate corresponding rhythms based on the number of times and frequency of user operations on the physical object in the video information. This allows for the acquisition of information corresponding to multiple operations through different methods, while also determining the appropriate rhythm, thus improving the artistry of the music.

[0053] For children's teaching scenarios, a terminal device can capture video footage of children selecting classroom teaching aids in sequence, such as a triangle, a book, and a desk. The server can analyze the video data and determine the selected item, such as a teaching aid, based on the specified actions performed by the characters in the video, such as picking up and putting down the object. Based on the characteristics of the teaching aid and the order of operations, a corresponding sound effect sequence is generated. Continuing the previous example, a music sequence of "metallic knocking sound - page turning sound - wooden knocking sound" is generated and automatically arranged into a three-beat dance tune with piano accompaniment. The server can also determine the corresponding target element based on preset operations performed by the people in the video. Specifically, this could be determined by picking up or putting down the target element, clicking on the target element with a finger or a specified shaped object, or by the people in the video speaking a specified voice command for the target element, such as "select this" or "select a book," which contains preset voice information.

[0054] Step 204: For any one of the multiple operations, determine the element features of the target element corresponding to that operation in the image.

[0055] The element features include at least one of the following: material, category, position in the image, and scene information of the image.

[0056] In the embodiments of this specification, elements can be object information contained in an image, such as trees, rivers, fountains, seats, glasses, etc. Target elements can represent the elements corresponding to the location of user operation in the image.

[0057] Element features can be extracted from images using multimodal models, such as MaterialPalette, Qwen2.5VL, MobileNet, and ResNet, to extract features from elements contained in an image. Material can represent the material category of an element; for example, a seat can be made of wood, iron, alloy, plastic, etc., or it can represent physical properties such as the object's hardness, density, elasticity, and structural form. Element category can represent the category to which the element belongs, such as teaching tools, scenery, etc.; or a more granular category, such as tree, grass, fruit, person, car, etc.

[0058] The position of an element in an image can be represented by the position of a corresponding point operated by the user; or it can be represented by the position of a corresponding region operated by the user. The position of an element in an image can also be represented by the position of a single point within the element; or it can be represented by the position of a region within the element, and so on.

[0059] Scene information can be obtained by identifying images using scene recognition models, such as teaching scenes, therapeutic scenes, entertainment scenes, event scenes, and so on. Scene recognition models can capture key elements in an image and determine the matching scene category based on these elements. For example, if an image contains teaching-related elements such as a blackboard and desks, it can be identified as a teaching scene; if the image contains soft furnishings such as sofas and green plants, it can be identified as a therapeutic scene. Scene information can also be obtained by the user selecting the corresponding scene tag on the terminal device's interface.

[0060] In one or more embodiments of this specification, optionally, determining the element features of the target element corresponding to any one of the multiple operations in the image may include: for any one of the multiple operations, performing material recognition on the region image at the operation position in the image based on the operation position of the operation to determine the material information of the target element at the operation position; and / or, for any one of the multiple operations, performing category recognition on the region image at the operation position in the image based on the operation position of the operation to determine the category information of the target element at the operation position.

[0061] In the embodiments of this specification, the server can perform image cutout based on the user's operation location, and then identify the material or category of the cutout area. For example, an image region can be obtained by expanding outwards from the user's operation location at a preset distance, or by combining or based on pixel changes to determine the image region, and then performing element identification within that image region to obtain the material or category information of the elements contained therein. Alternatively, element identification can be performed first to determine the names of the elements contained within the image region, and then the corresponding material or category can be determined based on a material or category information database. Alternatively, a computer model for material or category identification can be used to identify the material or category of the determined image region to obtain the material information, category information, and other element features of the elements contained within the image region. This allows for the determination of the subsequent identification region by cutting out the entire image based on the user's operation, eliminating the need for element feature identification of the entire image and improving identification efficiency.

[0062] In the embodiments of this specification, the region image at the operation location can be a region image of the element corresponding to the operation location recorded by the server during or after the user's operation on the image. Alternatively, after acquiring the image, before the user performs an operation, the server can identify each element in the image, determine the position corresponding to the outline of each element, and obtain a position set containing the position information of each element; then, establish the correspondence between the region images of each element and each position; then, based on the user's operation on the image, determine the position information that includes the operation location in the position set, and determine the element corresponding to that position based on the correspondence. For example, the horizontal coordinate of the operation location is greater than the minimum horizontal coordinate in the position set and less than the maximum horizontal coordinate in the position set; the vertical coordinate of the operation location is less than the maximum vertical coordinate in the position set and greater than the minimum vertical coordinate in the position set. Alternatively, a model can be used to identify the region image containing the operation location from the image based on the operation location.

[0063] In the embodiments of this specification, the material corresponding to a region image can be determined based on the color information and texture information contained in the region image, thereby obtaining material information. For example, if the texture is irregular and layered, the edges may have wood pores or a rough feel, the reflective area is large and soft, and the color is brown, then the material can be determined to be wood; if the texture is fine and uniform, some parts have artificial textures such as brushed or frosted, the reflective area is small and bright, and the color is silver, then the material can be determined to be metal.

[0064] In the embodiments of this specification, a sample library of materials and corresponding sound effects can be pre-built. The sample library contains original sound effect samples corresponding to various materials. For example, wood corresponds to sound effects of different actions such as knocking, dragging, and stepping; metal corresponds to sound effects of different actions such as collision and scratching. Thus, it is possible to select any sound effect corresponding to the material from the sample library based on the material and determine it as the sound effect segment corresponding to the target element.

[0065] As another implementation method, a material physical parameter library can be pre-established. For example, metals have high density, high hardness, and high sound absorption coefficient; wood has medium density, moderate hardness, and a higher sound absorption coefficient than metal. The server can obtain the corresponding physical parameters from the parameter library based on the material information in the element characteristics; convert the physical parameters into acoustic indicators that can be used for sound effect adjustment. For example, the fundamental frequency of the sound effect particles can be calculated using density and hardness. The higher the hardness and density, the higher the fundamental frequency is usually; the particle length can be calculated based on the sound absorption coefficient. The smaller the sound absorption coefficient, the slower the sound effect decays, and the longer the particle length; the vibration of the object can be simulated using the wave equation or the finite element method to determine the basic waveform of the element, and the basic waveform can be adjusted using acoustic indicators to generate the corresponding sound effect segment.

[0066] In the embodiments of this specification, the category can be pre-set category information used to distinguish object types or categories. The category can include broad categories, such as teaching aids, scenery, and animals; it can also include subcategories, such as teaching aids including books, desks, blackboards, and set squares; and scenery including fountains, waterfalls, trees, and landmarks. The object name information represented by the target element can be identified based on the region image; the corresponding category can be determined based on the object name information to obtain the category information of the target element. The target elements of each operation belong to the same level of category. For example, if element 1 in the region image is classified as trees, then element 2 is classified as a waterfall, not scenery.

[0067] As one implementation method, a sample library of element categories and their corresponding sound effects can be pre-built. This sample library can pre-collect the sounds emitted by each element in different actions within each element category. For example, a tree can correspond to the rustling sound of leaves rubbing together, or the sound of leaves colliding. A blackboard can correspond to the sound of friction, or the sound of knocking. After determining the category of the target element, the corresponding sound effect segment can be determined from this sample library.

[0068] As another implementation method, corresponding visual features can be identified based on element categories, and corresponding sound effect segments can be generated based on these visual features. For example, if the element category is a waterfall, the extracted visual features are rapid water flow and a large drop; the corresponding acoustic parameters can be determined based on these visual features, such as a fundamental frequency concentrated in the mid-low frequency range, high amplitude peaks, a gentle decay rate, and weak reverberation, generating the sound of water splashing. Similarly, if the element category is a tree, the extracted visual feature is large-scale leaf swaying; the corresponding acoustic parameters can be determined based on these visual features, such as a mid-low frequency base superimposed with irregular mid-frequency fluctuations, amplitude varying with the swaying of branches and leaves, and a moderate decay rate, generating the rustling sound of leaves rubbing together. A correspondence between element categories and visual feature types can be pre-established, and corresponding visual features can be extracted based on element categories to obtain corresponding acoustic parameters, and sound effect segments can be generated based on these acoustic parameters.

[0069] In practical applications, scene recognition can also be performed on images to obtain scene information of target elements. Based on the element features and scene information, the corresponding sound effect segments can be determined. In this way, the sound effect segments generated based on element features can conform to the scene information, improving the adaptability of the generated music to the usage scenario.

[0070] In practical applications, preset rules can be used to identify the element features of target elements. For example, the correspondence between the visual features of each element and its material can be preset, thereby determining the corresponding material based on the visual features obtained from the image. Alternatively, neural network models or large models can be used to identify the element features of target elements. Specifically, a model with the function of identifying materials or categories from images can be used to identify target elements and obtain element features. This can be set according to actual needs and is not a specific limitation.

[0071] Step 206: Generate a sound effect segment that matches the element features of the target element in the image based on the element features of any one operation, and obtain multiple sound effect segments.

[0072] In the embodiments of this specification, a sound effect clip can also be generated for a single operation, and this sound effect clip can be matched with the element characteristics of the element targeted by the operation. For example, if the element corresponding to a certain operation is a fountain, the sound effect clip can contain the sound of the fountain spraying water; if the element corresponding to a certain operation is a bird, the sound effect clip can contain the sound of birds chirping; if the element corresponding to a certain operation is a tree with branches and leaves, the sound effect clip can contain the sound of leaves rustling in the wind.

[0073] Multiple user actions on an image can generate multiple sound effect clips. In practical applications, the server or terminal device can store the correspondence between sound effect clips and element features. Based on this correspondence, the sound effect clip corresponding to the target element for each action can be determined. Alternatively, sound effect clips can be generated by processing element features using a model. For example, the model can match element features with sound effect features and output the sound effect clip corresponding to the matching sound effect features.

[0074] The aforementioned multiple sound effect segments can be generated one segment at a time according to the user's operation sequence. After generating the sound effect segment for one operation, a sound effect segment for the next operation is then generated. Alternatively, multiple sound effect segments can be generated in parallel or partially parallel. For example, after determining the element features corresponding to each operation, the process of generating sound effect segments based on element features can be executed synchronously or in parallel. Alternatively, the process of determining element features based on operations and the process of generating sound effect segments based on element features can be parallel. After identifying the element features corresponding to an operation, the process of generating sound effect segments is executed based on those element features. While executing the process of generating sound effect segments based on those element features, the process of identifying the element features corresponding to the next operation can be executed simultaneously, and so on, to obtain the sound effect segments corresponding to each operation.

[0075] In one or more embodiments of this specification, optionally, generating a sound effect segment that matches the element features of the target element corresponding to the image based on the element features of the target element in any one operation may include: determining a basic timbre that matches the element features based on the element features of the target element; and performing music art style transfer on the basic timbre to obtain the sound effect segment.

[0076] In the embodiments of this specification, the basic timbre can represent the sound emitted by an element or the sound produced by the vibration of an element after performing certain actions on it. It exhibits unique characteristics in terms of waveform; different objects vibrate with different characteristics. Different materials and structures produce different timbres.

[0077] Musical style transfer can refer to transferring the sound characteristics of a particular instrument or the characteristics of a particular music genre from the original basic timbre of the element. For example, the basic timbre of a glass is the sound of glass being struck; musical style transfer can refer to transferring the style of a harp from the sound of glass being struck to obtain a sound effect fragment. Music genres can include jazz, rock, country, ballads, and so on.

[0078] Different elemental characteristics can correspond to different basic timbres. For example, if the elemental characteristic is a material characteristic, then the basic timbres of different materials can be different.

[0079] Regarding the determination of basic timbres, specifically, a pre-established timbre library containing the correspondence between element features and basic timbres can be created. After determining the element features of the target element of the user's operation, basic timbres that have a mapping relationship with the element features can be obtained from the pre-established timbre library. Alternatively, a pre-trained model capable of generating timbres can be used to process the element features to obtain basic timbres that match the element features output by the model. The model can be a pre-trained neural network model or a pre-trained large model. If it is a large model, the server can generate corresponding prompts based on the extracted element features, and call the large model to generate basic timbres that match the element features based on the prompts, without requiring the user to provide corresponding prompts, thus improving the user experience. The embodiments in this specification can quickly and accurately obtain basic timbres that match element features through the above methods, facilitating subsequent music generation.

[0080] In the embodiments of this specification, based on the spectral characteristics contained in the basic timbre, an instrument with similar timbre characteristics to the spectral characteristics can be identified. The playing techniques, harmonic structure, dynamic expression, and other stylistic features of this instrument are then fused with the basic timbre to obtain a sound effect fragment with a new timbre that combines the characteristics of both the instrument and the target element. Thus, by fusing the musical characteristics of the instrument with the basic timbre, the sound effect fragment generated based on the target element can not only be distinguished from the instrument's conventional timbre, but also possess greater musicality and aesthetic appeal than the target element's own basic timbre, creating an innovative auditory experience and enhancing the artistry of the sound effect fragment.

[0081] In practical applications, artificial intelligence tools can be used to perform musical style transfer on basic timbres. For example, the Artist Brains neural style transfer tool can extract the texture and spectral characteristics of a sound and blend it with another instrument to obtain a new timbre. Other artificial intelligence tools, such as Combobulator, Time Domain Neural Audio Style Transfer, and Stability Audio 2.0, can also be used to perform musical style transfer on basic timbres. The appropriate tool can be selected based on actual needs; no specific limitations are made here.

[0082] In practical applications, the basic timbre of different elements can be transferred to the artistic style of the same instrument; conversely, the basic timbre of one element can also be transferred to the artistic style of different instruments. Based on user needs and the basic timbre of elements in different segments, musical artistic styles can be transferred from the basic timbre. For example, if the element is a glass and the user selects the "lyrical" tag, it can be determined that the user wants to generate lyrical music, thus transferring the musical style of the harp; if the user selects the "strong rhythm" tag, it can be determined that the user wants to generate music with a strong rhythm, thus transferring the musical style of the xylophone. If the element is a ceramic bowl and the user selects the "lyrical" tag, it can be determined that the user wants lyrical music, thus transferring the musical style of the violin.

[0083] In one or more embodiments of this specification, optionally, determining the basic timbre that matches the element features based on the element features of the target element may include: determining the basic timbre that matches the element features from a preset timbre library based on the element features of the target element; the preset timbre library includes multiple element features and corresponding audio information.

[0084] In the embodiments of this specification, the preset timbre library can be set based on expert experience. Different audio information can correspond to different frequency characteristics; frequency characteristics can represent the frequency composition of sound waves, such as the sound wave components of different frequencies and the intensity distribution of each frequency component. The preset timbre library can store the correspondence between element characteristics and audio information. Different element characteristics correspond to different audio information, for example, wood corresponds to 200-800 Hz; rain corresponds to 200-800 Hz. The server can use the audio information matching the element characteristics as the base timbre. If the element characteristic is lightning, the base timbre can be thunder of 50-200 Hz. Thus, the base timbre matching the element characteristics can be quickly obtained through the preset timbre library, and artistic sound effect segments can be generated based on the base timbre.

[0085] In practical applications, the frequency characteristics corresponding to the material can be determined based on the material features of the element, and the corresponding audio information can be determined based on the frequency characteristics. The correspondence between the audio information and the element features is then stored in a preset timbre library. For example, stainless steel plates are metals with high density, high hardness, and good elasticity. When struck, they produce a sound rich in high-frequency components and with rapid vibration decay. Oak wood has medium density, relatively weak elasticity, and a structure containing many fiber gaps. When struck, it produces a sound with more low-frequency components, fewer high-frequency components, and slow vibration decay. The timbre library can be obtained by pre-collecting the sounds generated by different actions performed on each element, such as friction sounds, knocking sounds, and falling sounds for glass; or it can be extracted from an existing sound library, and then a timbre library containing sound effect segments that correspond to the element features can be established.

[0086] In practical applications, large-scale models or pre-trained neural network models can be used to generate matching basic timbres based on elemental features. Both large-scale models and pre-trained neural network models are obtained by training with elemental feature samples and corresponding sample timbres. This can be achieved after one round of training or after multiple rounds of iterative training.

[0087] Specifically, for model training, feature samples of various elements and their corresponding standard timbres can be collected. Structured extraction of element features is performed; audio feature transformation is applied to the timbres; element feature vectors are bound one-to-one with their corresponding timbre feature vectors to form "element-timbre" sample pairs; outlier samples are removed from these pairs to obtain training samples, such as distorted timbres caused by recording errors or feature data so blurred that elements are unrecognizable; the model is then trained using these training samples to obtain a trained model capable of generating basic timbres. This allows the generation of basic timbres that match the element features, and the generation of artistically styled sound effect segments based on these basic timbres.

[0088] Step 208: According to the order in which the user performs the multiple operations, the multiple sound effect segments are spliced ​​together to obtain music that matches the image.

[0089] In the embodiments of this specification, the order may be determined by the server or terminal device based on the time of each operation performed by the user. Alternatively, when the user performs each operation, the server or terminal device may display a marker indicating the order of operations on the element corresponding to each operation, and determine the order of multiple operations based on the marker. For example, if the user performs the first operation on the textbook in the image, the server or terminal device marks the textbook as 1; if the user performs the second operation on the blackboard in the image, the server or terminal device marks the blackboard as 2, and so on.

[0090] In the embodiments of this specification, the server can determine the operation order corresponding to the sound effect segments, determine the splicing order of the sound effect segments, and splice the sound effect segments according to the splicing order. For example, the elements operated in operation order 1 and the corresponding sound effect segments are: trees - rustling sound; the elements operated in operation order 2 and the corresponding sound effect segments are: stream - babbling sound; the elements operated in operation order 3 and the corresponding sound effect segments are: birds - birdsong, then the splicing order is rustling sound - babbling sound - birdsong.

[0091] In one or more embodiments of this specification, a musical rhythm can be added to the sound effect segments to generate rhythmic music and enhance the artistic effect of the music. Optionally, the step of splicing the multiple sound effect segments to obtain music that matches the image may include: splicing the multiple sound effect segments to obtain a sound effect sequence; determining a matching musical rhythm based on the audio characteristics of each sound effect segment in the sound effect sequence; and merging the musical rhythm with the sound effect sequence to obtain music that matches the image.

[0092] In the embodiments of this specification, the sound effect sequence can be generated by splicing sound effect segments based on the order of operations performed on each target element in the image. The sound effect sequence is a complete sequence carrying the sound effect segments corresponding to all operations. Audio features are parameters that can characterize the acoustic properties of sound effect segments. These parameters mainly include time-domain features, frequency-domain features, and rhythm-related features. Time-domain features can include audio duration, loudness, peak amplitude, etc. Frequency-domain features can include frequency distribution such as the range of high, mid, and low frequencies, and spectral centroid such as the frequency region where audio energy is concentrated. Rhythm-related features can include beat such as the intensity pattern of the sound effect, and tempo such as the number of beats per minute. Musical rhythm can represent regular and periodic audio beats, which can give audio a sense of order and smoothness. If the rhythmic regularity represented by the audio features is strong, the server can add drum beats to the sound effect segment; if the rhythmic regularity represented by the audio features is weak, the server can add harmony to the sound effect segment.

[0093] In the embodiments of this specification, when splicing together various sound effect segments, transition effects such as fade-in / fade-out or volume gradients can be added at the joints of the segments to avoid abrupt transitions between segments.

[0094] Specifically, regarding the fade-in / fade-out method, sound effect segments can be stored in the form of sampling points. The amplitude of each sampling point determines the volume. The amplitude of the sampling points of the previous sound effect segment gradually decreases, while the amplitude of the sampling points of the next sound effect segment gradually increases. The total volume remains smooth after the two are superimposed. For example, the first segment (audio1) is 3 seconds long, the second segment (audio2) is 2 seconds long, and the transition duration is set to 0.1 seconds. The "end transition zone" of the first segment (e.g., the last 0.1 seconds of the sample) and the "beginning transition zone" of the second segment (e.g., the first 0.1 seconds of the sample) are taken. Linear interpolation is used to generate weighted weights. The weight of the first segment's transition zone decreases linearly from 1 to 0. For example, at sample point 4410, the first sample point has a weight of 1.0, and the 4410th sample point has a weight of 0.0, achieving a fade-out. The weight of the second segment's transition zone increases linearly from 0.0 to 1.0, achieving a fade-in. The weighted sum of the first segment's transition zone and the fade-out weights yields the first transition zone. The weighted sum of the second segment's transition zone and the fade-in weights yields the second transition zone. The first and second transition zones are added together to obtain the transition blend zone. The first segment (excluding the transition zone), the transition blend zone, and the second segment (excluding the transition zone) are then spliced ​​together to obtain the complete music. Volume transitions can be achieved by adjusting the volume of two segments, for example, the volume of the first segment goes from 0 to 100, and the volume of the second segment goes from 100 to 0. This will not be elaborated on further here.

[0095] The server can use audio processing tools to extract features from the sound effect sequence and obtain audio features. For determining the music rhythm, the server can obtain it from a pre-established music rhythm library. This library can contain pre-established music rhythms of different styles and parameters to meet the needs of different scenarios.

[0096] In the embodiments of this specification, for the fusion of music rhythm and sound effect sequence, the server can use the music rhythm as a background color and overlay the sound effect sequence on top of the rhythm to ensure that the sound effects are clearly distinguishable and the rhythm provides rhythmic support. Alternatively, the music rhythm can be dynamically adjusted based on the audio characteristics of the sound effect sequence; for example, when the loudness of the sound effect sequence increases, the rhythm intensity increases synchronously; when the duration of the sound effect sequence shortens, the rhythm beat accelerates. Alternatively, music rhythm segments can be inserted between the sound effect sequences to avoid interference between the sound effects and the rhythm, such as inserting a rhythmic melody between dialogue-type sound effect segments. This allows for the fusion of music rhythm and sound effect sequence to obtain music that matches the image, making the auditory effect of the music more pronounced and rhythmic.

[0097] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps and does not represent the only possible execution order. The order of some steps may be adjusted according to actual needs, or some steps may be omitted. When the claims involve method steps, changes in the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.

[0098] Figure 2 The method described above acquires multiple user actions performed on an image displayed on a terminal device. For any given action, it determines the element features of the target element in the image corresponding to that action. Based on these element features, it generates sound effect segments that match the target element's features, resulting in multiple sound effect segments. These segments are then concatenated according to the order in which the user performed the actions to generate music that matches the image. Users interact with the image displayed on the terminal device through actions such as clicking or double-clicking, and the server can generate music based on these interactions. This eliminates the need for users to possess advanced technical skills or specialized tools, allowing for simple image interaction to generate corresponding music. It lowers the barrier to entry for music generation, making it highly universal. Furthermore, the server can generate multiple music segments based on multiple user actions and concatenate them according to the user's action sequence to generate personalized music that matches the image, thus meeting user needs.

[0099] based on Figure 2 In addition to the method described herein, this specification also provides some improved implementation methods, which will be described below.

[0100] In one or more embodiments of this specification, the corresponding music rhythm and sound effect sequence can also be determined and fused based on the user's preferences. Optionally, the method may further include: obtaining user characteristic information; determining the user's preferred music rhythm based on the user characteristic information; and splicing the multiple sound effect segments to obtain music that matches the image, which may include: splicing the multiple sound effect segments to obtain a sound effect sequence; and adding the user's preferred music rhythm to the sound effect sequence to obtain music that matches the image.

[0101] In the embodiments of this specification, user feature information can be feature data reflecting attributes such as user identity, behavior, and interests. User features can include static features, dynamic behavioral features, and scene-related features. Static features can include features such as occupation and geographical location. Dynamic behavioral features can include historical music listening records, such as folk, punk, artistic, or classical music; audio and video creation habits, such as upbeat or steady rhythms; and interactive feedback features, such as collecting, sharing, or liking a certain type of rhythmic music.

[0102] In the embodiments of this specification, user characteristic information can be obtained from channels such as user registration information, preference questionnaire information, automatic recording of user behavior, or third-party platforms authorized by the user. The server can determine the user's preferred music rhythm based on preset rules using this user characteristic information. For example, a mapping rule between user characteristics and rhythm preferences can be pre-established, and the user's preferred music rhythm can be determined based on this mapping rule. The preset rules can be based on historical data or expert experience. For example, if a user's characteristic is "18-25 years old, with a historical listening history of 82% electronic music," then the user's preference for a strong and energetic music rhythm can be determined, and a mapping relationship between the two can be established. Similarly, if a user's characteristic is "50-65 years old, with a historical listening history of 70% drum music," then the user's preference for a weaker and more soothing music rhythm can be determined, and a mapping relationship between the two can be established. Thus, by using user characteristics, the user's preferred music rhythm can be determined, and then music can be generated based on the fusion of music rhythms and sound effect sequences that match the user's preferences. This improves the auditory effect of the music and its fit with the image, while also enhancing the user experience.

[0103] In one or more embodiments of this specification, a matching music style can also be determined based on scene information, making the generated music more suitable for the application scenario. Optionally, the method may further include: acquiring application scene information; the application scene information includes at least one of a scene representing children's education, a scene representing psychological healing, and a scene representing mass entertainment; determining a music style that matches the application scene information based on the application scene information; the music style includes at least one of music rhythm, harmony, drum beat position, and background sound.

[0104] Correspondingly, the above-mentioned method of splicing the multiple sound effect segments to obtain music that matches the image may include: splicing the multiple sound effect segments to obtain a sound effect sequence; and adding a music style that matches the application scenario information to the sound effect sequence to obtain music that matches the image.

[0105] In the embodiments of this specification, the application scenario information can represent the specific use case category served by the music generation. Different scenarios correspond to different music needs. For example, a children's education scenario can represent the need to generate music that suits children's auditory preferences, such as simple, cheerful, and rhythmically clear music, and can provide children with early cognitive and fun guidance. A psychological healing scenario can represent the need to generate music with calming auditory attributes, such as soothing and non-stimulating auditory attributes, and can help users relax and unwind. A mass entertainment scenario can represent the need to generate music with a wide range of applicability, such as cheerful, energetic, and lyrical attributes, and can serve as a form of leisure and atmosphere enhancement. Application scenarios can be selected by the user from the selectable scene labels displayed on the terminal device's display interface (such as controls for selecting scenes in the image display interface, or setting scenes through the terminal application's settings page, etc.); or, they can be the application's initial default scene, which the user can change using scene change controls; or, they can be based on image content recognition; or they can be based on user characteristics to determine the user's preferred scene as the application scenario, without any specific limitations here.

[0106] In the embodiments of this specification, music style can represent a comprehensive set of auditory characteristics of music. Music rhythm can represent the beat and tempo of music. Harmony can represent the sound effect based on the combination of multiple notes according to certain rules. Drum beat position can represent the distribution of percussion elements on the music timeline. Background sound can represent auxiliary audio to enhance the atmosphere. Different application scenarios correspond to different music styles. Thus, it is possible to determine the appropriate music style based on the obtained application scenario, and generate music that fits the functional requirements of the application scenario based on the music style, avoiding a disconnect between the music style and the scenario function. For example, a music style corresponding to mass entertainment cannot be used in a psychological healing scenario to avoid stimulating the user.

[0107] In the embodiments described in this specification, by adding a music style that matches the application scenario information to the sound effect sequence, the generated music not only fits the image but also serves the core function of the scenario, greatly improving the usability of the music and the user experience.

[0108] In one or more embodiments of this specification, the method may optionally include: displaying a visual acoustic particle effect in the terminal device.

[0109] In the embodiments of this specification, the visualized sound wave particle effect can be a visual effect that transforms the acoustic characteristics of audio into dynamic particle movement. The visualized sound wave particles can have various shapes, such as dots, light spots, lines, or irregular shapes; they can also have different types of colors, such as monochrome, gradient colors, or dynamic colors that change with the audio. The size of the visualized sound wave particles can be fixed, or it can be a dynamic size that scales according to the audio characteristics.

[0110] In practical applications, the state of visualized sound wave particles can change with the changes in audio characteristics. For example, visualized sound wave particles corresponding to low-frequency audio move slowly and are larger in size; visualized sound wave particles corresponding to high-frequency audio move rapidly and are smaller in size; visualized sound wave particles with heavy rhythm may exhibit a flickering state, and so on.

[0111] In practical applications, the area displaying the visual sound particles can be overlaid on the image, such as displaying them at the bottom or top of the image. Alternatively, the area displaying the visual sound particles can be separated from the image, such as displaying the image at the top and the visual sound particles at the bottom, or vice versa. Furthermore, the area displaying the visual sound particles can overlap with the image; for example, adjusting the transparency of the visual sound particles and overlaying them on top of the image. This allows users to visually perceive changes in audio while listening to music, enhancing the user experience and content appeal.

[0112] In one or more embodiments of this specification, the method may optionally include: for the multiple operations, the terminal device displays information indicating the order of operations at the area in the image where the user performs each operation.

[0113] In the embodiments of this specification, the information indicating the order of operations can be sequential Arabic numerals, such as 1, 2, 3, etc.; or sequential letters, such as a, b, c, etc.; or other sequential characters can be used for identification. The terminal device or server can sense and record the order in which the user clicks on each area, and then the terminal device can display the corresponding information indicating the order. Therefore, by displaying the information indicating the operation order in the image, the operation order is visually presented to the user, avoiding situations where the user forgets the operation order and has to repeat the operation, or repeatedly perform multiple operations on the same area.

[0114] In one or more embodiments of this specification, the method may optionally include: if a user performs a deletion operation on information indicating the order of operations for any one operation, then the information indicating the order of operations for any one operation displayed at the location of the operation is deleted from the area where the user performs the operation; and the sound effect segment corresponding to the operation is deleted.

[0115] In the embodiments of this specification, the deletion operation can be a preset operation performed by the user on information indicating the order of any operation, such as long-pressing, clicking, double-clicking, or swiping on displayed sequential numbers. Alternatively, after the user performs a long-press, click, double-click, or swipe operation, an operation menu pops up on the page, containing a deletion option. The user can then operate on the deletion control in the operation menu to complete the deletion operation. Alternatively, the deletion operation can also be performed by the user operating on an editing control displayed on the terminal device interface, entering an editing page, such as a pop-up window or overlay containing an editing box, where the user can delete information.

[0116] After a user performs a deletion operation, the server or terminal device can adjust the remaining information indicating the order of operations and display the adjusted order. For example, the adjusted order can be obtained according to the order before the deletion operation. If the user initially selected four items in the image, and then deleted the third selected item, the remaining three items can be arranged in a way that maintains the original order. For instance, the item selected the first time will still display the number indicating first, the item selected the second time will still display the number indicating second, but because the third selected item was deleted, the number indicating third will disappear and will no longer be displayed. The item selected the fourth time will then change from its original number indicating fourth to its original number indicating third.

[0117] In the embodiments of this specification, the deletion operation may be performed before the sound effect segment of the element corresponding to the deletion operation is generated, or it may be performed after the corresponding sound effect frequency band is generated.

[0118] In practical applications, users can also delete generated sound effect segments. For example, the user terminal's display page can visually display each sound effect segment, and the user can delete the visual information displayed on the page.

[0119] In one implementation, a user can perform a deletion operation on the information indicating the operation order displayed in the image. The server or terminal device can then adjust the remaining information indicating the operation order displayed in the image. Alternatively, the server or terminal device can choose not to adjust the remaining information indicating the operation order displayed in the image. During processing, sound effect segments can be generated for the elements marked with information indicating the operation order. When splicing the sound effect segments according to the operation order, the splicing requirement for the deleted operation can be skipped. For example, taking the numbers 1, 2, 3, and 4 as information indicating the operation order, 1, 2, 3, and 4 are displayed on elements a-b-c-d in the image, respectively. If the information indicating the operation order for element b is deleted, the remaining order can be determined as 1, 3, and 4. Furthermore, the sound effect segments corresponding to elements a-c-d in the image can be spliced ​​together.

[0120] As another implementation, each sound effect segment displays information indicating the order of operations. Users can delete the sound effect segments based on this information. After receiving the user's deletion request, the server can delete the sound effect segments corresponding to the order of operations. The server can then splice the remaining sound effect segments together according to the order of operations they correspond to, thus obtaining music.

[0121] After generating music, users can also delete music segments within a preset time period. The server can determine the part of the music before and after the deleted segment, and then splice the two parts together to obtain music that meets the user's needs. This allows users to delete elements that have already been deleted and obtain music that meets their requirements, improving the user experience.

[0122] In practical applications, if the deletion operation does not target the last operation in a series of operations, the server can adjust the information representing the order of the other operations, and then display the adjusted information on the terminal device so that the user can view the adjusted information. Continuing the example above, if the number 2 is deleted, the displayed number 3 can be adjusted to the number 2, and the number 4 can be adjusted to the number 3, resulting in the order 1, 2, 3, corresponding to elements a-element c-element d.

[0123] In one or more embodiments of this specification, to meet the user's personalized needs, the user can also adjust the operation sequence. Optionally, the method may further include: obtaining a sequence adjustment operation performed by the user based on displayed information indicating the order of each operation; determining the adjusted order of each of the target elements based on the sequence adjustment operation; and splicing the multiple sound effect segments according to the order in which the user performs the multiple operations to obtain music matching the image, which includes: splicing the sound effect segments corresponding to each of the target elements according to the adjusted order of each of the target elements to obtain music matching the image.

[0124] In the embodiments of this specification, the sequence adjustment operation can represent an operation that changes the information indicating the order of operations, so that after the sequence adjustment operation adjusts the information indicating the order of operations according to the order executed by the user, the page displays new information indicating the order of operations, and the new information indicating the order of operations is different from the original information indicating the order of operations.

[0125] Sequence adjustment operations can be performed by dragging and rearranging information indicating the order of operations. For example, numbers can represent the order: number 1 indicates the user performs the first operation on element 1 of the image; number 2 indicates the user performs the second operation on element 2; and number 3 indicates the user performs the third operation on element 3. If the user drags number 3 to element 1, numbers 1 and 3 can be swapped, resulting in the operation order of element 3-element 2-element 1. Alternatively, the order of element 3 can be moved forward while other orders remain unchanged, resulting in the operation order of element 3-element 1-element 2. The specific dragging and adjustment method is determined based on actual needs. Alternatively, sequence adjustment operations can also be performed by editing information indicating the order of operations. Continuing the previous example, number 1 corresponding to element 1 can be changed to number 2, and number 2 corresponding to element 2 can be changed to number 1. The adjustment process can display the adjustment effect; for example, dragging can show the effect of moving the information; modifying can show text boxes containing information and the cursor position for editing.

[0126] In the embodiments described in this specification, after adjusting the order, the server can splice the various sound effect segments according to the new order to obtain music that matches the image. Continuing the previous example, the user initially selected element 1-element 2-element 3 in that order, and later adjusted it to element 3-element 1-element 2. For each operation, assuming the sound effect segments generated in the original order are sound effect segment 1, sound effect segment 2, and sound effect segment 3; and after the order adjustment, they are sound effect segment 3, sound effect segment 1, and sound effect segment 2; then, the sound effect segments 3, sound effect segment 1, and sound effect segment 2 can be spliced ​​according to the new order to obtain the music. This allows the user to change the operation order without having to perform the operation again, improving the user experience.

[0127] In practical applications, if a user is not satisfied with the music generated in the adjusted order and prefers the music generated in the original order, they can perform an undo operation. Specifically, the undo control is displayed on the music generation interface before the music is generated, or on the interface containing the complete music after the music is generated.

[0128] In practical applications, the music playback interface of a terminal device can display a play button, allowing users to control the playback or pause of music. A synthesis progress bar can also be displayed, allowing users to visually track the music's progress. Specifically, the music playback interface can display an image as a background; alternatively, it can generate a matching cartoon character based on the music and image, displaying the cartoon character as a background.

[0129] In practical applications, terminal devices can also display controls for adding lyrics. Users can interact with these controls to input their own lyrics. The server can then use a large-scale model to process the music and lyrics, outputting a song with lyrics that perfectly match the music. The lyrics and music should have a unified emotional tone. The lyrics can be user-written, and users can choose whether to modify their own lyrics using optional controls. If the user chooses to modify, the large-scale model can adapt the original lyrics during music and lyric processing to make the modified lyrics more compatible with the music. If the user chooses not to modify, the large-scale model will use the original lyrics and blend them with the music to obtain the song.

[0130] In practical applications, since users often lack specialized knowledge, controls can be added to the lyrics for manipulation. Users can input words or sentences describing the requirements for generating lyrics. The server can then input this description into a large model, which can generate lyrics based on this information, integrate them with the music, and output a song containing both lyrics and music. The requirements for generating lyrics can be emotional terms, such as sad music, soothing music, or celebratory music; they can also be scene-related terms, such as wedding scenes, educational scenes, healing scenes, work scenes, homework scenes, or gaming scenes; and they can even be terms representing the objects described in the lyrics, such as paternal love, maternal love, friendship, the sea, or the sky.

[0131] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0132] According to the above explanation, Figure 3 This is a swimlane diagram of an image-based music generation method provided in the embodiments of this specification. For example... Figure 3 As shown, the process includes a user operation stage, a data processing stage, and a music generation stage. This explanation uses the user's interaction with the images displayed on the terminal device's screen as an example. The process may include:

[0133] The terminal device can display the image for which music needs to be generated based on the user's preset operation, and executes step 302: display the image.

[0134] Figure 4 This is a schematic diagram illustrating the display of an image on a terminal device, as provided in an embodiment of this specification. Figure 4 As shown, the image may include elements such as benches, fountains, and trees.

[0135] Images can be captured using a terminal device; or obtained from the terminal device's local storage; or obtained from the network; or provided by an application, etc.

[0136] Step 304: Obtain the multiple operations performed by the user on the image.

[0137] continue Figure 4 As shown in the image, assuming the user clicks on the tree the first time, the fountain the second time, and the bench the third time, information indicating the order of the operations can be displayed at the corresponding elements in the image. Figure 5 This is a schematic diagram illustrating multiple operations performed by a user on an image, provided as an embodiment of this specification. For example... Figure 5As shown, the terminal device's image can display information indicating the order of operations, such as labeling a tree as 1, a fountain as 2, and a bench as 3.

[0138] The server can provide user operation information, such as the location of the operation and the displayed image information. The server executes step 306: Based on the multiple operations performed by the user on the image, it determines the element features of the target element in the image corresponding to each operation.

[0139] by Figure 5 For example, we can determine that the element category of the target element in the first operation is tree, the element category of the target element in the second operation is fountain, and the element category of the target element in the third operation is bench, thus obtaining the element characteristics representing the category. Alternatively, we can determine that the element material of the target element in the first operation is wood, the element material of the target element in the second operation is water, and the element material of the target element in the third operation is wood, thus obtaining the element characteristics representing the material.

[0140] Step 308: Determine the basic timbre that matches the element characteristics based on the element characteristics of the target element.

[0141] A base timbre matching the element's characteristics can be determined from a preset timbre library. Specifically, the base timbre can be determined based on the element's category. If the image is blurry or the area clicked by the user is out of focus, and the category can only be identified as a tree, then a general tree sound, such as the rustling sound of leaves rubbing, can be used as the base timbre. If the specific category of the tree, such as an oak tree, can be identified, a deep creaking or groaning sound can be added to the rustling sound.

[0142] Step 310: Transfer the musical art style of the basic timbre to obtain sound effect fragments.

[0143] In the embodiments of this specification, a large model or a neural network model can be used to determine the musical art style that matches the basic timbre; or it can be a musical art style selected by the user. The musical art style can include styles representing music genres, such as jazz, classical, rock, etc.; or it can include styles representing instrument types, such as harp, guzheng, guitar, drum kit, etc.

[0144] like Figure 5As shown, after generating sound effect clips, the server can also display visual sound wave particle effects on the image displayed on the terminal device in a floating area with a certain degree of transparency. The corresponding sequence information is marked on the visual sound wave particle effects for each element. The floating area can also display a play button; after the user clicks the play button, the terminal device can play the various sound effect clips in sequence. The dotted lines in the floating area are for more clearly distinguishing the lengths of the sound effect clips corresponding to different elements; in practical applications, the terminal device may not display the dotted lines.

[0145] Step 312: Segment the various sound effect segments to obtain a sound effect sequence.

[0146] In the embodiments described in this specification, the sound effect segments of each target element can be spliced ​​together according to the operation order for each target element to obtain a sound effect sequence.

[0147] Step 314: Determine the matching music rhythm based on the audio characteristics of each sound effect segment in the sound effect sequence.

[0148] In the embodiments of this specification, the music rhythm preferred by the user can also be determined based on user characteristics, or the music rhythm can be determined based on information such as image scene and image style.

[0149] In this embodiment of the specification, the time information of the user's operation can also be obtained, and the corresponding festival can be determined based on the time information. Based on the atmosphere of the festival, the appropriate music rhythm for the festival can be determined. For example, if the festival is Qingming Festival, the music rhythm can be a low and sad rhythm; if the festival is Spring Festival, the music rhythm can be a high and cheerful rhythm.

[0150] Step 316: Combine the music rhythm with the sound effect sequence to obtain music that matches the image.

[0151] The terminal device can receive the music generated by the server and execute step 318: display the visual sound wave particle effects corresponding to the music.

[0152] In the embodiments described in this specification, visual sound wave particle effects that match the music can be generated and displayed on the terminal device based on the audio characteristics of the music, such as tempo and rhythm intensity. The sound effect segments can be generated after the user has performed all operations, or a sound effect segment corresponding to each operation can be generated, etc. The specific timing of the generation of the sound effect segments is not limited here.

[0153] Figure 6 This is a schematic diagram illustrating a music playback page displaying visual sound wave particle effects, as provided in an embodiment of this specification. Figure 6As shown, the page can display visual sound wave particle effects corresponding to the music; it can also display the name of the music, such as "Morning Music in the Park"; it can also display the creator, such as "Author: Ms. Wang"; it can also display an image as a background; and it can also display buttons to control the music to play or pause. Figure 6 The visual sound wave particle effects displayed in the image are... Figure 5 The visual sound wave particle effects displayed can be the same or similar.

[0154] Step 320: Play music based on user interaction.

[0155] like Figure 6 As shown, a music playback page can contain playback controls, allowing the terminal to play music based on the user's actions on these controls.

[0156] In the embodiments described in this specification, users can also adjust the order of operations, such as deletion, addition, or changing the order of operations, and can generate corresponding music based on the adjusted operations.

[0157] In the embodiments of this specification, some steps are the same as or similar to those in the foregoing embodiments. Please refer to the foregoing embodiments for further details.

[0158] Through these methods, users can generate music that matches images without needing professional music knowledge, simply by manipulating objects within the image. Furthermore, they can customize music to suit different application scenarios. For example, in a psychological healing scenario, natural sound effects can be generated; in a children's education scenario, clicking on historical war events in a textbook can generate historical war sound effects, and so on. This satisfies diverse user needs in different scenarios, improves user experience, lowers the barrier to music creation, and increases user enthusiasm for creation.

[0159] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.

[0160] Figure 7 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a device for generating music based on images.

[0161] like Figure 7 As shown, the device may include:

[0162] The operation acquisition module 702 is used to acquire multiple operations performed by the user on the image displayed on the terminal device;

[0163] The element feature determination module 704 is used to determine the element features of the target element corresponding to any one of the multiple operations in the image; the element features include at least one of material, category, position in the image, and scene information of the image;

[0164] The sound effect segment determination module 706 is used to generate a sound effect segment that matches the element features of the target element corresponding to the image in any one operation, thereby obtaining multiple sound effect segments.

[0165] The music generation module 708 is used to splice together the multiple sound effect segments according to the order in which the user performs the multiple operations to obtain music that matches the image.

[0166] based on Figure 7 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.

[0167] Optionally, the operation includes at least one of a single-click operation, a double-click operation, a long-press operation, and a selection operation; or, the image displayed by the terminal device includes at least one of an image scanned by the terminal device's camera, an image captured by the terminal device, and an image locally saved by the terminal device; or, the multiple operations include at least one of multiple operations by the user on a single frame of an image displayed by the terminal device and multiple operations by the user on multiple frames of an image displayed by the terminal device.

[0168] Optionally, the element feature determination module may be specifically used to: for any one of the multiple operations, perform material recognition on the region image at the operation position in the image based on the operation position of the operation, and determine the material information of the target element at the operation position; and / or, for any one of the multiple operations, perform category recognition on the region image at the operation position in the image based on the operation position of the operation, and determine the category information of the target element at the operation position.

[0169] Optionally, the sound effect segment determination module can be specifically used to: determine a basic timbre that matches the element characteristics based on the element characteristics of the target element; and perform music art style transfer on the basic timbre to obtain a sound effect segment.

[0170] Optionally, the sound effect segment determination module can be specifically used to: determine a basic timbre that matches the element features from a preset timbre library based on the element features of the target element; the preset timbre library includes multiple element features and corresponding audio information.

[0171] Optionally, the music generation module can be used to: splice the multiple sound effect segments to obtain a sound effect sequence; determine a matching music rhythm based on the audio characteristics of each sound effect segment in the sound effect sequence; and merge the music rhythm with the sound effect sequence to obtain music that matches the image.

[0172] Optionally, the device may include a user feature processing module, specifically used for: acquiring user feature information; determining the user's preferred music rhythm based on the user feature information; and splicing the multiple sound effect segments to obtain music matching the image, including: splicing the multiple sound effect segments to obtain a sound effect sequence; and adding the user's preferred music rhythm to the sound effect sequence to obtain music matching the image.

[0173] Optionally, the device may further include a scene processing module, specifically used for: acquiring application scene information; the application scene information includes at least one of a scene representing children's education, a scene representing psychological healing, and a scene representing mass entertainment; determining a music style matching the application scene information based on the application scene information; the music style includes at least one of music rhythm, harmony, drum beat position, and background sound; and splicing the multiple sound effect segments to obtain music matching the image includes: splicing the multiple sound effect segments to obtain a sound effect sequence; and adding a music style matching the application scene information to the sound effect sequence to obtain music matching the image.

[0174] Optionally, the terminal device displays a visual acoustic particle memory effect.

[0175] Optionally, for the multiple operations, the terminal device displays information indicating the order of operations in the area of ​​the image where the user performs each operation.

[0176] Optionally, the device may further include an operation adjustment module, which may be used to: if a user performs a deletion operation on information indicating the operation sequence of any operation, then delete the information indicating the operation sequence of any operation displayed at the location of the operation from the area where the user performs the operation; and delete the sound effect segment corresponding to the operation.

[0177] Optionally, the operation adjustment module can be specifically used to: obtain the order adjustment operation performed by the user based on the displayed information indicating the order of each operation; determine the adjusted order of each of the target elements based on the order adjustment operation; and splice the multiple sound effect segments according to the order in which the user performs the multiple operations to obtain music that matches the image, including: splicing the sound effect segments corresponding to each of the target elements according to the adjusted order of each of the target elements to obtain music that matches the image.

[0178] Optionally, the operation acquisition module may be used to: acquire multiple operations performed by the user on the image displayed on the terminal device's display interface; or, acquire video information containing the user's selection operations on several physical objects; and parse the multiple operations performed by the user on each physical object from the video information.

[0179] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.

[0180] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0181] The above is an illustrative scheme of an image-based music generation device according to this embodiment. It should be noted that the technical solution of this image-based music generation device and the technical solution of the image-based music generation method described above belong to the same concept. For details not described in detail in the technical solution of the image-based music generation device, please refer to the description of the technical solution of the image-based music generation method described above.

[0182] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0183] Figure 8 A structural block diagram of a computing device provided according to an embodiment of this specification is shown.

[0184] The computing device 800 includes:

[0185] Memory 810 and processor 820;

[0186] The memory 810 is used to store computer programs / instructions, and the processor 820 is used to execute the computer programs / instructions, which, when executed by the processor 820, implement the steps of the image-based music generation method.

[0187] Specifically, the components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and the database 850 is used to store data.

[0188] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0189] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0190] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.

[0191] The processor 820 executes the computer instructions to implement the steps of the image-based music generation method.

[0192] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the image-based music generation method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the image-based music generation method described above.

[0193] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the image-based music generation method described above.

[0194] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the image-based music generation method described above. Details not described in detail in the technical solution of the storage medium can be found in the description of the image-based music generation method described above.

[0195] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for generating music based on images.

[0196] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described method for generating music based on images belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described method for generating music based on images.

[0197] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments of apparatus, devices, media, and products, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The apparatus, devices, media, and products provided in the embodiments of this specification correspond to the methods; therefore, the apparatus, devices, media, and products also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, devices, media, and products will not be repeated here.

[0198] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0199] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0200] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0201] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0202] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0203] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0204] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0205] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0206] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0207] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0208] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0209] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0210] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0211] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating music based on an image, comprising: obtaining a plurality of operations performed by a user on an image displayed by a terminal device; determining, for each operation of the plurality of operations, an element feature of a target element in the image corresponding to the operation; the element feature comprises at least one of a material, a category, a position in the image, and scene information of the image; generating, according to the element feature of the target element in the image corresponding to the operation, a sound effect segment matching the element feature, to obtain a plurality of sound effect segments; splicing the plurality of sound effect segments in a sequence in which the user performs the plurality of operations, to obtain music matching the image.

2. The method of claim 1, wherein the operation comprises at least one of a single-click operation, a double-click operation, a long-press operation, and a circle selection operation; or the image displayed by the terminal device comprises at least one of an image scanned by a camera of the terminal device, an image captured by the terminal device, and an image stored locally by the terminal device; or the plurality of operations comprises at least one of a plurality of operations performed by the user on one frame of the image displayed by the terminal device, and a plurality of operations performed by the user on a plurality of frames of the image displayed by the terminal device.

3. The method of claim 1, wherein the determining, for each operation of the plurality of operations, an element feature of a target element in the image corresponding to the operation comprises: for each operation of the plurality of operations, performing material recognition on a region image at an operation position of the operation in the image according to the operation position, to determine material information of the target element at the operation position; and / or for each operation of the plurality of operations, performing category recognition on the region image at the operation position of the operation in the image according to the operation position, to determine category information of the target element at the operation position.

4. The method of claim 1, wherein the generating, according to the element feature of the target element in the image corresponding to the operation, a sound effect segment matching the element feature comprises: determining, according to the element feature of the target element, a basic tone matching the element feature; and performing music artistic style transfer on the basic tone to obtain the sound effect segment.

5. The method of claim 4, wherein the determining, according to the element feature of the target element, a basic tone matching the element feature comprises: determining, according to the element feature of the target element, a basic tone matching the element feature from a preset tone library; the preset tone library comprises a plurality of element features and corresponding audio information.

6. The method of claim 1, wherein the splicing the plurality of sound effect segments to obtain music matching the image comprises: splicing the plurality of sound effect segments to obtain a sound effect sequence; determining a matching music rhythm according to audio features of each sound effect segment in the sound effect sequence; and fusing the music rhythm with the sound effect sequence to obtain music matching the image. ​ ​ ​ 7. The method of claim 1, further comprising: obtaining user feature information of the user; determining a user preferred music rhythm according to the user feature information; wherein the splicing the plurality of sound effect segments to obtain the music matching the image comprises: splicing the plurality of sound effect segments to obtain a sound effect sequence; adding the user preferred music rhythm to the sound effect sequence to obtain the music matching the image.

8. The method of claim 1, further comprising: obtaining application scenario information; wherein the application scenario information comprises at least one of a scenario representing children education, a scenario representing psychological healing, and a scenario representing mass entertainment; determining a music style matching the application scenario information according to the application scenario information, wherein the music style comprises at least one of a music rhythm, a harmony, a drum position, and a background sound; wherein the splicing the plurality of sound effect segments to obtain the music matching the image comprises: splicing the plurality of sound effect segments to obtain a sound effect sequence; adding the music style matching the application scenario information to the sound effect sequence to obtain the music matching the image.

9. The method of claim 1, further comprising: displaying a visual sound wave particle special effect in the terminal device.

10. The method of claim 1, wherein for the plurality of operations, the terminal device displays information representing an operation sequence of each operation in a region where the user performs each operation in the image.

11. The method of claim 10, further comprising: if a deletion operation performed by the user on the information representing the operation sequence of any operation is obtained, deleting the information representing the operation sequence of the any operation displayed at the region where the user performs the any operation; deleting a sound effect segment corresponding to the any operation.

12. The method of claim 10, further comprising: obtaining a sequence adjustment operation performed by the user based on the displayed information representing the operation sequence of each operation; determining an adjusted sequence of each of the target elements based on the sequence adjustment operation; wherein the splicing the plurality of sound effect segments to obtain the music matching the image according to the sequence in which the user performs the plurality of operations comprises: splicing a sound effect segment corresponding to each of the target elements according to the adjusted sequence of each of the target elements to obtain the music matching the image.

13. The method of claim 1, wherein the obtaining the plurality of operations performed by the user on the image displayed by the terminal device comprises: obtaining a plurality of operations performed by the user on an image displayed by a display interface of the terminal device; or, obtaining video information containing a selection operation performed by the user on a plurality of real objects; obtaining a plurality of operations performed by the user on each of the real objects from the video information.

14. An apparatus for generating music based on an image, comprising: an operation obtaining module configured to obtain a plurality of operations performed by a user on an image displayed by a terminal device. An element feature determination module is configured to determine, for any one of the multiple operations, an element feature of a target element corresponding to the any one of the multiple operations in the image, wherein the element feature comprises at least one of a material, a category, a position in the image, and scene information of the image. An audio effect segment determination module is configured to generate an audio effect segment matched with the element feature according to the element feature of the target element corresponding to the any one of the multiple operations in the image, to obtain multiple audio effect segments. A music generation module is configured to splice the multiple audio effect segments according to a sequence in which the user performs the multiple operations, to obtain music matched with the image.

15. A computing device, comprising: a memory and a processor; the memory is configured to store a computer program or instructions, and the processor is configured to execute the computer program or instructions, and the computer program or instructions, when executed by the processor, implement the steps of the method in any one of claims 1 to 13.