Music dialogue method and system, medium, computing device and program product

By obtaining the target keywords of the image in the music dialogue system to generate search terms, recalling songs that meet the description, solving the problem of insufficient dialogue requirements for image input in the prior art, and realizing personalized music recommendations.

CN120296198APending Publication Date: 2025-07-11HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323600.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing music dialogue system enters pictures, it fails to effectively meet the user's dialogue needs and cannot effectively use image content for song recommendations.

Method used

By obtaining the target keywords of the input image, generating the target search terms, recalling songs that match the description from the music library, providing personalized music recommendations.

Benefits of technology

It realizes music that matches image content based on user input image recommendations, broadens the application scenarios of music recommendations, simplifies the complexity of image content, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296198A_ABST
    Figure CN120296198A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a music dialogue method and system, a medium, computing equipment and a program product. The method comprises the following steps: for each round of dialogue in a dialogue interface, acquiring an input image input in the round of dialogue under the condition that the type of a content problem input in the round of dialogue is an image; and determining a target keyword associated with the input image, and generating a target search word based on the target keyword. And displaying songs recalled from a music library in an output reply of the dialogue interface, wherein the songs accord with the description of the target search word.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology. More specifically, embodiments of the present disclosure relate to a music dialogue method, system, medium, computing device, and program product. Background Art

[0002] This section aims to provide background or context for embodiments of the present disclosure. The descriptions herein are not admitted to be prior art merely because they are included in this section.

[0003] Conversational AI is a broad field that includes various applications such as voice assistants, chatbots, and intelligent customer service. Users can ask questions, give commands, or express needs using voice or text, and the conversational AI interprets and responds based on semantic understanding technology.

[0004] Related technologies innovatively expand a music dialogue system based on existing conversational AI. It combines traditional dialogue systems with specific requirements in the music field to provide users with a more convenient and intelligent way to interact in music. Summary of the Invention

[0005] However, existing methods mainly focus on the need for playlist recommendations for text questions input by users. When a user inputs a picture in the music dialogue system, existing methods still adopt the processing method of traditional dialogue systems, that is, presenting the content interpretation of the input picture to the user. In the music field, this processing method fails to effectively meet the dialogue needs of the user's input picture.

[0006] Therefore, there is a great need for an improved music dialogue method to effectively meet the dialogue needs of users' input pictures.

[0007] In this context, embodiments of the present disclosure are expected to provide a music dialogue method, system, medium, computing device, and program product.

[0008] In a first aspect of the embodiments of the present disclosure, a music dialogue method is provided. The method includes:

[0009] For each round of dialogue in the dialogue interface, when the type of the content question input in this round of dialogue is an image, obtain the input image input in this round of dialogue;

[0010] Determine the target keyword associated with the input image, and generate a target retrieval term based on the target keyword;

[0011] Display in the output reply of the dialogue interface the songs recalled from the music library that meet the description of the target retrieval term.

[0012] Optionally, determining the target retrieval term associated with the input image includes: determining a target keyword associated with the input image, and generating a target retrieval term based on the target keyword.

[0013] Optionally, determining the target keyword associated with the input image includes: when there is a face in the input image, extracting the face feature from the input image, and determining the target keyword based on the face feature; when there is no face in the input image, extracting the global feature of the input image, and determining the target keyword based on the global feature.

[0014] Optionally, extracting the face feature from the input image and determining the target keyword based on the face feature includes: if the number of faces is greater than 1, extracting the face feature of each face from the input image respectively, and determining the corresponding target keyword based on each face feature respectively.

[0015] Optionally, determining the target keyword based on the face feature includes: retrieving the registered face feature similar to the face feature from the face feature library, and obtaining the face keyword corresponding to the registered face feature; the face feature library stores the mapping relationship between the registered face feature, the registered face feature and the corresponding face keyword; determining the target keyword based on the global feature includes: retrieving the keyword feature similar to the global feature from the keyword feature library, and obtaining the image keyword corresponding to the keyword feature; the keyword feature library stores the mapping relationship between the keyword feature, the keyword feature and the corresponding image keyword.

[0016] Optionally, the method further includes: obtaining the operation behavior data of the output reply for any round of conversation, and increasing the weight of the keyword used in any round of conversation in the keyword library when the operation behavior data indicates positive feedback; determining the target keyword associated with the input image, and generating a target retrieval term based on the target keyword includes: obtaining the target keyword associated with the input image from the keyword library, and generating a target retrieval term based on the sorted target keyword; wherein, the greater the weight of the target keyword, the more forward the sorting.

[0017] In the second aspect of the embodiments of the present disclosure, a music dialogue system is provided, and the system includes:

[0018] A content input module, configured to obtain the input content, and the type of the content includes text and image;

[0019] An image semantic understanding module, configured to obtain the input image input by the content input module and determine a target retrieval term associated with the input image when the type of the input content is an image;

[0020] An intent recognition module, configured to obtain the text content input by the content input module and recognize the problem intent of the text content, and call a response module that matches the problem intent to reply to the text question when the type of the input content is text; when the type of the input content is an image, obtain the target retrieval term determined by the image semantic understanding module, and call a response module for recommending songs to reply to the target retrieval term;

[0021] A response module, in response to the call of the intent recognition module, outputs a corresponding response.

[0022] Optionally, determining the target retrieval term associated with the input image includes: determining a target keyword associated with the input image, and generating a target retrieval term based on the target keyword.

[0023] Optionally, the image semantic understanding module includes: an image keyword generation sub-module, configured to extract the global features of the input image and determine image keywords based on the global features; a face keyword generation sub-module, configured to extract the face features of the input image and determine face keywords based on the face features; a retrieval term generation sub-module, configured to obtain the image keywords and the face keywords; when the face keywords exist, generate a target retrieval term based on the face keywords, and when the face keywords do not exist, generate a target retrieval term based on the image keywords.

[0024] Optionally, the retrieval term generation sub-module is further configured to generate a guiding term based on the keywords for display in the response of the current round of conversation; wherein the keywords include face keywords and image keywords.

[0025] In a third aspect of the embodiments of the present disclosure, a medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the music dialogue method described in any one of the embodiments in the first aspect is implemented.

[0026] In a fourth aspect of the embodiments of the present disclosure, a computing device is provided, including:

[0027] A processor;

[0028] A memory for storing processor-executable instructions;

[0029] Among them, the processor realizes the music dialogue method described in any embodiment of the first aspect above by running the executable instructions.

[0030] In the fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program and / or instructions, and when the computer program and / or instructions are executed by a processor, the music dialogue method described in any embodiment of the first aspect above is realized.

[0031] According to the music dialogue method of the embodiments of the present disclosure, when the content input by the user is an image, a target retrieval term associated with the input image can be determined, and a song that conforms to the description of the target retrieval term can be displayed in the output reply.

[0032] In this way, first of all, in the music field, a recommendation mechanism for recommending music that matches the image content according to the image input by the user is provided, which broadens the application scenarios of music recommendation. Secondly, this solution converts the content of the input image into a text-type target retrieval term, and recalls songs from the music library based on the target retrieval term. By converting the image content into a target retrieval term, it is easier to express the core features of the image content and simplify the complexity of the image content. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where:

[0034] Figure 1 Schematically shows a schematic diagram of a music dialogue interface.

[0035] Figure 2 Schematically shows a flowchart of a music dialogue method.

[0036] Figure 3 Schematically shows a processing framework diagram of picture semantic understanding.

[0037] Figure 4 Schematically shows another schematic diagram of a processing framework for picture semantic understanding.

[0038] Figure 5 Schematically shows a schematic diagram of a music dialogue interface for a user to input a portrait image.

[0039] Figure 6 Schematically shows a schematic diagram of a music dialogue interface for a user to input a landscape image.

[0040] Figure 7 Schematically shows a schematic diagram of a music dialogue system.

[0041] Figure 8 Schematically shows a schematic diagram of a music dialogue interface with poems and guiding words in the reply.

[0042] Figure 9 Schematically shows a schematic diagram of a medium.

[0043] Figure 10 Schematically shows a block diagram of a music dialogue device.

[0044] Figure 11 Schematically shows a schematic diagram of a computing device.

[0045] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners

[0046] The principles and spirit of the present disclosure will be described below with reference to several exemplary implementation manners. It should be understood that these implementation manners are provided only to enable those skilled in the art to better understand and then implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these implementation manners are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0047] Those skilled in the art know that the implementation manners of the present disclosure can be implemented as a system, a device, an apparatus, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0048] According to the implementation manners of the present disclosure, a music dialogue method, a system, a medium, a computing device, and a program product are provided.

[0049] In this article, the number of any element in the drawings is used for illustration rather than limitation, and any naming is only used for distinction and does not have any limiting meaning.

[0050] The principles and spirit of the present disclosure will be elaborated below with reference to several representative implementation manners of the present disclosure. Summary of the Invention

[0052] The inventors found that the existing methods mainly focus on the need for playlist recommendation for text problems input by users. When a user inputs a picture in a music dialogue system, the existing methods still adopt the processing manner of a traditional dialogue system, that is, presenting the content interpretation of the input picture to the user, and in the music field, this processing manner fails to effectively meet the dialogue needs of the user's input picture.

[0053] To solve the above problems, the present disclosure provides a music dialogue method, system, medium, computing device, and program product. The music dialogue method includes: for each round of dialogue in the dialogue interface, when the type of the content problem input in this round of dialogue is an image, obtain the input image input in this round of dialogue. Determine the target keyword associated with the input image, and generate a target retrieval term based on the target keyword. Display the songs recalled from the music library in the output reply of the dialogue interface, and the songs conform to the description of the target retrieval term.

[0054] In this way, first of all, in the music field, a new recommendation mechanism for recommending music that matches the image content according to the user input image can be provided, which broadens the application scenarios of music recommendation. Secondly, this solution converts the input image into a text-type target retrieval term, and recalls songs from the music library based on the target retrieval term. By converting the image content into a target retrieval term, it is easier to express the core features of the image content and simplify the complexity of the image content.

[0055] After introducing the basic principles of the present disclosure, the following specifically introduces various non-limiting implementation manners of the present disclosure.

[0056] Overview of Application Scenarios

[0057] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0058] Figure 1 Schematically shows a schematic diagram of a music dialogue interface. As Figure 1 shown, an exemplary music dialogue interface 10 opened on an electronic device is shown. The music dialogue interface 10 is provided with an input box 14, and the user can input information such as pictures, texts, and voices in the input box 14. The music assistant 13 responds to the information input by the user and can present the response result in the music dialogue interface 10. Among them, the interaction scenarios between the user and the music assistant 13 include but are not limited to song recommendation, knowledge Q&A, and chatting, etc. For example, the user can input "Recommend some songs suitable for running in the morning" in the input box 14, and the music assistant 13 can respond to the question input by the user and output the songs in the morning running scenario in the music dialogue interface 10. For another example, the user can input "Basic information of ** singer" in the input box 14, and the music assistant 13 can respond to the question input by the user and output the basic information of the ** singer in the music dialogue interface 10. For another example, the user can also have some chat interactions with the music assistant 13 that are not limited to the music field.

[0059] The present disclosure does not limit the above-mentioned some conventional Q&A functions that the music assistant 13 can implement. The present disclosure also provides a design in which, when the content input by the user is only an image, the music assistant 13 makes music recommendations for the input image, and this design can be integrated with the Q&A function shown above into the music assistant 13. According to the music dialogue method of the embodiments of the present disclosure, when the type of the content input in the current round of dialogue is an image, obtain the input image input in the current round of dialogue. Determine the target keyword associated with the input image, and generate a target search term based on the target keyword. Display the song recalled from the music library in the output reply of the dialogue interface, and the song conforms to the description of the target search term.

[0060] In practical applications, when the type of the content input by the user in the music dialogue interface 11 is an image, the music assistant 13 can output a song that matches the content of the input image. For example, when the content of the input image 11 includes mountains and waters, the theme of the song 12 output by the music assistant 13 can include mountains and waters. Another example is that when the content of the input image 11 contains the head portrait of a singer, the author of the song 12 output by the music assistant 13 can be the singer. Through image input, the user does not need to describe with complex language or text. The music assistant 13 can extract key information from the image and automatically infer the user's song-finding needs, reducing the burden on the user's input.

[0061] Exemplary Method

[0062] The following combines Figure 1 the application scenario of Figure 2 to describe a music dialogue method according to an exemplary embodiment of the present disclosure. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0063] Figure 2 Schematically shows a flowchart of a music dialogue method. As Figure 2 shown, it may include S201 - S203:

[0064] S201: When the type of the content input in the current round of dialogue is an image, obtain the input image input in the current round of dialogue.

[0065] S202: Determine the target keyword associated with the input image, and generate a target search term based on the target keyword.

[0066] S203: Display the song recalled from the music library in the output reply of the dialogue interface, and the song conforms to the description of the target search term.

[0067] In an exemplary application scenario, a user can have multi-turn conversations in a music conversation interface. For any turn of the multi-turn conversations, the voice assistant will perform corresponding response operations for the questions or instructions input by the user. Among them, the current turn of conversation described in this method can be any turn of the multi-turn conversations.

[0068] Suppose the user inputs a picture of a small town street scene full of a retro style in the current turn of conversation, then a target retrieval term associated with the picture can be determined. For example, the target retrieval term can be "Recommend retro-style songs suitable for strolling slowly under ancient buildings". The target retrieval term can be a high-level understanding and expression of the content of the input image. By extracting the theme, emotion, and scene information of the input image, these abstract features are transformed into specific music recommendation requirements.

[0069] Specific objects of the input image, such as "buildings", "streams", "vehicles", etc., can be extracted, and the target retrieval term can be constructed based on the extracted specific objects. It is also possible to analyze the emotion contained in the image based on the visual features of the input image, generate an emotion label, and construct the target retrieval term based on the emotion label. This specification does not limit the method of determining the target retrieval term associated with the input image.

[0070] Songs are recalled from the music library based on the target retrieval term, and the recalled songs can conform to the description of the target retrieval term. For example, the recalled song can be "jazz", whose melody and arrangement style are retro and conform to the "retro" style described by the target retrieval term; for example, it can also include "piano music", whose rhythm characteristics conform to the "leisurely" scene described by the target retrieval term. This specification does not limit the method of recalling songs according to the target retrieval term.

[0071] This embodiment provides a recommendation mechanism for recommending music that matches the image content according to the image input by the user. In the field of music recommendation, it can meet the user's need to obtain music recommendations through simple image input.

[0072] In one embodiment, when determining the target retrieval term associated with the input image, a target keyword associated with the input image can be determined, and the target retrieval term can be generated based on the target keyword.

[0073] Exemplarily, when determining the target keywords associated with the input image, an image classification model can be trained. This image classification model can assign different scene labels to the input image, and these scene labels can be used as the target keywords. Exemplarily, the OCR technology can be used to extract the text of the image from the input image, and the extracted text can be used as the target keywords. By extracting the text in the input image and generating the target retrieval term based on the extracted text, the generated target retrieval term can be more in line with the connotation of the image and can capture the intention of the user's input image more accurately. Exemplarily, the input image can be analyzed by a sentiment analysis model and the sentiment keywords can be generated as the target keywords. Of course, the above exemplary methods can be combined to generate the target keywords that synthesize sentiment, scene, and image text. This specification does not limit the method of determining the target keywords associated with the input image.

[0074] In one embodiment, when determining the target keywords associated with the input image, the global features of the input image can be extracted, and the target keywords can be determined based on the global features, and the target retrieval term can be generated based on the target keywords. Specifically, as Figure 3 shown, the input image can be input into the feature extraction module 30. The feature extraction module 30 outputs the image features to the keyword generation module 31. The keyword generation module 31 outputs the target keywords determined according to the image features to the retrieval term generation module 32, and the retrieval term generation module generates the target retrieval term according to the obtained target keywords.

[0075] In another embodiment, when determining the target keywords associated with the input image, when there is a face in the input image, the face features can be extracted from the input image, and the target keywords can be determined based on the face features. When there is no face in the input image, the global features of the input image are extracted, and the target keywords are determined based on the global features.

[0076] In the actual application scenario, when the user inputs an image containing a face, it usually means that they hope to recommend music related to the person in the image. For example, if the user uploads a photo of a singer, the target keyword can be the name of the singer, and thus the target retrieval term determined based on the target keyword can retrieve the music related to the singer.

[0077] When the user uploads an image without a face, such as images of scenery, buildings, street scenes, etc., the user's needs are usually related to the context or atmosphere expressed by the image. These images are usually taken casually by the user, and the user hopes that the music assistant can recommend music that matches the current environment or activity according to the visual content of the image (such as scene, sentiment, etc.).

[0078] The target keyword determined based on the facial features can be the name of the singer that matches the face determined through face recognition technology. The target keyword determined based on the global features of the input image can be a keyword that matches the image scene.

[0079] In this embodiment, if the image contains a face, recommendations related to the person are preferred; if the image does not contain a face, recommendations related to the scene are preferred. Specifically, when determining the target keyword, by classifying the user's needs into two categories according to whether the image contains a face and determining the corresponding target keyword respectively, the target keyword can be more in line with the user's current personalized needs.

[0080] In one embodiment, when there is a face in the input image, the facial features are extracted from the input image, and the target keyword is determined based on the facial features. Exemplarily, if the number of faces is greater than 1, the facial features of any one face can be extracted from the input image, and the target keyword is determined based on the facial features of the any one face. Among them, the any one face can be the face with the largest area and / or located in the middle area among all the faces in the input image, which is beneficial to focusing on the main person in the input image and reducing the interference of irrelevant people. For example, the user uploads a group photo of a music concert. In the photo, multiple singers are standing on the stage. If singer A is in the center of all the singers or the facial area of singer A is significantly larger than that of other singers, then singer A may be the focus of this concert, that is, the user's favorite singer A.

[0081] Exemplarily, if the number of the faces is greater than 1, the facial features of each face can be extracted from the input image respectively, and the corresponding target keywords are determined based on the facial features of each face. Still taking the group photo of the music concert as an example, assuming that the facial features of singer A, singer B, and singer C are extracted from the input image respectively, the corresponding target keywords determined based on the facial features of each face can be the names of singer A, B, and C. The target retrieval term generated based on the target keyword can contain the names of singer A, B, and C at the same time. For example, "Recommend songs of singer A, B, and C".

[0082] In this embodiment, by the above method, songs of all the people in the recommended input image can be generated, improving the comprehensiveness of the recommended content and avoiding missing the user's preferences.

[0083] In one embodiment, when extracting face features from an input image, a face region can be extracted from the input image, and the face region can be input into a feature extraction model to extract the face features of the face region. Compared with the method of directly inputting the input image into the feature extraction model for face feature extraction, the present disclosure only extracts the face region in the input image and inputs the face region into the feature extraction model, enabling the feature extraction model to focus on facial features and avoiding interference from other elements such as the background of the input image, thereby improving the accuracy of facial feature extraction.

[0084] Exemplarily, the RetinaFace model is used to extract the face region in the input image and output the face region to the ArcFace model, and the ArcFace model is used to extract the face features of the face region in the input image.

[0085] Specifically, an input image containing a face is obtained, and the input image is normalized to adjust the pixel values of the image to a specific range, such as [0, 1] or [-1, 1], to eliminate differences in brightness, contrast, etc. between different images. At the same time, according to the input requirements of the RetinaFace model, the image is scaled in size and adjusted to a fixed size suitable for model processing, such as [112, 112] pixels.

[0086] Load the pre-trained RetinaFace model, which is based on a deep learning architecture and includes multiple convolutional layers, pooling layers, and specific multi-task learning modules. Input the pre-processed image data into the RetinaFace model, and the model extracts the feature information in the image through convolutional operations to form feature maps at different levels. These feature maps contain different feature representations of the face in the image, from low-level edge and texture features to high-level semantic features.

[0087] Based on the extracted feature maps, the RetinaFace model simultaneously performs face box prediction, facial key point localization, and preliminary analysis of face attributes (such as pose, expression, etc.) through a multi-task learning mechanism. In the face box prediction task, the model predicts the possible positions and bounding box sizes of the face in the image; for facial key point localization, the model determines the specific coordinate positions of the key parts of the face (such as eyes, nose, mouth, etc.). Through the collaborative processing of these multi-tasks, precise spatial localization of the face is achieved.

[0088] Post-process and optimize the face localization results output by the model. Use the non-maximum suppression algorithm to remove the predicted results of face bounding boxes with too high overlap, and retain the most representative face bounding boxes. At the same time, according to the position information of the facial key points, fine-tune the position and angle of the face bounding box to further improve the accuracy of face localization. Finally, output the localization results including the accurate face position, the coordinates of the facial key points, and the relevant face attribute information, providing a high-quality data basis for subsequent face recognition, analysis, and other operations.

[0089] After successfully localizing the face, enter the stage of portrait feature extraction. During this process, the algorithm can use the facial features as anchor points to extract portrait features. This is because the facial features are the most distinguishable parts of the face and can effectively represent a person's facial characteristics. By accurately extracting the features such as the shape, relative position, and proportion of the facial features, a feature vector that can represent this face is formed.

[0090] Obtain the face image extracted by the RetinaFace model. First, perform normalization processing on it, mapping the image pixel values to a specific range, such as [0,1], to eliminate the influence caused by differences in illumination, contrast, etc. between different images. Then, according to the input requirements of the ArcFace model, adjust the size of the image. For example, uniformly adjust it to 112×112 pixels. At the same time, to enhance the generalization ability of the model, data augmentation operations such as random flipping and rotation can be performed on the image.

[0091] Build an ArcFace model based on a deep convolutional neural network (CNN). This model usually includes multiple components such as convolutional layers, pooling layers, and fully connected layers. The convolutional layer extracts the local features of the image through convolutional kernels, the pooling layer is used to reduce the data dimension, and the fully connected layer integrates the extracted features. Subsequently, load the pre-trained ArcFace model parameters on a large-scale face dataset. These parameters contain rich face feature information.

[0092] Input the preprocessed face image into the loaded ArcFace model. The model gradually extracts the low-level to high-level features of the image through forward propagation in the convolutional layer and pooling layer, from basic features such as edges and textures to high-level features with semantic information. After the fully connected layer, use the Additive Angular Margin Loss (AAM loss) function for optimization. This loss function introduces an angular margin between the feature vector and the classification boundary, making the features learned by the model more compact for the same class and more dispersed for different classes in the feature space. Finally, the model outputs a highly discriminative face feature vector, which can effectively represent the unique features of the face.

[0093] The extracted face feature vectors are normalized to have a unified scale and distribution, further enhancing the stability and comparability of the features.

[0094] In one embodiment, when determining target keywords based on face features, registered face features similar to the face features can be retrieved from a face feature library, and face keywords corresponding to the registered face features can be obtained. The face feature library stores the mapping relationship between registered face features, registered face features, and corresponding face keywords.

[0095] The face feature library may store registered face features of celebrities and the mapping relationship between registered face features and corresponding face keywords. This mapping relationship can be used to find face keywords corresponding to the registered face features, where the face keywords can be the names of the celebrities. Specifically, for the same celebrity, multiple facial photos taken from different angles can be collected, and the registered face features of the multiple facial photos can be extracted and stored in the face feature library. Over time, for newly emerging celebrities, their registered face features can also be added to the face feature library. This incremental update can ensure that the registered face feature library can adapt to the changing celebrity environment.

[0096] When retrieving registered face features similar to the face features from the face feature library, the face features extracted from the input image can be calculated for similarity with the registered face features stored in the face feature library, and the top k registered face features with the highest similarity can be extracted. Based on the mapping relationship between these top k registered face features and face keywords, the top k face keywords are determined, and the target retrieval term is determined based on the top k face keywords.

[0097] Through the above method, the person in the image can be associated with a specific celebrity, and then songs related to the celebrity can be recommended based on the target retrieval term. For example, songs sung by the celebrity, songs themed on the celebrity, etc.

[0098] In one embodiment, when determining target keywords based on global features, keyword features similar to the global features can be retrieved from a keyword feature library, and image keywords corresponding to the keyword features can be obtained. The keyword feature library stores the mapping relationship between keyword features, keyword features, and corresponding image keywords.

[0099] Among them, the keyword feature library can be the corresponding keyword features extracted for each image keyword included. The sources of the included image keywords can cover multiple fields such as music, film and television, figures, and culture. For example, in the music field, the image keywords can include various music styles (such as classical, pop, rock, etc.), the names of famous musicians and bands. In the film and television field, there are the names of popular movies and TV dramas, directors and actors, etc. In the figure field, it includes the names of celebrities from all walks of life. In the culture field, there are relevant words such as various cultural phenomena and art genres. The above-mentioned keywords that are included can be stored in the keyword library. Through this extensive collection and collation of keywords, the keyword library can provide comprehensive retrieval clues for different types of pictures and rich materials for accurately finding music related to the pictures.

[0100] When retrieving keyword features similar to the global feature from the keyword feature library, the similarity between the global feature and each keyword feature in the keyword feature library can be calculated, and the top k keyword features with the highest similarity can be obtained. The value of k can be defined according to actual needs. When obtaining the image keyword corresponding to the keyword feature, the corresponding image keyword can be obtained from the keyword library according to the mapping relationship between the keyword feature and the image keyword.

[0101] In one embodiment, the present disclosure also provides an embodiment of constructing a keyword library:

[0102] Obtain a publicly available picture dataset. For each picture in the picture dataset, generate a text description for each picture, and the text description is used to describe the content included in the picture. For example, for a picture of a beach, a text description such as "There are several people sunbathing on the beautiful beach, and the waves are hitting the shore" may be generated.

[0103] Obtain the text descriptions of all pictures and convert the text descriptions into candidate keywords. For example, a large language model can be used to generate topic words from the text descriptions as candidate keywords. For example, topic words such as "beach", "sun", and "waves" are generated from "There are several people sunbathing on the beautiful beach, and the waves are hitting the shore" as candidate keywords. It should be noted that the above process of extracting topic words may not be simply extracting keywords from the text description, but topic words that can be generated for the content summary of the text description.

[0104] In order to improve the quality and accuracy of the keywords, the candidate keywords can be screened, and the screened keywords can be stored in the keyword library. Among them, the screening conditions can be to screen out unqualified candidate keywords such as sensitive words, words that deviate too much from the music recommendation scenario, and incomplete words.

[0105] In this embodiment, text descriptions are generated for the pictures in the publicly available picture dataset, and the topic words are extracted from the text descriptions as keywords. Compared with the method of extracting words from a large number of web pages, this method is more efficient. Moreover, by generating keywords from the text descriptions of pictures, the generated keywords can be more centered around the theme scope of the pictures.

[0106] In addition, based on the above embodiment, it is also possible to obtain the words used by the user during the interaction with the music assistant and supplement the words into the keyword library to ensure the dynamic update of the keyword library.

[0107] In one embodiment, it is also possible to obtain the operation behavior data of the output response for any round of conversation. When the operation behavior data indicates positive feedback, increase the weight of the keywords used in the any round of conversation in the keyword library. When determining the target keyword associated with the input image and generating the target retrieval term based on the target keyword, the target keyword associated with the input image can be obtained from the keyword library, and the target retrieval term is generated based on the sorted target keywords. Among them, the greater the weight of the target keyword, the more forward it is sorted. Of course, it is also possible to reduce the weight of the keywords used in the any round of conversation in the keyword library when the operation behavior data indicates negative feedback.

[0108] In an exemplary application scenario, after the user inputs an image in any round of conversation, the output response can be a playlist that matches the input image. The operation behavior data for the user's response to the output can be that the user clicks on the playlist recommended in the output response, such as playing any song in the playlist recommended in the output response; the user performs an evaluation operation on the recommended result of the output response, such as clicking like, dislike, favorite, etc. In the case where the operation behavior data indicates positive feedback, the weight of the keywords used in any round of conversation in the keyword library can be increased. For example, in the output response of this round of conversation, if the user plays the song recommended in the output response or performs a positive feedback operation on the recommended result of the output response, such as clicking like or favorite, then the weight of the target keyword used in this round of conversation in the keyword library can be increased. For instance, if the playlist corresponding to the keyword "outdoor adventure music" is selected by the user more frequently, the weight of this keyword in the keyword library will gradually increase. Of course, in the case where the operation behavior data indicates negative feedback, the weight of the keywords used in any round of conversation in the keyword library can be decreased. For example, in the output response of this round of conversation, if the user does not click on the song recommended in the output response but starts the next round of conversation to continue asking the music assistant to recommend songs, or performs a negative feedback operation on the recommended result of the output response, such as clicking dislike, then the weight of the target keyword used in this round of conversation in the keyword library can be decreased. When generating the target retrieval term based on the target keyword, the target keywords can be sorted from largest to smallest according to the weight of each target keyword, and then the target retrieval term can be generated based on the sorted target keywords, so that when recalling songs from the music library, the songs corresponding to the keywords with greater weight are recalled first.

[0109] In this embodiment, dynamically adjusting the weight of keywords in the keyword library based on the user's feedback on the response result can enable the songs corresponding to the keywords with greater weight to be recalled first when recalling music from the music library based on the keywords. In other words, the songs corresponding to the keywords that the user is more interested in are recalled first.

[0110] In one embodiment, Figure 4 Schematically shows another schematic diagram of the processing framework for picture semantic understanding. As Figure 4As shown, when the type of the content input in this round of conversation is an image, the input image input in this round of conversation can be obtained, and the input image is input into the feature extraction module 30. The feature extraction module 30 includes a global feature extraction module 301 and a face feature extraction module 302. Specifically, the input image is respectively input into the global feature extraction module 301 and the face feature extraction module 302. Among them, the global feature extraction module 301 is used to extract the global features of the input image, while the face feature extraction module 302 is used to extract face features from the input image. The global feature extraction module 301 inputs the global features into the image keyword generation module 311. The image keyword generation module 311 is used to retrieve image keyword features similar to the global features from the image keyword feature library, obtain the image keywords corresponding to the image keyword features from the image keyword library, and input the image keywords into the retrieval term generation module 32. The face keyword generation module 312 is used to retrieve the registered face features similar to the face features from the face feature library, obtain the face keywords corresponding to the registered face features from the face keyword library, and input the face keywords into the retrieval term generation module 32. If the retrieval term generation module 32 obtains both face keywords and image keywords at the same time, a target retrieval term is generated based on the face keywords; if the retrieval term generation module 32 only obtains image keywords, a target retrieval term is generated based on the image keywords. Exemplarily, the global feature extraction module 301 and the face feature extraction module 302 can adopt different feature extraction models. For example, the global feature extraction module 301 can adopt the CLIP model, while the face feature extraction module 302 can adopt the RetinaFace model and the ArcFace model shown in the foregoing embodiments. Of course, other feature extraction models can also be adopted, and this specification does not impose any restrictions on this.

[0111] It should be noted that Figure 4 The storage method is only exemplary. The face feature library and the image keyword feature library can be the same feature library, that is, this feature library stores both face features and image keyword features at the same time. Similarly, the face keyword library and the image keyword library can also be the same keyword library, that is, this keyword library stores both face keywords and image keywords at the same time. Of course, other storage methods can also be adopted, such as storing image features and keywords in the same database at the same time. The present disclosure does not impose any restrictions on the storage methods of face features and keywords.

[0112] The processing framework given in this embodiment first determines the face keywords and image keywords of the input image in parallel, which can take into account the overall environment of the image and the character information to generate retrieval terms. If the image contains character information, it is assumed that the user is more inclined to retrieve songs related to that character; if the image does not contain characters, it is assumed that the user is more inclined to retrieve songs related to the environmental atmosphere of the image. This parallel processing method can flexibly adjust the song recommendation strategy according to different image contents, making the recommendation results more accurately respond to the personalized needs of the pictures uploaded by users. Secondly, a global feature extraction module 301 and a face feature extraction module 302 are respectively set for the feature extraction module 30, and the fine-grained semantic extraction of the face can also be optimized separately in the face feature extraction module 302, so as to improve the accuracy of face feature extraction. Finally, the processing framework given in this embodiment also satisfies the advantage of pluggable update of picture modality understanding. Those skilled in the art can add / modify processing sub-modules in different modules on the basis of this processing framework without affecting the working process of the overall framework. For example, other face feature extraction models can be replaced in the face feature extraction module 302. Another example is to add new feature extraction logic in the global feature extraction module 301 or the face feature extraction module 302, such as adding an OCR extraction sub-module in the global feature extraction module 301 to extract the text information in the input image.

[0113] Next, it will be described in combination with actual application scenarios:

[0114] Figure 5 Schematically shows a schematic diagram of a music dialogue interface for a user to input a portrait image. As Figure 5 shown, in the music dialogue interface 50, the user inputs a portrait image 51 through the input box 54, and the reply 52 output by the music assistant 53 contains songs related to the person recognized from the portrait image 51, such as the songs sung by this person.

[0115] Figure 6 Schematically shows a schematic diagram of a music dialogue interface for a user to input a landscape image. As Figure 6 shown, in the music dialogue interface 60, the user inputs a landscape image 61 through the input box 64, and the reply 62 output by the music assistant 63 contains songs related to the content of the landscape image 61. If the input image is a landscape painting, most of the themes of the recommended songs are **mountains and waters**, **looking at the sea**, **great rivers**, **ink painting**, etc.

[0116] Combined with Figure 5 and Figure 6 's embodiments, it can be seen that the song recommendation results generated by this disclosure can flexibly adapt to different types of image contents and adjust the recommendation strategy accordingly, so as to provide personalized recommendations for users.

[0117] Figure 7 Schematically shown is a schematic diagram of a music dialogue system. As Figure 7 shown, the system includes:

[0118] A content input module 71 for obtaining the input content, and the types of the content include text and images. Among them, the text can be the original text input by the user or the text converted from the voice input by the user.

[0119] An image semantic understanding module 3, which is used to obtain the input image input by the content input module 71 and determine the target keyword associated with the input image when the type of the input content is an image. The exemplary processing framework logic of the image semantic understanding module 3 can refer to the foregoing Figure 3 and Figure 4 .

[0120] An intention recognition module 72, which is used to obtain the text content input by the content input module 71 and recognize the problem intention of the text content when the type of the input content is text, and call the reply module 73 that matches the problem intention to reply to the text question. When the type of the input content is an image, obtain the target retrieval term determined by the image semantic understanding module 3, and call the reply module 73 for recommending songs to reply to the target retrieval term.

[0121] The reply module 73 outputs a corresponding reply in response to the call of the intention recognition module 72. In an actual application scenario, the reply module may include agents for different purposes. For example, a knowledge Q&A agent, a song-finding agent, a chatting agent, etc. When the intention recognition module 72 recognizes that the problem intention of the text content is a knowledge Q&A, the knowledge Q&A agent can be called to reply to the text question. When the intention recognition module 72 obtains the target retrieval term determined by the image semantic understanding module 3, the song-finding agent can be called to reply to the target retrieval term, specifically, songs that meet the description of the target retrieval term can be recalled from the music library and displayed in the output reply.

[0122] In an embodiment, when determining the target retrieval term associated with the input image, the image semantic understanding module 3 may determine the target keyword associated with the input image and generate the target retrieval term based on the target keyword.

[0123] In one embodiment, the image semantic understanding module 3 may specifically include an image keyword generation sub-module for extracting the global features of the input image and determining image keywords based on the global features; a face keyword generation sub-module for extracting the face features of the input image and determining face keywords based on the face features; and a retrieval term generation sub-module for obtaining the image keywords and the face keywords, and generating a target retrieval term based on the face keywords when the face keywords exist, and generating a target retrieval term based on the image keywords when the face keywords do not exist.

[0124] In one embodiment, the retrieval term generation sub-module is further configured to generate a guiding term based on the keywords for display in the current conversation. The keywords include face keywords and image keywords. In this embodiment, by generating a guiding term based on the keywords, clearer guidance can be provided to the user to guide the user to further interact with the music dialogue system.

[0125] In one embodiment, the retrieval term generation sub-module is further configured to generate a poem based on the keywords for display in the current conversation. The keywords include face keywords and image keywords. In this embodiment, generating a poem based on the keywords enables the expression of the poem to revolve around the content of the image.

[0126] Figure 8 Schematically shows a schematic diagram of a music dialogue interface with a poem and a guiding term in the reply. As Figure 8 shown, after the user inputs an image in the input box, the music assistant outputs a reply 82 that matches the content of the input image. And a poem generation area 821 is also displayed in the reply 82. The poem generation area 821 is used to display the poem generated according to the keywords used in the current reply. For example, assuming that the content of the image input by the user describes a garden scene, and the corresponding generated keywords are flowers, sunlight, and gentle breeze, the poem generated based on these keywords can be "The fragrance of flowers fills the prosperous scene, whispers softly, and the gentle spring blooms in the melody." Another example, assuming that the content of the image input by the user is the headshot of a **singer, and the corresponding generated keyword is the **singer's name, the poem generated based on the **singer's name can be "Listen to the voice of the **singer, and every sentence is the flowing of emotions."

[0127] A guiding term 84 can also be displayed in the reply 82. For example, the guiding term 84 can be "Take a look at the content related to [**]". Since the guiding term 84 is generated based on the keywords, it is related to the image content, which to a certain extent ensures that the guiding direction does not deviate from the content of the image input by the user, and improves the relevance between the guiding term and the input image.

[0128] Exemplary Medium

[0129] After introducing the methods of the exemplary embodiments of the present disclosure, next, reference will be made to Figure 9 describe the media of the exemplary embodiments of the present disclosure.

[0130] In this exemplary embodiment, the above method can be implemented by a program product. For example, a portable compact disc read-only memory (CD-ROM) can be used and includes program code, and this memory can run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable medium 90 can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.

[0131] The program product can adopt any combination of one or more readable media. The readable medium 90 can be a readable signal medium or a readable medium. The readable medium 90 can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable medium include: an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0132] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component.

[0133] The program code contained on the readable medium 90 can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0134] Program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the C language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0135] Exemplary Device

[0136] After introducing the medium of the exemplary embodiments of the present disclosure, next, reference is made to Figure 10 illustrate the device of the exemplary embodiments of the present disclosure. Regarding the following device, the specific manner in which each functional module performs operations and the specific functions achieved after performing the operations have been described in detail in the foregoing embodiments of the music dialogue method, and will not be elaborated herein again.

[0137] Figure 10 A block diagram of a music dialogue device according to an embodiment of the present disclosure is schematically shown. The music dialogue device may include:

[0138] An image acquisition module 1002, configured to acquire an input image input in the current round of dialogue when the type of the content problem input in the current round of dialogue is an image.

[0139] A target retrieval term generation module 1004, configured to determine a target keyword associated with the input image and generate a target retrieval term based on the target keyword.

[0140] A reply output module 1006, configured to display a song recalled from a music library in the output reply of the dialogue interface, and the song conforms to the description of the target retrieval term.

[0141] Optionally, the target retrieval term generation module 1004 is specifically configured to determine a target keyword associated with the input image and generate a target retrieval term based on the target keyword.

[0142] Optionally, the target retrieval term generation module 1004 is specifically configured to extract face features from the input image and determine target keywords based on the face features when a face exists in the input image; and extract global features of the input image and determine target keywords based on the global features when no face exists in the input image.

[0143] Optionally, if the number of faces is greater than 1, the target retrieval term generation module 1004 is specifically configured to extract face features of each face from the input image and determine corresponding target keywords based on the face features of each face.

[0144] Optionally, the target retrieval term generation module 1004 is specifically configured to retrieve registered face features similar to the face features from a face feature library and obtain face keywords corresponding to the registered face features; the face feature library stores the mapping relationship between the registered face features, the registered face features and the corresponding face keywords. Retrieve keyword features similar to the global features from a keyword feature library and obtain image keywords corresponding to the keyword features; the keyword feature library stores the mapping relationship between the keyword features, the keyword features and the corresponding image keywords.

[0145] Optionally, the device further includes a keyword weight update module 1008, configured to obtain operation behavior data of an output reply for any round of conversation, and increase the weight of the keywords used in the any round of conversation in a keyword library when the operation behavior data indicates positive feedback. The target retrieval term generation module 1004 is specifically configured to obtain target keywords associated with the input image from the keyword library and generate a target retrieval term based on the sorted target keywords; wherein, the greater the weight of the target keyword, the more forward it is sorted.

[0146] Exemplary Computing Device

[0147] After introducing the methods, media and devices of the exemplary embodiments of the present disclosure, next, reference is made to Figure 11 to describe the computing device of the exemplary embodiments of the present disclosure.

[0148] Figure 11 The computing device 110 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0149] Such as Figure 11As shown, the computing device 110 is presented in the form of a general-purpose computing device. The components of the computing device 110 may include, but are not limited to: at least one of the above-mentioned processing units 1101, at least one of the above-mentioned storage units 1102, and a bus 1103 that connects different system components (including the processing unit 1101 and the storage unit 1102).

[0150] The bus 1103 includes a data bus, a control bus, and an address bus.

[0151] The storage unit 1102 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 11021 and / or cache memory 11022, and may further include a readable medium in the form of non-volatile memory, such as read-only memory (ROM) 11023.

[0152] The storage unit 1102 may also include a program / utility 11025 having a set (at least one) of program modules 11024. Such program modules 11024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0153] The computing device 110 may also communicate with one or more external devices 1104 (such as a keyboard, a pointing device, etc.).

[0154] Such communication may be carried out through an input / output (I / O) interface 1105. Also, the computing device 110 may further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1106. As Figure 11 shown, the network adapter 1106 communicates with other modules of the computing device 110 through the bus 1103. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computing device 110, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0155] It should be noted that although several units / modules or sub-units / modules of the music dialogue device are mentioned in the above detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned units / modules may be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above may be further divided and embodied by multiple units / modules.

[0156] Moreover, although the operations of the method of the present disclosure are depicted in the drawings in a particular order, this is not a requirement or an implication that the operations must be performed in that particular order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.

[0157] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. Such division is only for convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A music dialogue method, characterized in that The method includes: When the type of the content input in the current round of conversation is an image, obtaining the input image input in the current round of conversation; Determining a target retrieval term associated with the input image; Displaying in the output response a song recalled from a music library that meets the description of the target retrieval term.

2. The method according to claim 1, wherein The determining a target retrieval term associated with the input image includes: Determining a target keyword associated with the input image and generating a target retrieval term based on the target keyword.

3. The method according to claim 2, wherein The determining a target keyword associated with the input image includes: When there is a human face in the input image, extracting the human face feature from the input image and determining a target keyword based on the human face feature; When there is no human face in the input image, extracting the global feature of the input image and determining a target keyword based on the global feature.

4. The method according to claim 3, wherein The extracting the human face feature from the input image and determining a target keyword based on the human face feature includes: If the number of human faces is greater than 1, extracting the human face feature of each human face from the input image respectively and determining the corresponding target keyword based on each human face feature respectively.

5. The method according to claim 3, wherein The determining a target keyword based on the human face feature includes: Retrieving from a human face feature library a registered human face feature similar to the human face feature and obtaining a human face keyword corresponding to the registered human face feature; the human face feature library stores a mapping relationship between the registered human face feature, the registered human face feature and the corresponding human face keyword; The determining a target keyword based on the global feature includes: Retrieving from a keyword feature library a keyword feature similar to the global feature and obtaining an image keyword corresponding to the keyword feature; the keyword feature library stores a mapping relationship between the keyword feature, the keyword feature and the corresponding image keyword.

6. The method according to claim 2, wherein The method further includes: Obtaining operation behavior data of the output response for any round of conversation, and when the operation behavior data indicates positive feedback, increasing the weight of the keyword used in the any round of conversation in the keyword library; The determining a target keyword associated with the input image and generating a target retrieval term based on the target keyword includes: Obtaining from the keyword library a target keyword associated with the input image and generating a target retrieval term based on the sorted target keyword; wherein, the greater the weight of the target keyword, the more forward the sorting.

7. A music dialogue system, characterized in that, The system includes: A content input module for obtaining the input content, and the type of the content includes text and image; An image semantic understanding module for, when the type of the input content is an image, obtaining the input image input by the content input module and determining a target retrieval term associated with the input image; An intent recognition module, configured to, when the type of the input content is text, obtain the text content input by the content input module, recognize the problem intent of the text content, and call a response module that matches the problem intent to respond to the text problem; when the type of the input content is an image, obtain the target retrieval term determined by the image semantic understanding module, and call a response module for recommending songs to respond to the target retrieval term; A response module, in response to the call of the intent recognition module, outputs a corresponding response.

8. A medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of claims 1-6 is implemented.

9. A computing device, comprising: A processor; A memory for storing instructions executable by the processor; wherein, the processor implements the steps of the method described in any one of claims 1-6 by running the executable instructions.

10. A computer program product comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, the steps of the method described in any one of claims 1-6 are implemented.