Music recommendation method and apparatus

By acquiring user viewpoint information and attention duration, and determining attention patterns, music can be accurately recommended, solving the problem of low music matching in complex environments and improving user experience.

CN114930319BActive Publication Date: 2025-12-12HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080092641.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-31
Publication Date
2025-12-12
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

Existing music recommendation methods cannot accurately match users' interests in complex environments, resulting in low music matching accuracy.

Method used

By acquiring users' viewpoint information and attention duration, we can determine their attention patterns, accurately identify the content they are paying attention to, and thus recommend suitable music.

Benefits of technology

It improves the matching accuracy of music recommendations, ensuring that the recommended music matches the things and behaviors that users are truly interested in, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930319B_ABST
    Figure CN114930319B_ABST
Patent Text Reader

Abstract

A music recommendation method and device determine the attention mode of a user in a complex environment through the viewpoint information of the user, and more accurately match music. In a first aspect, a music recommendation method is provided, which includes: receiving visual data of a user (S501); obtaining at least one attention unit and an attention duration of the at least one attention unit according to the visual data (S502); determining an attention mode of the user according to the attention duration of the at least one attention unit (S503); and determining recommended music information according to the attention mode (S504).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and more particularly, to a music recommendation method and device. BACKGROUND

[0002] Personalized music recommendation technology can improve the user's music experience. The traditional method is to recommend music based on the user's historical music playing information through data mining technology. This method cannot take into account the user's current state information. Some current methods can collect the user's current state information through different sensors, such as sensing environmental information, including location, weather, time, season, environmental sound and environmental picture information, to recommend related music; or by measuring the user's current state information, such as by collecting brain waves to analyze the user's current psychological state, or collecting the pictures the user sees, or obtaining the user's heart rate, etc. to recommend related music.

[0003] In the current method, music is recommended after the image seen by the user is collected by shooting, involving the matching process of music and image. In a real scene, the environment may contain many scenes, and if music is recommended only according to the overall image, the music matching degree is reduced. SUMMARY

[0004] The present application provides a music recommendation method and device, which determines the user's attention mode in a complex environment through the user's viewpoint information, and more accurately matches music.

[0005] In a first aspect, a music recommendation method is provided, the method comprising: receiving visual data of a user; obtaining at least one attention unit and an attention duration of the at least one attention unit according to the visual data; determining an attention mode of the user according to the attention duration of the at least one attention unit; and determining recommended music information according to the attention mode.

[0006] The music recommendation method of the present application embodiment can more accurately determine the user's attention content according to the user's visual information, thereby recommending more suitable music, so that the recommended music conforms to the user's real interests and conforms to the user's real behavior state, improving the user's use experience.

[0007] In combination with the first aspect, in a possible implementation manner of the first aspect, the visual data comprises viewpoint information of the user and picture information viewed by the user, and the viewpoint information comprises a position of a viewpoint and an attention duration of the viewpoint.

[0008] With reference to the first aspect, in a possible implementation form of the first aspect, the at least one attention unit and the attention duration of the at least one attention unit are acquired according to the visual data, including: acquiring the at least one attention unit according to the picture information; and acquiring a sum of the attention durations of the viewpoints in the at least one attention unit as the attention duration of the at least one attention unit.

[0009] The music recommendation method provided in the embodiments of the present application determines an initial attention unit according to acquired picture information, and determines the duration of each attention unit according to viewpoint information of the user. Compared with the prior art of recommending music only according to the entire picture viewed by the user, the viewpoint information can accurately indicate the location of the attention content interested by the user, so that the recommended music can be more in line with the needs of the user.

[0010] With reference to the first aspect, in a possible implementation form of the first aspect, the at least one attention unit and the attention duration of the at least one attention unit are acquired according to the visual data, further including: judging a similarity between a first attention unit and a second attention unit in the at least one attention unit, the first attention unit and the second attention unit being attention units at different time instants; and if the similarity is greater than or equal to a first threshold value, the attention duration of the second attention unit is equal to a sum of the attention duration of the first attention unit and the attention duration of the second attention unit.

[0011] In the music recommendation method provided in the embodiments of the present application, the first attention unit and the second attention unit can be attention units in frame images at different time instants within a preset time, or can be attention units in a historical library and newly acquired attention units respectively.

[0012] With reference to the first aspect, in a possible implementation form of the first aspect, the attention mode of the user is determined according to the attention duration of the at least one attention unit, including: if a standard deviation of the attention duration of the at least one attention unit is greater than or equal to a second threshold value, determining that the attention mode of the user is staring; and if the standard deviation of the attention duration of the at least one attention unit is less than the second threshold value, determining that the attention mode of the user is scanning.

[0013] With reference to the first aspect, in a possible implementation form of the first aspect, the music information is determined according to the attention mode, including: if the attention mode is scanning, determining the music information according to the picture information; and if the attention mode is staring, determining the music information according to an attention unit with the highest attention degree in the attention unit.

[0014] In the music recommendation method of the embodiments of the present application, after the attention mode of the user is determined, the music information suitable for being recommended to the user in the preset time period can be determined according to the attention mode of the user in the preset time period. When the attention mode of the user is scanning, it is considered that the user mainly perceives the environment in the preset time period, and the music can be recommended according to the picture information (environment); when the attention mode of the user is staring, it is considered that the user mainly perceives the interesting thing in the preset time period, and the music can be recommended according to the attention unit (interesting thing) with the highest attention degree.

[0015] With reference to the first aspect, in a possible implementation of the first aspect, the music information is determined according to the attention mode, and the method further includes: determining the behavior state of the user at each time in the first time period according to the attention mode; determining the behavior state of the user in the first time period according to the behavior state at each time; and determining the music information according to the behavior state in the first time period.

[0016] In the music recommendation method of the embodiments of the present application, after the attention content is determined according to the attention mode of the user in a preset time period, the music information can not be determined first, but the behavior state of the user in the preset time period is determined, and then the total behavior state of the user in the first time period is determined according to the behavior state in a plurality of preset time periods, so that the actual behavior state of the user can be more accurately determined, and the music is recommended according to the total behavior state, so that the recommended music is more consistent with the actual behavior state of the user.

[0017] The second aspect provides a device for music recommendation, which includes: a transceiver module configured to receive visual data of a user; a determination module configured to acquire at least one attention unit and an attention duration of the at least one attention unit according to the visual data; the determination module is further configured to determine an attention mode of the user according to the attention duration of the at least one attention unit; and the determination module is further configured to determine recommended music information according to the attention mode.

[0018] The embodiments of the present application provide a device for music recommendation, which is used to implement the music recommendation method in the first aspect.

[0019] With reference to the second aspect, in a possible implementation of the second aspect, the visual data includes viewpoint information of the user and picture information viewed by the user, and the viewpoint information includes a position of a viewpoint and an attention duration of the viewpoint.

[0020] With reference to the second aspect, in a possible implementation of the second aspect, the determining module acquires at least one attention unit and an attention duration of the at least one attention unit according to the visual data, including: acquiring the at least one attention unit according to the picture information; and acquiring a sum of the attention durations of the viewpoints in the at least one attention unit as the attention duration of the at least one attention unit.

[0021] With reference to the second aspect, in a possible implementation of the second aspect, the determining module acquires at least one attention unit and an attention duration of the at least one attention unit according to the visual data, including: acquiring the at least one attention unit according to the picture information; and acquiring a sum of the attention durations of the viewpoints in the at least one attention unit as the attention duration of the at least one attention unit.

[0022] With reference to the second aspect, in a possible implementation of the second aspect, the determining module determines the attention mode of the user according to the attention duration of the at least one attention unit, including: if a standard deviation of the attention durations of the at least one attention unit is greater than or equal to a second threshold value, determining that the attention mode of the user is staring; or if the standard deviation of the attention durations of the at least one attention unit is less than the second threshold value, determining that the attention mode of the user is scanning.

[0023] With reference to the second aspect, in a possible implementation of the second aspect, the determining module is configured to determine the music information according to the attention mode, including: if the attention mode is scanning, determining the music information according to the picture information; or if the attention mode is staring, determining the music information according to an attention unit with the highest attention degree in the attention unit.

[0024] With reference to the second aspect, in a possible implementation of the second aspect, the determining module determines the music information according to the attention mode, including: determining a behavior state of the user at each time point in a first time period according to the attention mode; determining a behavior state of the user in the first time period according to the behavior state at each time point; and determining the music information according to the behavior state in the first time period.

[0025] The third aspect provides a computer readable storage medium, in which a program instruction is stored, and when the program instruction is run by a processor, the method of the first aspect and any implementation manner of the first aspect is implemented.

[0026] In a fourth aspect, a computer program product is provided, which includes computer program codes, when the computer program codes are run on a computer, to implement the method of the first aspect and any implementation manner of the first aspect.

[0027] In a fifth aspect, a music recommendation system is provided, which includes a data collection device and a terminal device, the terminal device includes a processor and a memory, the memory stores one or more programs, the one or more computer programs include instructions, wherein the data collection device is configured to collect visual data of a user; when the instructions are executed by the one or more processors, the terminal device executes the method of the first aspect and any implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 is a system architecture to which the music recommendation method of the embodiments of the present application is applied;

[0029] Figure 2 is a schematic block diagram of a first wearable device in the system architecture to which the music recommendation method of the embodiments of the present application is applied;

[0030] Figure 3 is a schematic block diagram of a terminal device in the system architecture to which the music recommendation method of the embodiments of the present application is applied;

[0031] Figure 4 is a schematic block diagram of a second wearable device in the system architecture to which the music recommendation method of the embodiments of the present application is applied;

[0032] Figure 5 is a schematic flowchart of the music recommendation method of the embodiments of the present application;

[0033] Figure 6 is a schematic block diagram of the music recommendation method of the embodiments of the present application;

[0034] Figure 7 is a schematic block diagram of the music recommendation device of the embodiments of the present application;

[0035] Figure 8 is a schematic block diagram of the music recommendation device of the embodiments of the present application. DETAILED DESCRIPTION

[0036] The terminology used in the following description merely for the purpose of describing particular embodiments and is not intended to limit the application. As used in this description and the accompanying claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "at least one of" followed by a list of two or more items, means that at least one of the listed items is present at the minimum, but that one or more such items may be present either singly or in multiple amounts. The term "and / or" used in the context of describing the relationship between two or more objects is intended to mean that either one object can be used alone or in combination with one or more other objects. The term "and / or" is used in the context of describing the relationship between two or more objects is intended to mean that either one object can be used alone or in combination with one or more other objects.

[0037] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrases "in one embodiment" or "in some embodiments" in various places throughout this specification are not necessarily all referring to the same embodiment, unless otherwise specified. The terms "including", "containing", "having" and variations thereof mean "including but not limited to", unless expressly specified otherwise.

[0038] The technical solutions in the application will be described below with reference to the drawings.

[0039] The existing image and application matching method mainly includes: one is to extract the traditional bottom features of two modalities of music and image, and then establish the relationship between the two through a relationship model. This method has low matching degree of recommended music and image. Two is to collect matching pair data of music and image first, and automatically learn the matching model of music and image based on deep neural network. This method can recommend appropriate music in a simple scene.

[0040] However, in a real scene, the environment may contain many scenes and different style elements. The above existing methods do not consider the user's interest in the current environment, which reduces the matching degree of music. For example, when the user focuses on the clouds in the scene, the matching music should be different from when the user focuses on the animals in the scene.

[0041] Therefore, the embodiment of the application provides a music recommendation method, which obtains the attention area of the user in the complex environment by obtaining the viewpoint information of the user, so as to know the real interest of the user in the current environment and improve the matching degree of music.

[0042] Figure 1 The system architecture to which the music recommendation method of the embodiments of the present application is applied is shown in FIG. 1, which includes a first wearable device, a second wearable device, and a mobile terminal device. Figure 1 As shown in FIG. 1, the first wearable device is a wearable device that can collect visual data of a user and record head movement data of the user, such as smart glasses and the like, which is installed with an advanced photo system (APS) camera, a dynamic vision sensor (DVS) camera, an eye tracker, and an inertial measurement unit (IMU) sensor. The second wearable device is a wearable device that can play music, such as earphones and the like. The mobile terminal device can be a mobile phone, a tablet computer, a wearable device (for example, a smart watch), a vehicle-mounted device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like. The terminal device of the embodiments of the present application can include a touch screen for displaying service content to a user. The specific type of the terminal device is not limited in the embodiments of the present application.

[0043] It should be understood that the above is only an example of the devices in the embodiments of the present application and does not constitute a limitation on the embodiments of the present application. In addition to the devices exemplified above, the devices in the embodiments of the present application can also be other devices that can achieve the same functions. Figure 1 Figure 1 It should be understood that the above is only an example of the devices in the embodiments of the present application and does not constitute a limitation on the embodiments of the present application. In addition to the devices exemplified above, the devices in the embodiments of the present application can also be other devices that can achieve the same functions.

[0044] In the application of the music recommendation method of the embodiments of the present application, the mobile terminal device sends a data collection instruction to the first wearable device. After receiving the instruction, the first wearable device collects full-frame data at a certain frequency, records picture change data, and simultaneously records user viewpoint data and local picture data, as well as acceleration and angle data of head rotation, and continuously sends them to the mobile terminal device. After receiving the data, the mobile terminal device judges the attention area and attention mode of the user, extracts corresponding features according to the attention mode and attention area of the user, and matches music. The mobile terminal device sends audio data to the second wearable device, and the second wearable device plays music.

[0045] Figure 2 The modules included in the first wearable device in the application of the music recommendation method of the embodiments of the present application are shown in FIG. 2.

[0046] ​The wireless module is configured to establish a wireless link and communicate with other nodes. The wireless communication can be implemented by using a communication mode such as WiFi, Bluetooth, and cellular network.

[0047] The video frame acquisition module is configured to drive an APS camera on the first wearable device to acquire a video frame describing an environment.

[0048] The viewpoint acquisition module is configured to drive an eye tracker on the glasses to acquire viewpoint data. The viewpoint data includes a viewpoint position, an acquisition time, a fixation time, and a pupil diameter.

[0049] The head movement acquisition module is configured to drive an IMU module on the glasses to acquire a speed and an acceleration of head rotation.

[0050] The picture change capture module is configured to drive a DVS camera on the glasses to acquire picture change data.

[0051] The data receiving module is configured to receive data sent by the mobile terminal device.

[0052] The data sending module is configured to send the acquired data to the mobile terminal device.

[0053] Figure 3 The modules included in the mobile terminal device when the music recommendation method of the embodiment of the application is applied are shown.

[0054] The wireless module is configured to establish a wireless link and communicate with other nodes. The wireless communication can be implemented by using a communication mode such as WiFi, Bluetooth, and cellular network.

[0055] The attention mode determination module is configured to calculate an attention area and an attention mode according to the data acquired by the glasses.

[0056] The feature extraction and music matching module is configured to extract a feature and match music according to the attention mode category.

[0057] The data receiving module is configured to receive data sent by the first wearable device.

[0058] The data sending module is configured to send audio data and a play instruction of the music to the second wearable device.

[0059] Figure 4 The modules included in the second wearable device when the music recommendation method of the embodiment of the application is applied are shown.

[0060] The wireless module is configured to establish a wireless link and communicate with other nodes. The wireless communication can be implemented by using a communication mode such as WiFi, Bluetooth, and cellular network.

[0061] The data receiving module is configured to receive audio data and a playing instruction sent by the mobile terminal device.

[0062] The audio playing module is configured to play music according to the audio data and the playing instruction sent by the mobile terminal device.

[0063] Figure 5 A schematic flowchart of a music recommendation method according to an embodiment of the present application is shown, which includes steps 501 to 504, which will be described in detail below. Among them, Figure 5 The music recommendation method in the method can be executed by the terminal device in the method. Figure 1 The music recommendation method in the method can be executed by the terminal device in the method.

[0064] S501, receiving visual data of a user.

[0065] Specifically, the terminal device can receive visual data of a user sent by a first wearable device, and the first wearable device can collect visual data of the user within a preset time (for example, 1 second). The visual data of the user includes viewpoint information of the user and picture information viewed by the user. The viewpoint information includes position coordinates (x, y) of the viewpoint and attention duration of the viewpoint. The picture information includes a video frame image collected by an APS camera and picture change data collected by a DVS camera.

[0066] S502, obtaining at least one attention unit and attention duration of the at least one attention unit according to the visual data.

[0067] Specifically, the at least one attention unit can be obtained according to the picture information. For example, a macroblock in the video frame image can be taken as an attention unit. The macroblocks can be overlapped or non-overlapped. Alternatively, one or more object rectangular frames can be extracted as attention units according to an algorithm (for example, an objectness algorithm) for quantifying whether an object exists in a region. Alternatively, motion rectangular frames at different time points can be obtained according to the picture change data, and then the motion rectangular frames can be taken as attention units. Each attention unit can take image data at the same position in a frame image at the most recent time point as content of the attention unit.

[0068] When the picture viewed by the user is static or in a frame image, at this time, the DVS camera does not collect picture change data. The attention duration of the attention unit can be obtained by voting of all viewpoints on the attention unit. When a viewpoint is located in an attention unit, the attention duration of the viewpoint is added to the attention duration of the attention unit.

[0069] Optionally, according to the visual data, the at least one attention unit and the attention duration of the at least one attention unit are acquired, and the method further comprises: when the scene viewed by the user is changing, the DVS camera captures scene change data, and the attention unit in the frame image is still voted according to the above method, so that the attention duration of the attention unit in each frame image can be obtained. For the attention units in any two adjacent images, taking an attention unit in the latter image as an example, which is named as a second attention unit, N attention units in the former image are found, which have a distance less than a preset value from the second attention unit, wherein the distance between the attention units is the Euclidean distance of the center coordinates of the two attention units, and N can be a value artificially specified or the maximum number of attention units meeting the condition. Taking one of the N attention units as an example, which is named as a first attention unit, the similarity of the first attention unit and the second attention unit is judged, that is, the features of the first attention unit and the second attention unit are matched, wherein the feature matching method of the first attention unit and the second attention unit can be any existing image feature matching method, and the embodiments of the present application are not limited here. If it is determined that the features of the first attention unit and the second attention unit are similar, that is, the similarity of the first attention unit and the second attention unit is greater than or equal to a first threshold, it is considered that the first attention unit and the second attention unit are the same object at different times, and the attention duration of the second attention unit is equal to the sum of the attention duration of the first attention unit and the attention duration of the second attention unit, and then the attention duration of the first attention unit is zero. If it is determined that the features of the first unit and the second unit are not similar, that is, the similarity of the first attention unit and the second attention unit is less than the first threshold, the attention durations of the first unit and the second unit are retained. The above method is used to determine the attention units in any two adjacent images.

[0070] Optionally, the acquiring the at least one attention unit and the attention duration of the at least one attention unit according to the visual data further comprises: establishing a history library of attention units, the size of the history library is fixed, for example, only 10 attention units can be stored. The similarity between the newly acquired attention unit and the attention unit in the history library is determined, for example, the similarity between the second attention unit and the first attention unit in the history library is determined. The visual features of the first attention unit and the second attention unit can be extracted respectively, and then the similarity between the visual features is calculated. If it is determined that the features of the first attention unit and the second attention unit are similar, that is, the similarity between the first attention unit and the second attention unit is greater than or equal to a third threshold value, the attention duration of the second attention unit is equal to the sum of the attention duration of the first attention unit and the attention duration of the second attention unit, and then the second attention unit is used to replace the first attention unit and is stored in the history library. If it is determined that the features of the first unit and the second unit are not similar, that is, the similarity between the first attention unit and the second attention unit is less than the third threshold value, the first attention unit in the history library is retained. In this way, the attention units in the history library and the attention time of each attention unit can be obtained within a predetermined time, for example, 1 second, and then the attention mode of the user within the 1 second is determined according to the method in S503. The attention unit in the history library that has existed for more than 1 second and has an attention duration of less than 600 milliseconds is deleted, and the newly acquired attention unit is supplemented.

[0071] S503, determining the attention mode of the user according to the attention duration of the at least one attention unit.

[0072] Within a predetermined time, if the standard deviation of the attention duration of all attention units is greater than or equal to a second threshold value, it is determined that the attention mode of the user is staring; if the standard deviation of the attention duration of all attention units is less than the second threshold value, it is determined that the attention mode of the user is scanning.

[0073] S504, determining the recommended music information according to the attention mode.

[0074] If the attention mode of the user is scanning, the frame image captured by the APS camera is directly taken as the attention content of the user; if the attention mode of the user is staring, the attention unit with the highest attention degree in all attention units in a preset time period is taken as the attention content of the user. The attention degree can be determined according to the attention duration, for example, the attention unit with the longest attention duration is taken as the attention unit with the highest attention degree; or the attention degree can be determined according to the pupil dilation degree of the user, for example, the attention unit with the largest pupil dilation degree of the user is taken as the attention unit with the highest attention degree; or the attention degree can be determined according to the re-reading times of the user, for example, if the user re-reads an attention unit multiple times after staring at it and the re-reading times are greater than a preset value, the attention unit is taken as the attention unit with the highest attention degree; or the attention degree of the attention unit can be estimated by considering the three factors simultaneously, for example, the attention degree is the product of the pupil dilation degree of the user, the attention duration and the re-reading times.

[0075] Then the music information is determined according to the attention content. The method of determining the music information according to the attention content can be an existing method of matching music according to an image, for example, the attention content (frame image or attention unit with the highest attention degree) is taken as the input of a neural network model, and the music category with the largest probability value output by the neural network model is taken as the judgment result, for example, when the probability value is greater than 0.8, it is considered that the matching degree of the image and the music is high enough.

[0076] Optionally, after the attention content is determined according to the attention mode of the user in a preset time period, the music information can also not be determined, but the behavior state of the user in the preset time period is determined. The method of determining the behavior state according to the attention content can adopt an existing classification machine learning method, for example, the attention content is taken as the input of a neural network model, and then the behavior state category with the largest probability value output by the neural network model is taken as the judgment result. The behavior state includes driving, learning, traveling, exercising, etc. Thus, the behavior states of the user in multiple preset time periods in a first time period can be determined, for example, the first time period is 10 seconds, and the preset time period is 1 second, so the behavior states of the user in 10 seconds can be determined. The 10 behavior states are voted, for example, 7 of the 10 behavior states are determined to be learning, 2 are determined to be exercising, and 1 is determined to be traveling, so the behavior state of the user in the 10 seconds is considered to be learning. Finally, the music is matched according to the behavior state of the user in the first time period. The method of matching music according to the behavior state can be an existing method, for example, the music can be matched according to the label information of the behavior state, which is not limited in the embodiments of the present application.

[0077] After the music information is determined, the terminal device can send a music playing instruction to the second wearable device according to the music information, and the second wearable device plays the specified music. Or the terminal device can also play music according to the music information.

[0078] The music recommendation method of the embodiment of the present application can determine the attention mode of the user according to the visual information of the user, can more accurately determine the attention content of the user, and thus can recommend more suitable music, so that the recommended music conforms to the things that the user is really interested in and conforms to the real behavior state of the user, and improves the use experience of the user.

[0079] The music recommendation method of the embodiment of the present application is described in detail below according to specific examples. The first wearable device is taken as an example of smart glasses, the second wearable device is taken as an example of earphones, and the mobile terminal device is taken as an example of a mobile phone.

[0080] Figure 6 A schematic block diagram of the music recommendation method provided by the embodiment of the present application is shown in FIG. 1. Figure 6 As shown in FIG. 1, the music recommendation method comprises the following steps.

[0081] 1. Data acquisition

[0082] The mobile phone sends a data acquisition instruction to the smart glasses. After receiving the data acquisition instruction sent by the mobile phone, the smart glasses start to acquire data and continuously transmit the acquired data to the mobile phone. The acquired data comprises:

[0083] (1) Frame data: frame data of the whole image that can be seen by the user through the smart glasses is acquired at a certain frequency (for example, 30 Hz);

[0084] (2) Viewpoint data: the viewpoint position coordinates (x, y) of the user, the pupil diameter, the acquisition time and the gaze time are recorded;

[0085] (3) Head movement data: the angle and acceleration of the head rotation;

[0086] (4) Picture change data: the number of events acquired by the DVS camera.

[0087] 2. Analysis based on the acquired data and extraction of feature matching music

[0088] I. Determine one or more attention units of the user in a period of time and the attention time corresponding to each attention unit.

[0089] Specifically, the above period of time can be 1 second. An APS frame is taken at the beginning of 1 second, picture change and eye movement data are started to be recorded, the data at the end of this time is analyzed, and the feature matching music is extracted. If the situation changes in this 1 second, for example, the head of the user rotates greatly at 500 milliseconds, then only the data of 500 milliseconds can be analyzed, but if the above period of time is less than 100 milliseconds, which is not enough to generate a gaze point, then the data is discarded.

[0090] Wherein, the attention unit can be a macro block, an object rectangular frame or a motion rectangular frame. When the attention unit is a macro block, the macro blocks can be overlapped or non-overlapped; when the attention unit is an object rectangular frame, the attention unit at the initial time can be extracted by an algorithm (such as an objectness algorithm) quantifying whether an object exists in a region to obtain one or more object rectangular frames as the attention unit; when the attention unit is a motion rectangular frame, the motion rectangular frame at each time can be obtained based on the event data collected by the DVS camera. Specifically, at each time, the event data collected by the DVS camera is first represented as frame data, i.e. the gray value of the pixel position of the event is 255, and the gray value of the remaining pixel positions is 0, then the motion region is obtained by first eroding and then dilating the frame data, and finally the smallest rectangular frame that can cover the entire connected motion region is taken as the attention unit.

[0091] When the user's head is not moving (the head rotation angle is less than or equal to 5 degrees) and the DVS camera has no local output in this 1 second, i.e. the user sees a static picture:

[0092] (1) When a fixation point is located in an attention unit, the attention duration of the fixation point is accumulated to the attention duration of the current attention unit.

[0093] (2) The attention unit with an attention duration of 0 is removed, and the attention units with highly overlapped areas are removed according to the non maximum suppression (NMS) method.

[0094] When the user's head is not moving (the head rotation angle is less than or equal to 5 degrees) and the DVS camera has local output in this 1 second, i.e. the user sees a changing picture, and the pursuit behavior can occur:

[0095] (1) At the same time, when a fixation point is located in an attention unit, the attention duration of the fixation point is accumulated to the attention duration of the current attention unit.

[0096] (2) The attention unit with an attention duration of 0 at each time is removed, and the attention units with highly overlapped areas are removed according to the NMS method.

[0097] (3) In two adjacent time, for a attention unit A in the later time, find N attention units which are the nearest to the attention unit A in the former time, N is a positive integer greater than or equal to 1, the distance between two attention units is the Euclidean distance of the center coordinates of the two attention units. Perform feature matching between each of the N attention units and the attention unit A respectively, if the feature of the attention unit B in the former time is similar to the feature of the attention unit A, it is considered that the two attention units are the same object in different time, then delete the attention unit B in the former time, and accumulate the attention time of the attention unit B in the former time to the attention unit A; if the feature of the attention unit B in the former time is not similar to the feature of the attention unit A, keep the two attention units.

[0098] The music recommendation method of the embodiment of the present application is suitable for the case when the user's head is not moving, if the user's head is moving, the music matching is not performed at this time, and the music recommendation method of the embodiment of the present application is executed again when the user's head is not moving.

[0099] II, determine the attention mode and the attention content.

[0100] Attention mode:

[0101] According to the above determination method:

[0102] (1) If the number of attention units is 0, it is determined that the attention mode is "scanning look";

[0103] (2) If the number of attention units is not 0, and the mean square deviation of the attention time of different attention units is greater than a preset value, for example, 100 ms, it is determined that the attention mode is "staring look", otherwise, it is "scanning look".

[0104] Attention content:

[0105] (1) When the attention mode is "scanning look", it is considered that the user is mainly perceiving the environment at this time, and therefore the APS frame image is taken as the attention content;

[0106] (2) When the attention mode is "staring look", it is considered that the user is perceiving the interesting object at this time, and the attention unit with the highest attention degree is taken as the attention content.

[0107] III, extract features and match music according to the attention mode and the attention content.

[0108] The embodiment of the present application provides two methods of extracting features and matching music according to the attention mode and the attention content.

[0109] (1) Short term strategy

[0110] Directly match the visual features of the attention content and the audio features of the music in the current period. For example, using a classification machine learning method, taking the attention content as the input of a deep convolutional neural network, and taking the class with the maximum probability value in the output of the neural network as the judgment result, for example, when the probability value is greater than 0.8, it is judged that the matching degree of the visual features and the music is high, and the music is in line with the current perception of the user. The process of matching the music according to the image can be according to any existing method of matching the music according to the image, and the embodiments of the present application are not limited here.

[0111] (2) Long-term strategy

[0112] Determine the state category to which the content of the user's attention area at each moment belongs, associate the state category information at different moments, obtain the state of the user in a period of time, and match the music according to the label information of the state. Among them, the state category can be "driving", "learning", "traveling", "exercise" and other high-frequency scenarios of listening to music. The process of determining the user state category according to the content of the user's attention area at a certain moment can use a classification machine learning method, for example, taking the attention content as the input of a deep convolutional neural network, and taking the class with the maximum probability value in the output of the network as the judgment result. Associating the state category information at different moments can use a time-independent voting method to get the highest, or a time-dependent time weighting method, for example, dividing a period of time into ten moments, among which the user is judged to be learning for 8 moments and exercising for 2 moments, then the state of the user in this period of time can be obtained as learning.

[0113] 3、Earphone end plays music

[0114] After the earphone end receives the audio data and the play instruction sent by the mobile phone end, the corresponding music is played.

[0115] Optionally, in the music recommendation method shown in Figure 6 The embodiments of the present application also provide another method of analyzing and extracting features to match music based on the collected data. The following describes this another method.

[0116] I. Determine one or more attention units of the user in a period of time and the attention duration corresponding to each attention unit.

[0117] A history library of attention units is established, and the size of the history library is fixed, for example, set to 10 attention units. The history library is empty at the beginning, and the user-generated attention units are put into the history library until the history library is full. The attention duration of the attention units can be determined according to the viewpoint voting in the above method. After the history library is full, each newly generated attention unit is matched with each attention unit in the history library. The attention duration of the newly generated attention unit can also be determined according to the viewpoint voting in the above method. If the similarity between the attention unit A in the history library and the newly generated attention unit B is the highest, the attention time corresponding to the attention unit A is accumulated to the attention time corresponding to the attention unit B, then the attention unit A is deleted, and the attention unit B is put into the history library. The process of matching the similarity of different attention units is as follows: the visual features of different units are extracted respectively, and the similarity between different unit features is calculated according to the algorithm of speeded up robust features (SURF). If there is an attention unit in the history library that has an existence time of more than 1 second and an attention time of less than 600 milliseconds, the attention unit is deleted, and a newly generated attention unit is randomly filled in.

[0118] II. According to the attention units and attention times in the history library, the attention mode and attention content are determined.

[0119] Every 1 second, the attention allocation balance degree of different attention units is quantified according to the attention units and attention times in the history library.

[0120] When the user's head rotation angle is greater than 90 degrees and less than 270 degrees, that is, the user's view angle changes greatly, the history library of attention units is emptied. When the user's head is stationary, the history library is re-accumulated, and the balance degree of the attention units is quantified again after 1 second.

[0121] Attention mode:

[0122] According to the above determination method:

[0123] (1) If the number of attention units in the history library is 0, it is determined that the attention mode is "scanning look";

[0124] (2) If the number of attention units in the history library is not 0, and the mean square deviation of the attention times of different attention units is greater than a preset value, for example, 100 ms, it is determined that the attention mode is "staring look", otherwise it is "scanning look".

[0125] Attention content:

[0126] (1) When the attention mode is "scanning look", it is considered that the user is mainly perceiving the environment at this time, so the APS frame image is taken as the attention content;

[0127] (2) When the attention mode is "staring", it is considered that the user is perceiving the object of interest, and the attention unit with the highest attention is taken as the attention content.

[0128] III. Extracting features according to the attention mode and the attention content, and matching music.

[0129] The embodiments of the present application provide two methods of extracting features according to the attention mode and the attention content, and matching music.

[0130] (1) Short-term strategy

[0131] Directly match the visual features of the attention content in the current period and the audio features of the music. For example, a classification machine learning method is used, the attention content is taken as the input of a deep convolutional neural network, and the class with the maximum probability value in the output of the neural network is taken as the judgment result. For example, when the probability value is greater than 0.8, it is judged that the matching degree of the visual features and the music is high, and the music is consistent with the current perception of the user. The process of matching music according to the image can be any existing method of matching music according to the image, and the embodiments of the present application are not limited here.

[0132] (2) Long-term strategy

[0133] Determine the state category to which the content of the user's attention area at each moment belongs, associate the state category information at different moments, obtain the state of the user in a period of time, and match music according to the label information of the state. The state category can be "driving", "learning", "traveling", "exercise", and the like. The process of determining the state category of the user according to the content of the user's attention area at a moment can use a classification machine learning method, for example, taking the attention content as the input of a deep convolutional neural network, and taking the class with the maximum probability value in the output of the network as the judgment result. The association of state category information at different moments can use a time-independent voting method to obtain the highest value, or a time-dependent time weighting method, for example, dividing a period of time into ten moments, in which the user is judged to be learning for 8 moments and exercising for 2 moments, and the state of the user in this period of time is learning.

[0134] The method of data collection and the method of playing music at the earphone end are the same as those in the previous music recommendation method, and for the sake of brevity, the embodiments of the present application will not be repeated here.

[0135] The music recommendation method according to the embodiments of the present application recommends different music according to different contents to which the user pays attention, thereby providing better music experience. The music recommendation method according to the embodiments of the present application judges the current attention mode of the user by acquiring the viewpoint data, head movement data and environment data of the user, and selects a full frame image or a local attention region as the basis for matching music according to the judgment result.

[0136] The music recommendation method according to the embodiments of the present application is introduced above, and the music recommendation device according to the embodiments of the present application is introduced below.

[0137] Figure 7 A schematic block diagram of the music recommendation device according to the embodiments of the present application is shown in FIG. 7, which includes a transceiver module 710 and a determination module 720. The functions of the transceiver module 710 and the determination module 720 are introduced below respectively. Figure 7

[0138] The transceiver module 710 is configured to receive visual data of a user.

[0139] The determination module 720 is configured to acquire at least one attention unit and an attention duration of the at least one attention unit according to the visual data.

[0140] The determination module 720 is further configured to determine an attention mode of the user according to the attention duration of the at least one attention unit.

[0141] The determination module 720 is further configured to determine recommended music information according to the attention mode.

[0142] Optionally, the visual data includes viewpoint information of the user and picture information viewed by the user, and the viewpoint information includes a position of a viewpoint and an attention duration of the viewpoint.

[0143] Optionally, the determination module 720 acquires at least one attention unit and an attention duration of the at least one attention unit according to the visual data, including: acquiring the at least one attention unit according to the picture information; and acquiring a sum of the attention durations of the viewpoints in the at least one attention unit as the attention duration of the at least one attention unit.

[0144] Optionally, the determination module 720 acquires at least one attention unit and an attention duration of the at least one attention unit according to the visual data, further including: judging a similarity of a first attention unit and a second attention unit in the at least one attention unit, the first attention unit and the second attention unit being attention units at different time instants; and if the similarity is greater than or equal to a first threshold value, the attention duration of the second attention unit is equal to a sum of the attention duration of the first attention unit and the attention duration of the second attention unit. ​

[0145] Optionally, the determining module 720 determines the attention mode of the user according to the attention time length of the at least one attention unit, including: if the standard deviation of the attention time length of the at least one attention unit is greater than or equal to a second threshold, determining that the attention mode of the user is staring; if the standard deviation of the attention time length of the at least one attention unit is less than the second threshold, determining that the attention mode of the user is scanning.

[0146] Optionally, the determining module 720 is configured to determine the music information according to the attention mode, including: if the attention mode is scanning, determining the music information according to the picture information; if the attention mode is staring, determining the music information according to the attention unit with the highest attention degree in the attention unit.

[0147] The determining module 720 determines the music information according to the attention mode, and further includes: determining the behavior state of the user at each time in the first time period according to the attention mode; determining the behavior state of the user in the first time period according to the behavior state at each time; and determining the music information according to the behavior state in the first time period.

[0148] It should be understood that the transceiver module 710 in the music recommendation device 700 of the embodiment of the present application can be used to execute the method of S501 in the embodiment of the present application. Figure 5 The determining module 720 can be used to execute the method of S502 to S504 in the embodiment of the present application. Figure 5 The specific description can refer to the introduction of the above-mentioned embodiment, and for the sake of brevity, the embodiment of the present application will not be described here. Figure 5

[0149] Figure 8 is a schematic block diagram of the music recommendation device 800 of the embodiment of the present application. The music recommendation device 800 can be used to execute the music recommendation method provided by the above-mentioned embodiment, and for the sake of brevity, will not be described here. The music recommendation device 800 includes: a processor 810, the processor 810 is coupled with a memory 820, the memory 820 is used to store computer programs or instructions, and the processor 810 is used to execute the computer programs or instructions stored in the memory 820, so that the method in the above-mentioned method embodiment is executed.

[0150] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores program instructions, when the program instructions are run by a processor, to implement the music recommendation method of the embodiment of the present application.

[0151] The embodiment of the present application further provides a computer program product, characterized in that the computer program product includes computer program codes, when the computer program codes are run on a computer, to implement the music recommendation method of the embodiment of the present application.​

[0152] The embodiment of the present application further provides a music recommendation system, characterized in that the system comprises a data acquisition device and a terminal device, the terminal device comprises a processor and a memory, the memory stores one or more programs, and the one or more computer programs comprise instructions, wherein the data acquisition device is configured to acquire visual data of a user; and when the instructions are executed by the one or more processors, the terminal device executes the method for music recommendation of the embodiment of the present application.

[0153] Those skilled in the art can understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or in combination of computer software and electronic hardware. Whether the functions are realized in hardware or software mode depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0154] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0155] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0156] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application.

[0157] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.

[0158] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0159] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A music recommendation method characterized by, The method comprises: receiving visual data of a user; acquiring at least one attention unit and an attention duration of the at least one attention unit according to the visual data; determining an attention mode of the user according to the attention duration of the at least one attention unit; determining a behavior state of the user at each time point in a first time period according to the attention mode; determining a behavior state of the user in the first time period according to the behavior state at each time point; determining music information according to the behavior state in the first time period.

2. The method of claim 1, wherein, The visual data comprises viewpoint information of the user and picture information viewed by the user, and the viewpoint information comprises a position of a viewpoint and an attention duration of the viewpoint.

3. The method of claim 2, wherein, The acquiring at least one attention unit and an attention duration of the at least one attention unit according to the visual data comprises: acquiring the at least one attention unit according to the picture information; acquiring a sum of the attention durations of the viewpoint in the at least one attention unit as the attention duration of the at least one attention unit.

4. The method of claim 3, wherein, The acquiring at least one attention unit and an attention duration of the at least one attention unit according to the visual data further comprises: judging a similarity between a first attention unit and a second attention unit in the at least one attention unit, the first attention unit and the second attention unit being attention units at different time points; if the similarity is greater than or equal to a first threshold value, the attention duration of the second attention unit is equal to a sum of the attention duration of the first attention unit and the attention duration of the second attention unit.

5. The method according to any one of claims 1 to 4, characterized in that, The determining an attention mode of the user according to the attention duration of the at least one attention unit comprises: if a standard deviation of the attention durations of the at least one attention unit is greater than or equal to a second threshold value, determining that the attention mode of the user is staring; if the standard deviation of the attention durations of the at least one attention unit is less than the second threshold value, determining that the attention mode of the user is scanning.

6. The method of claim 5, wherein, The determining music information according to the attention mode comprises: if the attention mode is scanning, determining music information according to the picture information; if the attention mode is staring, determining music information according to an attention unit with the highest attention degree in the attention unit.

7. An apparatus for music recommendation, characterized by The method comprises: a transceiving module configured to receive visual data of a user; a determining module configured to acquire at least one attention unit and an attention duration of the at least one attention unit according to the visual data; the determining module is further configured to determine an attention mode of the user according to the attention duration of the at least one attention unit; the determining module is further configured to determine a behavior state of the user at each time point in a first time period according to the attention mode; a behavior state of the user in the first time period is determined according to the behavior state at each time point; and music information is determined according to the behavior state in the first time period.

8. The apparatus of claim 7, wherein, The visual data comprises viewpoint information of the user and picture information viewed by the user, and the viewpoint information comprises a position of a viewpoint and an attention duration of the viewpoint.

9. The apparatus of claim 8, wherein, The determining module acquires at least one attention unit and an attention duration of the at least one attention unit according to the visual data, and comprises: According to the picture information, the at least one attention unit is obtained; A sum of attention durations of the viewpoints in the at least one attention unit is obtained as an attention duration of the at least one attention unit.

10. The apparatus of claim 9, wherein, The determining module obtains at least one attention unit and an attention duration of the at least one attention unit according to the visual data, and further includes: Similarities of a first attention unit and a second attention unit in the at least one attention unit are judged, the first attention unit and the second attention unit being attention units at different time points; If the similarity is greater than or equal to a first threshold value, the attention duration of the second attention unit is equal to a sum of the attention duration of the first attention unit and the attention duration of the second attention unit.

11. The apparatus of any one of claims 7 to 10, wherein, The determining module determines an attention mode of the user according to the attention duration of the at least one attention unit, and includes: If a standard deviation of the attention duration of the at least one attention unit is greater than or equal to a second threshold value, it is determined that the attention mode of the user is staring; If the standard deviation of the attention duration of the at least one attention unit is less than the second threshold value, it is determined that the attention mode of the user is scanning.

12. The apparatus of claim 11, wherein, The determining module is configured to determine music information according to the attention mode, and includes: If the attention mode is scanning, the music information is determined according to the picture information; If the attention mode is staring, the music information is determined according to an attention unit with the highest attention degree in the attention unit.

13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program instructions, and when the program instructions are run by the processor, the method in any one of claims 1 to 6 is implemented.

14. A computer program product, characterised in that, The computer program product includes computer program code, and when the computer program code is run on the computer, the method in any one of claims 1 to 6 is implemented.

15. A music recommendation system characterized by, The system includes a data acquisition device and a terminal device, the terminal device includes a processor and a memory, the memory stores one or more programs, and the one or more computer programs include instructions, wherein, The data acquisition device is configured to acquire visual data of a user; When the instructions are executed by the one or more processors, the terminal device executes the method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information acquisition method and terminal

    CN109151176A

  • Information processing method and device, computer system and medium

    CN111241385A