House recommendation method and device, electronic equipment and storage medium
By obtaining the user's visual focus characteristics and emotional parameters, using a multi-scale time domain feature extraction network and a self-attention feature extraction network to determine the recommended feature images, solving the problems of low efficiency and low accuracy of house recommendation in the prior art, and achieving efficient and accurate house recommendations.
Patent Information
- Application Number
- CN202510576043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the prior art, house recommendations are inefficient and low in accuracy. Users' behaviors such as filling out forms, browsing and scoring cannot accurately reflect users' spontaneous preferences, and a large number of questionnaires, browsing records and user scores cannot be collected in a short period of time.
By obtaining the visual focus features when the user is viewing the feature house set information, the target feature image of interest to the user is determined, and the emotional parameters of the user are obtained when the user is viewing the target feature image, the multi-scale time domain feature extraction network, self-attention feature extraction network, and emotion estimation network determine the target feature image whose emotional parameters meet the preset recommendation conditions as the recommended feature image, and then determine the recommended house based on the recommended feature image.
It realizes automatic house recommendation, improves recommendation efficiency and accuracy, and comprehensively uses users' visual focus information and user emotions to recommend houses.
Smart Images

Figure CN120541249A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a housing recommendation method, device, electronic device, and storage medium. Background Art
[0002] Home recommendation is a key technology supporting the real estate industry. It aims to recommend homes that meet users' expectations based on their personalized needs and preferences, thereby improving the conversion rate of home recommendations and the success rate and efficiency of users finding satisfactory homes. Accurately capturing user preferences is a key factor in determining the effectiveness of home recommendations. Currently, questionnaires, browsing history analysis, and user rating analysis are the main methods used to quantify user preferences. However, user behaviors such as completing forms, browsing, and ratings do not accurately reflect users' spontaneous preferences and are also inefficient, making it difficult to collect large amounts of questionnaires, browsing history, and user ratings in a short period of time.
[0003] Therefore, there is an urgent need for a housing recommendation system that can efficiently obtain users' personalized and spontaneous preferences to solve the above problems. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a housing recommendation method, device, electronic device, and storage medium to solve the problems of low efficiency and low accuracy of housing recommendation in the prior art.
[0005] A first aspect of an embodiment of the present application provides a housing recommendation method, comprising:
[0006] Display characteristic house set information;
[0007] Track the user's visual focus when viewing the feature house set and determine the target feature image that the user is interested in;
[0008] Acquiring emotional parameters of the user when viewing the target feature image, and determining the target feature image whose emotional parameters meet the preset recommendation conditions as the recommended feature image;
[0009] Determine recommended houses based on the recommended feature images.
[0010] In certain embodiments of the present application, tracking the visual focus of a user when viewing a set of characteristic housing information and determining a target characteristic image of interest to the user includes:
[0011] Determine the user's visual focus at the target moment;
[0012] In the image viewed by the user at the target moment, an image of a preset size is cropped with the visual focus as the center to obtain a target feature image;
[0013] The target moment is any moment when the user views the feature house set.
[0014] In certain embodiments of the present application, determining the visual focus of the user at the target moment includes:
[0015] Get the user's eye rotation angle at the target moment;
[0016] Determine the ray direction vector in the eye coordinate system based on the eye rotation angle;
[0017] According to the calibration relationship between the eyeball coordinate system and the display coordinate system, the ray direction vector is transformed into the display coordinate system to obtain the line of sight equation of the display coordinate system;
[0018] The intersection of the sight line equation and the plane showing the characteristic house set is determined as the visual focus.
[0019] In some embodiments of the present application, obtaining the emotional parameters of the user when viewing the target feature image includes:
[0020] Obtaining EEG data of the user when viewing the target feature image;
[0021] Remove the baseline data from the EEG data to obtain valid EEG data;
[0022] The effective EEG data is input into the pre-trained emotion perception model to obtain emotion parameters.
[0023] In certain embodiments of the present application, the pre-trained emotion perception model includes a multi-scale temporal feature extraction network, a self-attention feature extraction network, and an emotion estimation network;
[0024] The multi-scale time-domain feature extraction network includes multiple convolutional layers and a multi-scale fusion module. The multiple convolutional layers extract features from the effective EEG data respectively. The multi-scale fusion module fuses the features extracted by each convolutional layer to obtain multi-scale features.
[0025] The self-attention feature extraction network generates an attention feature matrix based on multi-scale features. The attention feature matrix is used to represent the relationship between different time steps and different electrodes for obtaining EEG data.
[0026] The emotion estimation network maps the attention feature matrix into a two-dimensional vector, uses the growth curve function to process the two-dimensional vector to obtain the preference arousal value, and uses the hyperbolic tangent function to process the two-dimensional vector to obtain the preference efficacy value. The emotion parameters include the preference arousal value and the preference efficacy value.
[0027] In certain embodiments of the present application, the self-attention feature extraction network generates an attention feature matrix based on multi-scale features, including:
[0028] Determine the feature dimension and number of sequences of multi-scale features. The number of sequences is the product of the number of electrodes and the number of time steps for collecting EEG data.
[0029] Embed the multi-scale features based on the feature dimension and sequence number to obtain the input feature matrix;
[0030] Add position encoding to the input feature matrix;
[0031] The multi-head self-attention mechanism is used to extract features from the input feature matrix with position encoding added to obtain the attention feature matrix.
[0032] In certain embodiments of the present application, determining whether the emotion parameter satisfies a preset recommendation condition includes:
[0033] Obtaining the preferred arousal value and preferred efficacy value in the emotion parameters;
[0034] It is determined that the preference arousal value is greater than a first preset threshold value, and the preference efficacy value is greater than a second preset threshold value, and it is determined that the emotion parameter meets the preset recommendation condition.
[0035] A second aspect of the embodiments of the present application provides a housing recommendation device, comprising:
[0036] a display module configured to display characteristic house set information;
[0037] a visual feature extraction module configured to track the visual focus features of the user when viewing the feature house set and determine a target feature image that the user is interested in;
[0038] An emotional feature extraction module is configured to obtain emotional parameters of a user when viewing a target feature image, and determine a target feature image whose emotional parameters meet a preset recommendation condition as a recommended feature image;
[0039] The recommendation module is configured to determine a recommended house based on the recommended feature image.
[0040] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0041] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0042] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the embodiments of the present application obtain the visual focus characteristics of the user when viewing the characteristic house set information, determine the target characteristic image that the user is interested in, and obtain the emotional parameters of the user when viewing the target characteristic image, determine the target characteristic image whose emotional parameters meet the preset recommendation conditions as the recommended characteristic image, and then determine the recommended house based on the recommended characteristic image, thereby realizing automatic house recommendation and improving recommendation efficiency; at the same time, comprehensively utilizing the user's visual focus information and user emotions to recommend houses, thereby improving recommendation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] Figure 1 This is a flowchart of a housing recommendation method provided in an embodiment of the present application.
[0045] Figure 2 This is a flow chart of a method for determining a target feature image of interest to a user, provided in an embodiment of the present application.
[0046] Figure 3 This is a flow chart of a method for determining a user's visual focus at a target moment provided in an embodiment of the present application.
[0047] Figure 4 This is a flow chart of a method for obtaining emotional parameters of a user when viewing a target feature image, provided in an embodiment of the present application.
[0048] Figure 5 This is a flow chart of a method for generating an attention feature matrix based on multi-scale features using a self-attention feature extraction network provided in an embodiment of the present application.
[0049] Figure 6 It is a flowchart of a method for determining whether an emotion parameter meets a preset recommendation condition provided in an embodiment of the present application.
[0050] Figure 7 It is a block diagram of a system for implementing the housing recommendation method provided in an embodiment of the present application.
[0051] Figure 8 This is a structural block diagram of the emotion perception model provided in the embodiment of the present application.
[0052] Figure 9This is a schematic diagram of obtaining an image of interest in a head-mounted display screen based on a user visual focus tracking system provided in an embodiment of the present application.
[0053] Figure 10 This is a schematic diagram of a house recommendation device provided in an embodiment of the present application.
[0054] Figure 11 Schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0056] A housing recommendation method and apparatus according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0057] As mentioned above, currently, questionnaires, browsing history analysis, and user rating analysis are mainly used to quantify user preferences. However, users' behaviors such as filling out forms, browsing, and rating cannot accurately reflect their spontaneous preferences. Moreover, they are inefficient and cannot collect a large number of questionnaires, browsing history, and user ratings in a short period of time.
[0058] In view of this, an embodiment of the present application provides a house recommendation method, which obtains the visual focus characteristics of the user when viewing the characteristic house set information, determines the target characteristic image that the user is interested in, and obtains the emotional parameters of the user when viewing the target characteristic image, determines that the target characteristic image whose emotional parameters meet the preset recommendation conditions is the recommended characteristic image, and then determines the recommended house based on the recommended characteristic image, which can realize automatic house recommendation and improve the recommendation efficiency; at the same time, the user's visual focus information and user emotions are comprehensively utilized to recommend houses, thereby improving the recommendation accuracy.
[0059] Figure 1 This is a flow chart of a housing recommendation method provided by an embodiment of the present application. Figure 1 As shown, the method includes the following steps:
[0060] In step S101, characteristic house set information is displayed.
[0061] In step S102, the visual focus features of the user when viewing the feature house set are tracked to determine the target feature image that the user is interested in.
[0062] In step S103, the emotional parameters of the user when viewing the target feature image are obtained, and the target feature image whose emotional parameters meet the preset recommendation conditions is determined as the recommended feature image.
[0063] In step S104, a recommended house is determined based on the recommended feature image.
[0064] In some embodiments of the present application, the method may be executed by a terminal device with certain processing capabilities. The terminal device may include or be connected to a display device, and the display device is used to display the characteristic housing set information to the user.
[0065] In one example, the characteristic housing set information may be a pre-set housing information set, such as housing images or housing videos of different types of housing. A universal housing information set may be set, or different housing information sets may be set for different users based on user requirements.
[0066] In some embodiments of the present application, the terminal device may further include or be connected to a user eye tracking module. The eye tracking module is configured to track the user's eye movements when viewing the featured housing set information, thereby obtaining the user's visual focus characteristics when viewing the featured housing set, and further determining a target featured image that the user is interested in. The target featured image is the housing image that the user is visually focused on when viewing the featured housing set information.
[0067] In certain embodiments of the present application, the terminal device may also include or be connected to a brain-computer interface that can collect EEG data from the user. The terminal device can extract the EEG data of the user viewing the target feature image and determine the user's emotional parameters when viewing the target feature image based on the EEG data. The emotional parameters are used to represent the user's level of interest in the target feature image.
[0068] If the emotion parameter meets the preset recommendation condition, that is, the user's interest in the target feature image is greater than a preset threshold, the target feature image can be determined as a recommended feature image, and the target feature image can be added to a recommended feature image set.
[0069] When a user views the feature house set information, the operations of obtaining a target feature image, determining the emotional parameters of the user when viewing the target feature image, and adding the target feature image to the recommended feature image set when it is determined that the emotional parameters meet the preset recommendation conditions can be iteratively performed until the user has viewed all the content in the feature house set information, or the number of images in the recommended feature image set reaches a preset threshold.
[0070] The terminal device can determine the recommended house based on the recommended feature image. In one example, similar features of each image in the recommended feature image set can be extracted, and the house recommended to the user can be determined based on the similar features.
[0071] For example, a multimodal housing database can be configured on the terminal device, containing all images, text descriptions, and quantitative data of the listings. Alternatively, the multimodal housing database can be configured on a cloud server, and the terminal device can communicate with the cloud server to obtain data from the multimodal housing database.
[0072] For any house, all its pictures can be input into the visual Transformer model, and the visual feature set of each salient element in the picture can be obtained {f nm}. In this visual feature set {f nm} to determine the recommended house by retrieving the similar features of each image in the extracted recommended feature image set. Alternatively, the similar features of each image in the extracted recommended feature image set are calculated and compared with the visual feature set {f nm The distance between each feature is used to determine the recommended house. In one example, the visual Transformer model can use the pre-trained DinoV2 model.
[0073] According to the technical solution provided in the embodiments of the present application, by obtaining the visual focus features of the user when viewing the feature house set information, determining the target feature image that the user is interested in, and obtaining the emotional parameters of the user when viewing the target feature image, determining the target feature image whose emotional parameters meet the preset recommendation conditions as the recommended feature image, and then determining the recommended house based on the recommended feature image, automatic house recommendation can be achieved, thereby improving the recommendation efficiency; at the same time, the user's visual focus information and user emotions are comprehensively utilized to recommend houses, thereby improving the recommendation accuracy.
[0074] Figure 2 This is a flow chart of a method for determining a target feature image of interest to a user provided in an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0075] In step S201 , the visual focus of the user at the target moment is determined.
[0076] In step S202, in the image viewed by the user at the target moment, an image of a preset size is cropped with the visual focus as the center to obtain a target feature image.
[0077] The target moment is any moment when the user views the feature house set.
[0078] In some embodiments of the present application, tracking the visual focus characteristics of the user when viewing the feature house set information and determining the target feature image that the user is interested in can be done by first determining the visual focus of the user at the target moment, and then, from the image viewed by the user at the target moment, cropping an image of a preset size with the visual focus as the center to obtain the target feature image.
[0079] Figure 3 FIG is a flow chart of a method for determining the user's visual focus at a target moment provided by an embodiment of the present application. Figure 3 As shown, the method includes the following steps:
[0080] In step S301, the eye rotation angle of the user at the target moment is obtained.
[0081] In step S302, the ray direction vector in the eye coordinate system is determined based on the eye rotation angle.
[0082] In step S303 , according to the calibration relationship between the eyeball coordinate system and the display coordinate system, the ray direction vector is transformed into the display coordinate system to obtain the sight line equation of the display coordinate system.
[0083] In step S304, the intersection of the sight line equation and the plane displaying the characteristic house set is determined as the visual focus.
[0084] In some embodiments of the present application, when determining the visual focus of the user at the target moment, the eye rotation angle of the user at the target moment can be first obtained. In one example, the eye rotation angle can be, for example, a two-dimensional rotation angle (θ xt ,θ yt ).
[0085] The ray direction vector in the eyeball coordinate system can be calculated based on the eyeball rotation angle. If the eyeball coordinate system is {E}, the ray direction vector from the origin of the coordinate system {E} to the point (θ xt ,θ yt ) is the ray direction vector v t The origin of the coordinate system {E} can also be recorded as the virtual center position of the eyeball P EC,t .
[0086] The ray direction vector v can be t and the virtual center position of the eyeball P EC,t Convert to the display coordinate system {D}, the transformation relationship between the eye coordinate system {E} and the display coordinate system {D} is obtained by calibration. t After converting to the display coordinate system {D}, we can obtain the line of sight equation in the display coordinate system {D}, that is, the ray direction in the display coordinate system {D}. Finally, the intersection of this line of sight equation and the plane displaying the feature house set is determined as the visual focus.
[0087] In some implementations, the intersection point P of the sight line equation and the plane showing the characteristic house set may also be E,t After distortion correction, the user visual focus P in the real-time image is obtained. F,t , the P F,t As the visual focus, it is the local position that the user's eyes are focused on at the current moment.
[0088] In order to improve the stability of the user's visual focus, a Kalman filter can also be used to F,t Perform filtering.
[0089] When the preset size image is cropped with the visual focus as the center to obtain the target feature image, P F,t As the center, cut out the image of interest of preset size, such as 256×256 pixels, from the real-time screen image. I,t .
[0090] Figure 4 FIG. 1 is a flow chart of a method for obtaining the emotional parameters of a user when viewing a target feature image provided by an embodiment of the present application. Figure 4 As shown, the method includes the following steps:
[0091] In step S401 , EEG data of a user viewing a target feature image is obtained.
[0092] In step S402, baseline data in the EEG data is removed to obtain valid EEG data.
[0093] In step S403, the effective EEG data is input into a pre-trained emotion perception model to obtain emotion parameters.
[0094] In some embodiments of the present application, obtaining the emotional parameters of a user when viewing a target feature image can be accomplished by first obtaining EEG data of the user when viewing the target feature image, and then removing the baseline data from the EEG data to obtain the effective EEG data. This is because the baseline data in the EEG data represents the user's EEG signals at rest, which vary greatly from person to person. Therefore, reliable emotional perception can be achieved based on the difference between the real-time EEG data and the EEG signals at rest.
[0095] Next, the effective EEG data can be input into the pre-trained emotion perception model to obtain emotion parameters.
[0096] In some implementations, the pre-trained emotion perception model may include a multi-scale temporal feature extraction network, a self-attention feature extraction network, and an emotion estimation network.
[0097] Among them, the multi-scale time domain feature extraction network includes multiple convolutional layers and multi-scale fusion modules. Multiple convolutional layers extract features from effective EEG data respectively, and the multi-scale fusion module fuses the features extracted by each convolutional layer to obtain multi-scale features.
[0098] In one example, the EEG data x t and baseline data x b Difference x t -x b As the input of the emotion perception model, the x t -x b The dimensions may be, for example, 32×128×1.
[0099] The multi-scale temporal feature extraction network can include four convolutional layers and a multi-scale fusion module. The convolution kernel size of the first convolutional layer can be 1×7, the number of convolution kernels can be 64, the time span can be 2, and each convolutional layer is followed by a batch normalization layer and a ReLU layer, outputting the first-layer feature h1 with a dimension of 32×64×64.
[0100] The convolution kernel size of the second convolution layer can be 1×5, the number of convolution kernels is 64, the time span is 2, each convolution layer is followed by a batch normalization layer and a ReLU layer, and the output dimension is 32×32×64 second layer feature h2.
[0101] The convolution kernel size of the third convolution layer can be 1×3, the number of convolution kernels is 128, the stride is 2, and each convolution layer is followed by a batch normalization layer and a ReLU layer, and the output dimension is the third layer feature h3 of 32×16×128.
[0102] The convolution kernel size of the 4th convolution layer can be 1×3, the number of convolution kernels is 128, the stride is 2, and each convolution layer is followed by a batch normalization layer and a ReLU layer, and the output dimension is the 4th layer feature h4 of 32×8×128.
[0103] The multi-scale fusion module can first average pool the 4th layer features to obtain a global feature h0 with a dimension of 32×1×128. Then, the second-dimensional features of h0, h1, h2, h3, and h4 are bilinearly interpolated, aligned to a size of 32×64, and channel cascaded to obtain a multi-scale feature h with a dimension of 32×64×512. MS Next, we use two convolution layers with a kernel size of 1×1, 64 kernels, and a span of 1 to perform multi-scale feature fusion processing. Each convolution layer is followed by a batch normalization layer and a ReLU layer. Finally, the multi-scale fusion module outputs a 32×64×64-dimensional multi-scale feature h MF .
[0104] The self-attention feature extraction network generates an attention feature matrix based on multi-scale features. The attention feature matrix is used to characterize the relationship between different time steps and different electrodes for acquiring EEG data.
[0105] Figure 5 This is a flow chart of a method for generating an attention feature matrix based on multi-scale features from a self-attention feature extraction network provided in an embodiment of the present application. Figure 5 As shown, the method includes the following steps:
[0106] In step S501 , the feature dimension and sequence number of the multi-scale feature are determined.
[0107] The number of sequences is the product of the number of electrodes collecting EEG data and the number of time steps.
[0108] In step S502, the multi-scale features are embedded and encoded based on the feature dimension and the number of sequences to obtain an input feature matrix.
[0109] In step S503, position encoding is added to the input feature matrix.
[0110] In step S504, a multi-head self-attention mechanism is used to extract features from the input feature matrix with position encoding added to obtain an attention feature matrix.
[0111] In some embodiments of the present application, when the self-attention feature extraction network generates an attention feature matrix based on multi-scale features, the feature dimension and number of sequences of the multi-scale features can be first determined, and then the multi-scale features can be embedded and encoded based on the feature dimension and number of sequences to obtain an input feature matrix. The number of sequences is the product of the number of electrodes and the number of time steps used to collect EEG data.
[0112] In an example, if the feature dimension is 64, the number of electrodes collecting EEG data is 32, and the time step is 64, the number of sequences is 2048. Then the above multi-scale feature h MF Encode the data to a 2048×64 dimensional format to fit the Transformer.
[0113] Next, we can add positional encodings to the input feature matrix. For example, we can use a linear transformation to map the input to the Transformer's embedding space, mapping the 64-dimensional features to 32 dimensions, resulting in a 2048×32 input matrix. We then add positional encodings to generate a 2048×32 positional encoding matrix and superimpose it on the input matrix.
[0114] Finally, the multi-head self-attention mechanism is used to extract features from the input feature matrix with position encoding added to obtain the attention feature matrix.
[0115] In one example, the input matrix can be divided into multiple heads, and each head calculates self-attention. First, the attention score is calculated by the query, key, and value. Then, the values are weighted and summed according to the attention score to obtain the output of each head. After splicing and linear transformation, the outputs of all heads are spliced and linearly transformed to obtain the final output. Finally, the feedforward neural network further processes the output of the self-attention mechanism through two fully connected layers to form a 64×64 dimension output matrix h. out .
[0116] In some embodiments of the present application, the emotion estimation network of the pre-trained emotion perception model maps the attention feature matrix into a two-dimensional vector, uses the growth curve function to process the two-dimensional vector to obtain a preference arousal value, and uses the hyperbolic tangent function to process the two-dimensional vector to obtain a preference efficacy value, and the emotion parameters include preference arousal value and preference efficacy value.
[0117] The emotion estimation network can be composed of two layers of fully connected networks, which are used to transform the output matrix h of the self-attention feature extraction network into out Mapped into a two-dimensional vector. The first dimension of this two-dimensional vector is subjected to a Sigmoid (growth curve) function to obtain the preferred arousal value A, which can range from [0, 1] and represents the degree of arousal of the preferred emotion. The second dimension is subjected to a tanh (hyperbolic tangent) function to obtain the preferred efficacy value V, which can range from [-1, 1] and represents the preferred emotion of like or dislike.
[0118] In some embodiments of the present application, the emotion perception model may include two stages during training: pre-training and fine-tuning. In the pre-training stage, a large amount of general training data T may be collected first. G , which aims to enable the model to learn emotion-related EEG signal features from a large amount of general training data.
[0119] In one example, when multiple users watch a specific video material, the EEG signal sample x of each user can be collected. train,i , score each user's emotional arousal value and preference value, and obtain the training label y Atrain,i and y Vtrain,i . Using the above-mentioned EEG signal samples and training labels, the above-mentioned user emotion perception model is supervised trained to minimize the error between the model output label and the training label, thereby obtaining a pre-trained general model.
[0120] In the fine-tuning phase, individual training data T can be collected F , which aims to enable the model to further learn EEG features and emotion estimation related to individual users.
[0121] In one example, when a specific user is watching a specific video material, an EEG signal sample x of the specific user may be collected. ft,i , score the emotional arousal value and preference value of the specific user and obtain the training label y Aft,i and y Vft,i . Using the above EEG signal samples and training labels, the above general model is fine-tuned to minimize the error between the model output label and the training label, and a personalized emotion perception model for the specific user is obtained.
[0122] Figure 6 Schematic diagram of the flow of the method for determining whether the emotion parameter meets the preset recommendation condition provided by the embodiment of the present application. Figure 6 As shown, the method includes the following steps:
[0123] In step S601 , the preferred arousal value and the preferred efficacy value in the emotion parameters are obtained.
[0124] In step S602 , it is determined that the preference arousal value is greater than a first preset threshold, and the preference efficacy value is greater than a second preset threshold, and it is determined that the emotion parameter meets the preset recommendation condition.
[0125] In some embodiments of the present application, when determining whether an emotion parameter meets a preset recommendation condition, the preferred arousal value and preferred efficacy value of the emotion parameter may be first obtained. A determination is then made as to whether the preferred arousal value is greater than a first preset threshold and whether the preferred efficacy value is greater than a second preset threshold. If it is determined that the preferred arousal value is greater than the first preset threshold and the preferred efficacy value is greater than the second preset threshold, then the emotion parameter is determined to meet the preset recommendation condition.
[0126] The first preset threshold may be set to, for example, 0.5, and the second preset threshold may be set to, for example, 0.75.
[0127] Figure 7 is a block diagram of a system for implementing the housing recommendation method provided in the embodiment of the present application. Figure 7 As shown, the system may include a head-mounted display (HMD), which includes a brain-computer interface module, an eye tracking module, and a display module. Based on the HMD, a user can immersively watch video images, and the user's EEG and eye movement physiological signals can be synchronously collected.
[0128] The brain-computer interface module can have 32 surface-mount electrodes, which can synchronously collect brain waves with a sampling rate of 1024 Hz. The original brain wave signal is filtered by a 1-75 Hz bandpass filter and downsampled to 128 Hz with a time window of 1 second to obtain brain wave data x t , with dimensions of 32×128.
[0129] The eye tracking module consists of an infrared light source and a micro camera, which can capture pupil images and obtain the two-dimensional rotation angle of the eye (θ x ,θ y ).
[0130] The display module is composed of microOLED and lens modules, which can project the video image into the user's left and right eyes respectively to form a three-dimensional display. The image at time t is I t .
[0131] The brain-computer interface module outputs EEG data to the user emotion perception model, which calculates the user's preference arousal and preference efficacy for the currently viewed image based on the EEG data. The eye tracking module outputs eye direction to the user's visual focus tracking system, which outputs the image of interest in the video frame displayed on the display module.
[0132] Based on the preference arousal and preference validity values, we can determine whether the image of interest is a recommended positive sample. If so, we save it to the recommended positive sample library and use the top K positive samples in the recommended positive sample library to train the visual Transformer to obtain the retrieval visual feature set. Here, K is a positive integer.
[0133] By performing distance measurement and conditional screening on the retrieval visual feature set and the data in the multimodal housing database, recommended houses can be obtained.
[0134] That is, when recommending a house, the user can wear the head mounted display and use the display module to play a video with house content to the user. During the user's viewing process, the image of interest output by the user's visual focus tracking system is obtained. I,t , and the current preference arousal A output by the user emotion perception model t and preference utility value V t .
[0135] When the user's preference awakening degree A t Greater than the first preset threshold 0.5, and the preference value V t When it is greater than the second preset threshold value of 0.75, it is considered that the current image of interest I I,t The image of interest can be stored in the recommended positive sample library P.
[0136] When the number of images in the recommended positive sample library is greater than the third preset threshold (for example, 10), the visual Transformer model can be used to extract features from the 10 images with the highest preference value in the positive sample library, and obtain the retrieval visual feature set {f nm'}.
[0137] By retrieving the visual feature set {f nm The system then uses the cosine distance metric with the visual feature set in the multimodal housing database to retrieve recommended properties for the user. Finally, based on other filtering information provided by the user, such as budget and population, it selects valid properties from the recommended listings and pushes them to the user.
[0138] Figure 8 This is a structural diagram of the emotion perception model provided by the embodiment of this application. Figure 8 As shown in Figure 3, the emotion perception module includes 4 convolutional layers, a multi-scale fusion module, a multi-head self-attention module, an emotion estimation network, a Sigmoid function, and a tanh function.
[0139] The four convolutional layers are Convolution 1, Convolution 2, Convolution 3, and Convolution 4. The multi-scale fusion module is used to perform multi-scale fusion on the features output by the four convolutional layers. The multi-head self-attention module may include six heads and further extract features from the features output by the multi-scale fusion module. The emotion estimation network estimates the output features of the multi-head self-attention module, outputs the preferred arousal value through a sigmoid function, and outputs the preferred efficacy value through a tanh function.
[0140] Figure 9 Schematic diagram of obtaining an image of interest in a head-mounted display screen based on a user visual focus tracking system provided by an embodiment of the present application. Figure 9 As shown, the overall image may be a head-mounted display screen. After the visual focus is determined according to the eye tracking module, the image of interest may be cropped with the visual focus as the center.
[0141] The technical solution provided in the embodiments of the present application provides a housing recommendation method that can efficiently obtain users' personalized and spontaneous preferences. Through the combination of advanced technologies such as head-mounted displays, brain-computer interfaces and artificial intelligence, it helps to improve the service level and user experience of the real estate market.
[0142] The head-mounted display's brain-computer interface and eye-tracking module enable immersive viewing while accurately capturing EEG and eye movement signals, improving data quality. Combining EEG signals with deep learning models accurately perceives user emotional preferences, enabling personalized housing recommendations and improving user satisfaction. Leveraging a visual Transformer model and a multimodal housing database, the system rapidly extracts and matches features, effectively recommending properties that meet user preferences and needs.
[0143] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0144] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0145] Figure 10 Schematic diagram of a house recommendation device provided in an embodiment of the present application. Figure 10 As shown, the device includes:
[0146] The display module 1001 is configured to display characteristic house set information.
[0147] The visual feature extraction module 1002 is configured to track the visual focus features of the user when viewing the feature house set, and determine the target feature image that the user is interested in.
[0148] The emotion feature extraction module 1003 is configured to obtain emotion parameters of the user when viewing the target feature image, and determine the target feature image whose emotion parameters meet the preset recommendation conditions as the recommended feature image.
[0149] The recommendation module 1004 is configured to determine a recommended house based on the recommended feature image.
[0150] According to the technical solution provided in the embodiments of the present application, by obtaining the visual focus features of the user when viewing the feature house set information, determining the target feature image that the user is interested in, and obtaining the emotional parameters of the user when viewing the target feature image, determining the target feature image whose emotional parameters meet the preset recommendation conditions as the recommended feature image, and then determining the recommended house based on the recommended feature image, automatic house recommendation can be achieved, thereby improving the recommendation efficiency; at the same time, the user's visual focus information and user emotions are comprehensively utilized to recommend houses, thereby improving the recommendation accuracy.
[0151] In some embodiments, tracking the visual focus characteristics of a user when viewing the feature house set information and determining the target feature image that the user is interested in includes: determining the visual focus of the user at a target moment; cropping an image of a preset size centered on the visual focus from the image viewed by the user at the target moment to obtain the target feature image; wherein the target moment is any moment when the user views the feature house set.
[0152] In some embodiments, determining the visual focus of the user at the target moment includes: obtaining the eye rotation angle of the user at the target moment; determining the ray direction vector in the eye coordinate system based on the eye rotation angle; transforming the ray direction vector to the display coordinate system according to the calibration relationship between the eye coordinate system and the display coordinate system to obtain the line of sight equation of the display coordinate system; determining the intersection of the line of sight equation and the plane displaying the feature house set as the visual focus.
[0153] In some embodiments, obtaining emotional parameters of a user when viewing a target feature image includes: obtaining EEG data of the user when viewing the target feature image; removing baseline data from the EEG data to obtain effective EEG data; and inputting the effective EEG data into a pre-trained emotion perception model to obtain emotional parameters.
[0154] In some embodiments, the pre-trained emotion perception model includes a multi-scale time domain feature extraction network, a self-attention feature extraction network and an emotion estimation network; the multi-scale time domain feature extraction network includes multiple convolutional layers and a multi-scale fusion module, the multiple convolutional layers respectively extract features of the effective EEG data, and the multi-scale fusion module fuses the features extracted by each convolutional layer to obtain multi-scale features; the self-attention feature extraction network generates an attention feature matrix based on the multi-scale features, and the attention feature matrix is used to characterize the relationship between different time steps and different electrodes for obtaining EEG data; the emotion estimation network maps the attention feature matrix into a two-dimensional vector, uses the growth curve function to process the two-dimensional vector to obtain a preference arousal value, and uses the hyperbolic tangent function to process the two-dimensional vector to obtain a preference efficacy value, and the emotion parameters include preference arousal value and preference efficacy value.
[0155] In some embodiments, the self-attention feature extraction network generates an attention feature matrix based on multi-scale features, including: determining the feature dimension and number of sequences of the multi-scale features, where the number of sequences is the product of the number of electrodes and the number of time steps for collecting EEG data; embedding and encoding the multi-scale features based on the feature dimension and the number of sequences to obtain an input feature matrix; adding position coding to the input feature matrix; and using a multi-head self-attention mechanism to extract features from the input feature matrix with position coding added to obtain an attention feature matrix.
[0156] In some embodiments, determining whether the emotional parameters meet the preset recommendation conditions includes: obtaining the preference arousal value and the preference efficacy value in the emotional parameters; determining that the preference arousal value is greater than a first preset threshold and the preference efficacy value is greater than a second preset threshold, and determining that the emotional parameters meet the preset recommendation conditions.
[0157] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0158] Figure 11 Schematic diagram of an electronic device provided in an embodiment of the present application. Figure 11As shown, the electronic device 11 of this embodiment includes: a processor 1101, a memory 1102, and a computer program 1103 stored in the memory 1102 and executable by the processor 1101. When the processor 1101 executes the computer program 1103, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 1101 executes the computer program 1103, the functions of the modules / units in the above-described device embodiments are implemented.
[0159] The electronic device 11 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 11 may include but is not limited to a processor 1101 and a memory 1102. Those skilled in the art will appreciate that Figure 11 The electronic device 11 is merely an example and does not limit the electronic device 11 , and may include more or fewer components than shown in the figure, or different components.
[0160] The processor 1101 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0161] The memory 1102 may be an internal storage unit of the electronic device 11, such as a hard disk or memory of the electronic device 11. The memory 1102 may also be an external storage device of the electronic device 11, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 1102 may also include both an internal storage unit of the electronic device 11 and an external storage device. The memory 1102 is used to store computer programs and other programs and data required by the electronic device.
[0162] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0163] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0164] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A housing recommendation method, characterized in that: include: Display characteristic house set information; Tracking the visual focus features of the user when viewing the feature house set, and determining the target feature image that the user is interested in; Acquiring an emotional parameter of the user when viewing the target feature image, and determining the target feature image whose emotional parameter meets a preset recommendation condition as a recommended feature image; A recommended house is determined based on the recommended feature image.
2. The method according to claim 1, characterized in that Tracking the visual focus features of the user when viewing the feature house set information, and determining the target feature image that the user is interested in, including: Determine the user's visual focus at the target moment; In the image viewed by the user at the target moment, cropping an image of a preset size with the visual focus as the center to obtain the target feature image; The target moment is any moment when the user views the feature house set.
3. The method according to claim 2, characterized in that Determine the user's visual focus at the target moment, including: Get the user's eye rotation angle at the target moment; Determining a ray direction vector in an eyeball coordinate system based on the eyeball rotation angle; transforming the ray direction vector to the display coordinate system according to the calibration relationship between the eye coordinate system and the display coordinate system to obtain a sight line equation of the display coordinate system; An intersection point of the sight line equation and a plane displaying the characteristic house set is determined as the visual focus.
4. The method according to claim 1, wherein Acquiring the emotional parameters of the user when viewing the target feature image, including: Acquiring EEG data of a user when viewing the target feature image; removing baseline data from the EEG data to obtain valid EEG data; The effective EEG data is input into a pre-trained emotion perception model to obtain the emotion parameters.
5. The method according to claim 4, characterized in that The pre-trained emotion perception model includes a multi-scale time-domain feature extraction network, a self-attention feature extraction network and an emotion estimation network; The multi-scale time domain feature extraction network includes multiple convolutional layers and a multi-scale fusion module. The multiple convolutional layers respectively extract features from the effective EEG data. The multi-scale fusion module fuses the features extracted by each convolutional layer to obtain multi-scale features. The self-attention feature extraction network generates an attention feature matrix based on the multi-scale features, wherein the attention feature matrix is used to represent the relationship between different time steps and different electrodes for obtaining the EEG data; The emotion estimation network maps the attention feature matrix into a two-dimensional vector, uses a growth curve function to process the two-dimensional vector to obtain a preference arousal value, and uses a hyperbolic tangent function to process the two-dimensional vector to obtain a preference efficacy value. The emotion parameters include the preference arousal value and the preference efficacy value.
6. The method according to claim 4, characterized in that The self-attention feature extraction network generates an attention feature matrix based on the multi-scale features, including: Determining a feature dimension and a number of sequences of the multi-scale feature, where the number of sequences is the product of the number of electrodes and the number of time steps for collecting the EEG data; Performing embedding encoding on the multi-scale features based on the feature dimension and the number of sequences to obtain an input feature matrix; Adding position encoding to the input feature matrix; A multi-head self-attention mechanism is used to extract features from the input feature matrix with position encoding added to obtain the attention feature matrix.
7. The method according to any one of claims 1 to 6, characterized in that Determine whether the emotion parameters meet the preset recommendation conditions, including: Obtaining a preferred arousal value and a preferred efficacy value in the emotion parameters; It is determined that the preference arousal value is greater than a first preset threshold, and the preference efficacy value is greater than a second preset threshold, and it is determined that the emotion parameter meets a preset recommendation condition.
8. A house recommendation device, characterized in that: include: a display module configured to display characteristic house set information; a visual feature extraction module configured to track the visual focus features of the user when viewing the feature house set, and determine a target feature image that the user is interested in; An emotional feature extraction module is configured to obtain emotional parameters of a user when viewing the target feature image, and determine a target feature image whose emotional parameters meet a preset recommendation condition as a recommended feature image; The recommendation module is configured to determine a recommended house based on the recommended feature image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Interest data recommendation method and device, computer equipment and storage medium
CN117056574A
Systems, Methods, And Devices to Curate and Present Content and Physical Elements Based on Personal Biometric Identifier Information
US20240361827A1