House retrieval method and device
By building a floor plan diagram analysis model and a graphic matching model, the problem of low house search accuracy in the existing technology is solved, and more accurate and flexible house search is achieved to meet the personalized needs of users.
Patent Information
- Application Number
- CN202411998821.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the house search accuracy is low and cannot meet the personalized needs of users, because the method based on keyword matching has limitations on inaccurate labeling information and coarse-grained retrieval.
Build a floor plan analytical model and a graphic matching model, including a multi-scale feature extraction network, a multi-scale feature fusion network, a room key point extraction network and a room type semantic key point construction network, as well as a graph coded feature branch, a user coded feature branch and a matching degree calculation network. Through these models, the floor plan and user description are converted into semantic matching degrees, and houses that meet user needs are recommended.
It improves the accuracy of house search, can more accurately meet users' personalized needs, and enhances the flexibility and accuracy of search.
Smart Images

Figure CN119938967A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of search technology, and in particular to a housing search method and device. Background Art
[0002] House type retrieval refers to automatically retrieving house types that meet the personalized needs of users from a massive house type database, helping customers and intermediaries to find houses quickly and accurately, and playing an important role in the real estate market. Existing house type retrieval methods mainly use keywords for screening, such as matching and screening by parameters such as the number of rooms, area, and orientation. The keyword matching-based method relies on manual keyword annotation of house types, which may have the problem of inaccurate annotation information. In addition, the keyword matching-based method can only be used for coarse-grained retrieval and cannot meet the flexible and changeable personalized needs of different users, because the original expression of the user description is natural language text, which describes several demand points in detail, and several keywords cannot cover all the information in the original natural language text. Summary of the invention
[0003] In view of this, embodiments of the present application provide a housing search method, device, electronic device, and computer-readable storage medium to solve the problem of low housing search accuracy in the prior art.
[0004] According to a first aspect of an embodiment of the present application, a housing retrieval method is provided, including: constructing a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch and a matching degree calculation network; obtaining floor plans of each house and a user description of a target user describing the requirements for the house; inputting each floor plan into the floor plan parsing model, and outputting a set of floor plan semantic key points of each floor plan; inputting the set of floor plan semantic key points of each floor plan and the user description into the graph-text matching model, and outputting a matching degree between each floor plan and the user description; and recommending a house to the target user based on the matching degree between each floor plan and the user description.
[0005] According to a second aspect of an embodiment of the present application, a housing search device is provided, including: a construction module, configured to construct a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network and a floor plan semantic key point construction network, and the graph-text matching model includes a graph encoding feature branch, a user encoding feature branch and a matching degree calculation network; an acquisition module, configured to obtain floor plans of each house and a user description of the target user describing the requirements of the house; a first processing module, configured to input each floor plan into the floor plan parsing model, and output a set of floor plan semantic key points of each floor plan; a second processing module, configured to input the set of floor plan semantic key points of each floor plan and the user description into the graph-text matching model, and output the matching degree of each floor plan with the user description; a recommendation module, configured to recommend houses to the target user based on the matching degree of each floor plan with the user description.
[0006] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0007] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0008] Compared with the prior art, the embodiments of the present application have the following beneficial effects: constructing a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network, and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch, and a matching degree calculation network; obtaining the floor plan of each house and the user description of the target user describing the house requirements; inputting each floor plan into the floor plan parsing model, and outputting the floor plan semantic key point set of each floor plan; inputting the floor plan semantic key point set of each floor plan and the user description into the graph-text matching model, and outputting the matching degree of each floor plan with the user description; and recommending houses to the target user based on the matching degree of each floor plan with the user description. The above technical means can solve the problem of low house retrieval accuracy in the prior art, thereby improving the house retrieval accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0010] Figure 1 It is a flowchart of a housing search method provided in an embodiment of the present application;
[0011] Figure 2 It is a flowchart of another housing search method provided in an embodiment of the present application;
[0012] Figure 3 It is a structural schematic diagram of a house search device provided in an embodiment of the present application;
[0013] Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0014] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0015] A housing search method and device according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0016] Figure 1 It is a flowchart of a housing search method provided in an embodiment of the present application. Figure 1 The housing search method can be executed by a computer or a server, or by software on the computer or the server. Figure 1 As shown, the housing search method includes:
[0017] S101, constructing a floor plan parsing model and a picture-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network, and a floor plan semantic key point construction network, and the picture-text matching model includes a picture coding feature branch, a user coding feature branch, and a matching degree calculation network;
[0018] S102, obtaining floor plans of each house and a user description of the target user's requirements for the house;
[0019] S103, inputting each floor plan into a floor plan parsing model, and outputting a set of floor plan semantic key points of each floor plan;
[0020] S104, inputting the semantic key point set of each floor plan and the user description into a picture-text matching model, and outputting the matching degree between each floor plan and the user description;
[0021] S105, recommending houses to target users based on the matching degree between each floor plan and the user description.
[0022] The set of semantic key points of apartment type includes the coordinates and category values of each key point. The category value indicates the semantic category of the key point, including wall endpoint, door endpoint, window endpoint, etc. User description is the text description of the target user's needs, such as "a one-bedroom apartment, I expect its layout to be compact and practical. There is a small corridor between the living room and the door. The bathroom must be larger. The bedroom must have a south-facing balcony."
[0023] According to the technical solution provided in the embodiment of the present application, a floor plan parsing model and a graph-text matching model are constructed, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch and a matching degree calculation network; the floor plans of each house and the user description of the target user describing the house requirements are obtained; each floor plan is input into the floor plan parsing model, and the floor plan semantic key point set of each floor plan is output; the floor plan semantic key point set of each floor plan and the user description are input into the graph-text matching model, and the matching degree of each floor plan and the user description is output; based on the matching degree of each floor plan and the user description, the house is recommended to the target user. The above technical means can solve the problem of low house retrieval accuracy in the prior art, thereby improving the house retrieval accuracy.
[0024] Furthermore, each floor plan is input into a floor plan parsing model, and a set of floor plan semantic key points of each floor plan is output, including: within the floor plan parsing model: extracting the scale feature map of each floor plan through a scale feature extraction network; processing the scale feature map of each floor plan through a multi-scale feature fusion network to obtain a fusion feature map of each floor plan; converting the fusion feature map of each floor plan into a key point confidence map and a key point category map of each floor plan through a floor plan key point extraction network; processing the key point confidence map and the key point category map of each floor plan through a floor plan semantic key point construction network to obtain a set of floor plan semantic key points of each floor plan.
[0025] A multi-scale feature extraction network is constructed, which is implemented by a deep convolutional neural network. The original floor plan is used as input, and the deep convolutional neural network is forward propagated to obtain the intermediate layer features and the final output features to form multiple multi-scale feature maps.
[0026] Preferably, ResNet-50 is used as the multi-scale housing feature extraction network, and the size of the final output feature map is 8×8 pixels. In addition, four intermediate layer feature maps with sizes of 128×128, 64×64, 32×32, and 16×16 are selected, and a total of 5 multi-scale feature maps are obtained.
[0027] Furthermore, the scale feature maps of each house plan are processed by a multi-scale feature fusion network to obtain fused feature maps of each house plan, including: inside the multi-scale feature fusion network: processing the scale feature maps of each house plan by a deformable convolution layer to obtain the convolution feature maps of each house plan; processing the convolution feature maps of each house plan by a bilinear interpolation layer to obtain the interpolation feature maps of each house plan; processing the interpolation feature maps of each house plan by a splicing layer to obtain the cascade feature maps of each house plan; processing the cascade feature maps of each house plan by a multi-layer convolution layer to obtain the fused feature maps of each house plan.
[0028] First, the feature map of each scale is processed by deformable convolution to enhance the network's sensitivity to shape information. Then, the processed multi-scale feature map is scaled by bilinear interpolation to achieve spatial size alignment. The multi-scale feature maps after spatial size alignment are then channel cascaded (concatenation layer). Finally, the multi-scale feature maps after channel cascading are fused through multi-layer convolution operations to obtain a fused feature map.
[0029] Preferably, the convolution kernel size of the deformable convolution is 3×3, and the number of convolution kernels is the same as the number of channels of the input feature map. The spatial size of all multi-scale feature maps is aligned to 128×128 pixels, and the multi-layer convolution operation is specifically 3 layers of 3×3 convolution, with the number of convolution kernels being 512, 256, and 128 respectively.
[0030] Furthermore, the fused feature maps of each floor plan are converted into key point confidence maps and key point category maps of each floor plan through the floor plan key point extraction network, including: inside the floor plan key point extraction network: processing the fused feature maps of each floor plan through the confidence branch to obtain a preliminary key point confidence map of each floor plan, wherein the confidence branch includes multiple layers of convolutional layers and Sigmoid layers; processing the fused feature maps of each floor plan through the classification branch to obtain a preliminary key point category map of each floor plan, wherein the classification branch includes multiple layers of convolutional layers and Softmax; processing the preliminary key point confidence map and the preliminary key point category map of each floor plan respectively through the bilinear interpolation layer to obtain the key point confidence map and the key point category map of each floor plan.
[0031] The fused feature map is transformed into the key point confidence map C and the key point category map T. The key point confidence map is a single channel, its size is the same as the input image size, and the value range of each pixel is 0 to 1. The key point category map T has 3 channels, its size is the same as the input image size, and the value range of each pixel is 0 to 1.
[0032] Preferably, the fused feature map is input into the confidence branch and the classification branch respectively. The confidence branch is composed of 2 layers of 1×1 convolutional layers and Sigmoid output function, and the number of convolution kernels of the 2 layers of 1×1 convolutional layers is 64 and 1 respectively. The classification branch is composed of 2 layers of 1×1 convolutional layers and Softmax output function, and the number of convolution kernels of the 2 layers of 1×1 convolutional layers is 64 and 3 respectively. The outputs of the Sigmoid function and the Softmax function are both resized by bilinear interpolation to restore the spatial size to 256×256 pixels.
[0033] Furthermore, the key point confidence map and key point category map of each floor plan are processed through the floor plan semantic key point construction network to obtain the floor plan semantic key point set of each floor plan, including: using a non-maximum suppression algorithm to extract multiple key points of the floor plan and the coordinates of each key point from the key point confidence map of each floor plan; extracting the category value of each key point of the floor plan from the key point category map of the floor plan based on the coordinates of each key point of each floor plan; combining the coordinates and category values of each key point of each floor plan to obtain the floor plan semantic key point set of each floor plan.
[0034] First, N key points are extracted from the key point confidence map C using the non-maximum suppression algorithm with the lowest threshold, with two-dimensional coordinates {(ui,vi)}. Then, according to the two-dimensional coordinates of the N key points, the maximum confidence semantic category ci of the N key points is extracted from the key point category map T. Finally, the semantic key point set {Pi} (i = 1, 2, ..., N) of the floor plan I is obtained.
[0035] Training the floor plan parsing model. First, collect floor plan images, preprocess the image size and quality, and manually annotate the semantic key points of each floor plan image to obtain an annotated floor plan image dataset. Then, optimize the neural network parameters of each module of the floor plan parsing model using the gradient descent method. The training loss function consists of key point detection loss and key point classification loss.
[0036] Preferably, the key point detection loss is the Dice coefficient loss between the key point confidence map C and its true value. The key point classification loss is the cross entropy loss between the category {ci} marked by the key point of the apartment and the category of the corresponding position on the key point category map T of the apartment.
[0037] Furthermore, the semantic key point set of each floor plan and the user description are input into the image-text matching model, and the matching degree of each floor plan and the user description is output, including: within the image-text matching model: processing the semantic key point set of each floor plan through the image coding feature branch to obtain the image coding features of each floor plan; processing the user description through the user coding feature branch to obtain the user coding features; and calculating the matching degree between the image coding features and the user coding features of each floor plan through the matching degree calculation network.
[0038] The image-text matching model is implemented using a dual-branch architecture based on Transformer. The input is the semantic key point set {Pi} of the floor plan I and the natural language description {ti} of the user description R. The output is the matching degree m between the floor plan I and the user description R.
[0039] The first branch of the image-text matching model is the graph encoding feature branch, which converts the semantic key point set {Pi} of the floor plan I into a D-dimensional graph encoding feature Etype. First, the N key points are used as word units, each of which is linearly mapped to a D-dimensional key point embedding vector. The key point embedding vector is then input into the graph encoding feature branch composed of multiple layers of Transformer layers, and the D-dimensional embedding vector of the last Transformer layer is used as the graph encoding feature Etype of the floor plan I.
[0040] Preferably, the maximum number of key points N is 128. If the number of key points N is less than 128, zero padding is performed. The number of Transformer layers is 12, and each Transformer layer consists of layer normalization, multi-head self-attention, and multi-layer perceptron structure, which can capture the spatial and structural relationships between key points through the self-attention mechanism. The dimension D is set to 256 dimensions.
[0041] The second branch of the image-text matching model is the user encoding feature branch, which converts the natural language description {ti} of the user description R into a D-dimensional user encoding feature Ereq. First, the natural language description {ti} of the user description R is merged with the standard prompt to obtain a prompt suitable for a large language model, and then the prompt is input into the pre-trained large language model. The embedding vector Ereq0 output by the encoder of the large language model is linearly mapped to a D-dimensional vector, and then continued to be input into the user encoding feature branch composed of multiple layers of Transformer layers, and the D-dimensional embedding vector of the last layer of Transformer layer is used as the user encoding feature Ereq.
[0042] Preferably, the standard prompt is designed as “Based on the above description, analyze and summarize the quantitative key points of housing type requirements.” The large language model directly uses an existing mature model that has completed large-scale pre-training, and there is no need to optimize or adjust the parameters of the large language model during use.
[0043] Preferably, the number of Transformer layers is 12, and each Transformer layer consists of layer normalization, multi-head self-attention, and a multi-layer perceptron structure. The self-attention mechanism can be used to parse the semantic expression of the apartment structure from the implicit semantic expression of the apartment type.
[0044] The matching degree calculation network takes the graph encoding feature Etype and the user encoding feature Ereq as input, and calculates the similarity between the graph encoding feature Etype and the user encoding feature Ereq. The higher the similarity, the more the user description matches the apartment structure. Conversely, the lower the similarity, the less the user description matches the apartment structure.
[0045] Preferably, cosine similarity is used to measure the similarity between the graph encoding feature Etype and the user encoding feature Ereq, and the value range of the matching degree is [-1, 1].
[0046] Train the image-text matching model. First, collect the paired data of floor plans and corresponding text descriptions, convert each floor plan into a set of semantic key points of the floor plan in advance, and remove samples with exactly the same floor plan structure. Then, in each round of training, randomly select a batch of semantic key points of the floor plan and paired text descriptions, input them into the image-text matching model for forward propagation, calculate the matching ranking loss within the batch, and then optimize the parameters of all Transformer layers in the image-text matching model through gradient backpropagation.
[0047] Figure 2 FIG. 1 is a flow chart of another housing search method provided in an embodiment of the present application. Figure 2 As shown, the method includes:
[0048] S201, processing each floor plan through the floor plan parsing model and the image coding feature branch in the image-text matching model in turn to obtain the image coding feature of each floor plan;
[0049] S202, constructing a housing type code database using the image coding features of each housing type map;
[0050] S203, when receiving the user description, the user description is processed by the user coding feature branch in the image-text matching model to obtain the user coding feature;
[0051] S204, calculating the similarity between each bar code feature in the housing type code database and the user code feature;
[0052] S205, recommending houses to the target user based on the similarity between each floor plan and the user's description.
[0053] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.
[0054] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0055] Figure 3 Schematic diagram of a housing search device provided in an embodiment of the present application. Figure 3 As shown, the house search device comprises:
[0056] A construction module 301 is configured to construct a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network, and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch, and a matching degree calculation network;
[0057] The acquisition module 302 is configured to acquire the floor plan of each house and the user description of the target user's requirements for the house;
[0058] The first processing module 303 is configured to input each floor plan into a floor plan parsing model and output a set of floor plan semantic key points of each floor plan;
[0059] The second processing module 304 is configured to input the semantic key point set of each floor plan and the user description into the image-text matching model, and output the matching degree between each floor plan and the user description;
[0060] The recommendation module 305 is configured to recommend houses to the target user based on the matching degree between each floor plan and the user description.
[0061] According to the technical solution provided in the embodiment of the present application, a floor plan parsing model and a graph-text matching model are constructed, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch and a matching degree calculation network; the floor plans of each house and the user description of the target user describing the house requirements are obtained; each floor plan is input into the floor plan parsing model, and the floor plan semantic key point set of each floor plan is output; the floor plan semantic key point set of each floor plan and the user description are input into the graph-text matching model, and the matching degree of each floor plan and the user description is output; based on the matching degree of each floor plan and the user description, the house is recommended to the target user. The above technical means can solve the problem of low house retrieval accuracy in the prior art, thereby improving the house retrieval accuracy.
[0062] In some embodiments, the first processing module 303 is configured to: extract the scale feature map of each house plan through a scale feature extraction network; process the scale feature map of each house plan through a multi-scale feature fusion network to obtain a fused feature map of each house plan; convert the fused feature map of each house plan into a key point confidence map and a key point category map of each house plan through a house type key point extraction network; process the key point confidence map and the key point category map of each house plan through a house type semantic key point construction network to obtain a house type semantic key point set of each house plan.
[0063] In some embodiments, the first processing module 303 is configured to: process the scale feature map of each house plan through a deformable convolution layer to obtain the convolution feature map of each house plan; process the convolution feature map of each house plan through a bilinear interpolation layer to obtain the interpolation feature map of each house plan; process the interpolation feature map of each house plan through a splicing layer to obtain the cascade feature map of each house plan; process the cascade feature map of each house plan through a multi-layer convolution layer to obtain the fusion feature map of each house plan.
[0064] In some embodiments, the first processing module 303 is configured to: process the fused feature map of each floor plan through a confidence branch inside the house plan key point extraction network to obtain a preliminary key point confidence map of each floor plan, wherein the confidence branch includes multiple layers of convolutional layers and Sigmoid layers; process the fused feature map of each floor plan through a classification branch to obtain a preliminary key point category map of each floor plan, wherein the classification branch includes multiple layers of convolutional layers and Softmax; process the preliminary key point confidence map and the preliminary key point category map of each floor plan respectively through a bilinear interpolation layer to obtain a key point confidence map and a key point category map of each floor plan.
[0065] In some embodiments, the first processing module 303 is configured to use a non-maximum suppression algorithm to extract multiple key points of the floor plan and the coordinates of each key point from the key point confidence map of each floor plan; extract the category value of each key point of the floor plan from the key point category map of the floor plan based on the coordinates of each key point of each floor plan; combine the coordinates and category values of each key point of each floor plan to obtain a set of semantic key points of each floor plan.
[0066] In some embodiments, the second processing module 304 is configured to: process the set of semantic key points of each floor plan through the graph coding feature branch to obtain the graph coding features of each floor plan; process the user description through the user coding feature branch to obtain the user coding features; and calculate the matching degree between the graph coding features and the user coding features of each floor plan through the matching degree calculation network.
[0067] In some embodiments, the recommendation module 304 is configured to process each floor plan in turn through the floor plan parsing model and the image coding feature branch in the image-text matching model to obtain the image coding features of each floor plan; construct a floor plan code database using the image coding features of each floor plan; when receiving a user description, process the user description through the user coding feature branch in the image-text matching model to obtain the user coding features; calculate the similarity between each image coding feature and the user coding feature in the floor plan code database; and recommend houses to the target user based on the similarity between each floor plan and the user description.
[0068] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0069] Figure 4 Schematic diagram of an electronic device 4 provided in an embodiment of the present application. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0070] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4The electronic device 4 is merely an example and does not limit the electronic device 4 , and may include more or less components than those shown in the figure, or different components.
[0071] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0072] The memory 402 may be an internal storage unit of the electronic device 4, for example, a hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0073] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0074] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.
[0075] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A housing search method, characterized in that: include: Constructing a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network, and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch, and a matching degree calculation network; Obtain floor plans of each house and user descriptions of the target user's house requirements; Input each floor plan into the floor plan parsing model, and output a set of floor plan semantic key points of each floor plan; Inputting the semantic key point set of each floor plan and the user description into the image-text matching model, and outputting the matching degree between each floor plan and the user description; Based on the matching degree between each floor plan and the user description, a house is recommended to the target user.
2. The method according to claim 1, characterized in that Input each floor plan into the floor plan parsing model, and output a set of floor plan semantic key points of each floor plan, including: Inside the floor plan parsing model: Extracting scale feature maps of each floor plan through a scale feature extraction network; The scale feature maps of each floor plan are processed through a multi-scale feature fusion network to obtain fusion feature maps of each floor plan; The fusion feature map of each floor plan is converted into a key point confidence map and a key point category map of each floor plan through the floor plan key point extraction network; The key point confidence map and key point category map of each floor plan are processed through the apartment semantic key point construction network to obtain the apartment semantic key point set of each floor plan.
3. The method according to claim 2, characterized in that The scale feature maps of each floor plan are processed through a multi-scale feature fusion network to obtain the fusion feature maps of each floor plan, including: Inside the multi-scale feature fusion network: The scale feature map of each floor plan is processed by a deformable convolution layer to obtain a convolution feature map of each floor plan; The convolution feature map of each floor plan is processed through a bilinear interpolation layer to obtain an interpolation feature map of each floor plan; The interpolation feature maps of each floor plan are processed through the splicing layer to obtain the cascade feature maps of each floor plan; The cascade feature maps of each floor plan are processed through multiple convolutional layers to obtain the fused feature maps of each floor plan.
4. The method according to claim 2, characterized in that: The fusion feature map of each floor plan is converted into a key point confidence map and a key point category map of each floor plan through the key point extraction network, including: Inside the key point extraction network of the house type: Processing the fusion feature map of each floor plan through the confidence branch to obtain a preliminary confidence map of the key points of each floor plan, wherein the confidence branch includes multiple convolutional layers and Sigmoid layers; Processing the fusion feature map of each floor plan through a classification branch to obtain a preliminary map of key point categories of each floor plan, wherein the classification branch includes multiple convolutional layers and Softmax; The preliminary key point confidence map and the preliminary key point category map of each floor plan are processed separately through the bilinear interpolation layer to obtain the key point confidence map and the key point category map of each floor plan.
5. The method according to claim 1, characterized in that The key point confidence map and key point category map of each floor plan are processed through the apartment semantic key point construction network to obtain the apartment semantic key point set of each floor plan, including: A non-maximum suppression algorithm is used to extract multiple key points of the floor plan and the coordinates of each key point from the key point confidence map of each floor plan; According to the coordinates of each key point of each floor plan, extract the category value of each key point of the floor plan from the key point category map of the floor plan; The coordinates and category values of each key point of each floor plan are combined to obtain a set of semantic key points of each floor plan.
6. The method according to claim 1, characterized in that Inputting the semantic key point set of each floor plan and the user description into the image-text matching model, and outputting the matching degree between each floor plan and the user description, including: Inside the image-text matching model: Processing the semantic key point set of each floor plan through the graph coding feature branch to obtain the graph coding features of each floor plan; Processing the user description through a user coding feature branch to obtain a user coding feature; The matching degree between the image coding features and user coding features of each floor plan is calculated through the matching degree calculation network.
7. The method according to claim 1, characterized in that The method further comprises: Processing each floor plan through the floor plan parsing model and the image coding feature branch in the image-text matching model in turn to obtain the image coding feature of each floor plan; Using the image coding features of each floor plan to build a floor plan coding database; When receiving a user description, the user description is processed by the user coding feature branch in the image-text matching model to obtain a user coding feature; Calculate the similarity between the coding features of each image in the housing coding database and the coding features of the user; Based on the similarity between each floor plan and the user's description, a house is recommended to the target user.
8. A house search device, characterized in that: include: A construction module is configured to construct a floor plan parsing model and a graph-text matching model, wherein the floor plan parsing model includes a multi-scale feature extraction network, a multi-scale feature fusion network, a floor plan key point extraction network, and a floor plan semantic key point construction network, and the graph-text matching model includes a graph coding feature branch, a user coding feature branch, and a matching degree calculation network; An acquisition module is configured to acquire a floor plan of each house and a user description of the target user's requirements for the house; A first processing module is configured to input each floor plan into the floor plan parsing model and output a set of floor plan semantic key points of each floor plan; The second processing module is configured to input the semantic key point set of each floor plan and the user description into the image-text matching model, and output the matching degree between each floor plan and the user description; The recommendation module is configured to recommend a house to the target user based on the matching degree between each floor plan and the user description.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Space generation method based on natural language input driving and related apparatus
CN122594450A