Method and device for searching window for specific landscape

By building a three-dimensional model and visual domain database of the target area, combined with multimodal query technology, the problem of low efficiency in finding specific landscape windows in the existing technology is solved, and users' independent and accurate window retrieval is realized, and query efficiency and data interoperability are improved.

CN120219636AActive Publication Date: 2025-06-27HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510677803.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-27
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

In the prior art, it is inefficient to find windows with specific exterior window landscapes, making it difficult for users to find windows that meet exterior window landscapes in an independent system.

Method used

By obtaining the building information of the target area and a multi-view image set, a three-dimensional model is constructed, and the geometric spatial characterization data of the window is determined through image analysis. Based on the preset field angle and outer normal vector, the view cone of the window is generated and projected to the three-dimensional model to obtain the visual field data. This data is integrated to build a visual database, and users can find windows that meet the specific exterior landscape in the database through multimodal query information.

Benefits of technology

The efficiency of finding windows with specific exterior landscapes is improved. Users can independently and accurately retrieve windows that meet specific exterior landscape needs, breaking through the independent closure of traditional systems, and realizing the interconnection and sharing of cross-industry data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219636A_ABST
    Figure CN120219636A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for searching a window for a specific landscape, and the method comprises the steps: obtaining building information and a multi-view image set in a target region, and constructing a three-dimensional model of the target region according to the building information and the image set; analyzing frame images in the image set to determine a window data set; determining a view cone of the window according to a preset view field angle and an outer normal vector in the geometric space representation data, and determining view field data of the window generated when the view cone is projected to the three-dimensional model; according to the building information, the geometric space representation data of each window and the visual field data, constructing a visual field database of all windows in the target area; and according to the multi-modal query information uploaded by the user, searching a window conforming to the multi-modal query information in the visual field database. The searching efficiency of the window with the specific landscape outside the window is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer systems, and particularly to a method and device for finding windows for a specific landscape. Background Art

[0002] With the rapid advancement of urbanization and the continuous improvement of people's living standards, in the fields related to architecture, whether it is in scenarios such as residential, commercial, tourism, or urban planning, people's requirements for the spatial environment are becoming increasingly refined and diverse. In the scenarios of residential and hotel accommodation, residents and travelers not only pay attention to the facilities and layouts inside the building, but also expect the space they are in to have a high-quality visual experience, so as to enjoy the desired landscape or a good natural environment from the window. For example, hotel guests expect to book a room with a direct view of a specific landmark landscape, homebuyers focus on visual field indicators such as the greening rate outside the house window, and urban planners also need to evaluate the sky view and green coverage outside the windows of the block.

[0003] However, in the prior art, various systems such as hotels, real estate, and urban landscapes are independent, and it is difficult for ordinary users to search for the desired window view in multiple independent systems. For example, if a guest wants to book a hotel room with a view of a specific landmark, they need to first determine the hotel by themselves and then determine the location of the room through the hotel staff. In this way, the efficiency of users searching for windows that meet the window view is very low. Another example is that it will be even more difficult for a user to rent a house with a green view rate and a water view rate of 40% outside the window. Summary of the Invention

[0004] This application provides a method and device for finding windows for a specific landscape to solve the problem of low efficiency in finding windows with a specific window view in the prior art.

[0005] In a first aspect, this application provides a method for finding windows for a specific landscape, the method including: Obtain building information and a multi-perspective image set in a target area, and construct a three-dimensional model of the target area according to the building information and the image set; Determine a window data set by analyzing the frame images in the image set, where the window data set includes geometric space representation data of each window in the target area; Determine the frustum of the window according to a preset field of view angle and the outward normal vector in the geometric space representation data, and determine the visible area data of the window generated by projecting the frustum onto the three-dimensional model, where the visible area data is used to indicate the window view; Construct a visible area database of all windows in the target area according to the building information, the geometric space representation data of each window, and the visible area data; According to the multi-modal query information uploaded by the user, search for windows in the visual field database that match the multi-modal query information, where the multi-modal query information is used to query windows with specific outdoor views.

[0006] Optionally, by analyzing the frame images in the image set, it is determined that the window data set includes: By back-projecting the two-dimensional window frame corner points in the frame image onto the three-dimensional model, a three-dimensional window frame contour under the image path is obtained, where the three-dimensional window frame contour is used to indicate the shape contour of the window in three-dimensional space; Determine the window point clusters under the point cloud path according to the dense point cloud of the target area, and determine the geometric space characterization data of the window according to the window point clusters, where the window point clusters are used to indicate the point cloud data of multiple windows; Determine the overlap degree of the three-dimensional window frame contour under the image path and the window point clusters under the point cloud path on the two-dimensional plane; If the overlap degree is greater than the preset overlap degree threshold, determine that the three-dimensional window frame contour and the window point clusters are the same window, and use the geometric space characterization data under the point cloud path as the geometric space characterization data of the window; Construct the window data set according to the geometric space characterization data of multiple windows.

[0007] Optionally, by back-projecting the two-dimensional window frame corner points in the frame image onto the three-dimensional model, obtaining the three-dimensional window frame contour under the image path includes: Perform semantic segmentation on the frame image through a semantic segmentation model to determine the initial window frame mask in the frame image; Determine the final window frame mask of the window by tracking the same window in multiple consecutive frame images; Extract the two-dimensional window frame corner points in the final window frame mask; According to the camera parameters corresponding to the frame image, back-project the two-dimensional window frame corner points onto the three-dimensional model to obtain the three-dimensional window frame contour of the window under the image path, where the camera parameters include shooting parameters and camera pose parameters.

[0008] Optionally, determining the geometric space characterization data of the window according to the window point clusters includes: Repeat the following operations: Fit an initial plane according to a part of the points randomly selected from the window point clusters, and determine the distance from each point in the window point clusters to the initial plane; Determine the number of points with a distance less than the preset distance threshold, where the initial plane is used to indicate a wall with multiple windows; After reaching the set number of repeated operations, use the initial plane with the largest number of points as the target plane; Determine the geometric space representation data of each window according to the distribution characteristics of the window point clusters in the target plane.

[0009] Optionally, determining the overlap degree between the three-dimensional window frame contour under the image path and the window point clusters under the point cloud path in the two-dimensional plane includes: Project the three-dimensional window frame contour under the image path onto the two-dimensional plane to obtain a polygon contour; Project the window point clusters under the point cloud path onto the two-dimensional plane to obtain a projected point cluster; Select the internal point cluster that falls within the polygon contour from the projected point cluster; Calculate the overlap degree according to the area of the region formed by the internal point cluster and the area of the region formed by the polygon contour.

[0010] Optionally, constructing a visual field database for all windows in the target area according to the building information, the geometric space representation data of each window, and the visual field data includes: Determine the cone-of-vision images of the visual field data, and construct a cone-of-vision database according to the cone-of-vision images of multiple windows; Generate visual vectors according to the cone-of-vision images of the windows, and construct a vector database according to the visual vectors of multiple windows; Determine the building floor data where the window is located according to the preset building information, determine the visual environment index according to the visual field data, and construct an attribute database by combining the building floor data where multiple windows are located, the geometric space representation data, and the visual environment index; Construct a knowledge graph database, where the knowledge graph database uses entities, buildings, blocks, and windows as nodes, and uses the visible entities, the buildings where the windows are located, and the blocks where the windows are located as edges; Among them, the cone-of-vision database, the vector database, the attribute database, and the knowledge graph database all include window identifiers.

[0011] Optionally, searching for windows that meet the multimodal query information in the visual field database according to the multimodal query information uploaded by the user includes: Obtain the text to be recognized and the image to be recognized uploaded by the user, where the text to be recognized and the image to be recognized contain the outdoor landscape that the user expects to see; Determine the image vector of the image to be recognized, and search for the first set of window identifiers that match the image vector in the vector database. If the image vector contains landmark entities, add landmark keywords to the text to be recognized; Search for the second set of window identifiers that meet the text to be recognized in the first set of window identifiers in the attribute database; Search for target nodes and target edges that satisfy the second window identifier set in the knowledge graph database, and construct a third window identifier set based on the target nodes and the target edges; Determine the cone image to be recognized of the image to be recognized, and search for the final window identifier set that matches the cone image to be recognized in the third window identifier set of the cone database; Determine all window identifiers in the final window identifier set.

[0012] In a second aspect, the present application provides a device for finding windows for a specific landscape, the device includes: An acquisition and construction module, configured to acquire building information and an image set with multiple perspectives in a target area, and construct a three-dimensional model of the target area according to the building information and the image set; A first determination module, configured to determine a window data set by analyzing frame images in the image set, where the window data set includes geometric space representation data of each window in the target area; A second determination module, configured to determine a viewing frustum of the window according to a preset field of view angle and an outward normal vector in the geometric space representation data, and determine visual field data of the window generated by projecting the viewing frustum onto the three-dimensional model, where the visual field data is used to indicate the landscape outside the window; A construction module, configured to construct a visual field database of all windows in the target area according to the building information, geometric space representation data of each window, and visual field data; A search module, configured to search for windows that meet the multimodal query information in the visual field database according to multimodal query information uploaded by a user, where the multimodal query information is used to query windows with a specific landscape outside the window.

[0013] In a third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.

[0014] In a fourth aspect, the present application further provides a computer storage medium, storing computer-executable instructions, where the computer-executable instructions are used to execute the method for finding windows for a specific landscape according to any one of the above in the present application.

[0015] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: First, a three-dimensional model is constructed through building information and a multi-view image set, and the geometric space representation data of the window is determined by combining image analysis. Then, a frustum is generated based on a preset field of view angle and an outward normal vector, and the visible area data is obtained by projecting it onto the three-dimensional model, converting the abstract field of view into structured and quantifiable information. Furthermore, a visible area database is constructed by integrating building information, the geometric space representation data of the window, and the visible area data, and combined with multi-modal query technology, a quick response to the complex query requirements of users is achieved. The present application realizes the interconnection and sharing of cross-industry data by integrating multi-source heterogeneous data. With the help of the multi-modal query function, users can break through the query limitations of traditional systems that are independent and closed, and without relying on multi-platform switching and manual consultation, they can independently and accurately retrieve windows that meet specific outdoor landscape requirements, improving the user's query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0019] Figure 1 It is the system schematic diagram for finding windows for a specific landscape provided by the embodiments of the present application; Figure 2 It is the flowchart of a method for finding windows for a specific landscape provided by the embodiments of the present application; Figure 3 It is the window field of view rendering diagram of the frustum image provided by the embodiments of the present application; Figure 4 It is the window field of view semantic diagram provided by the embodiments of the present application; Figure 5 It is the window field of view depth of field diagram provided by the embodiments of the present application; Figure 6 It is the knowledge graph schematic diagram provided by the embodiments of the present application; Figure 7Schematic structural diagram of a device for finding windows for a specific landscape provided by an embodiment of the present application; Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0022] To solve the problem of low efficiency in finding windows with a specific outdoor landscape mentioned in the background art, the embodiment of the present application infers the visible field data of the windows based on the three-dimensional model of the target area, and then searches for windows that meet the specific landscape according to the user input, thereby improving the user's search efficiency.

[0023] The application fields of the present application include but are not limited to: hotel reservation, real estate transaction, urban planning and environmental assessment, tourism and cultural guide, and building design and renovation, etc.

[0024] Optionally, in the embodiment of the present application, the above method for finding windows for a specific landscape may be applied to a hardware environment composed of a terminal 101 and a server 103 as shown in Figure 1 As shown in Figure 1 As shown, the server 103 is connected to the terminal 101 through a network. The server constructs a visible field database according to the visible field data of each window in the target area. The user inputs the text to be recognized and the picture to be recognized through the terminal to query the windows with a specific outdoor landscape. The server searches for the windows that meet the conditions in the visible field database according to the user input and feeds them back to the user. The database 105 may be set on the server or independently of the server to provide data storage services for the server 103. The above network includes but is not limited to: wide area network, metropolitan area network or local area network. The terminal 101 includes but is not limited to PC, mobile phone, tablet computer, etc.

[0025] Next, in combination with specific implementation manners, a method for finding windows for a specific landscape provided by an embodiment of the present application will be described in detail. Taking the application to a server as an example, as Figure 2 shown, the specific steps are as follows: Step 201: Obtain building information and a multi-view image set in a target area, and construct a three-dimensional model of the target area according to the building information and the image set; Step 202: Determine a window data set by analyzing the frame images in the image set. Among them, the window data set includes geometric space representation data of each window in the target area; Step 203: Determine the frustum of a window according to a preset field of view angle and the outward normal vector in the geometric space representation data, and determine the visible area data of the window generated by projecting the frustum onto the three-dimensional model. The visible area data is used to indicate the landscape outside the window; Step 204: Construct a visible area database of all windows in the target area according to the building information, the geometric space representation data of each window, and the visible area data; Step 205: Search for windows that meet the multi-modal query information in the visible area database according to the multi-modal query information uploaded by the user. The multi-modal query information is used to query windows with a specific landscape outside the window.

[0026] First, some terms mentioned in the embodiments of the present application will be explained.

[0027] Multi-view image set: A series of image collections obtained by a user using devices such as a mobile phone, a drone, or a panoramic camera to capture a target area from different angles and positions. These images cover information in all directions of the target area and are the basic data for constructing a three-dimensional model.

[0028] Three-dimensional model: A three-dimensional digital presentation of a target area generated based on a multi-view image set and building information, accurately showing the structural relationship and respective spatial positions among buildings, water surfaces, and green plants in the target area.

[0029] Window data set: A set containing detailed information of each window in the target area. The geometric space representation data therein specifically describes geometric features such as the position, size, and shape of the window in three-dimensional space.

[0030] Geometric space representation data: Data used to accurately describe the geometric features of a window in three-dimensional space, including window frame vertex coordinates, window center point coordinates, outward normal vector (representing the orientation of the window), and window area, etc. These data can accurately determine the position and form of the window in space.

[0031] View frustum: A conical spatial range determined with the center point of the window as the viewing point and the outer normal vector of the window as the main axis direction, based on a preset horizontal field of view and vertical field of view. The light projection situation within this spatial range is used to determine the visible range and visible field data of the window.

[0032] Visible field data: The data result generated by projecting the view frustum onto a 3D model, which contains various information about the outdoor landscape, such as the proportion of different objects outside the window (vegetation, sky, buildings, etc.), whether a specific landmark (such as the Eiffel Tower) is visible and the visibility probability, etc., intuitively reflecting the characteristics of the external scene that the window can see.

[0033] Visible field database: Integrates the visible field data and geometric space representation data of all windows in the target area, as well as the building information in the target area. By centrally storing and managing these data, it facilitates quick query and acquisition of visible field-related information for specific windows, and supports various visible field-based analysis and query operations.

[0034] Multi-modal query information: The query instructions input by the user in various ways, such as text description, voice question, uploading pictures, etc., to find windows with a specific outdoor landscape. The system will perform corresponding processing and queries based on this different forms of information.

[0035] In step 201, the user uses devices such as mobile phones, drones, panoramic cameras, etc. to collect an image set with multiple perspectives in the target area, and then uploads the uploaded image set to the server. In addition, the server also supports importing building model files, parsing and extracting key information such as window geometric dimensions, floor height, building number, etc. The image set can be collected in the following ways.

[0036] 1. Mobile device collection: The user uses a mobile phone terminal to quickly upload a single photo or short video. When the data is uploaded, the system synchronously records metadata such as the shooting time, GPS (Global Positioning System) positioning coordinates, and mobile phone attitude (pitch angle, roll angle, heading angle), which is suitable for quickly capturing and instant querying of local scenes, meeting the user's fragmented data collection needs.

[0037] 2. Drone aerial survey data collection: For large-scale scene requirements, a drone is used to periodically collect image data according to a preset route. During the collection process, the real-time trajectory information recorded by the inertial measurement unit of the aircraft is archived synchronously with attitude parameters such as pitch angle, roll angle, and heading angle, ensuring the acquisition of high-precision spatial data and providing a basis for 3D modeling and visible field analysis of large areas.

[0038] 3. Panoramic and indoor image supplementary acquisition: For areas that are difficult to observe on the building facade or indoor spaces, a 360-degree panoramic camera is used for supplementary acquisition. By taking circumferential shots, complete spatial images are obtained, filling the blind spots of traditional perspectives and constructing a more comprehensive building visual field dataset, which is especially suitable for capturing details of buildings with complex structures.

[0039] The system constructs a 3D model of the target area using the frame images in the image set and building information. For example, when analyzing a certain urban block, images are collected by various devices and imported into a CAD (Computer Aided Design) file, and then a 3D model of the block is constructed. The 3D model can clearly present the layout and structure of the buildings in the block.

[0040] In step 202, the system analyzes the frame images in the image set, identifies the relevant information of the windows from the frame images, and then obtains the geometric space representation data of each window, such as the coordinates of the window frame vertices, the center point coordinates, the outer normal vector, and the area, etc., forming a window dataset containing these detailed information. The geometric space representation data provides an accurate data basis for the generation of the frustum in step 203 on the one hand, and also constructs an accurate data basis for building the attribute database on the other hand.

[0041] In step 203, based on the three-dimensional coordinates of the center point and the outer normal vector of each window in the window dataset, combined with the user-defined or scene-adapted horizontal field of view and vertical field of view, the system constructs a set of light rays with direction attributes through the ray tracing algorithm. These light rays start from the center point of the window and diverge evenly at a specific angle, jointly forming a frustum with a spatial geometric shape. In the 3D model space, the system performs projection operations on the frustum, and uses spatial collision detection technology to real-time track the propagation path of each light ray, accurately collecting multi-dimensional data such as the intersection coordinates of the light ray and the object in the 3D model, the category label of the hit object (such as vegetation, sky, building, water body, landmark building, etc.), and the occlusion degree. Through the integration and analysis of these data, detailed window visual field data is finally generated. The visual field data can not only intuitively present the specific composition of the landscape outside the window, but also accurately quantify the proportion and distribution of various landscape elements. For example, when analyzing the windows of a hotel, information such as the visible probability of each window looking at the landmark building and the greening rate outside the window can be clearly obtained.

[0042] In the embodiments of the present application, the abstract and vague concept of vision is transformed into structured and quantifiable visual field data. By using ray tracing and spatial collision detection technologies, the projection of the frustum of each window is accurately calculated, and multi-dimensional information such as the intersection coordinates of light with various objects in the 3D model, the category of the hit object, and the visibility probability of landmarks is detailedly recorded. After being standardized, these quantitative data can be efficiently integrated into the visual field database, forming a data storage system with clear logic and distinct levels. This data storage system improves the efficiency of data retrieval and analysis. Even in the face of extremely detailed query requirements from users (such as the visibility of specific landmarks, the exact value of the vegetation coverage rate outside the window, the change of the vision at different time periods, etc.), the system can rely on the structured visual field data to quickly locate and match relevant information and respond to the requirements with high precision.

[0043] In step 204, based on the building information of the target area (including building number, floor height where each window is located, etc.), the high-precision geometric space representation data of each window (coordinates of window frame vertices, center point coordinates, outer normal vector, area, etc.), and the quantified visual field data (intersection of light with objects, category of the hit object, visibility probability of landmarks, etc.), the system constructs a visual field database for all windows in the target area. The visual field database facilitates the management and query of a large number of windows and provides data support for multi-modal queries.

[0044] In step 205, the user inputs query information through multi-modal methods such as text, voice, and pictures. After receiving it, the system converts the voice to text, analyzes the query intention by combining technologies such as picture feature extraction, transforms it into query conditions, then searches for windows that meet the conditions in the visual field database according to specific query algorithms and logics, and finally filters and sorts the found windows and feeds back the results to the user. The multi-modal query function meets the diverse needs of users and improves the query efficiency and accuracy. For example, when the user is in a certain scenic area, uploads a photo containing the landmark of the scenic area and inputs the occlusion rate condition, the system can quickly screen out the windows that meet the conditions and display their positions on the mobile terminal.

[0045] In this application, first, a three-dimensional model is constructed through building information and a multi-view image set. The geometric spatial representation data of the window is determined by combining image analysis. Then, a frustum is generated based on a preset field of view angle and an outward normal vector, and projected onto the three-dimensional model to obtain visible field data, converting the abstract field of view into structured and quantifiable information. Then, the building information, the geometric spatial representation data of the window, and the visible field data are integrated to construct a visible field database, and combined with multi-modal query technology to achieve a rapid response to the complex query requirements of users. This application realizes the interconnection and sharing of cross-industry data by integrating multi-source heterogeneous data. With the help of the multi-modal query function, users can break through the independent and closed query limitations of traditional systems, and without relying on multi-platform switching and manual consultation, independently and accurately retrieve windows that meet specific outdoor landscape requirements, improving the user query efficiency.

[0046] As an alternative implementation, in step 201, constructing a three-dimensional model of the target area according to the image set includes the following steps.

[0047] Step S11: Perform feature point matching on the frame images in the image set, and use a preset method to recover the camera parameters used when the frame images were taken. Among them, the camera parameters include shooting parameters and camera pose parameters; Step S12: Generate a sparse three-dimensional point cloud of the target area according to the matched feature point pairs, camera parameters, and building information; Step S13: Process the sparse three-dimensional point cloud by density reconstruction to generate a dense point cloud, where the density of the dense point cloud is higher than that of the sparse three-dimensional point cloud; Step S14: Use the Poisson reconstruction algorithm to convert the dense point cloud into a three-dimensional model in a unified coordinate system.

[0048] The system extracts feature points from the frame images in the image set, identifies key points with uniqueness in each frame image, and then uses a method based on descriptor matching to complete the feature point matching of cross-frame images, generating matched feature point pairs. At the same time, the system uses preset algorithms such as bundle adjustment to solve and recover the camera parameters when each frame image was taken, including shooting parameters such as focal length and distortion coefficient, and camera pose parameters such as position and rotation angle.

[0049] Based on the matched feature point pairs and accurate camera parameters, combined with data such as geographical coordinates and structural layout in the building information, the system uses the principle of triangulation to calculate the three-dimensional spatial coordinates of the feature points, generating a sparse three-dimensional point cloud of the target area. The geographical coordinates in the building information are used to determine the position of the point cloud in the real-world coordinate system, and the structural layout information helps to verify the rationality of the point cloud distribution, ensuring that the point cloud initially outlines the spatial framework of the target area.

[0050] To improve the detailed expression ability of the point cloud, the system adopts density reconstruction technology (such as multi-view stereo algorithm) to process the sparse three-dimensional point cloud. By analyzing the perspective differences between adjacent images and using the texture and depth information of the images, more points are interpolated on the basis of the sparse point cloud to increase the point cloud density and form a dense point cloud, which can present the surface details of the target area more finely.

[0051] For the dense point cloud, the system adopts the Poisson reconstruction algorithm to convert the discrete point cloud data into a continuous triangular mesh model by means of implicit surface fitting. During the reconstruction process, if the building information contains global coordinates, the iterative closest point algorithm is used to align the model to the world coordinate system; if the global coordinates are missing, the coordinate calibration is performed through a small number of ground control points, and finally a three-dimensional model with clear texture and accurate geometry in the unified coordinate system is generated.

[0052] In this application, feature point matching and camera parameter recovery ensure the accuracy of three-dimensional reconstruction; generating sparse point clouds in combination with building information improves the geographical positioning accuracy and spatial structure rationality of the model; density reconstruction enhances the detail expressiveness of the model; Poisson reconstruction and coordinate calibration ensure the geometric integrity and coordinate system consistency of the model. The generated high-precision three-dimensional model can truly restore the spatial form of the target area.

[0053] As an alternative implementation, in step 202, by analyzing the frame images in the image set, it is determined that the window data set includes the following content.

[0054] Step S21: By back-projecting the two-dimensional window frame corner points in the frame image onto the three-dimensional model, the three-dimensional window frame contour under the image path is obtained, where the three-dimensional window frame contour is used to indicate the shape contour of the window in three-dimensional space; Step S22: Determine the window point clusters under the point cloud path according to the dense point cloud of the target area, and determine the geometric space representation data of the window according to the window point clusters, where the window point clusters are used to indicate the point cloud data of multiple windows; Step S23: Determine the overlap degree of the three-dimensional window frame contour under the image path and the window point clusters under the point cloud path on the two-dimensional plane; Step S24: If the overlap degree is greater than the preset overlap degree threshold, determine that the three-dimensional window frame contour and the window point clusters are the same window, and use the geometric space representation data under the point cloud path as the geometric space representation data of the window; Step S25: Construct a window data set according to the geometric space representation data of multiple windows.

[0055] In step S21, on the image path, the system uses a pre-trained semantic segmentation model (such as the DeepLabv3+ model) to perform pixel-level analysis on the frame image, quickly identify the window area and generate a window frame mask. Subsequently, the two-dimensional window frame corner points are extracted through a corner detection algorithm. Combining accurate camera parameters (such as pose, intrinsic matrix), the principle of triangulation is used to back-project them into the three-dimensional model, thereby constructing a high-precision three-dimensional window frame contour. The advantage of this path is that it can utilize the rich texture and semantic information in the image to accurately capture the boundary features of the window, especially suitable for the contour extraction of windows with complex appearances. Even in scenarios with light and shadow changes or local occlusions, the window position can be quickly located.

[0056] In step S22, in terms of the point cloud path, the system uses a density-based clustering algorithm to deeply mine the dense point cloud data in the target area. Since the window structure presents a unique point cloud aggregation form in space, by setting optimized density thresholds and neighborhood parameters, the window point cloud can be efficiently separated from the overall point cloud to form independent window point clusters. For each window point cluster, the system further calculates geometric parameters such as the vertex coordinates, center point coordinates, and outer normal vector of the window frame to generate accurate geometric space representation data. The core advantage of this path lies in relying on the three-dimensional spatial characteristics of the point cloud data, which can accurately describe the three-dimensional geometric shape of the window and has strong ability to capture details such as the depth and concave-convex structure of the window, especially suitable for scenarios where the building facade is complex and it is difficult to directly judge the structure through images.

[0057] In step S23, to verify whether the three-dimensional window frame contour under the image path and the window point cluster under the point cloud path correspond to the same window, the system projects the three-dimensional window frame contour and the window point cluster onto the two-dimensional plane respectively. By calculating the overlapping area between the two on the two-dimensional plane and comparing it with the total area of the two, the overlapping degree value is obtained, and this overlapping degree reflects the degree of coincidence of the window information determined by the two methods.

[0058] Although the image path can accurately identify the window contour, it is difficult to accurately obtain the spatial depth information of the window due to limitations such as the image perspective and occlusion problems; although the point cloud path can provide reliable three-dimensional geometric space representation data, point cloud clustering errors may occur in complex texture backgrounds. In the embodiments of this application, by projecting the three-dimensional window frame contour and the window point cluster onto the two-dimensional plane, calculating the overlapping degree and comparing it with a preset threshold, the advantages of the two paths can be effectively integrated and the errors of a single data source can be eliminated. Only when the overlapping degree meets the standard can it be ensured that the obtained window geometric space representation data has both the contour accuracy of the image path and the spatial accuracy of the point cloud path, improving the accuracy of the geometric space representation data.

[0059] In step S24, the system preset an overlap threshold in advance. If the calculated overlap is greater than the threshold, it indicates that the window information obtained from the two different paths of the image and the point cloud has high consistency. Thus, it is determined that the three-dimensional window frame contour and the window point cluster represent the same window. At this time, the geometric space representation data determined under the point cloud path is used as the final geometric space representation data of the window because the point cloud data has higher accuracy and reliability in describing spatial geometric features.

[0060] In step S25, after the system performs the above processing on all windows in the target area, it integrates the geometric space representation data of each window and constructs a complete window data set according to a certain data structure and storage format. This data set contains detailed and accurate geometric information of all windows in the target area, providing a data basis for subsequent operations such as window visibility analysis and database construction.

[0061] In this application, the image path relies on semantic segmentation and camera parameter inversion to accurately outline the window contour using image texture details for boundary recognition in complex appearance and occlusion scenarios; the point cloud path relies on the density clustering algorithm to accurately capture the three-dimensional structure of the window based on the three-dimensional spatial distribution characteristics. Through the collaborative operation of the image and point cloud dual paths and cross-verification, the image path can accurately identify the appearance features of the window, and the point cloud path can accurately describe the spatial geometric structure of the window, enabling the selection of window information with accurate contours and geometries. The verified geometric space representation data combines the visual accuracy of the image and the spatial depth of the point cloud, improving data reliability.

[0062] As an optional implementation manner, in step S21, obtaining the three-dimensional window frame contour under the image path by back-projecting the two-dimensional window frame corner points in the frame image into the three-dimensional model includes the following content.

[0063] Step S211: Perform semantic segmentation on the frame image through a semantic segmentation model to determine the initial window frame mask in the frame image; Step S212: Determine the final window frame mask of the window by tracking the same window in multiple consecutive frame images; Step S213: Extract the two-dimensional window frame corner points from the final window frame mask; Step S214: Back-project the two-dimensional window frame corner points into the three-dimensional model according to the camera parameters corresponding to the frame image to obtain the three-dimensional window frame contour of the window under the image path, where the camera parameters include shooting parameters and camera pose parameters.

[0064] In step S211, the system calls a semantic segmentation model (such as DeepLabv3+, U-Net, etc.) trained with a large amount of window image data to perform pixel-level semantic segmentation on the input frame image. Based on deep learning algorithms, this model can learn the texture, shape and other feature patterns of windows in the image. Through forward propagation calculation, it assigns class labels to each pixel in the image, separates the window area from the complex background environment, and generates an initial window frame mask. This mask is presented in the form of a binary image, where the white area represents the window and the black area represents the background, initially outlining the position and approximate contour of the window in the image.

[0065] In step S212, to improve the accuracy of the window frame mask, the system uses a multi-frame object tracking algorithm to dynamically track the same window in multiple consecutive frame images. The algorithm analyzes the motion trajectory and appearance feature changes of the window between adjacent frames, establishes the correspondence of the window between frames, and excludes misidentifications or missed detections caused by factors such as occlusion and lighting changes. After comprehensive processing and optimization of multiple-frame data, the final window frame mask of the window is finally determined, which can more accurately depict the true contour and boundary of the window.

[0066] In step S213, after obtaining the final window frame mask, the system uses an edge detection algorithm to extract the edge of the mask image to obtain the pixel-level contour of the window edge. Subsequently, a corner detection algorithm is used to locate two-dimensional window frame corner points with obvious features on the contour.

[0067] In step S214, based on the precise camera parameters corresponding to each frame image (including shooting parameters such as focal length and distortion coefficient, and camera pose parameters such as position and rotation angle), the system uses the principle of triangulation to back-project the two-dimensional window frame corner points from the image plane into the constructed three-dimensional model space. Specifically, the system converts the image coordinates into coordinates in the camera coordinate system through the internal parameter matrix of the camera, and then combines the external parameter matrix (i.e., the camera pose parameters) to convert the coordinates in the camera coordinate system into the world coordinate system, thereby calculating the position of the two-dimensional window frame corner points in the three-dimensional space. Finally, the three-dimensional window frame contour of the window under the image path is constructed, intuitively presenting the shape, size and position information of the window in the three-dimensional space.

[0068] In this application, the semantic segmentation model realizes automatic and pixel-level accurate recognition of the window area, improving the accuracy and stability of window recognition; the multi-frame tracking algorithm further eliminates the interference of environmental factors on the recognition results, ensuring that the window frame mask can accurately fit the actual contour of the window. The system accurately extracts the two-dimensional window frame corner points, and then uses precise camera parameters and the principle of triangulation for two-dimensional to three-dimensional conversion, so that the three-dimensional window frame contour obtained from the image path can highly restore the shape and position of the window in the real space.

[0069] As an alternative implementation, in step S22, determining the geometric space representation data of the window based on the window point cluster includes the following content.

[0070] Step S221: Repeat the following operations: Fit an initial plane according to a randomly selected part of the points from the window point cluster, and determine the distance from each point in the window point cluster to the initial plane; Determine the number of points whose distance is less than a preset distance threshold, where the initial plane is used to indicate a wall surface with multiple windows. Step S222: After reaching the set number of repeated operations, take the initial plane with the largest number of points as the target plane. Step S223: Determine the geometric space representation data of each window according to the distribution characteristics of the window point cluster in the target plane.

[0071] In step S221, the system adopts an iterative random sampling and plane fitting strategy for each window point cluster. First, randomly extract 20%-30% of the points from the window point cluster, and use the least squares method to fit an initial plane, which can approximately represent the wall containing multiple windows. Subsequently, based on a preset distance threshold (such as 5 cm), the system calculates the perpendicular distance from each point in the point cluster to the initial plane through vector operations, determines the points with a distance less than the threshold as "inliers", representing valid data closely fitting the wall plane; the points with a distance greater than the threshold are "outliers", representing invalid data caused by wall surface unevenness, noise interference, etc. The system counts the number of inliers under each fitted plane, and generates multiple groups of different initial planes and their corresponding inlier distribution results through multiple random samplings.

[0072] In step S222, when the iterative operation reaches 50-100 set times, the system comprehensively evaluates all the generated initial planes. Due to the complex shapes such as convex decorations on the building wall and concave window frames, the fitting degrees of different initial planes to the real wall vary due to sampling randomness. The system selects the plane with the largest number of inliers as the target plane by comparing the number of inliers of each initial plane. This is because the more inliers there are, the higher the degree of fitting of the plane to the real shape of the wall, which can filter out the interference points caused by wall surface unevenness and data noise to the greatest extent, accurately restore the spatial relationship between the wall and the window, and provide a reliable basis for subsequent calculation of geometric space representation data. This method is especially suitable for processing buildings with complex facades such as relief decorations, curved window frames, and special-shaped curtain walls, effectively avoiding the problem of plane misjudgment in the case of irregular wall scenes by the traditional fixed threshold method, and improving the accuracy and robustness of geometric space representation data extraction.

[0073] In step S223, after determining the target plane, the system further analyzes the distribution characteristics of the window point clusters within the target plane. By calculating the convex hull boundary of the point clusters, the coordinates of the window frame vertices are identified; the centroid algorithm is used to determine the coordinates of the window center point; by calculating the normal vector of the point clusters, the external normal vector of the window is obtained to determine the orientation of the window; finally, based on the area of the region covered by the point clusters, the actual area of the window is calculated. These operations comprehensively apply point cloud processing algorithms to accurately extract the geometric space representation data of each window from the target plane, and fully depict the shape, position, and orientation information of the window in three-dimensional space.

[0074] In this application, through the method of random sampling and iterative fitting, the system can adaptively process complex and variable wall forms. Whether it is a regular planar wall or an irregular wall with local concavities and convexities, it can accurately fit the target plane representing the wall. Compared with the traditional planar fitting method with fixed parameters, this method improves the adaptability to different building structures and reduces the calculation error of geometric space representation data caused by wall form differences.

[0075] Based on the reliable target plane, the subsequent analysis of the distribution characteristics of the window point clusters can more accurately determine the geometric space representation data of the windows. The screening mechanism with a preset distance threshold can effectively filter out the noise points and outliers in the point cloud data, avoiding the influence of these interfering data on planar fitting and parameter calculation. At the same time, the strategy of multiple iterative samplings enhances the robustness of the algorithm. Even when there are local missing or error in the point cloud data, it can still stably output reliable target planes and geometric space representation data.

[0076] As an optional implementation manner, in step S23, determining the overlap degree between the three-dimensional window frame contour under the image path and the window point clusters under the point cloud path in the two-dimensional plane includes the following contents.

[0077] In step S231, the three-dimensional window frame contour under the image path is projected onto the two-dimensional plane to obtain a polygonal contour; In step S232, the window point clusters under the point cloud path are projected onto the two-dimensional plane to obtain a projected point cluster; In step S233, the internal point clusters that fall within the polygonal contour are selected from the projected point cluster; In step S234, based on the area of the region formed by the internal point clusters and the area of the region formed by the polygonal contour, the overlap degree is calculated.

[0078] In step S231, the system maps the three-dimensional window frame contour constructed under the image path onto a two-dimensional plane by means of orthographic projection or perspective projection. During the projection process, according to the coordinate system of the three-dimensional model and the preset projection rules, the vertex coordinates of the window frame contour in three-dimensional space are transformed, and finally a two-dimensional polygon contour is formed. This polygon contour completely retains the shape and boundary information of the window under the image path. The number of its vertices corresponds to the number of corner points of the three-dimensional window frame contour, and the connection relationship of the edges also matches the window frame structure in three-dimensional space.

[0079] In step S232, for the window point clusters under the point cloud path, the system also performs a projection operation. The points in each window point cluster in three-dimensional space are mapped onto the two-dimensional plane one by one according to the same projection rules and coordinate system transformation methods as the three-dimensional window frame contour, forming projected point clusters. The distribution of these projected points on the two-dimensional plane reflects the morphological characteristics and positional relationships of the window point cloud in three-dimensional space and retains the discrete characteristics of the point cloud data.

[0080] In step S233, the system detects each point in the projected point cluster through the inclusion relationship between the point and the polygon. Specifically, each point is traversed from the projected point cluster to determine whether it falls inside the polygon contour generated in step S231. If the point is inside the polygon, it is included in the internal point cluster; otherwise, it is excluded. Through this operation, the point cloud data that has a spatial overlap relationship with the window frame contour under the image path can be accurately screened out.

[0081] Step S234: The system calculates the area of the region formed by the internal point cluster and the area of the region formed by the polygon contour respectively. For the internal point cluster, a convex hull algorithm or a grid-based area calculation method is used to obtain the area of the region it covers on the two-dimensional plane; for the polygon contour, the polygon area calculation formula is directly used to obtain the area. Finally, the area of the internal point cluster region is divided by the area of the polygon contour region to obtain the overlap degree value of the two. This value intuitively reflects the spatial coincidence degree of the window information obtained from the image path and the point cloud path in the form of a percentage. When the calculated overlap degree value reaches a preset threshold (such as 80%), it means that the three-dimensional window frame contour under the image path and the window point cluster under the point cloud path are highly coincident in spatial form and position, and the verification process passes smoothly.

[0082] In this application, by uniformly projecting the three-dimensional window frame contour and the window point cluster onto a two-dimensional plane, an intuitive comparison and quantitative analysis of image data and point cloud data in the same dimension are realized. In a building scenario, the image data may have contour deviations due to occlusion and perspective problems, and the point cloud data may have errors due to noise and sampling density. By calculating the overlap degree, the system can quantitatively evaluate the consistency of the information obtained by the two paths, effectively eliminate the wrong matches caused by data errors or ambiguities, improve the reliability of window recognition, and avoid misjudgment caused by the limitations of a single data source. Only when the overlap degree meets the preset threshold, the geometric space representation data under the point cloud path is used as the final information of the window, ensuring the high quality of the window data set. This method has strong adaptability to complex building scenarios. Whether it is a window with a regular or irregular shape, whether it is image data with light changes and occlusion, or point cloud data with noise and holes, effective screening and verification can be achieved through two-dimensional projection and overlap degree calculation, enhancing the stability and reliability of the system under different data qualities and complex environments.

[0083] As an alternative implementation manner, in step 204, according to the building information, the geometric space representation data of each window, and the visual field data, a visual field database of all windows in the target area is constructed, including the following contents.

[0084] Step S31: Determine the frustum image of the visual field data, and construct a frustum database according to the frustum images of multiple windows; Step S32: Generate a visual vector according to the frustum image of the window, and construct a vector database according to the visual vectors of multiple windows; Step S33: Determine the building floor data where the window is located according to the preset building information, determine the visual environment index according to the visual field data, and construct an attribute database by combining the building floor data, the geometric space representation data, and the visual environment index of multiple windows; Step S34: Construct a knowledge graph database, where the knowledge graph database uses entities, buildings, blocks, and windows as nodes, and the visible entities, the buildings where the windows are located, and the blocks where the windows are located as edges; Among them, the frustum database, the vector database, the attribute database, and the knowledge graph database all include window identifiers.

[0085] In step S31, in the frustum generation stage, the system traverses all windows in the target area, collects and organizes the frustum images of each window, and constructs a frustum database. During the construction of the frustum database, to ensure the accuracy and integrity of the data, the frustum images are standardized, and parameters such as the image resolution and color space are unified. At the same time, the indexing of the frustum images is optimized so that the corresponding frustum images can be quickly retrieved according to the window ID, improving the data query efficiency.

[0086] Figure 3 is the window field of view rendering of the cone image, from Figure 3 It can be seen that the visible field of view of the window includes but is not limited to: city skyline, waterfront landscape, high-rise buildings, urban density and visual corridors. The scene description of this image is: looking out through the window, you can see the high-density waterfront city landscape, and the nearby industrial or office low-rise buildings. The line of sight gradually extends to the dense high-rise buildings and distant mountains on both sides of Victoria Harbour, forming a rich depth of field. The mountains and the sea blend in the scene, forming a symbiotic visual corridor of nature and the city. High-rise towers and coastal skylines constitute the visual focus, with good openness and depth. In addition, the image also has the corresponding final semantic vector (visual vector and text vector).

[0087] In step S32, the system reprojects the cone image into a fixed-resolution image snapshot, which is used as a carrier of visual information and input into the visual pre-trained model for vectorized encoding. The model snapshot generates a set of floating-point vectors of the same length, namely, visual vectors, for the cone image of each window by extracting and transforming the features of the image.

[0088] Figure 4 It is a window view semantic map generated based on the image snapshot, in which each color represents a kind of semantic information. The window view semantic map uses different colors to distinguish semantic information such as buildings, corridors, sky and water, so that users can intuitively know the category composition of the scenery outside the window.

[0089] Figure 5 The window field of view depth map is generated based on the image snapshot. The window field of view depth map is used to display the depth information of the scenery outside the window and present the far and near hierarchical relationship of each part of the scenery.

[0090] The window field of view semantic map and window field of view depth map can provide more image feature information for visual vector generation, so that the generated vector can more accurately and comprehensively describe the scene outside the window.

[0091] Preferably, the system can also obtain prompts input by the user, which are used to describe the scene outside the window of the image snapshot. The system feeds the constructed prompts to the large language model. The large language model, with its powerful language understanding and generation capabilities, conducts in-depth analysis and processing of the input information, and then generates a brief Chinese summary, which is then converted into a numerical vector representation, that is, a text vector.

[0092] At this time, the system already has the visual vector and text vector corresponding to each window. Since the visual vector and text vector are dimensionally compatible, they can be concatenated bit by bit. The resulting final semantic vector integrates information from both the visual and text modalities, and can describe the outdoor scene more comprehensively and accurately. Finally, the final semantic vector is written into the vector database. This application takes into account both visual information and text information, improving the accuracy of the vector database. The vector database uses the window ID as an index to find the corresponding final semantic vector.

[0093] In step S33, the system determines the building floor data where each window is located according to the preset building information, such as the floor structure and storey height data of the building. At the same time, based on the visual field data obtained from the visual field analysis, visual environment indicators such as greening rate, occlusion rate, and sky visibility are calculated. Combining the geometric space representation data of the window (including window frame vertex coordinates, center point coordinates, outer normal vector, area, etc.), the system constructs an attribute database.

[0094] In the construction of the attribute database, the window view semantic map and the window view depth map can jointly assist in determining the visual environment indicators. Analyze the greening rate and occlusion rate through the window view semantic map, and use the window view depth map to judge the spatial depth relationship to improve the attribute data.

[0095] When constructing the attribute database, a suitable spatial database management system is selected to store the data in a structured manner. The various attribute data of the window are classified and stored according to different fields, and a unique identifier (window ID) is established for each window to realize the associated query of the various attribute data through the window ID.

[0096] The attribute database integrates the building floor data, geometric space representation data, and visual environment indicators of the window, providing rich data support for in-depth analysis of the characteristics of the window and its surrounding environment. For example, in the real estate field, these data can be used to evaluate value factors such as the lighting and view of the house; in urban planning, it can be used to analyze the landscape visibility and environmental quality of different regions.

[0097] In step S34, the system constructs a knowledge graph database, with entities (such as landmark buildings, green areas, etc.), buildings, blocks, and window IDs as nodes, and the visible entities, the buildings they belong to, and the blocks they are located in as edges, to construct a complex relationship network. For example, "window - visible - landmark" represents the relationship between the window and the visible landmark, "window - belongs to - building" represents the relationship between the window and the building it belongs to, and "building - located in - block" represents the relationship between the building and the block it is located in.

[0098] The present application can also enrich the semantic information of the knowledge graph by defining the attributes of nodes and edges. For example, attributes such as visibility probability and occlusion degree are added to the "window - visible - landmark" edge, enabling the knowledge graph to more accurately describe the relationship between the window and the surrounding environment. The embodiments of the present application can periodically update and expand the knowledge graph. With the collection and analysis of new data, the structure and content of the knowledge graph are continuously improved to adapt to the changing actual situation.

[0099] Figure 6 For the four - level knowledge graph structure of "window - entity - building - block", each node and edge in the figure represents the complex semantic relationship between the window and the visible elements around it. For example, window 7 belongs to Jiangfeng Villa, and at the same time, window 7 has the feature: medium green view rate, and Jiangfeng Villa is located in Dapu.

[0100] The knowledge graph database visually displays the relationship between the window and the surrounding entities, buildings, and blocks in a graphical way. This relationship model is not only easy to understand and visualize but also supports complex relationship reasoning. For example, through the knowledge graph, it is possible to quickly query all visible windows within a certain range around a landmark building, or the window information of specific types of buildings within a certain block.

[0101] Table 1 shows the multiple sub - databases included in the visible field database and the main content of each sub - database.

[0102] Table 1

[0103] The present application integrates multi - source heterogeneous data related to windows by constructing a frustum database, a vector database, an attribute database, and a knowledge graph database, establishing associations with the window identifier as the link, and realizing the unified management and efficient storage of data. This structured data organization method enables quick positioning and acquisition of the required information when querying and analyzing data, reducing data processing time and improving the system operation efficiency.

[0104] As an alternative implementation, in step 205, finding the windows that match the multi - modal query information in the visible field database according to the multi - modal query information uploaded by the user includes the following content.

[0105] Step S41: Obtain the text to be recognized and the image to be recognized uploaded by the user, where the text to be recognized and the image to be recognized contain the outdoor landscape that the user expects to see; Step S42: Determine the image vector of the image to be recognized, and find the first set of window identifiers that match the image vector in the vector database. If the image vector contains a landmark entity, add the landmark keyword to the text to be recognized; Step S43: Search for the second window identification set that meets the text to be recognized in the first window identification set of the attribute database; Step S44: Search for the target nodes and target edges that meet the second window identification set in the knowledge graph database, and construct the third window identification set based on the target nodes and target edges; Step S45: Determine the cone image to be recognized of the image to be recognized, and search for the final window identification set that matches the cone image to be recognized in the third window identification set of the cone database; Step S46: Determine all window identifications in the final window identification set.

[0106] In step S41, the system provides a convenient interaction interface to support the user in uploading the text to be recognized and the image to be recognized that contain the information of the expected view outside the window. The user can upload the text description of the expected view outside the window (such as being able to see the sea and the beach) and the relevant pictures (such as pictures of the sea and the beach) through the mobile terminal or the web terminal. The system will perform format verification and preliminary preprocessing on the uploaded files to ensure the integrity and availability of the data.

[0107] In step S42, the system uses an image feature extraction algorithm to deeply analyze the image to be recognized and extract its image vector. This image vector is a numerical representation of the image content and contains key feature information such as the color, texture, and shape of the image. After the extraction is completed, the system will perform a search in the vector database. By calculating the similarity (such as cosine similarity) between the image vector and the visual vectors stored in the database, it finds the window identifications corresponding to the vectors with higher similarity to form the first window identification set.

[0108] If landmark entities (such as the Eiffel Tower, the Oriental Pearl Tower, etc.) are recognized in the extracted image vector through the target detection algorithm, the system will automatically extract keywords from the names of the landmark entities and add these keywords to the text to be recognized. For example, if the Eiffel Tower is recognized in the image, the system will add the keyword "Eiffel Tower" to the original text to be recognized, further enriching the query conditions and improving the accuracy of the search.

[0109] In step S43, for the first window identification set, the system filters in the attribute database. The system uses natural language processing technology to parse the text to be recognized, extracts key information such as greening rate, occlusion rate requirements, landmark names, etc. Then, the system queries the corresponding window information in the attribute database according to these conditions. For numerical conditions such as greening rate and occlusion rate, the system directly compares the visual environment indicators of the windows in the attribute database; for text conditions such as landmark names, the system performs text matching operations. Only the identifications corresponding to the windows that meet all the extracted conditions will be included in the second window identification set, thus further narrowing the query scope.

[0110] In step S44, based on the second window identification set, the system conducts in-depth mining in the knowledge graph database. The knowledge graph database constructs a complex relationship network with entities, buildings, blocks, and windows as nodes and the visible entities, the buildings where the windows are located, and the blocks where the windows are located as edges. The system will search for the relevant target nodes and target edges in the knowledge graph according to the window identifications. For example, searching for the "visible" relationship edge between the window and the landmark entity, as well as the relationships between the window and the building and block to which it belongs. Through the sorting and screening of these relationships, the system can obtain window information more relevant to the query conditions and construct the third window identification set. This process makes full use of the rich semantic relationships in the knowledge graph database, further improving the accuracy and relevance of the query results.

[0111] In step S45, based on the information such as the shooting perspective and position of the image to be recognized (if there is relevant metadata), the system simulates and generates the cone image to be recognized. This cone image represents the scene range and content that can be observed from the perspective of the user's captured image. Then, in the cone database, for each window in the third window identification set, the system compares the similarity between the corresponding cone image and the cone image to be recognized. The similarity calculation can use methods such as Hausdorff distance comparison. Only the identifications corresponding to the windows whose cone image similarity reaches a certain threshold will be included in the final window identification set, thus ensuring that the query results highly match the user's expectations in terms of visual effects and visible range.

[0112] In step S46, the system obtains all the window identifications in the final window identification set. For these windows, a learning-to-rank model is used for comprehensive scoring. The learning-to-rank model comprehensively considers multiple features such as perspective similarity, landmark confidence, and inverse occlusion rate. Among them, the perspective similarity is determined by comparing the similarity between the cone image to be recognized and the actual cone image of the window; the landmark confidence is evaluated according to the probability and accuracy of the visible landmark entity of the window; the inverse occlusion rate is to take the reciprocal of the occlusion rate of the window to ensure that the window with a lower occlusion rate has more advantages in the scoring.

[0113] By learning the calculations of the ranking model, the system ranks these windows and generates a final result list. At the same time, for the convenience of users' understanding and use, the system attaches Chinese interpretable explanations to each result in the list, elaborating in detail the reasons why the window meets the query conditions, as well as a confidence value, which is used to represent the matching degree between the result and the user's query conditions, enabling users to intuitively understand the reliability and relevance of each query result.

[0114] Exemplarily, a user plans to rent a house near a tourist scenic area and hopes to see green mountains and a landmark ancient temple outside the window, with an open view and few obstructions. The user uploads a visible (image to be recognized) containing green mountains and the ancient temple to the system and enters "There are green mountains outside the window and the view is open" (text to be recognized). The system first extracts the image vector, finds the first set of window identifiers corresponding to similar visual vectors in the vector database. Since the image contains the ancient temple landmark, the system adds "ancient temple" to the text to be recognized. Then, in the attribute database, it filters out the second set of window identifiers with a low obstruction rate and visible green mountains and ancient temples according to the new text. Then, in the knowledge graph database, it searches and constructs the third set of window identifiers with the triples (window - visible - green mountain, window - visible - ancient temple). Then, it generates the image of the visual cone to be recognized by simulating the photo, and finds the final set of window identifiers that match it in the visual cone database. Finally, the system comprehensively scores and ranks these windows according to features such as the similarity of the viewing angle, the confidence of the landmark, and the inverse of the obstruction rate, and generates a result list with Chinese explanations and confidence values, such as "Window W1, confidence 85%, can clearly see green mountains and the ancient temple, with an open view; Window W5, confidence 70%, can see part of the green mountains and the ancient temple, with a relatively open view", etc., to help users quickly understand the matching degree of each window with their needs.

[0115] In this application, the system integrates information in two modalities, text and image, for retrieval. When the user holds a mobile device and aims it at the target building, the system quickly analyzes the captured image, accurately identifies the candidate windows. When the user inputs the desired window scene by voice, the system can respond quickly and accurately, making full use of the semantic description of the text to be recognized and the intuitive visual information of the image to be recognized. Through hierarchical screening and matching, it obtains relevant information from different types of databases, achieving the instant interaction effect of "what you shoot is what you get". Users do not need to wait for a long time and can obtain the required window information instantly when they raise the device, improving the accuracy and efficiency of the retrieval.

[0116] This application provides an overall process for finding windows for specific landscapes, including the following steps.

[0117] 1. Data preparation: Collect the building information of the target area, and at the same time obtain a multi - perspective image set, such as from drone aerial photography, ground shooting, etc. Organize this data to prepare for building a 3D model.

[0118] 2. Construct a 3D model: Extract feature points from multi-view images and match them. Use the "structure from motion" method to recover camera parameters and obtain a sparse 3D point cloud. Then use a multi-view density recovery algorithm to densify the point cloud and generate a textured 3D model through Poisson reconstruction, which is integrated with the building information.

[0119] 3. Determine the window dataset (image path): Process the frame images with a semantic segmentation model to obtain an initial window frame mask, and track consecutive frames to determine the final window frame mask. Extract the 2D window frame corner points and back-project them onto the 3D model in combination with the camera parameters to obtain the 3D window frame contour under the image path.

[0120] 4. Determine the window dataset (point cloud path): Input the dense point cloud of the target area into a point cloud instance segmentation network to obtain window point clusters. Randomly select points to fit an initial plane, and after repeating the operation multiple times, determine the target plane. Determine the window geometric space representation data based on the distribution of window point clusters on the target plane.

[0121] 5. Integrate and determine the window dataset: Project the 3D window frame contour of the image path and the window point clusters of the point cloud path onto a 2D plane and calculate the overlap degree. If the overlap degree exceeds the preset threshold, use the geometric space representation data of the point cloud path as window data to construct a window dataset.

[0122] 6. Determine the visible field data and construct a database: Determine the viewing frustum according to the outward normal vector of the window and the preset field of view angle, and project it onto the 3D model to obtain the visible field data. Generate a viewing frustum image based on the visible field data and construct a viewing frustum database; extract visual vectors to construct a vector database; integrate various data to construct an attribute database and a knowledge graph database.

[0123] 7. Multi-modal query processing: Obtain the text to be recognized and the image to be recognized uploaded by the user, extract the image vector of the image to be recognized and search for a first set of window identifiers that match in the vector database, then filter in the attribute database to obtain a second set of window identifiers, then process in the knowledge graph database to obtain a third set of window identifiers, and finally find the final set of window identifiers in the viewing frustum database to complete the processing through hierarchical query.

[0124] Based on the same technical concept, this application provides a device for finding windows for a specific landscape, as Figure 7 shown, the device includes: An acquisition and construction module 701, configured to acquire building information and a multi-view image set in a target area, and construct a 3D model of the target area according to the building information and the image set; A first determination module 702, configured to determine a window dataset by analyzing the frame images in the image set, where the window dataset includes geometric space representation data of each window in the target area; The second determination module 703 is configured to determine a viewing frustum of a window according to a preset field of view angle and an outer normal vector in the geometric space representation data, and determine visible region data of the window generated by projecting the viewing frustum onto a three-dimensional model, where the visible region data is used to indicate the landscape outside the window; The construction module 704 is configured to construct a visible region database of all windows in the target area according to the building information, the geometric space representation data of each window, and the visible region data; The search module 705 is configured to search for windows that match the multimodal query information in the visible region database according to the multimodal query information uploaded by the user, where the multimodal query information is used to query windows with a specific landscape outside the window.

[0125] Optionally, the first determination module 702 is configured to: Back-project the two-dimensional window frame corner points in the frame image onto the three-dimensional model to obtain a three-dimensional window frame contour under the image path, where the three-dimensional window frame contour is used to indicate the shape contour of the window in the three-dimensional space; Determine a window point cluster of the window under the point cloud path according to the dense point cloud of the target area, and determine the geometric space representation data of the window according to the window point cluster, where the window point cluster is used to indicate the point cloud data of multiple windows; Determine the overlap degree of the three-dimensional window frame contour under the image path and the window point cluster under the point cloud path on the two-dimensional plane; If the overlap degree is greater than a preset overlap degree threshold, determine that the three-dimensional window frame contour and the window point cluster are the same window, and use the geometric space representation data under the point cloud path as the geometric space representation data of the window; Construct a window data set according to the geometric space representation data of multiple windows.

[0126] Optionally, the first determination module 702 is configured to: Perform semantic segmentation on the frame image through a semantic segmentation model to determine an initial window frame mask in the frame image; Determine a final window frame mask of the window by tracking the same window in multiple consecutive frame images; Extract the two-dimensional window frame corner points in the final window frame mask; Back-project the two-dimensional window frame corner points onto the three-dimensional model according to the camera parameters corresponding to the frame image to obtain a three-dimensional window frame contour of the window under the image path, where the camera parameters include shooting parameters and camera pose parameters.

[0127] Optionally, the first determination module 702 is configured to: Repeat the following operations: Fit an initial plane according to some randomly selected points from the window point cluster, and determine the distance from each point in the window point cluster to the initial plane; Determine the number of points with a distance less than a preset distance threshold, where the initial plane is used to indicate a wall with multiple windows; After reaching the set number of repeated operations, the initial plane with the largest number of points is used as the target plane; According to the distribution characteristics of the window point clusters in the target plane, the geometric space representation data of each window is determined.

[0128] Optionally, the first determination module 702 is used for: Project the three-dimensional window frame contour under the image path onto a two-dimensional plane to obtain a polygon contour; Project the window point clusters under the point cloud path onto a two-dimensional plane to obtain a projected point cluster; Select the internal point cluster that falls within the polygon contour from the projected point cluster; Calculate the overlap degree according to the area of the region formed by the internal point cluster and the area of the region formed by the polygon contour.

[0129] Optionally, the construction module 704 is used for: Determine the frustum image of the visible field data, and construct a frustum database according to the frustum images of multiple windows; Generate visual vectors according to the frustum images of the windows, and construct a vector database according to the visual vectors of multiple windows; Determine the building floor data where the window is located according to the preset building information, determine the visible environment index according to the visible field data, and construct an attribute database by combining the building floor data, geometric space representation data, and visible environment index where multiple windows are located; Construct a knowledge graph database, where the knowledge graph database has entities, buildings, blocks, and windows as nodes, and the visible entities, the buildings where they are located, and the blocks where they are located of the windows as edges; Among them, the frustum database, the vector database, the attribute database, and the knowledge graph database all include window identifiers.

[0130] Optionally, the search module 705 is used for: Obtain the text to be recognized and the image to be recognized uploaded by the user, where the text to be recognized and the image to be recognized contain the outdoor landscape that the user expects to see; Determine the image vector of the image to be recognized, and search for the first set of window identifiers that match the image vector in the vector database. If the image vector contains landmark entities, add landmark keywords to the text to be recognized; Search for the second set of window identifiers that meet the text to be recognized in the first set of window identifiers in the attribute database; Search for the target nodes and target edges that meet the second set of window identifiers in the knowledge graph database, and construct the third set of window identifiers according to the target nodes and target edges; Determine the cone image to be recognized in the image to be recognized, and search for the final window identification set that matches the cone image to be recognized in the third window identification set of the cone database; Determine all window identifications in the final window identification set.

[0131] As Figure 8 shown, an embodiment of the present application provides an electronic device, including a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 complete mutual communication through the communication bus 804.

[0132] The memory 803 is used to store computer programs.

[0133] In an embodiment of the present application, when the processor 801 executes the program stored on the memory 803, it implements the method for finding windows for a specific landscape provided by any one of the foregoing method embodiments.

[0134] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method for finding windows for a specific landscape provided by any one of the foregoing method embodiments.

[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0137] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order described or illustrated, unless an execution order is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0138] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for finding windows for a specific landscape, characterized in that, The method includes: Obtaining building information and an image set with multiple perspectives in a target area, and constructing a three-dimensional model of the target area based on the building information and the image set; Determining a window data set by analyzing frame images in the image set, where the window data set includes geometric space characterization data of each window in the target area; Determining a viewing frustum of the window according to a preset field of view angle and an outward normal vector in the geometric space characterization data, and determining visible area data of the window generated by projecting the viewing frustum onto the three-dimensional model, where the visible area data is used to indicate the landscape outside the window; Constructing a visible area database of all windows in the target area according to the building information, geometric space characterization data of each window, and visible area data; Searching for windows that match the multimodal query information in the visible area database according to the multimodal query information uploaded by the user, where the multimodal query information is used to query windows with a specific landscape outside the window.

2. The method according to claim 1, characterized in that Determining the window data set by analyzing frame images in the image set includes: Back-projecting two-dimensional window frame corner points in the frame image onto the three-dimensional model to obtain a three-dimensional window frame contour in the image path, where the three-dimensional window frame contour is used to indicate the shape contour of the window in three-dimensional space; Determining a window point cluster in the point cloud path according to the dense point cloud of the target area, and determining geometric space characterization data of the window according to the window point cluster, where the window point cluster is used to indicate point cloud data of multiple windows; Determining the overlap degree between the three-dimensional window frame contour in the image path and the window point cluster in the point cloud path on a two-dimensional plane; If the overlap degree is greater than a preset overlap degree threshold, determining that the three-dimensional window frame contour and the window point cluster are the same window, and using the geometric space characterization data in the point cloud path as the geometric space characterization data of the window; Constructing the window data set according to the geometric space characterization data of multiple windows.

3. The method according to claim 2, wherein Back-projecting two-dimensional window frame corner points in the frame image onto the three-dimensional model to obtain a three-dimensional window frame contour in the image path includes: Performing semantic segmentation on the frame image through a semantic segmentation model to determine an initial window frame mask in the frame image; Determining a final window frame mask of the window by tracking the same window in multiple consecutive frame images; Extracting two-dimensional window frame corner points from the final window frame mask; Back-projecting the two-dimensional window frame corner points onto the three-dimensional model according to camera parameters corresponding to the frame image to obtain the three-dimensional window frame contour of the window in the image path, where the camera parameters include shooting parameters and camera pose parameters.

4. The method according to claim 2, wherein Determining geometric space characterization data of the window according to the window point cluster includes: Repeating the following operations: fitting an initial plane according to a part of points randomly selected from the window point cluster, and determining the distance from each point in the window point cluster to the initial plane; determining the number of points with a distance less than a preset distance threshold, where the initial plane is used to indicate a wall surface with multiple windows; After reaching the set number of repeated operations, the initial plane with the largest number of points is used as the target plane; According to the distribution characteristics of the window point clusters in the target plane, determine the geometric space representation data of each window.

5. The method according to claim 2, characterized in that, Determining the overlap degree between the three-dimensional window frame contour in the image path and the window point clusters in the point cloud path on a two-dimensional plane includes: Project the three-dimensional window frame contour in the image path onto a two-dimensional plane to obtain a polygon contour; Project the window point clusters in the point cloud path onto the two-dimensional plane to obtain a projected point cluster; Select the internal point cluster that falls within the polygon contour from the projected point cluster; Calculate the overlap degree based on the area of the region formed by the internal point cluster and the area of the region formed by the polygon contour.

6. The method according to claim 1, characterized in that, Constructing a visible area database for all windows in the target area according to the building information, the geometric space representation data of each window, and the visible area data includes: Determine the cone-of-vision images of the visible area data and construct a cone-of-vision database based on the cone-of-vision images of multiple windows; Generate visual vectors based on the cone-of-vision images of the windows and construct a vector database based on the visual vectors of multiple windows; Determine the building floor data where the window is located according to the preset building information, determine the visible environment indicators according to the visible area data, and construct an attribute database by combining the building floor data where multiple windows are located, the geometric space representation data, and the visible environment indicators; Construct a knowledge graph database, where the knowledge graph database uses entities, buildings, blocks, and windows as nodes, and uses the visible entities, the buildings where the windows are located, and the blocks where the windows are located as edges; Among them, the cone-of-vision database, the vector database, the attribute database, and the knowledge graph database all include window identifiers.

7. The method according to claim 6, characterized in that Searching for windows that match the multi-modal query information in the visible area database according to the multi-modal query information uploaded by the user includes: Obtain the text to be recognized and the image to be recognized uploaded by the user, where the text to be recognized and the image to be recognized contain the outdoor landscape that the user expects to see; Determine the image vector of the image to be recognized and search for a first set of window identifiers that match the image vector in the vector database. If the image vector contains landmark entities, add landmark keywords to the text to be recognized; Search for a second set of window identifiers that satisfy the text to be recognized in the first set of window identifiers in the attribute database; Search for target nodes and target edges that satisfy the second set of window identifiers in the knowledge graph database, and construct a third set of window identifiers based on the target nodes and the target edges; Determine the cone-of-vision image to be recognized of the image to be recognized and search for a final set of window identifiers that match the cone-of-vision image to be recognized in the third set of window identifiers in the cone-of-vision database; Determine all window identifiers in the final set of window identifiers.

8. A device for finding windows for a specific landscape, characterized in that, The device includes: An acquisition and construction module, configured to acquire building information and a multi-view image set in a target area, and construct a three-dimensional model of the target area according to the building information and the image set; A first determination module, configured to determine a window data set by analyzing frame images in the image set, where the window data set includes geometric space characterization data of each window in the target area; A second determination module, configured to determine a viewing frustum of the window according to a preset field of view angle and an outward normal vector in the geometric space characterization data, and determine visible area data of the window generated by projecting the viewing frustum onto the three-dimensional model, where the visible area data is used to indicate the landscape outside the window; A construction module, configured to construct a visible area database of all windows in the target area according to the building information, the geometric space characterization data of each window, and the visible area data; A search module, configured to search for windows that match the multimodal query information in the visible area database according to the multimodal query information uploaded by the user, where the multimodal query information is used to query windows with a specific landscape outside the window.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Multi-view visual field measuring and calculating method and system based on Cesium

    CN119205482A

  • Point cloud labeling method, apparatus, and system, device, and storage medium

    WO2021114884A1

Cited By

  • Low-altitude remote sensing image processing method and device, storage medium and computer equipment

    CN121032868A