An old-age home environment risk assessment method and system based on digital twinning
Patent Information
- Application Number
- CN202610802060.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-08-28
AI Technical Summary
[0005]为了至少解决上述背景技术中现有适老化评估中人工测量效率低下、主观性强,2D图像方法缺乏深度信息的缺陷,通过构建多任务微调模型,自动、精准地对室内环境进行空间几何、安全设施、活动可达性及细节隐患四大维度的合规性检测与可视化标注
A、全维度覆盖:本发明创造性地将适老化改造标准解构为四个功能模块,不仅覆盖了刚性的几何尺寸(模块一),还覆盖了语义层面的设施完整性(模块二)和复杂的空间逻辑(模块三、四),实现了国标规范的数字化映射。
Smart Images

Figure CN122657421A_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of digital twin technology. More specifically, this invention relates to a method and system for risk assessment of the home environment for the elderly based on digital twins. Background Technology
[0002] Potential risks in the home environment are one of the main factors leading to falls and other accidents among the elderly. Conducting scientific and accurate risk assessments of the elderly's home environment and making targeted modifications is a crucial step in ensuring their quality of life and safety.
[0003] Current environmental risk assessment mainly includes four types of models. The first type, traditional models, primarily includes methods such as pixel analysis, multi-criteria decision analysis, weight determination and data fusion, and mathematical modeling. These can only handle relatively simple mapping relationships and are difficult to extract deep semantic information. The second type, machine learning models, commonly use BP neural networks, Bayesian networks, and random forests for evaluation, and also combine statistical methods and decision trees for risk classification. Although they perform well in simple pattern recognition and classification problems, their algorithmic limitations may make them unsuitable for complex and dynamic environments. The third type, unimodal models in deep learning, has seen researchers propose hybrid neural networks such as CNN-LSTM, instance segmentation, and object detection. These models assess risk levels by analyzing the relative positions of people and dangerous objects. Unimodal models have good versatility and efficiency in safety assessment and object detection, but they can only handle one type of data, limiting their application potential in more complex or dynamic environments. The fourth category is multimodal models, such as those that use affine transformation techniques and fusion modules to combine infrared and RGB sensor data, align features, and reduce differences between sensors to achieve efficient target recognition. The integration of information from different perspectives can provide a more comprehensive and detailed view, significantly improving the overall assessment capability of scene safety and ensuring that the assessment results more accurately reflect the actual situation.
[0004] However, existing risk assessment methods for elderly people living in their homes have the following drawbacks: 1. Computational complexity prevents real-time online processing; there is a trade-off between voxel resolution and physical simulation accuracy. Physical simulations use a uniform material assumption and global friction coefficient, ignoring the differences in the actual material properties of objects. Furthermore, simplifying boundary shapes to regular geometric primitives introduces shape approximation errors, affecting the realism of mass distribution and collision detection. 2. Limited generalization: Due to factors such as small dataset size, simple scenarios, limited coverage of object categories and risk patterns, and lack of human activity, it is difficult to support dynamic risk assessment in real home environments, and its generalization ability is limited in complex environments. 3. Lack of reliable indoor positioning: GPS-based multimodal solutions rely on GPS coordinates in assessing the risk of elderly people getting lost. However, GPS signals are weak, inaccurate, or even malfunctioning in indoor environments, making it impossible to accurately capture the location information of elderly people indoors. Summary of the Invention
[0005] To address the shortcomings of existing age-friendliness assessment methods, such as low efficiency and high subjectivity of manual measurements, and the lack of depth information in 2D image methods, this invention proposes a digital twin-based method and system for risk assessment of elderly home environments. This system constructs a multi-task fine-tuning model to automatically and accurately perform compliance detection and visual annotation of the indoor environment across four dimensions: spatial geometry, safety facilities, activity accessibility, and potential hazards. In view of this, the invention provides solutions in the following aspects. The first aspect of this invention provides a method for risk assessment of elderly-friendly home environments based on digital twins, comprising: acquiring indoor 3D point cloud data; after preprocessing the data, inputting it into a point cloud large language model based on the Transformer architecture for structured scene understanding to obtain structured semantic output including the geometric layout of all walls in the scene, the position and size of doors and windows, and semantic category labels and 3D bounding box parameters of various indoor objects; based on the structured semantic output and the 3D point cloud data, performing the following safety checks: spatial geometric compliance check, based on computational geometry and rasterization analysis technology, used to process strongly quantitative indicators with clear numerical thresholds; safety facility integrity check, based on target detection and existence verification, used to determine whether specific age-friendly facilities are missing; spatial accessibility and activity space check, based on volume collision detection and virtual agent simulation, used to assess knee room and wheelchair turning; furniture and detail hazard check, used to detect unstructured, micro-scale, or visually perceptually dependent age-friendly safety hazards in the indoor environment; and visually overlaying the detection results of the safety checks in a 3D scene.
[0006] A second aspect of the present invention provides a digital twin-based home environment risk assessment system for the elderly, utilizing any of the aforementioned digital twin-based home environment risk assessment methods for the elderly.
[0007] The beneficial effects of this invention include: A. Full-dimensional coverage: This invention creatively deconstructs the standards for age-friendly renovation into four functional modules, which not only cover rigid geometric dimensions (module one), but also cover the semantic level of facility integrity (module two) and complex spatial logic (modules three and four), realizing the digital mapping of national standards.
[0008] B. High precision and anti-interference: Compared with traditional 2D image evaluation, point cloud-based methods are not affected by lighting conditions (such as accurate detection under a bed or in a bathroom in a dimly lit place) and can provide dimensional measurement accuracy at the centimeter or even millimeter level.
[0009] C. Intelligent semantic reasoning: The introduction of large model technologies such as PointGPT / SpatialLM (Module 4) solves the problem that traditional algorithms have difficulty in identifying abstract hidden dangers such as "messy lines" and "blurred step boundaries".
[0010] D. Automated Design Assistance: The system outputs not only risk reports, but can also directly mark the areas to be modified in the 3D model (such as highlighting the walls where handrails need to be installed), providing direct data support for subsequent decoration design. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 This demonstrates a method for assessing the risk of elderly people living in their homes. Figure 2 This shows the structure diagram of the algorithm model; Figure 3 This shows the system structure block diagram; Figure 4 This shows the method flow of modules one through four; Figure 5 This is a comparison chart showing the input / output effects; Figure 6 This shows the human-computer interaction interface. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this invention are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the specification and claims of this invention indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]." The first aspect of this invention provides a method for risk assessment of the home environment for the elderly. Specific embodiments of the invention will now be described in detail with reference to the accompanying drawings. Figure 1 The diagram shown is a flowchart of the evaluation method of the present invention. Figure 2 The diagram shown is a structural diagram of the algorithm model of this invention.
[0013] From the appendix Figure 1 , 2 As can be seen, the home environment risk assessment method of the present invention specifically includes steps S100-S300: Step S100: Obtain indoor 3D point cloud data. After preprocessing, the data is input into a point cloud big language model based on the Transformer architecture for structured scene understanding, so as to obtain structured semantic output including the geometric layout of all walls in the scene, the position and size of doors and windows, as well as the semantic category labels and 3D bounding box parameters of various indoor objects. Step S200: Based on the structured semantic output and the 3D point cloud data, perform the following security checks respectively: Spatial geometry compliance testing, based on computational geometry and rasterization analysis techniques, is used to process strongly quantitative indicators with clearly defined numerical thresholds; Safety facility integrity testing, based on target detection and existence verification, is used to determine whether specific age-friendly facilities are missing; Spatial accessibility and activity space detection, based on volume collision detection and virtual agent simulation, is used to evaluate knee space and wheelchair turning. Furniture and detail hazard detection is used to detect unstructured, micro-scale, or visually perceptible safety hazards in the indoor environment that may affect age-friendliness. Step S300: Visualize and overlay the detection results of the security detection in a three-dimensional scene.
[0014] The following is in conjunction with the appendix Figure 3 , 4 The apparatus described above is explained in detail below. It will be understood that the method and apparatus of the present invention have corresponding principles and implementation methods. (See attached diagram.) Figure 3 This is a system structure block diagram illustrating the present invention. Figure 4 This illustrates the method flow of modules one to four of the present invention.
[0015] This invention first inputs indoor 3D point cloud data ( Each point contains three-dimensional spatial coordinates (x, y, z) and corresponding color information (r, g, b). After preprocessing, the data is input into a spatial large language model based on the Transformer architecture for structured scene understanding.
[0016] In a preferred embodiment of the present invention, the data sources include a benchmark dataset and a proprietary dataset.
[0017] Benchmark datasets: Large-scale publicly available indoor 3D datasets (such as ScanNet V2 with 1,513 indoor scenes, S3DIS with 6 building regions, and Matterport3D with 90 buildings) were used for pre-training of the point cloud encoder. The pre-training method employed a self-supervised learning strategy: a large amount of unlabeled indoor point cloud data was input into the encoder, and the encoder was trained to learn general spatial geometric feature representations through self-supervised objective functions such as masked reconstruction and contrastive learning, without the need for manual annotation. After pre-training, the encoder possessed a basic perception ability for common structures such as walls, floors, and furniture in indoor spaces. Based on this, the encoder, projector, and large language model were combined into a complete Encoder-MLP-LLM architecture. End-to-end supervised fine-tuning was performed on a large-scale synthetic dataset containing structured annotations (wall coordinates, door and window coordinates, object bounding boxes, and category labels) to train the model's ability to generate structured scene descriptions.
[0018] Our proprietary dataset: We collected real-world home environment data of the elderly through MetaCam-EDU; we generated images of the indoor environment and specific items of the elderly’s homes through Flux-dev and collected them from the internet; and we converted the images into 3D point clouds using TRELLIS / Argus1.0.
[0019] Data types include point cloud data and unstructured text: Point cloud data: Format is... This includes spatial coordinates and RGB color information. Unstructured text (Text): A description of the age-friendly modification specifications for the corresponding scene, used as the Prompt input.
[0020] In a preferred embodiment of the present invention, the above data preprocessing includes: outlier removal, normalization, and image enhancement, specifically: 1. Outlier Removal: A statistical filter (Statistical Outlier Removal, SOR) is used. The specific steps are as follows: Step 1: For each point p_i in the point cloud, use the KD-Tree spatial index to search for its k nearest neighbor points (k is a preset parameter, typically 50).
[0021] Step 2: Calculate the sum of the Euclidean distances from point p_i to its k neighboring points, divide by k, and obtain the average neighborhood distance d_i = (1 / k)·Σ‖p_i-p_j‖2.
[0022] Step 3: Calculate the global mean μ and standard deviation σ of the average neighborhood distance {d_i} of all points.
[0023] Step 4: If the average neighborhood distance d_i of a point exceeds the range of [μ-α·σ,μ+α·σ] (α is a preset multiple, typically 2.0), then the point is identified as an outlier and removed.
[0024] The above steps can effectively filter out sparse flying point noise caused by reflection, multipath effect, etc. during the scanning process.
[0025] 2. Normalization: The specific steps are as follows: Step 1: Coordinate translation — Calculate the minimum values (x_min, y_min, z_min) of all points in the point cloud on the X, Y, and Z axes. Subtract the corresponding minimum value from the coordinates of each point so that the origin of the point cloud is moved to the boundary corner and all coordinate values become non-negative.
[0026] Step 2: Coordinate Quantization – The translated continuous coordinate values are discretized into integer values at a fixed resolution (e.g., 2.5cm / bin), i.e., q = round(coord / resolution), with a quantization range of [0, 1280]. The quantized integer coordinates can be directly used as text tokens for large language models.
[0027] Step 3: Inverse Transformation – The quantized integer coordinates output by the model are restored to continuous values in the original physical coordinate system through the inverse operation coord=q×resolution+offset.
[0028] 3. Image enhancement: Random rotation, scaling, jittering, and random discarding of color channels are applied to point cloud data to enhance the model's generalization ability under different lighting conditions and apartment layouts.
[0029] In one embodiment of the present invention, the spatial large language model adopted follows the standard multimodal architecture of "point cloud encoder-projector-large language model" (Encoder-MLP-LLM), and its processing flow is as follows: Step S1: Point Cloud Encoding. Input the preprocessed N six-dimensional points into a point cloud encoder (a network based on Transformer self-attention mechanism, such as the Sonata encoder or Point Transformer V3, can be used). The encoder adopts a five-level hierarchical structure, progressively reducing the spatial resolution of the point cloud through serialization, embedding, and multi-level encoding blocks (each level halves the resolution along each dimension through Grid Pooling), while progressively increasing the feature dimensions (e.g., from 48 dimensions to 96, 192, 384, and 512 dimensions), finally outputting K D-dimensional feature embedding vectors F∈ K×D, where K is much smaller than N (typically about 1 / 100 to 1 / 500 of the number of input points).
[0030] Step S2: Feature Projection. Using a two-layer multilayer perceptron (MLP) projector, the K feature embedding vectors output by the point cloud encoder are mapped from the point cloud feature space to the word embedding space of the large language model, so that the point cloud visual tokens and language tokens are aligned in the same semantic space.
[0031] Step S3: Autoregressive Text Generation. The projected visual token sequence is concatenated with text prompts (e.g., "Detect walls, doors, windows, bboxes") and input into a large language model (open-source pre-trained LLMs such as Qwen2.5 or Llama can be used). The model generates structured scene description text token by token in an autoregressive manner. The output is in the format of a general-purpose programming language (Python) data script, containing the following four types of structured elements: (a) Wall: includes the starting coordinates (a_x, a_y, a_z), the ending coordinates (b_x, b_y, b_z), and the height. (b) Door: Includes the ID of the wall it belongs to (wall_id), the coordinates of its center position (position_x, position_y, position_z), its width (width), and its height (height); (c) Window: Same structure as Door, including the wall it belongs to, location, width and height; (d) Object bounding box (Bbox): contains semantic category labels (class, such as bed, toilet, sofa, chair, etc., 59 indoor object categories), center position coordinates (position_x, position_y, position_z), rotation angle around the Z-axis (angle_z), and three-dimensional dimensions (scale_x, scale_y, scale_z).
[0032] All the coordinates mentioned above have been quantized, mapping continuous coordinate values to discrete integers (e.g., 1280 quantization intervals with a resolution of 2.5cm), which makes it easier for LLM to accurately express spatial location in the form of text tokens.
[0033] Through steps S1-S3, the system automatically obtains the geometric layout of all walls, the location and dimensions of doors and windows, and the semantic category labels and 3D bounding box parameters of various indoor objects in the scene. These structured outputs are then passed to the subsequent four functional detection modules as the spatial semantic basis for risk assessment. Specifically, the coordinates of Walls and Doors are used for geometric compliance detection in Module 1 (ground elevation difference positioning, door clear width measurement), the category labels and coordinates of Bboxes are used for facility integrity detection in Module 2 (association retrieval of anchor furniture and auxiliary facilities), and accessibility assessment in Module 3 (furniture positioning and virtual probe deployment).
[0034] In one embodiment of the present invention, based on the above-described structured semantic output and raw point cloud data, the system performs risk analysis using the following four dedicated functional detection modules: Module 1: Geometric Compliance Module This module is primarily based on computational geometry and rasterization analysis techniques to handle highly quantitative indicators with clearly defined numerical thresholds. It corresponds to the following clauses in the "General Requirements for Age-Friendly Home Environment Renovation for the Elderly" (MZ / T 218-2024): Clause 5.1.1.1 (The height difference between the bedroom and living room floors should not exceed 15mm), Clause 5.1.3.1 (The clear passage width of doors for wheelchair-bound elderly people after opening should not be less than 800mm), Clause 5.1.2.2 (The passage width of indoor corridors should comply with the relevant requirements of GB 50763 and GB 55019), Clause 5.1.11.3 (Height of washbasins and legroom), and Clause 5.1.12.3 (Bayside railing anti-fall measures; no gaps should be left within 350mm of the ground).
[0035] Preliminary localization: The spatial large language model (i.e., the "point cloud encoder-projector-large language model" architecture described in steps S1-S3 above, hereinafter referred to as "point cloud LLM") already includes the coordinates of Door elements (including the wall to which it belongs, center position (position_x, position_y, position_z), width, and height) and the coordinates of various object bounding boxes (Bboxes) (including center position and 3D dimensions) in the structured scene description output by autoregression in step S3. Among them, the Door coordinates are used for the preliminary localization of the clear width of the doorway; the area enclosed by each Bbox is the Zone (functional area), for example, the area where the Bbox with the category label "toilet" is located is the toilet area, and the area where the Bbox with the label "bed" is located is the bedroom sleeping area.
[0036] The criteria for point cloud LLM are as follows: The point cloud encoder encodes the spatial geometric features and color features of the original point cloud into high-dimensional semantic embedding vectors. After being mapped to the word space of a large language model by a projector, the LLM generates a sequence of structured parameters describing each building element and object in the scene autoregressively based on the spatial knowledge it has obtained by pre-training on a large-scale indoor scene dataset and the detection instructions in the text Prompt.
[0037] The post-processing module defines the Region of Interest (ROI) as follows: Using the Door coordinates or Bbox coordinates output by the LLM point cloud as the center, a pre-defined 3D bounding area is extracted from the original high-precision point cloud. The size of the ROI is determined based on the detection task: for door width measurement, the ROI is a rectangular area with the door center as the origin, extending outwards by one door width along both the door normal and the wall direction; for ground elevation difference detection, the ROI is an area with the Zone center as the origin, covering the entire Zone plane and extending outwards by 500mm. By performing detailed geometric analysis within the ROI, the computationally expensive point-by-point processing of the entire scene's point cloud is avoided.
[0038] Ground elevation difference calculation: The specific detection and calculation steps are as follows: (a) Ground height map construction: The point cloud within the ROI is projected onto the horizontal plane (XY plane) at a fixed grid resolution (e.g., 5cm×5cm). Each grid cell records the lowest Z coordinate value of all points within it as the ground elevation of that grid, generating a two-dimensional height map.
[0039] (b) Elevation difference detection: Apply the Sobel gradient operator or the Canny edge detection operator to the height map and calculate the absolute value of the elevation difference |ΔZ| between each grid cell and its adjacent grid cells.
[0040] (c) Threshold determination and risk marking: Grids whose |ΔZ| exceeds the preset threshold Threshold (according to Article 5.1.1.1 of MZ / T 218-2024, Threshold is set to 15mm) are marked as pixels with height difference risk.
[0041] (d) Morphological post-processing: Apply morphological closing operation (dilation followed by erosion) to the risk pixel map, connect adjacent risk grids, eliminate isolated noise points, and form a complete connected risk region.
[0042] (e) Quantitative index extraction: For each connected risk region, calculate the centroid coordinates (as spatial anchors for visualization annotation) and the maximum elevation difference Δh_max within the region (as a quantitative index of risk severity). If Δh_max > Threshold, the region is marked as elevation difference risk.
[0043] Channel / door clearance measurement: The specific algorithm steps are as follows: (a) Doorway location: Extract the coordinate parameters of the Door element from the output of step S3 of the point cloud LLM, and obtain the center position (px, py, pz), width w_door, height h_door, and the wall_id to which the door belongs. Calculate the normal vector n_wall of the wall containing the doorway based on the start and end coordinates of the Wall element associated with wall_id.
[0044] (b) Horizontal slice extraction: Extract horizontal point cloud slices near the height of the door center (e.g., within the range of z∈[pz, pz+h_door]) to reduce the dimensionality of the three-dimensional problem to two-dimensional planar analysis.
[0045] (c) Free space mask construction: Rasterize the horizontal slices on a two-dimensional plane at a fixed grid resolution (e.g., 1cm × 1cm). Mark the grids containing point clouds as "obstacles" (value 0) and the grids not containing point clouds as "free space" (value 1) to generate a two-dimensional free space mask.
[0046] (d) Euclidean Distance Transform (EDT): Apply an Euclidean distance transform to the free space mask and calculate the Euclidean distance d(i,j) from each free space pixel to the nearest obstacle pixel. The transformed value d(i,j) represents the distance of that point from the nearest wall or obstacle (unit: number of grids × resolution = mm).
[0047] (e) Calculation of net width: Within the free space area corresponding to the doorway, extract the maximum value d_max of the distance transformation along the direction of the wall normal vector n_wall, then the net width of passage at that location is calculated. = 2 × d_max. This is because the distance transformation value represents the distance to the nearest obstacle on one side, and multiplying it by 2 gives the total net width between the two opposing sides.
[0048] (f) Compliance determination: Compare with the standard threshold. According to Clause 5.1.3.1 of MZ / T 218-2024, if... If the width is less than 800mm, it is considered a risk of insufficient clearance for passage.
[0049] The above method is also applicable to the measurement of the net width of corridor passages. The difference is that the corridor section does not require door coordinate positioning. Instead, the distance transformation value is calculated segment by segment along the corridor axis to obtain the net width at the narrowest point.
[0050] Height logic verification: Fine-tuning point cloud LLM to identify point cloud clusters of balcony railings and washbasins, and calculating their highest points. With ground plane The difference is compared with the standard threshold (e.g., the railing needs to be >1100mm, and the point cloud density within the bottom 350mm needs to reach the threshold to prevent gaps).
[0051] Module 2: Safety Facility Integrity Module This module, based on target detection and existence verification, is used to determine whether specific age-friendly facilities are missing. It corresponds to the following clauses in the "General Requirements for Age-Friendly Home Environment Renovation for the Elderly" (MZ / T 218-2024): Clause 5.1.6 (Safety handrails should be installed according to the elderly person's physical condition, mode of movement, and route), Clause 5.1.9.3 (Guardrails should be installed at the foot of the bed, and handrails or combination handrails should preferably be installed around the bed), Clause 5.1.11.8 (Assistive handrails should be installed next to the toilet and in the shower area; the color of the handrails should contrast clearly with the bathroom wall color), Clause 5.1.7.5 (A shoe-changing stool should preferably be installed in the entryway; the shoe-changing stool should not be placed behind the door), and Clause 5.1.11.5 (A shower chair or shower stool should preferably be installed in an appropriate location; the height should meet the elderly person's requirements for sitting while bathing).
[0052] Preliminary localization: In the structured scene description output by Point Cloud LLM in step S3, each bounding box of an object is accompanied by a semantic category label (class field). The system parses this text output and filters out objects whose category labels belong to the predefined "anchor furniture" category set. For example, Point Cloud LLM may output the following structured text: bbox_5 = Bbox("toilet", 120, 85, 42, 0, 28, 30, 45) bbox_8 = Bbox("bed", 95, 160, 50, 90, 80, 100, 42) bbox_12 = Bbox("shower_room", 130, 78, 55, 0, 50, 48, 82) The first parameter, "toilet", "bed", and "shower_room", are semantic category labels. The system filters out objects from all output bounding boxes that belong to the anchor furniture categories {bed, toilet, shower, shower_room, tub, hand_sink}, and extracts their center coordinates as spatial anchors for subsequent related searches.
[0053] Related target retrieval: Within a preset neighborhood radius R of key furniture (e.g., 600mm around a toilet), search for the existence of point cloud clusters belonging to the categories of "armrest", "shower chair", and "shoe changing stool".
[0054] Missing Facility Alarm: If a critical area (Anchor Region) exists, but the corresponding auxiliary facility detection confidence score is less than α, it is judged as "facility missing risk". To avoid the impact of false detections (such as misidentifying pipelines as handrails) or missed detections by point cloud LLM on the results, this system adopts a dual-channel fusion strategy of "semantic detection + point cloud geometric cross-validation" to calculate the confidence score, as follows: Channel 1: Semantic Detection Score The system collects all bounding boxes (BBoxes) output from the point cloud LLM within the neighborhood radius R of the anchor furniture, matches their class fields with the set of associated age-friendly facility categories required by that anchor point, and counts the number of matches. To prevent Due to false positives and inflated values, a spatial rationality filter is applied to each matched Bbox: (a) Size verification – checking whether the three-dimensional dimensions of the Bbox fall within the reasonable physical dimensions of the target facility (e.g., the cross-sectional diameter of the handrail should be between 25mm and 45mm, and the length between 300mm and 900mm; those exceeding this range are considered false positives and are discarded); (b) Position verification – checking whether the spatial orientation of the Bbox relative to the anchor furniture conforms to the installation logic (e.g., the toilet handrail should be located on both sides of the toilet and at a height between 650mm and 750mm; a "handrail" appearing 2m directly above the toilet is considered a false positive). The number of valid matches after filtering is recorded as follows: The semantic score is The upper limit is truncated to 1.0 to eliminate interference from duplicate detections.
[0055] Channel 2: Point Cloud Geometric Verification Score This channel does not rely on the semantic output of LLM, but directly extracts geometric evidence from the original point cloud and performs independent cross-validation on the semantic detection results. Specifically, within the neighborhood of the anchor furniture, detection rules are constructed based on the typical geometric features of the target age-friendly facility. Taking handrail detection as an example: a subset of points in the neighborhood point cloud located near the wall (30mm-60mm from the wall) and with a height between 600mm-800mm is extracted; the eigenvalues of the local covariance matrix of this subset are calculated, and linearity indices are selected. Points (i.e., points exhibiting tubular / rod-like geometric features); DBSCAN clustering is performed on the filtered points; if a continuous linear point cloud cluster with a length greater than 300mm exists, it is considered geometric evidence of the presence of a handrail. Take the number of this type of cluster and The ratio is truncated to 1.0; if no point cloud clusters satisfy the condition, then... =0.
[0056] Final confidence level fusion: .
[0057] in Default value In other words, the geometric verification channel has a higher weight than the semantic detection channel to improve robustness against LLM false positives. The Score has the highest confidence level when the two channels reach the same conclusion (both high or low scores). When the two channels reach contradictory conclusions (e.g., the semantic channel detects it but the geometric channel has no evidence, suggesting a possible false positive; or the semantic channel does not detect it but the geometric channel has evidence, suggesting a possible missed detection), the system marks the detection item as "requiring manual review" and outputs the sub-scores for both channels in the report for assessors to judge. The threshold α is set to 0.5 by default; when the Score < α, it is judged as "facility missing risk".
[0058] Fine-tuning strategy: To enable the spatial large language model to distinguish between items with similar functions but different age-appropriate attributes (such as "ordinary chair" and "shower chair," "ordinary countertop" and "shoe-changing stool"), this invention performs domain adaptation fine-tuning on the model. The fine-tuning targets are the output layer of the large language model in step S3 (i.e., the classification head responsible for generating object semantic category label tokens) and the parameters of the projector MLP. The parameters of the point cloud encoder can be frozen or jointly updated with a lower learning rate. The specific steps are as follows: Step 1: Construct an age-friendly fine-grained labeled dataset. Based on standard indoor scene point cloud data, annotators further subdivide the original coarse-grained categories into age-friendly subcategories according to the functional attributes and spatial locations of items. For example, the general category "chair" is subdivided into "dining_chair", "shower_chair", and "stool"; "table" is subdivided into "dining_table" and "dressing_table", etc. The format of each labeled sample is consistent with the output format of Step S3 above, that is, the class field of the Bbox in the Python data class script uses the subdivided subcategory label.
[0059] Step 2: Construct fine-tuning training data. Organize each labeled scene into a triplet of "point cloud input + text prompt + target output". The prompt instruction is "Detect walls, doors, windows, bboxes" or a custom instruction for age-friendly tasks (such as "Detect age-friendly facilities and their categories in this scene"). The target output is a structured scene description text containing fine-grained category labels.
[0060] Step 3: Efficient Parameter Fine-Tuning. A LoRA (Low-Rank Adaptation) strategy is employed, inserting low-rank adaptation matrices (typically rank r=16) into the attention layer of the large language model. Only the parameters of these adaptation matrices and the projector MLP (approximately 1%–3% of the total parameters) are trained, while the remaining pre-trained parameters are frozen. The training objective is the standard language modeling cross-entropy loss, which maximizes the probability of the model generating the target structured text given a point cloud visual token and a prompt.
[0061] Step 4: After fine-tuning, the model can output fine-grained age-friendly item category labels during the inference phase, thereby supporting the accurate determination of auxiliary facility types in Module 2.
[0062] Module 3: Accessibility & Activity Space Detection Subnet This module, based on volume collision detection and virtual agent simulation, is used to assess knee space and wheelchair turning. It corresponds to the following clauses in the "General Requirements for Age-Friendly Home Renovation for the Elderly" (MZ / T 218-2024): Clause 5.1.7.3 (For small entryways, sufficient space should be reserved to facilitate the entry and exit of stretchers, wheelchairs, and other assistive devices); Clause 5.1.8.1 (Living rooms should be open and have space for easy wheelchair turning); Clause 5.1.8.2 (Space should be reserved under countertops for wheelchairs and other walking aids to approach or turn); Clause 5.1.11.3 (Knee and foot space should be reserved under washbasins); and Clause 5.1.10.2 (In kitchens used by elderly people in wheelchairs, knee and foot space should be reserved under the countertop).
[0063] ROI extraction: Locate key areas such as the washbasin, cabinets, and entryway.
[0064] Virtual Geometric Probe: Knee space: In accordance with the relevant clauses of the "Barrier-Free Design Code" (GB 50763) and ADA 2010 §306, the standard dimensional parameters for knee space are: clear height not less than 650mm (from the ground), clear knee depth not less than 280mm (from the front edge of the countertop inwards), clear toe height not less than 230mm, and clear toe depth not less than 480mm. Based on this, a parameterized L-shaped cross-section virtual geometry Vknee is constructed under the countertop (instead of a simple cube, to more accurately simulate the different clearance requirements for knee and toe spaces). The system automatically determines the accessible direction of the countertop (i.e., the open surface away from the wall), projects the virtual geometry in that direction, and then detects whether there is environmental point cloud intrusion within the Vknee area (such as pipes, cabinet base plates, supporting structures, etc. under the countertop). If the volume of the intruding point cloud accounts for the proportion of the total volume of Vknee (i.e., the intrusion rate) exceeds a preset threshold (e.g., 5%), or if the intruding point cloud causes the net height / net depth to not meet the above-mentioned size requirements, then the knee space is deemed to be substandard.
[0065] Wheelchair rotation: Project a virtual cylinder with a diameter of 1500mm in the center area of the living room and detect the point cloud density of obstacles within the cylinder.
[0066] Stretcher simulation: The specific simulation steps are as follows: (a) Based on the free space two-dimensional mask of the lobby area, construct a virtual cuboid probe according to the external dimensions of the stretcher (the standard emergency stretcher is about 2000 mm long and 550 mm wide).
[0067] (b) In free space, starting from the entrance door and ending in the direction of the main indoor passage, a grid-based path search algorithm (such as the A* algorithm) is used. During the search process, the virtual cuboid is translated along the path and rotated at the corners. The algorithm detects whether the virtual cuboid collides with the obstacle point cloud in each pose.
[0068] (c) If there is at least one collision-free path from the entrance door to the main interior space, the foyer is deemed to meet the minimum passage requirements for stretcher entry and exit; if all paths have collisions, they are deemed to be non-compliant, and the specific corner positions where the collisions occurred are marked.
[0069] Module 4: Furniture & Details Hazard Module This module is used to detect age-friendly safety hazards in indoor environments that are unstructured, small-scale, or rely on visual perception features. Unlike large-scale spatial geometric measurements, this module mainly focuses on the following clauses of the "General Requirements for Age-Friendly Home Renovation" (MZ / T 218-2024): Clause 5.1.5.1 (Electrical wiring should be renovated in rooms with messy wiring, aging and damaged wires, and potential safety hazards) and Clause 5.1.2.1 (Indoor stair treads should have clear boundaries and should not use black, dark-colored, or patterned finishing materials).
[0070] The system combines high-precision geometric feature calculation with multimodal semantic understanding. It uses local geometric features of point clouds to perform initial screening of the high sensitivity of small-scale objects (coarse screening), and then uses a large model for multimodal semantic confirmation.
[0071] Implementation Method: Ground-level minor obstacle and line detection (for 5.1.5.1): Based on the ground model built from the preceding modules, a height threshold is set. (Preferably 0mm to 150mm), extract point cloud slices that are close to the ground; calculate the local neighborhood covariance matrix and its eigenvalues for each point. Using linearity index Point cloud clusters exhibiting "linear extension" or "chaotic clump" geometric features are extracted using the rate of curvature change. Multi-view projection and image / text feature alignment are achieved by utilizing camera pose parameters (including the camera's extrinsic matrix [R|t] and intrinsic matrix K) recorded synchronously during point cloud data acquisition. The coordinates of candidate regions of interest (ROIs) in the 3D point cloud space are projected onto the 2D RGB image planes of each acquisition viewpoint using a pinhole camera model. The projection formula is: [u, v, 1]. = K · [R|t] ·[X, Y, Z, 1] Where (X,Y,Z) are the three-dimensional coordinates of points within the candidate region, and (u,v) are the two-dimensional pixel coordinates projected onto the image. The system selects the viewpoint image with the largest projected area and the least occlusion, and crops out local high-resolution image patches (Image Patches) containing the candidate region based on the pixel range of the projection.
[0072] Based on the semantic verification of the Visual-Language Model (VLM), prompts are constructed, such as: "Does the image contain scattered wires, power strips, or other miscellaneous objects that could trip someone?". Image patches and prompts are then input into the VLM for inference. If the model outputs a positive result with a confidence level higher than a preset threshold, the geometric candidate region is identified as a "tangle of wires risk point."
[0073] Staircase visual warning detection (for section 5.1.2.1): Using normal vector clustering or region growing algorithms, the staircase structure is identified, separating the horizontal tread from the vertical riser. The intersection line between the tread and riser is extracted and defined as the "warning edge area" extending inward from the intersection line by a certain width (e.g., 50mm). The center area of the tread is defined as the "background reference area".
[0074] Perceptual color difference calculation: The point cloud color information of the two regions mentioned above is converted from RGB space to CIELAB (CIE Lab) perceptual color space. The color mean vector of each region is calculated separately, and then... The color difference formula calculates the contrast ratio, and the calculated color difference value is compared. Compared with the visual thresholds recommended by the specifications (e.g.) If it is below the threshold, or the luminance component... If the difference is too small, it is judged as "the step boundary is blurred" and marked as high risk.
[0075] Furthermore, this invention also includes manual annotation. For example... Figure 5The image shown is a comparison of the input / output effects in this invention, with the original image on the far left, the manually annotated image in the middle, and the model prediction result on the far right.
[0076] Labeling guidelines: Strictly in accordance with the "General Requirements for Age-Friendly Renovation of Home Environments for the Elderly" (MZ / T 218-2024), "Barrier-Free Design Code" (GB 50763), and "Design Code for Residential Buildings for the Elderly" (GB 50340).
[0077] Annotated content: Semantic tags: room type, handrail, washbasin, toilet, regular chair, shower chair, etc.
[0078] Risk area labeling: Risk areas such as "elevation difference tripping points", "blurred steps", and "areas where it is impossible to turn around" are manually selected and used as the ground truth for model training.
[0079] (1) Obtaining semantic labels: For object-level semantic labels (such as room type, handrail, toilet, shower chair, etc.), the existing 3D object detection model is first used to automatically pre-label the scene to generate the initial 3D bounding box and category prediction. Available automated annotation models include: V-DETR (a 3D object detector based on the DETR framework, introducing 3D vertex relative position encoding, achieving state-of-the-art performance on the ScanNet dataset, source: Shen Yet et al., V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection, ICLR 2024), FCAF3D (a fully convolutional anchor-free 3D detector, source: Rukhovich D et al., FCAF3D: Fully Convolutional Anchor-Free 3D Object Detection, ECCV 2022), CAGroup3D (a point cloud 3D detector based on class-aware grouping, source: Wang H et al., CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds, NeurIPS 2022), and UniDet3D (an indoor 3D detector jointly trained on multiple datasets, source: Kolodiazhnyi M et al., UniDet3D: Multi-dataset Indoor). 3D Object Detection (AAAI 2025). The 3D bounding boxes and category labels output by the automated model are used as pre-annotation results.
[0080] (2) Manual verification and correction: The annotators load the pre-annotation results in the 3D point cloud visualization annotation tool (such as CloudCompare, Labelbox 3D, ScanNet annotation tool or self-developed annotation platform), check whether the object category is correct, whether the bounding box position and size are consistent, delete mis-detections (such as mislabeling pipelines as handrails), manually supplement missing detections, and correct category errors. For items that need to be distinguished by aging-appropriate fine-grained subcategories (such as subdividing "chair" into "dining_chair" and "shower_chair"), the annotators manually specify the subcategories based on the material appearance, functional attributes and spatial location of the items.
[0081] (3) Risk Area Labeling: After completing the object-level labeling, the labelers, referring to the clauses of the MZ / T 218-2024 standard, used tools such as CloudCompare to select risk areas (such as tripping points due to elevation differences, steps with blurred vision, areas where it is impossible to turn around, etc.) by using 3D bounding boxes or polygon region segmentation, and assigned a risk type label to each area. Finally, all labeling results underwent cross-quality inspection (independently reviewed by another labeler) to form Ground Truth labeling data that passed the quality inspection.
[0082] Furthermore, the algorithm model in this invention includes the following contents.
[0083] 1. Model Selection The backbone network employs a Sonata encoder / PointTransformer V3 or PointNet++ as the point cloud feature extractor (i.e., the point cloud encoder described in step S1 above), effectively capturing local geometric structure and global contextual information. The combination of the backbone network and the large language model follows the standard multimodal alignment architecture of "Encoder-MLP-LLM": the backbone network acts as the encoder, encoding the point cloud into K D-dimensional feature embedding vectors; subsequently, a two-layer MLP projector maps these feature embeddings from the point cloud feature space to the word embedding space of the large language model (i.e., achieving cross-modal alignment); the mapped visual token and the text prompt token are concatenated and input into the large language model. During training, the LLM learns to associate visual features with language instructions through language modeling loss (predicting the next token). During training, the parameters of both the encoder and the LLM can be updated (end-to-end fine-tuning), or only the projector and some adaptation layer parameters can be updated (efficient parameter fine-tuning).
[0084] Large model combination: This invention introduces SpatialLM or a finely tuned Point-Bind architecture to achieve "point cloud-language" modal alignment. Both are existing technologies. As mentioned earlier, SpatialLM can autoregressively generate structured Python script descriptions containing bounding boxes for walls, doors, windows, and 59 object classes from point clouds. Point-Bind, proposed by Guo Z et al. (Guo Z, Zhang R, Zhu X, et al. Point-Bind&Point-LLM: Aligning Point Cloud with Multi-modality for 3DUnderstanding, Generation, Instruction Following, and Editing. arXiv:2309.00615, 2023), is a multimodal alignment framework based on contrastive learning. It bridges point clouds with modalities such as text and images by aligning point cloud features with the joint embedding space of ImageBind. This invention can use either of the above models or other models with equivalent point cloud-language alignment capabilities.
[0085] The Multimodal Language Model (MLLM) described in Module 4 can be a multimodal pre-trained model based on the Transformer architecture, such as the chatGPT series, Gemini series, Claude series, Qwen-VL, LLaVA, SpatialLM, or other generative artificial intelligence models with similar image and point cloud semantic understanding and reasoning capabilities. This invention does not limit the specific model version used; any model capable of receiving images, point clouds, or other text inputs and outputting semantic judgment results is within the scope of this invention.
[0086] 2. Model Input and Output enter: Preprocessed indoor scene point cloud blocks, size: (Coordinates + Color).
[0087] A text prompt for a specific task.
[0088] Output: Instance Bounding Box: The 3D frame of critical infrastructure (doors, cabinets). .
[0089] Risk Heatmap: The probability that each bounding box belongs to a "non-compliant area". .
[0090] Text Diagnostics (VLM Output): Natural language description of potential hazards (e.g., "Detected obstacle inknee space").
[0091] 3. Innovations of the model Geometric-Semantic Dual-Stream Fusion Strategy: Unlike pure deep learning models, this solution designs a "rule calibration layer". That is, the semantic results predicted by the model (such as "This is a door") will be forced to undergo hard constraint verification (measuring net width) through the geometric calculation module (module 1) to ensure the absolute accuracy of hard indicators (such as 750mm net width), rather than relying solely on probabilistic prediction.
[0092] Virtual Geometric Probe: In Module 3, an innovative "virtual probe" algorithm is proposed, which dynamically generates virtual voxels representing ergonomics (such as a wheelchair rotating cylinder) in the point cloud. Accessibility is quantitatively evaluated by calculating the volume intersection-union ratio (IoU) between the virtual voxels and the environmental point cloud, which solves the problem that traditional algorithms have difficulty in handling the "space occupied" issue.
[0093] In one embodiment of the present invention, post-processing and visualization of the evaluation results are further included. Specifically, this includes the following: 1. Result Transformation Logic Threshold logic gate (LogicGate): The auxiliary facility detection confidence score output by Module 2 (Safety Facility Integrity Detection Module) (its meaning and calculation method are detailed in the "Missing Alarm" section of Module 2, and is determined by the semantic detection score) Point cloud geometric verification score The result is obtained by weighted fusion, with a value range of [0,1]. A threshold of T=0.75 is set, and a "suspected missing" alarm is triggered if the value is lower than this.
[0094] For the geometric measurement outputs of Module 1 (Spatial Geometric Compliance Detection Module) and Module 3 (Spatial Accessibility Detection Module) (including the ground elevation difference Δh and door width from Module 1) The height of the balcony railing, etc., and the wheelchair turning diameter D, knee space clear height / depth, etc. in Module 3, are compared with the standard values, such as... .like Calculate the deviation rate ,according to Size defines the risk level (low / medium / high).
[0095] Region Merging and Denoising: Utilizing the DBSCAN clustering algorithm (parameters being the neighborhood radius ε and the minimum number of points MinPts), scattered high-risk rasters or detection points output by each module are clustered into complete "risk regions," removing isolated false positives. A complete "problem point" refers to a spatially continuous, semantically belonging connected region of the same risk type—for example, a non-compliant staircase is marked as a single connected risk region containing multiple steps, rather than a few isolated pixels on the staircase. DBSCAN automatically groups spatially adjacent (distance ≤ ε) and sufficiently dense (number of points in the neighborhood ≥ MinPts) risk points into the same cluster based on density connectivity. Isolated points that do not meet the density condition are identified as noise and removed. Each cluster represents a complete problem point. The system calculates its bounding box, centroid coordinates, and area, serving as the spatial anchor for subsequent visualization annotations and report output.
[0096] 2. Visualization of medical-engineering integration (age-friendly integration) Augmented Reality Overlay Preview: The system visualizes and overlays the detection results of each module onto a 3D scene. The 3D bounding boxes are generated from two sources: first, the object bounding boxes directly output from point cloud LLM step S3 (used to label detected or missing facilities); and second, the bounding boxes of risk areas generated by each detection module (such as the circumscribed rectangle of the elevation difference connected region in module one, and the geometry of the virtual probe in module three). The text tag generation rules are as follows: the system generates structured text by combining [Detection Module Name] + [Non-compliant Item Name] + [Measured Value] + [Corresponding Standard Clause Number and Standard Value] + [Rectification Suggestion] according to a predefined tag template. Each tag contains: [Non-compliant Item Name] + [Measured Value] + [Standard Value] + [Rectification Suggestion]. For example: Tag: Insufficient Door Width | Measured: 680mm | Standard: ≥800mm | Suggestion: Remove door frame or enlarge opening.
[0097] 3. Human-computer interaction Age-Friendly Assessment Workstation (PC): A plugin integrated into CAD / BIM software. After designers import scanned point clouds, the system automatically generates an "Age-Friendly Diagnosis Layer," allowing designers to directly design renovation plans on this layer (e.g., directly capturing risk point coordinates to place handrails). Triggered Interaction: When assessors click on the red risk box on the screen, the system pops up a database of relevant original clauses from the relevant regulations and specific rectification suggestions. See the appendix for the human-computer interaction interface. Figure 6 .
[0098] 4. Automated report generation Based on the test results, the system automatically generates a "Home Environment Age-Friendly Risk Assessment Report" (PDF).
[0099] The report is structured and includes floor plans and a list of risk points.
[0100] While various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and intent of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover modular compositions, equivalents, or alternatives within the scope of these claims.
Claims
1. A method for risk assessment of the home environment for the elderly based on digital twins, characterized in that, include: The system acquires indoor 3D point cloud data. After preprocessing, the data is input into a point cloud big language model based on the Transformer architecture for structured scene understanding. This results in structured semantic output, including the geometric layout of all walls in the scene, the position and size of doors and windows, as well as the semantic category labels and 3D bounding box parameters of various indoor objects. Based on the structured semantic output and the 3D point cloud data, the following security checks are performed: Spatial geometry compliance testing, based on computational geometry and rasterization analysis techniques, is used to process strongly quantitative indicators with clearly defined numerical thresholds; Safety facility integrity testing, based on target detection and existence verification, is used to determine whether specific age-friendly facilities are missing; Spatial accessibility and activity space detection, based on volume collision detection and virtual agent simulation, is used to evaluate knee space and wheelchair turning. Furniture and detail hazard detection is used to detect unstructured, micro-scale, or visually perceptible safety hazards in the indoor environment that may affect age-friendliness. The detection results of the security inspection are visualized and overlaid in a three-dimensional scene.
2. The method for risk assessment of elderly home environment based on digital twins according to claim 1, characterized in that: The construction of the large language model based on the Transformer architecture includes: Point cloud encoding inputs the preprocessed N six-dimensional points into the point cloud encoder, gradually reducing the spatial resolution of the point cloud while progressively increasing the feature dimension, and finally outputs K D-dimensional feature embedding vectors, where K is much smaller than N; Feature projection maps the K feature embedding vectors output by the point cloud encoder from the point cloud feature space to the word embedding space of the large language model through a two-layer multilayer perceptron projector, so that the point cloud visual token and the language token are aligned in the same semantic space. Autoregressive text generation involves concatenating the projected visual token sequence with text prompts and inputting it into a large language model to generate structured scene description text token by token.
3. The method for risk assessment of elderly home environment based on digital twins according to claim 1, characterized in that: The spatial geometry compliance check includes: Preliminary localization: Based on the coordinates of the door or functional area zone output by the large language model, a three-dimensional enclosing region of a preset size is extracted from the original high-precision point cloud to define the region of interest (ROI). Ground elevation difference calculation: RANSAC plane fitting algorithm is performed on the original point cloud in the region of interest (ROI) to detect abrupt changes in the ground point cloud normal vector, and the vertical distance between adjacent plane patches is calculated. If the vertical distance is greater than the threshold, it is marked as elevation difference risk. Door / Aisle Clear Width Measurement: Based on the aforementioned large language model, the door frame and corridor walls are segmented into instances, bounding boxes are extracted, and the minimum Euclidean distance between opposite facades is calculated. And compare it with the corresponding threshold requirements in the specification; Height logic verification: Fine-tuning point cloud LLM to identify point cloud clusters of balcony railings and washbasins, and calculating their highest points. With ground plane The difference is compared with the standard threshold; the parts that do not conform to the standard are marked.
4. The method for risk assessment of elderly home environment based on digital twins according to claim 1, characterized in that: The integrity testing of the security facilities includes: Preliminary localization: By analyzing the semantic category labels of each bounding box of the object in the structured scene description output by the point cloud big language model, objects whose category labels belong to the predefined set of anchor point furniture categories are selected. Related target retrieval: Within a preset neighborhood radius R of key furniture, search for the existence of point cloud clusters belonging to the categories of armrests, shower chairs, and shoe-changing stools; Missing facility alarm: If a critical area exists, but the corresponding auxiliary facility detection confidence score is less than the preset value, it is judged as a facility missing risk.
5. The method for risk assessment of elderly home environment based on digital twins according to claim 4, characterized in that: The security facility integrity detection employs a dual-channel fusion strategy of semantic detection and point cloud geometric cross-validation to calculate confidence levels, including: Calculate the semantic detection score Ssem, and collect all bounding boxes output by the large language model within the neighborhood radius R of the anchor furniture; the number of bounding boxes is denoted as . Match its class field with the set of associated age-friendly facility categories required by the anchor point, and count the number of matches N_match; apply space rationality filtering to each matched Bbox; the number of valid matches after filtering is denoted as Semantic score: The upper limit is truncated to 1.0 to eliminate interference from duplicate detections; Calculate the point cloud geometric verification score S_geo. This channel does not rely on the semantic output of LLM, but directly extracts geometric evidence from the original point cloud and performs independent cross-validation on the semantic detection results. Final confidence score fusion: Score=w_1 S_sem+w_2 S_geo; Where w_1+w_2=1, with the default values of w_1=0.4 and w_2=0.6, the geometric verification channel has a higher weight than the semantic detection channel to improve robustness against LLM false detections. When the conclusions of the two channels are consistent, the Score has the highest credibility. When the conclusions of the two channels are contradictory, the system marks the detection item as "requires manual review" and outputs the sub-scores of both channels in the report for the evaluator to judge.
6. The method for risk assessment of elderly home environment based on digital twins according to claim 1, characterized in that: Spatial accessibility and activity space detection includes: Within the ROI, extract environmental point clouds of key areas such as the washbasin, cabinet, and foyer; and make the following judgments based on the environmental point clouds: Knee space: Construct a virtual cube V_knee below the platform, and check the intersection-union ratio (IU / UK) of this cube with the existing environmental point cloud. If any point cloud intrudes into the virtual volume, the knee space is deemed insufficient. Wheelchair rotation: Project a virtual cylinder with a diameter of 1500mm in the center area of the living room and detect the point cloud density of obstacles inside the cylinder; Stretcher simulation: Perform a cuboid path planning simulation in the lobby area to determine whether the minimum turning radius for stretcher entry and exit is met.
7. The method for risk assessment of elderly home environment based on digital twins according to claim 6, characterized in that: Furniture and detail hazard detection combines high-precision geometric feature calculation with multimodal semantic understanding. It uses the high sensitivity of local geometric features of point clouds to screen small-scale objects for the initial screening, and then uses a large model for multimodal semantic confirmation.
8. The method for risk assessment of elderly home environment based on digital twins according to claim 7, characterized in that: Furniture and detail inspection for potential hazards includes checking for minor obstacles on the floor and wiring, specifically including: Based on the acquired ground model, a height threshold is set, and point cloud slices close to the ground are extracted; the local neighborhood covariance matrix and its eigenvalues of each point are calculated, and point cloud clusters with "linear extension" or "random cluster" geometric features are extracted using linearity index and curvature change rate; multi-view projection and image feature alignment are used to map the extracted candidate region ROI back to the original acquired multi-view RGB image, and local high-resolution image patches containing the candidate region are cropped. Construct prompt words, input image patches and prompt words into a large model for inference. If the model outputs a positive result and the confidence level is higher than a preset threshold, then the geometric candidate area is identified as a risk point of messy lines.
9. The method for risk assessment of elderly home environment based on digital twins according to claim 8, characterized in that: Furniture and detail hazard inspection also includes visual warning inspection of stairs, specifically including: Using normal vector clustering or region growing algorithms, the stair structure is identified, the horizontal treads and vertical risers are separated, the intersection line between the treads and risers is extracted by defining the edge key area, the area extending inward from the intersection line with a certain width is defined as the warning edge area, and the center area of the tread is defined as the background reference area. The point cloud color information of the two regions mentioned above is converted from RGB space to CIELAB perceptual color space. The color mean vector of the two regions is calculated respectively, and the contrast is calculated using the ΔE color difference formula. The calculated color difference value ΔE is compared with the visual threshold recommended by the standard. If it is lower than the threshold, or the difference of the brightness component is too small, it is judged as blurry step boundary and marked as high risk.
10. The method for risk assessment of elderly home environment based on digital twins according to any one of claims 1-9, characterized in that: Visualize the output results.
11. A temperature measurement system for the internal flow channel of a heat exchanger, utilizing the digital twin-based risk assessment method for the home environment of the elderly as described in any one of claims 1-10.