Intelligent virtual fence generation method, device and equipment based on multi-modal large model
By combining a multimodal large model with monocular depth and RGB images, obstacle attributes and functions are identified. Combined with infant movement trajectory information, the adaptability problem of virtual fence boundaries and early warning methods in existing technologies is solved, enabling flexible and accurate assessment of infant safety monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO SIMSHINE INTELLIGENT TECH CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-06-02
AI Technical Summary
Existing infant and toddler safety monitoring technologies struggle to effectively understand the semantic attributes and functions of obstacles in complex home environments, cannot adaptively adjust virtual fence boundaries and warning methods, and lack the ability to dynamically assess the movement trajectories of infants and toddlers.
By combining multimodal large models with monocular depth information and RGB images, obstacle attributes and functions are identified. Combined with infant movement trajectory information, risk levels are dynamically assessed, and adaptive virtual fence boundaries and early warning methods are generated.
It enables semantic risk assessment of obstacles and dynamic monitoring of infant and toddler behavior in complex home environments, adaptively generates virtual fence boundaries and early warning methods, and improves the flexibility and accuracy of safety monitoring.
Smart Images

Figure CN122134919A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal large model technology, and particularly relates to a method, apparatus and equipment for generating intelligent virtual fences based on multimodal large models. Background Technology
[0002] With the development of society and the economy, parents are often too busy with work to adequately supervise their infants and young children, making it impossible for them to provide full-time care. Therefore, reducing the supervision burden on parents while ensuring the safety of infants and young children has become an urgent need in the field of infant and young child safety supervision.
[0003] Existing infant and toddler safety monitoring technologies mainly include video surveillance equipment, wearable positioning devices, and physical fences. While video surveillance systems can provide real-time images, they typically only have image acquisition and display functions, relying on continuous parental monitoring or post-event playback. They lack the ability to deeply understand the content of the images and cannot automatically assess potential safety risks in the scene. Wearable devices and positioning devices can provide location information or activity intensity data for infants and toddlers, but their perception dimensions are limited, making it difficult to reflect the spatial relationship between infants and toddlers and specific environmental targets, and they cannot conduct differentiated risk assessments for different types of obstacles. While physical fences can limit the range of movement for infants and toddlers to some extent, they suffer from inconvenient installation, fixed boundaries, interference with normal family use, and the possibility of infants and toddlers climbing or bypassing them, making them difficult to adapt to dynamically changing family scenarios.
[0004] With the development of intelligent technology, virtual fences are gradually being introduced into infant and toddler care scenarios. Virtual fences can dynamically set safety zones based on the location information of infants and toddlers, offering greater flexibility compared to physical fences. However, in real-world home environments, scenarios often feature numerous obstructions, complex object types, and varying lighting conditions. Obstacles are not only numerous but also possess different hazard attributes and functions, such as sharp table corners, climbable chairs, cabinets storing dangerous items, and easily tipped-over furniture. Existing virtual fence solutions are mostly based on two-dimensional image detection or simple distance threshold judgment, primarily focusing on the presence of obstacles without understanding what the obstacles are, what dangers they pose, or whether they can be used by infants and toddlers. They lack effective modeling of the semantic attributes and functions of obstacles. Furthermore, the analysis of infant and toddler behavior typically remains at the current location or instantaneous state, lacking the ability to predict movement trajectories, speed patterns, and approach trends. This makes it difficult for virtual fence boundaries, alert levels, and warning methods to adaptively adjust to changes in risk in complex environments.
[0005] In recent years, large-scale visual models (such as GPT-4V) have demonstrated powerful capabilities in multimodal information understanding, simultaneously processing image appearance information and spatial semantic information. They can perform attribute understanding, functional reasoning, and common-sense judgment on targets in a scene, providing new technical means for risk assessment in complex scenarios. However, a mature and effective technical solution is still lacking for how to effectively combine monocular depth information with RGB images, introduce large-scale visual models to semantically understand the attributes and functions of obstacles in a home setting, and further combine this with the movement trajectory information of infants and toddlers to accurately assess the current risk level, thereby adaptively determining virtual fence boundaries, fence alert levels, and warning methods.
[0006] Existing Chinese patent CN111340864A discloses a method and apparatus for 3D scene fusion based on monocular estimation. It obtains a target depth map by inputting the acquired first image into a target monocular depth estimation network, and combines this with a semantic segmentation map obtained after distortion correction to acquire the depth information of the target object. Furthermore, by combining the parameter information of the image acquisition device, it determines the position information of the target object in a preset static 3D scene, achieving fusion of the monitored object and the static 3D scene model, thereby improving fusion accuracy and reducing hardware costs. While this patent can recover the depth and position of the target object and achieve 3D scene fusion under monocular conditions, it focuses on the geometric fusion and positioning representation of the monitored object and the static 3D scene model. It lacks semantic understanding and differentiated modeling of obstacle hazard attributes and usage functions in infant and toddler care scenarios, and it does not dynamically assess risk levels based on infant and toddler movement trajectories. Furthermore, it is difficult to adaptively generate visualized virtual fence boundaries, fence warning levels, and early warning methods. Therefore, its proactive protection capability against potential safety risks in complex home environments remains limited.
[0007] Therefore, how to comprehensively understand the attributes and functions of obstacles in complex home environments based on the multimodal semantic understanding capabilities of monocular depth fields and large visual models, and how to combine the movement trajectory information of infants and young children to achieve risk level assessment and intelligent generation of virtual fences, are technical problems that urgently need to be solved. Summary of the Invention
[0008] In view of this, embodiments of the present invention provide a method, apparatus and device for generating intelligent virtual fences based on a multimodal large model, in order to solve the technical problem in the prior art that it is difficult to use a visual large model to fuse monocular depth and RGB information to accurately assess the semantic risk of obstacles and adaptively generate virtual fences.
[0009] In a first aspect, embodiments of the present invention provide a method for generating intelligent virtual fences based on a multimodal large model, the method comprising: Acquire several frames of RGB images in an infant care scenario; Monocular depth calculation and target tracking are performed on the RGB image to obtain a monocular depth image and the infant's motion trajectory information; The RGB image and the monocular depth image are input into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions. Based on the obstacle attributes, obstacle functions, and movement trajectory information, the current risk level is assessed. The fence warning level and warning method are determined based on the risk level assessment results. The real-time position of the infant and the position of the obstacle are weighted and fused according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
[0010] Preferably, the step of performing monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the infant's motion trajectory information includes: The RGB image is processed using a preset monocular depth field algorithm to obtain the depth information of each pixel. Based on the depth information, the monocular depth image is obtained; The RGB image is input into a preset target detection model to obtain the target detection box for infants and young children; The target detection box is tracked using a preset target tracking algorithm to obtain the motion trajectory information.
[0011] Preferably, the step of tracking the target detection box using a preset target tracking algorithm to obtain the motion trajectory information includes: In the consecutive frames of the RGB image, the target detection box and the target tracking algorithm are used to perform association matching on the infant target, and the tracking sequence box corresponding to each frame of the RGB image is output; Based on the tracking sequence frame, the real-time position of the infant in each frame is determined, and the real-time position includes the center coordinates; Within a preset time window, the real-time locations are sorted by time to form a sequence of the infant's movement trajectory; Based on the motion trajectory sequence, calculate the displacement difference of the infant in two adjacent RGB images; Based on the displacement difference and the inter-frame time interval, the infant's movement speed and direction information are obtained. The motion trajectory sequence, the movement speed information, and the movement direction information are used as the motion trajectory information.
[0012] Preferably, the step of inputting the RGB image and the monocular depth image into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions includes: Based on the RGB image and the monocular depth image, multimodal scene data is constructed for the input of the visual large model, wherein the multimodal scene data is used to simultaneously represent the appearance information and spatial depth information of the scene; The multimodal scene data is input into the visual big model to perform overall perception of the care scene, identify obstacle targets in the scene, and obtain the target type of each obstacle; For each identified obstacle target, semantic reasoning is performed on the physical characteristics of the obstacle based on the visual big model to obtain the obstacle attributes, wherein the obstacle attributes are used to characterize the sharpness, softness, stability or easyness to tip over of the obstacle; Based on the obstacle attributes, semantic reasoning is performed on the usage of each obstacle in the care scenario based on the visual big model to obtain the obstacle function, wherein the obstacle function is used to characterize whether the obstacle is climbable, whether it is used to store dangerous items, or whether it constitutes a passage obstruction.
[0013] Preferably, the step of assessing the current risk level based on the obstacle attributes, obstacle functions, and movement trajectory information, determining the fence alert level and warning method based on the risk level assessment results, and weighting and fusing the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual large model to obtain a safety threshold, and generating a virtual fence boundary based on the safety threshold includes: For each obstacle in the scene, based on the obstacle attributes and the obstacle functions, the visual big model is used to generate semantic risk parameters corresponding to each obstacle; Based on the motion trajectory information, the approach trend information and movement speed pattern information of infants and young children are extracted; Based on the semantic risk parameters, the proximity trend information, and the movement speed pattern information, the current risk level is assessed to obtain the risk level assessment result; The warning method is determined based on the fence alert level corresponding to the risk level assessment result. The warning method includes at least one of sound and light reminders and push notifications to the parent's terminal. Based on the weight coefficients corresponding to the structured parameters output by the visual big model, the real-time position of the infant and the position of the obstacle are weighted and fused to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
[0014] Preferably, the step of assessing the current risk level based on the semantic risk parameters, the proximity trend information, and the movement speed pattern information to obtain the risk level assessment result includes: Based on the proximity trend information, proximity state parameters of the infant relative to each obstacle are determined, including the proximity direction and the expected degree of proximity. Based on the movement speed pattern information, the speed state parameters of infants and young children are determined, including acceleration tendency and speed stability. Based on the semantic risk parameters, the proximity state parameters, and the velocity state parameters, a current risk score is generated; The current risk score is compared with a preset level threshold. Based on the comparison result, the current risk level is evaluated to obtain the risk level evaluation result.
[0015] Preferably, the method further includes: The care rule information input by the user is input into a preset visual big model for semantic parsing to obtain structured rule parameters, wherein the structured rule parameters include at least the target area, triggering conditions, triggering objects and reminder methods; Based on the structured rule parameters, the virtual fence boundary, the fence alert level, or the early warning method are constrained and updated to ensure that the virtual fence boundary, the fence alert level, or the early warning method meets the care rule information. Within a preset statistical period, the movement trajectory information, risk level assessment results, and early warning trigger records of infants and young children are statistically summarized to obtain the summary results; The summarized results are input into the visual big data model to generate an activity report, wherein the activity report is used to characterize the infant's activity trajectory, exploratory behavior and safety events.
[0016] Secondly, embodiments of the present invention provide an intelligent virtual fence generation device based on a multimodal large model, the device comprising: The image acquisition module is used to acquire several frames of RGB images in infant and toddler care scenarios. The depth calculation and tracking module is used to perform monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the movement trajectory information of the infant. The obstacle recognition and analysis module is used to input the RGB image and the monocular depth image into a preset visual large model, identify obstacles in the scene, and obtain obstacle attributes and obstacle functions; The risk assessment and fence alarm module is used to assess the current risk level based on the obstacle attributes, obstacle functions and the movement trajectory information, determine the fence alarm level and warning method based on the risk level assessment results, and perform weighted fusion of the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold, and generate a virtual fence boundary based on the safety threshold.
[0017] Thirdly, embodiments of the present invention provide an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect described above.
[0018] Fourthly, embodiments of the present invention provide a storage medium storing computer program instructions, which, when executed by a processor, implement the method of the first aspect described above.
[0019] In summary, the beneficial effects of the present invention are as follows: This invention provides a method, apparatus, and device for generating intelligent virtual fences based on a multimodal large model. The method includes: acquiring several frames of RGB images in an infant care scenario; performing monocular depth calculation and target tracking on the RGB images to obtain monocular depth images and the infant's motion trajectory information; inputting the RGB images and the monocular depth images into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions; assessing the current risk level based on the obstacle attributes, obstacle functions, and motion trajectory information; determining the fence warning level and warning method based on the risk level assessment results; and weighting and fusing the infant's real-time position and obstacle positions according to the weight coefficients corresponding to the structured parameters output by the visual large model to obtain a safety threshold; and generating a virtual fence boundary based on the safety threshold. This invention leverages the multimodal semantic understanding capabilities of a large visual model to upgrade the traditional virtual fence approach, which relies solely on two-dimensional image detection or simple distance thresholds, to dynamic protection based on semantic risk. First, it acquires several frames of RGB images from an infant care scenario and performs monocular depth calculation and target tracking on these images to obtain monocular depth images and the infant's motion trajectory information. This allows the system to simultaneously possess spatial structure information of the scene and the continuous motion characteristics of the infant. Then, the RGB images and monocular depth images are input into a pre-defined large visual model. Supported by multimodal input, the large visual model identifies obstacles in the scene and further outputs... By identifying obstacle attributes and functions, the system not only identifies the obstacles but also understands their potential hazards and interaction methods. Finally, the obstacle attributes, functions, and infant movement trajectory information are combined for current risk level assessment. This allows the risk assessment to simultaneously reflect the semantic risks of obstacles and the behavioral trends of infants. Based on this, the system adaptively determines the virtual fence boundary, fence alert level, and warning method, achieving differentiated protection and dynamic adjustment for different risk targets in complex home environments. This effectively solves the problem that existing technologies struggle to integrate monocular depth and large-scale visual model semantic understanding to accurately assess risks and intelligently generate virtual fences. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, and these are all within the protection scope of the present invention.
[0021] Figure 1 This is a schematic diagram of the overall process of the intelligent virtual fence generation method based on a multimodal large model in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the process of identifying obstacles in a scene and obtaining obstacle attributes and functions in Embodiment 1 of the present invention. Figure 3 This is a flowchart illustrating the process of assessing the current risk level and determining the virtual fence boundary, fence alert level, and early warning method based on the risk level assessment results in Embodiment 1 of the present invention. Figure 4 This is a schematic diagram of the intelligent virtual fence generation device based on a multimodal large model in Embodiment 2 of the present invention; Figure 5 This is a schematic diagram of the electronic device in Embodiment 3 of the present invention. Detailed Implementation
[0022] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details. The following description of the embodiments is merely intended to provide a better understanding of the present invention by illustrating examples of the invention.
[0023] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0024] It should be noted that all actions involving the acquisition of signals, information, or data in this invention are carried out in compliance with the relevant data protection laws and regulations of the locality and with authorization from the owner of the relevant device.
[0025] Example 1 Please see Figure 1 This invention provides a method, apparatus, and device for generating intelligent virtual fences based on a multimodal large model. The method includes: Acquire several frames of RGB images in an infant care scenario; Specifically, RGB images are the fundamental data carrier used to characterize the visible light appearance information of caregiving scenarios. They are typically acquired continuously by the camera of the caregiving device, including both the morphological features of the infant / toddler target in the image and the visual information of environmental elements such as furniture, doorways, and table corners. Setting them in a multi-frame format preserves the continuity of infant / toddler activity and scene changes over time, facilitating subsequent cross-frame correlation analysis. By using image frame sequences as input, this solution can establish a unified data entry point for subsequent depth calculations, target tracking, and risk inference without relying on additional wearable devices. This provides a reliable and reusable data foundation for the dynamic generation of virtual fences, thereby improving the system's adaptability and real-time performance in everyday family scenarios.
[0026] Monocular depth calculation and target tracking are performed on the RGB image to obtain a monocular depth image and the infant's motion trajectory information; Specifically, two-dimensional visual information is transformed into structured information that can be used for spatial and behavioral analysis: monocular depth calculation is used to deduce the relative depth distribution of the scene from the RGB images of a single camera, obtaining a monocular depth image, enabling the system to represent the distance relationship between infants and surrounding objects; target tracking is used to maintain the identity and continuously update the position of infant targets between consecutive frames, forming motion trajectory information, thereby characterizing the activity range, movement trend, and behavioral continuity of infants. In implementation, a depth estimate can first be generated for each frame image using a monocular depth model, and then the position information of the infant across frames can be updated by combining target detection and tracking algorithms, finally outputting depth and trajectory results consistent with the frame sequence. By simultaneously extracting spatial structure and motion behavior, this step provides key support for subsequent semantic understanding and risk assessment, improving the stability of risk judgment in care environments with occlusion or changing viewpoints, and providing more sufficient basis for the adaptive adjustment of virtual fence boundaries.
[0027] The RGB image and the monocular depth image are input into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions. Specifically, a large-scale visual model is introduced to perform higher-level semantic understanding of the caregiving scenario. This enables the system not only to know which obstacles exist in the scenario, but also to understand the potential hazards and interactive meanings of these obstacles. Obstacle attributes can be used to characterize the risk-related features of objects, such as sharpness, softness, stability, and fragility; obstacle functions can be used to characterize the interactive methods or potential uses of objects in the caregiving scenario, such as climbable furniture or cabinets that may store dangerous items. In implementation, the appearance texture information provided by RGB images and the spatial structure information provided by monocular depth images are used together as multimodal inputs, allowing the large-scale visual model to output semantic attributes and functional inference results related to obstacles while recognizing them. By elevating obstacle detection from simple target detection to semantic-level risk understanding, this step significantly enhances the ability to identify risk sources in complex home environments. This lays the foundation for upgrading subsequent risk level assessment from distance threshold judgment to semantic risk judgment, thereby improving the targeting and accuracy of virtual fence strategy generation.
[0028] Based on the obstacle attributes, obstacle functions, and movement trajectory information, the current risk level is assessed. The fence warning level and warning method are determined based on the risk level assessment results. The real-time position of the infant and the position of the obstacle are weighted and fused according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
[0029] Specifically, in the implementation process, the RGB image and monocular depth image are first input into a large visual model to identify obstacles in the scene. Semantic reasoning is then performed on the obstacle attributes and functions of each obstacle, outputting structured risk parameters for fence calculation. These structured risk parameters quantify the differences in the degree of danger of different obstacles in the care scenario and are further mapped to weight coefficients corresponding to each structured risk parameter, allowing for adaptive adjustment of the constraint strength of different obstacles on fence generation. Subsequently, the infant's real-time position is determined based on the infant's movement trajectory information. Combining the positional relationships of each obstacle, a weighted fusion calculation is performed on the infant's real-time position and obstacle positions according to the weight coefficients to obtain the safety threshold. This safety threshold characterizes the safety constraint strength or safety distance requirements that the fence boundary generation needs to meet. Based on this, spatial positions within the candidate fence area are filtered according to the safety threshold to determine boundary positions that meet the safety constraints. A continuous and closed set of boundary points or closed boundary lines are extracted as the geometric representation of the virtual fence boundary. Since the fence boundary is obtained through deterministic calculation, it avoids relying on a large model to directly output pixel-level fence graphics, thus ensuring that the fence boundary can be stably generated and used for visualization. After obtaining the virtual fence boundary, it is mapped to the image coordinate system of the RGB image, and the fence boundary line or semi-transparent fence area is overlaid on the display interface, allowing users to intuitively perceive the range of the virtual fence. Simultaneously, a risk level assessment result is obtained based on the obstacle attributes, obstacle functions, and the motion trajectory information. This risk level assessment result is mapped to the fence alert level and warning mode, increasing the alert level and triggering more timely audio-visual reminders or pushing notifications to the parent's terminal when the risk level is high, while reducing the warning frequency to minimize unnecessary interference when the risk level is low. By using the structured output of the large visual model to determine weight coefficients and combining it with the traditional location-weighted fusion fence generation logic, this step achieves adaptive fence generation and warning linkage output for care scenarios, reducing false alarms and missed alarms and improving the overall system availability, while ensuring that the fence is calculable, drawable, and interpretable.
[0030] Preferably, the step of performing monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the infant's motion trajectory information includes: The RGB image is processed using a preset monocular depth field algorithm to obtain the depth information of each pixel. Specifically, a pre-defined monocular depth field algorithm is used to process the RGB image to obtain the depth information of each pixel. This is to recover the spatial structure of the scene using only a monocular image, enabling subsequent determination of the distance relationship between infants and objects such as furniture. The monocular depth field algorithm can be understood as a computational model that infers three-dimensional distance from two-dimensional appearance. The output depth information typically provides relative or absolute distance values in pixels. In implementation, the input RGB image can first be normalized in size, corrected for distortion, and normalized in brightness before being fed into a depth estimation network or depth regression model to obtain a depth matrix consistent with the image resolution. For edge regions or low-texture regions, consistency constraints can be used for correction to reduce depth jumps. By directly generating pixel-level depth information, this step elevates the care scene from a purely two-dimensional observation to a spatially measurable representation, which is beneficial for improving the stability of distance judgment and the reliability of risk assessment in complex occlusion environments.
[0031] Based on the depth information, the monocular depth image is obtained; Specifically, obtaining a monocular depth image based on the aforementioned depth information essentially involves organizing the depth matrix obtained in the previous step into a depth representation that can be used for subsequent processing. This facilitates alignment with RGB images and participation in target association analysis. A monocular depth image can be understood as a depth distribution expressed in image form, where the value of each pixel corresponds to its depth information. It can be directly stored as a depth matrix or mapped to a grayscale or pseudo-color depth map for visualization and debugging. In implementation, the depth matrix can be scaled or quantized to unify the depth value range. If necessary, interpolation can be used to fill in gaps in the depth matrix, maintaining the same coordinate system and resolution as the RGB image. In cross-frame processing scenarios, temporal smoothing of the depth map can also be performed to suppress transient noise. By forming a standardized monocular depth image, this step provides a unified data interface for subsequent fusion of depth information with detection boxes and trajectory information, making spatial relationship calculations more direct and reducing error accumulation caused by inconsistent data formats.
[0032] The RGB image is input into a preset target detection model to obtain the target detection box for infants and young children; Specifically, the RGB image is input into a pre-defined target detection model to obtain target detection boxes for infants and toddlers. The purpose is to locate the two-dimensional range of the infant / toddler target in each frame of the image, thereby establishing a stable target starting point for subsequent tracking. Target detection boxes are typically rectangular region descriptions used to mark the infant / toddler's position and scale in the image. In implementation, the RGB image is fed into a pre-trained target detection network, which outputs candidate boxes and their confidence scores. Thresholding and overlap suppression are then used to retain the detection boxes corresponding to the infant / toddler. When multiple candidate boxes exist in the same frame, the box that best matches the infant / toddler's characteristics can be selected by combining class probability and scale constraints. This step clearly separates the infant / toddler from the complex background, providing repeatable observation input for the tracking algorithm, reducing the probability of target loss due to background interference, and making the trajectory generation process more stable.
[0033] The target detection box is tracked using a preset target tracking algorithm to obtain the motion trajectory information.
[0034] Specifically, a pre-defined target tracking algorithm is used to track the target detection boxes to obtain motion trajectory information. This is to maintain the consistency of the infant target's identity across consecutive frames and continuously update its positional changes, thereby forming a trajectory result that reflects the activity trend. The target tracking algorithm can perform cross-frame association based on the detection boxes, outputting the target position sequence corresponding to each frame, thus forming motion trajectory information. In implementation, detection boxes in adjacent frames can be matched. The matching criteria can include the spatial proximity of the boxes, the similarity of appearance features, and the motion prediction results. The trajectory sequence is updated after a successful match. When the detection in a certain frame is unstable or short-term missing, the target position can be continued based on the motion prediction of the previous moment, and the trajectory can be corrected after the target is detected again in subsequent frames. By connecting discrete detection results into a continuous trajectory, this step enables the system to continuously characterize the infant's movement direction and activity range, providing a behavioral basis for subsequent risk level assessment combined with obstacle semantic information, and reducing false alarms and delays caused by judging based on a single frame.
[0035] Preferably, the step of tracking the target detection box using a preset target tracking algorithm to obtain the motion trajectory information includes: In the consecutive frames of the RGB image, the target detection box and the target tracking algorithm are used to perform association matching on the infant target, and the tracking sequence box corresponding to each frame of the RGB image is output; Specifically, in consecutive RGB image frames, the system performs correlation matching based on target detection boxes and target tracking algorithms to output the corresponding tracking sequence boxes for each frame. This is to connect the detected infants in each frame into the same continuous target, preventing the same infant from being treated as different targets or incorrectly switched in different frames. The correlation matching here can be understood as the process of establishing cross-frame correspondences, while the tracking sequence boxes are a continuous sequence of boxes used to mark the location range of the infant in each frame. In implementation, the target detection boxes can be used as the observation input for the current frame, combined with the target position prediction results from the previous frame, to search for the most suitable candidate boxes in the neighborhood for matching. Matching criteria can include whether the changes in the detection box position are reasonable, whether the changes in the box size are continuous, and the consistency of the infant's appearance features. By outputting stable tracking sequence boxes, this step provides a continuous and reliable foundation for subsequent extraction of real-time position and formation of trajectory sequences, reducing the risk of trajectory breakage caused by occlusion, lighting changes, or background interference.
[0036] Based on the tracking sequence frame, the real-time position of the infant in each frame is determined, and the real-time position includes the center coordinates; Specifically, determining the infant's real-time position in each frame and providing the center coordinates based on the tracking sequence bounding boxes is to transform the box-level results into positional information that is easier to calculate and compare, allowing subsequent trajectory sorting, displacement calculation, and velocity direction derivation to be performed under a unified coordinate reference. Real-time position can be understood as the infant's position within a particular frame, while the center coordinates are the geometric center point of the tracking sequence bounding boxes in the image coordinate system, typically representing the target position in that frame. In implementation, the center coordinates can be directly calculated from the coordinates of the top-left and bottom-right corners of the tracking sequence bounding boxes, while retaining the box's width and height as auxiliary information to reflect the infant's scale changes within the image. When the detection box exhibits jitter, the center coordinates can be slightly smoothed to reduce instantaneous noise. By unifying the real-time position into a center coordinate sequence, this step simplifies the trajectory representation, facilitates subsequent displacement difference calculations, and improves the stability of motion trend judgment.
[0037] Within a preset time window, the real-time locations are sorted by time to form a sequence of the infant's movement trajectory; Specifically, within a preset time window, the real-time locations are sorted temporally to form a motion trajectory sequence. The purpose is to organize the positional information of discrete frames into a continuous sequence with temporal significance, thereby reflecting the activity path and change patterns of infants and toddlers over a period of time. The preset time window can be understood as a fixed duration interval or a fixed number of frames used for trajectory statistics. It controls the time scale of trajectory analysis, covering sufficient behavioral changes while avoiding response lag due to an excessively long window. In implementation, a collection timestamp or frame number can be attached to each real-time location. The positions are sorted chronologically within the window and cached as a trajectory list. When a new frame arrives, the latest real-time location is appended to the end of the trajectory sequence, while old locations outside the window are removed to maintain a constant window length. By forming a motion trajectory sequence, this step provides direct input for subsequent displacement difference and velocity direction calculations, elevating motion state extraction from single-frame judgment to temporally continuous behavioral analysis, making it more suitable for dynamic risk assessment and adaptive adjustment of virtual fences.
[0038] Based on the motion trajectory sequence, calculate the displacement difference of the infant in two adjacent RGB images; Specifically, the displacement difference of the infant in two adjacent RGB images is calculated based on the motion trajectory sequence. The purpose is to quantify the infant's positional changes over continuous moments into a fundamental quantity that can be directly used to derive velocity and direction. The displacement difference can be understood as the coordinate difference between the center coordinates of two adjacent frames, including horizontal and vertical change components. In implementation, the real-time center coordinates of the k-th and (k+1)-th frames in the trajectory sequence are read, and the differences in the x and y directions are calculated respectively to obtain the displacement vector between the frames. In the case of slight jitter in the trajectory sequence, the center coordinates can be smoothed with a short window before calculating the difference to reduce false displacements caused by detection box jitter. By converting continuous positional changes into displacement differences, this step transitions motion information from the position layer to the change layer, providing a direct calculation basis for subsequent judgment of the infant's speed and direction of movement, and improving the interpretability and consistency of trajectory analysis.
[0039] Based on the displacement difference and the inter-frame time interval, the infant's movement speed and direction information are obtained. Specifically, obtaining the infant's movement speed and direction information based on the displacement difference combined with the inter-frame time interval is to characterize the infant's movement state on a unified time scale, facilitating matching with subsequent risk assessments and alert strategies. The inter-frame time interval is typically determined by the video frame rate or the timestamps of adjacent frames. Movement speed information can be obtained by dividing the displacement amplitude by the time interval, while movement direction information is determined by the direction of the displacement vector. In implementation, the length of the displacement vector can be calculated first as the planar displacement, and then the planar displacement can be divided by the inter-frame time interval to obtain the speed value. Simultaneously, the movement direction is determined based on the sign and angle of the displacement vector. When a more stable speed and direction output is required, the speed of multiple frames can be averaged within a preset time window, or the direction can be consistently determined to avoid direction jumps caused by single displacement anomalies. By converting the displacement difference into speed and direction, this step can more closely reflect the real behavioral characteristics of infants, enabling the system not only to know where the infant is, but also to determine whether they are rapidly approaching a certain direction, providing key behavioral basis for subsequent risk level assessment and early warning method selection.
[0040] The motion trajectory sequence, the movement speed information, and the movement direction information are used as the motion trajectory information.
[0041] Specifically, using the motion trajectory sequence, movement speed information, and movement direction information as the motion trajectory information is to encapsulate the path shape and motion state elements of the trajectory into input data that can be directly used for subsequent decision-making, facilitating fusion analysis with obstacle semantic information. Here, the motion trajectory information is no longer just a list of location points, but includes the activity path described by the trajectory sequence, as well as the dynamic characteristics described by speed and direction, simultaneously reflecting the infant's activity range, approach trend, and movement intensity. In implementation, the trajectory sequence can be saved as a basic field, and the speed and direction corresponding to each adjacent frame can be used as additional fields aligned with the trajectory sequence, forming a structured trajectory dataset. During multiple updates, the motion trajectory information can be refreshed according to a preset time window to ensure it always maintains the latest behavioral characterization. Through this combined output, this step eliminates the need for subsequent risk assessments to repeatedly calculate motion features, enabling more efficient association between infant behavior and obstacle attributes, thereby improving the real-time performance and accuracy of virtual fence boundary adjustment and early warning strategy generation.
[0042] Preferably, please refer to Figure 2 The step of inputting the RGB image and the monocular depth image into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions includes: Based on the RGB image and the monocular depth image, multimodal scene data is constructed for the input of the visual large model, wherein the multimodal scene data is used to simultaneously represent the appearance information and spatial depth information of the scene; Specifically, multimodal scene data is constructed based on RGB images and monocular depth images for use as input to large-scale visual models. The aim is to organize appearance information and spatial depth information into a unified data representation, enabling subsequent large-scale model understanding to not only rely on color and texture but also utilize near and far structures to distinguish key areas such as doorways, table corners, and cabinet openings. Multimodal scene data can be understood as a combined representation of a scene at the same moment, including both the pixel appearance of the RGB images and the depth distribution of the monocular depth images. In implementation, RGB images and monocular depth images can be aligned at the frame level to maintain the same resolution and coordinate system. Then, depth information is encapsulated as model input through additional channels, feature concatenation, or paired input. Normalization can also be performed during the encapsulation process to ensure that the depth value range is adapted to the image numerical scale, thereby improving the model's sensitivity to spatial differences. By forming multimodal scene data, this step provides a more complete foundation of scene information for subsequent obstacle recognition and semantic reasoning, reducing misidentification caused by relying solely on two-dimensional appearance and improving robustness in complex home environments.
[0043] The multimodal scene data is input into the visual big model to perform overall perception of the care scene, identify obstacle targets in the scene, and obtain the target type of each obstacle; Specifically, multimodal scene data is input into a large-scale visual model to perceive the caregiving scene holistically and identify obstacle targets, obtaining the target type of each obstacle. This is to identify potential risk sources in the scene from a global perspective, forming the object set required for subsequent risk assessment. Target type can be understood as the category of an obstacle in the scene, such as a table, chair, cabinet, or doorway area, used to distinguish the basic semantics of different targets. In implementation, the large-scale visual model receives multimodal scene data and outputs a list of obstacle targets. Each target corresponds to its location range in the image and its type label. When there are multiple obstacles in the scene, the model can provide the type result for each target individually, enabling subsequent processing to perform attribute and functional reasoning on a target-by-target basis. By obtaining the target type first, this step elevates scene elements from chaotic pixel information into a structured target set, facilitating differentiated semantic understanding and risk assessment of different types of obstacles, and improving the targeting of virtual fence generation.
[0044] For each identified obstacle target, semantic reasoning is performed on the physical characteristics of the obstacle based on the visual big model to obtain the obstacle attributes, wherein the obstacle attributes are used to characterize the sharpness, softness, stability or easyness to tip over of the obstacle; Specifically, for each identified obstacle target, semantic reasoning is performed on the physical characteristics of the obstacle based on a large visual model to obtain obstacle attributes. These attributes characterize the sharpness, softness, stability, or susceptibility to tipping over, elevating obstacles from mere "existing objects" to risk semantic elements that can be used for safety assessment. Obstacle attributes reflect the potential for injury or accidents caused by the object itself; for example, the sharpness of a table corner affects the collision risk, while the susceptibility of a floor lamp or thin chair to tip over affects the risk of it falling. In implementation, the large visual model combines target type, appearance morphology, and depth structure features for inference: appearance texture and contour reflect edge morphology, and depth variations reflect the degree of protrusion and the stability of the supporting structure, thus outputting attribute conclusions corresponding to the obstacle. In multi-target scenarios, attribute information can be generated separately for each type of target, maintaining a one-to-one correspondence with the target. By introducing attribute semantic reasoning, this step enables risk assessment to incorporate judgment criteria closer to the actual level of danger, reducing false alarms and missed alarms caused by relying solely on distance thresholds, and improving the rationality of care strategies.
[0045] Based on the obstacle attributes, semantic reasoning is performed on the usage of each obstacle in the care scenario based on the visual big model to obtain the obstacle function, wherein the obstacle function is used to characterize whether the obstacle is climbable, whether it is used to store dangerous items, or whether it constitutes a passage obstruction.
[0046] Specifically, based on obstacle attributes, the visual big data model performs semantic reasoning on the usage of each obstacle in the caregiving scenario to obtain obstacle functions. These functions characterize whether the obstacle is climbable, whether it is used to store dangerous items, or whether it obstructs passage. This is to incorporate the potential interaction methods of the object into risk considerations, enabling the system to determine how infants or toddlers might come into contact with or use the object and potentially cause risks. Obstacle functions emphasize the use or interactivity within the context of the scene. For example, a climbable chair may pose a risk of falling, a cabinet that may store dangerous items implies a higher risk of approach, and obstruction of doorways or passageways may cause tripping or accidental entry into dangerous areas. In implementation, the visual big data model can infer the function of the obstacle based on the obtained target type and attributes, combined with the scene layout: the target's height relationship with the ground, accessibility, and spatial location can be supported by the depth structure, thereby outputting the functional judgment of the obstacle in the caregiving scenario and binding it to the corresponding obstacle target. By outputting obstacle functions, this step enables subsequent risk level assessments to simultaneously consider the inherent hazardous characteristics of objects and their interactive risks within the scene, thereby allowing for more precise determination of fence boundaries, alert levels, and warning methods, and enhancing the effectiveness and practicality of dynamic protection.
[0047] Preferably, please refer to Figure 3The process of assessing the current risk level based on the obstacle attributes, obstacle functions, and movement trajectory information; determining the fence alert level and warning method based on the risk level assessment results; and weighting and fusing the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual large model to obtain a safety threshold, and generating a virtual fence boundary based on the safety threshold includes: For each obstacle in the scene, based on the obstacle attributes and the obstacle functions, the visual big model is used to generate semantic risk parameters corresponding to each obstacle; Specifically, for each obstacle in the scene, semantic risk parameters are generated based on the obstacle's attributes and functions using a large visual model. This large visual model includes general models such as GPT and Deepseek, which are used to further transform the semantic understanding results of obstacles into quantitative or graded inputs that can be directly used in risk calculation, so that subsequent assessments no longer rely solely on simple distance thresholds. Semantic risk parameters can be understood as risk representation results corresponding to each obstacle, comprehensively reflecting the obstacle's degree of danger and interaction risk. They can reflect differences in attributes such as sharpness and fragility, as well as differences in functions such as climbability, dangerous storage, and obstruction of passage. In implementation, after outputting obstacle attributes and functions, the large visual model can generate risk weights, risk scores, or risk labels for each obstacle, maintaining a correspondence with the obstacle's target type and location. When multiple obstacles exist simultaneously, these semantic risk parameters can be formed into a list or mapping table according to a preset format, facilitating subsequent association with infant trajectory features. By merging semantic information into semantic risk parameters, this step makes risk assessment calculable and comparable, and can improve the accuracy of expressing the differences in risk of different obstacles in complex home environments.
[0048] Based on the motion trajectory information, the approach trend information and movement speed pattern information of infants and young children are extracted; Specifically, extracting approach trend information and movement speed pattern information from infants and toddlers based on motion trajectory information aims to determine, from a behavioral perspective, whether a risk is approaching and the urgency of its potential occurrence, thereby transforming static environmental risks into dynamic risks related to the infant's current behavior. Approach trend information reflects changes in the direction and degree of approach of the infant / toddler relative to obstacles, while movement speed pattern information reflects the changes in the infant's speed within a time window, such as continuous acceleration, stable movement, or deceleration and stillness. In implementation, the distance sequence from the infant to each obstacle can be calculated using the changes in the center coordinates of the motion trajectory sequence, combined with the positions of obstacles in the scene. The approach trend is then obtained based on the direction and magnitude of the changes in the distance sequence. Simultaneously, the speed sequence is obtained based on the displacement and time interval of adjacent frames in the trajectory, and the speed sequence is smoothed and its trend extracted within a preset time window to form movement speed pattern information. By outputting the approach trend and speed pattern, this step enables the system to distinguish between different states, such as an infant / toddler rapidly rushing towards a high-risk obstacle and lingering in a low-risk area, providing a more realistic behavioral basis for subsequent risk level assessment.
[0049] Based on the semantic risk parameters, the proximity trend information, and the movement speed pattern information, the current risk level is assessed to obtain the risk level assessment result; Specifically, assessing the current risk level based on semantic risk parameters, proximity trend information, and movement speed patterns aims to integrate environmental semantic risk with infant behavioral characteristics for decision-making. This ensures that the risk output reflects both the nature of the hazard and changes in its probability of occurrence. The risk level assessment result can be understood as a classification of the risk level at the current moment or within the current time window, driving the selection of subsequent fence boundaries, alert levels, and warning methods. In practice, the semantic risk parameters of each obstacle can be correlated with the infant's approach trend and speed patterns towards that obstacle to form a comprehensive risk quantity. This comprehensive risk quantity is then mapped to a risk level according to preset assessment rules. The assessment rules can employ weighted aggregation or hierarchical judgment methods, outputting a higher risk level for high semantic risk and rapid approach, and a lower risk level for low semantic risk and distance or hesitation. Through this integrated assessment, this step reduces biases caused by relying solely on distance or target type, making the risk level closer to actual care needs, thereby improving the accuracy of fence adjustments and warning triggers.
[0050] The warning method is determined based on the fence alert level corresponding to the risk level assessment result. The warning method includes at least one of sound and light reminders and push notifications to the parent's terminal. Based on the weight coefficients corresponding to the structured parameters output by the visual big model, the real-time position of the infant and the position of the obstacle are weighted and fused to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
[0051] Specifically, the risk level assessment results are first mapped to fence alert levels, and corresponding warning methods are selected accordingly. When the alert level is high, audio-visual alerts are prioritized and notifications are pushed to the parent's terminal. When the alert level is low, the alert intensity or push frequency is reduced, thus balancing timeliness and non-intrusiveness. Meanwhile, to ensure users can intuitively see the stably generated virtual fence in practical applications, this step does not require the visual model to directly output the fence graphic. Instead, the visual model outputs structured parameters for fence calculation, which are then converted into weighting coefficients to characterize the impact of different obstacle risk factors on the fence tightening degree. Subsequently, based on these weighting coefficients, the infant's real-time location and the locations of each obstacle are weighted and fused to obtain a safety threshold. This safety threshold is used to limit the safety constraint strength or safety distance requirements that the fence boundary should meet. Finally, based on the safety threshold, continuously closed boundary point sets or boundary lines are selected and extracted from the candidate area as virtual fence boundaries, and these are overlaid on the interface for display. This ensures that the warning strategy and the fence visualization range are consistent, improving the interpretability and usability of the protection strategy output in the care scenario.
[0052] Preferably, the step of assessing the current risk level based on the semantic risk parameters, the proximity trend information, and the movement speed pattern information to obtain the risk level assessment result includes: Based on the proximity trend information, proximity state parameters of the infant relative to each obstacle are determined, including the proximity direction and the expected degree of proximity. Specifically, determining the approach state parameters of infants and toddlers relative to various obstacles based on approach trend information is to transform trajectory-level distance changes into more direct risk assessment elements, enabling subsequent differentiation between different scenarios such as approaching versus moving away, and slow approach versus rapid approach. The approach direction in the approach state parameters reflects the infant's / toddler's direction of movement relative to the obstacle, while the expected approach level reflects the intensity or magnitude of the approach that the infant / toddler may achieve under the current trend. In implementation, the distance sequence from the infant / toddler to each obstacle can be calculated within a time window based on the motion trajectory sequence. If the distance continuously decreases, it is determined to be an approach direction; if the distance increases, it is determined to be a moving away direction. The expected approach level can be characterized by the distance decrease rate, the number of frames of decrease duration, or the minimum predicted distance, and the prediction results can be corrected by combining the current orientation of the trajectory. By summarizing the approach trend into approach direction and expected approach level, this step enables the approach relationship to have a structured expressive capability, providing a unified input for subsequent risk score generation and reducing accidental misjudgments caused by considering only the distance at a single moment.
[0053] Based on the movement speed pattern information, the speed state parameters of infants and young children are determined, including acceleration tendency and speed stability. Specifically, determining the speed state parameters of infants and toddlers based on movement speed patterns is to characterize the urgency and stability of their movements, enabling risk assessment to reflect the strength of behavioral changes rather than just positional changes. The acceleration tendency parameter characterizes whether the speed is trending upwards, while speed stability characterizes whether speed fluctuations are small or whether there are sudden changes. In implementation, a speed sequence can be obtained within a preset time window. The acceleration tendency is judged based on the slope of the speed sequence or continuously increasing frames, and speed stability is judged based on the fluctuation amplitude or variance of the speed sequence. When short-term speed spikes occur, smoothing processing can be used before calculating parameters to avoid misinterpreting jitter as acceleration. By outputting the acceleration tendency and speed stability, this step incorporates the intensity and pattern of infants and toddlers' movements into risk assessment, making a clear difference in scoring between rapidly rushing towards high-risk targets and slowly lingering in low-risk areas, thus improving the matching degree between risk assessment and early warning strategies.
[0054] Based on the semantic risk parameters, the proximity state parameters, and the velocity state parameters, a current risk score is generated; Specifically, generating a current risk score based on semantic risk parameters, proximity state parameters, and velocity state parameters aims to integrate the inherent danger of obstacles with the approach trends and movement intensity of infants and toddlers into a single, comparable quantitative result, facilitating risk classification and strategy triggering. The current risk score can be understood as a comprehensive risk level for the current moment or time window. Semantic risk parameters provide the prior danger of the obstacle, proximity state parameters provide the changing trend of the probability of risk occurrence, and velocity state parameters provide the urgency of the risk occurrence. In implementation, a comprehensive score can be calculated per obstacle. For example, semantic risk parameters can be used as the base weight, and approach scenarios can be enhanced based on the approach direction while away scenarios can be suppressed. The score amplitude can be adjusted based on the expected approach degree and acceleration tendency. Simultaneously, velocity stability can be used to robusten the score, preventing occasional fluctuations from causing abnormal score jumps. If multiple obstacles exist in the scene, the highest score can be taken, or multiple high-risk scores can be merged to form the final current risk score. By compressing multi-source information into a unified score, this step provides a basis for calculable, comparable, and iteratively updated risk assessment, which is beneficial for subsequent adaptive adjustments to alert levels and fence boundaries.
[0055] The current risk score is compared with a preset level threshold. Based on the comparison result, the current risk level is evaluated to obtain the risk level evaluation result.
[0056] Specifically, comparing the current risk score with a preset risk level threshold and assessing the current risk level accordingly is crucial for mapping continuous risk quantification results into discrete risk level outputs. This allows the system to drive fence alert levels and warning methods in a stable and controllable manner. The preset risk level threshold can be understood as a boundary standard for different risk levels, used to divide risk scores into low, medium, and high risk level ranges. In implementation, a multi-level threshold table can be pre-defined, and the current risk score is compared sequentially with each threshold. A higher risk level is output if the score meets a higher threshold, and a corresponding intermediate level is output if the score falls within the threshold range. In continuous frame assessment scenarios, a minimum hold period or a risk level fallback condition can be introduced into the risk level output to prevent frequent risk level jumps caused by score fluctuations around the threshold. Obtaining the risk level through threshold comparison makes the risk output more intuitive and easier to link with strategies. This ensures that high-risk situations can trigger stricter fences and stronger warning methods in a timely manner, while reducing unnecessary disturbances in low-risk situations, improving the usability and consistency of the monitoring system.
[0057] Preferably, the method further includes: The care rule information input by the user is input into a preset visual big model for semantic parsing to obtain structured rule parameters, wherein the structured rule parameters include at least the target area, triggering conditions, triggering objects and reminder methods; Specifically, the caregiving rule information input by the user is semantically parsed into a pre-defined visual model to obtain structured rule parameters. This allows parents to directly express their caregiving intentions in natural language, which are then transformed into control constraints that the system can execute. Caregiving rule information can be scene-related reminders, such as limiting reminders to a specific target area or under specific circumstances. Structured rule parameters break down natural language into explicit fields: target area refers to areas like the kitchen door or stairwell; trigger condition describes the logic of approaching or staying; trigger object limits the infant or specific behavior; and reminder method specifies audio-visual reminders or push notifications to the parent's device. In implementation, the rule text and current scene information can be provided to the visual model. The model extracts and standardizes the area names, action words, and reminder preferences in the text, mapping area names to corresponding locations in the scene when necessary. The final output is a set of parameters that can be directly read by the system. By obtaining structured rule parameters through semantic parsing, the system does not require parents to learn fixed command formats, reducing the setup threshold and providing clear input for subsequent personalized updates to the fence strategy.
[0058] Based on the structured rule parameters, the virtual fence boundary, the fence alert level, or the early warning method are constrained and updated to ensure that the virtual fence boundary, the fence alert level, or the early warning method meets the care rule information. Specifically, updating the virtual fence boundaries, alert levels, and warning methods based on structured rule parameters to meet caregiving rules is crucial for incorporating parents' personalized preferences into the fence generation and warning strategies. This allows the protection range and alert methods in the same scenario to be adapted to individual family needs. Constraint updates can be understood as imposing additional conditions on existing strategies, such as tightening the fence boundaries for the target area, increasing the alert level associated with the target area, or prioritizing push notifications over only audio-visual alerts when specific trigger conditions are met. In implementation, the target area is first located based on structured rule parameters and used as the constraint object for strategy updates. Then, rule constraints are overlaid on the basic risk level assessment strategy to form the final virtual fence boundary and warning method configuration. For example, when the trigger condition is proximity to the kitchen door and the alert method is push notification, even if the overall risk level is not high, a higher alert level can be set for that area and notifications can be enabled. By introducing rule constraints, this step allows the virtual fence to move beyond general risk assessment and adapt to different families' preferences for sensitive areas and alerts, improving the system's practicality and controllability.
[0059] Within a preset statistical period, the movement trajectory information, risk level assessment results, and early warning trigger records of infants and young children are statistically summarized to obtain the summary results; Specifically, the statistical summary of movement trajectory information, risk level assessment results, and warning trigger records within a preset statistical period aims to distill key events and behavioral characteristics scattered throughout the continuous operation into a data foundation for summarization and retrospection. The preset statistical period can be set to a day or several hours to unify the time scale for summary; trigger records reflect when and under what risk level the warning method was triggered. In practice, key segments or trajectory summaries of the movement trajectory can be continuously written within the statistical period, recording each risk level output and its duration, while also saving the trigger time and corresponding alert level of audio-visual alerts or push notifications. During the statistical summary phase, the trajectory can be segmented and summarized, extracting the activity range, hotspots, and the number of times approaching high-risk targets, and summarizing the risk level distribution and warning trigger frequency. By generating the summary results, this step provides structured material for activity report generation, reducing the need for repetitive analysis of raw data during report generation, and also allowing parents to quickly understand the general situation of risks occurring throughout the day.
[0060] The summarized results are input into the visual big data model to generate an activity report, wherein the activity report is used to characterize the infant's activity trajectory, exploratory behavior and safety events.
[0061] Specifically, the summarized results are input into a large-scale visual model to generate activity reports, which characterize the infants' and toddlers' activity trajectories, exploratory behaviors, and safety incidents. This transforms statistical data into a summary output that parents can easily understand, forming a more intuitive care feedback loop. The activity reports not only describe trajectories but also, by combining risk level assessment results and trigger records, present characteristics of exploratory behaviors and a summary of safety incidents, such as frequent lingering in certain areas or repeated approaches to sensitive areas that triggered alerts. In implementation, the summarized results can be organized into model input according to a preset template, including the statistical period, activity range summary, risk level time percentage, and alert trigger list. The large-scale visual model then generates a natural language report, highlighting high-risk periods or frequently triggered areas to facilitate targeted adjustments to care rules or environmental arrangements by parents. By generating activity reports through the large-scale model, this step extends the system from real-time protection to periodic review and communication, reducing the burden of continuous screen monitoring for parents and improving the interpretability and manageability of the care process.
[0062] Example 2 Please see Figure 4 This invention provides an intelligent virtual fence generation device based on a multimodal large model, the device comprising: The image acquisition module is used to acquire several frames of RGB images in infant and toddler care scenarios. The depth calculation and tracking module is used to perform monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the movement trajectory information of the infant. The obstacle recognition and analysis module is used to input the RGB image and the monocular depth image into a preset visual large model, identify obstacles in the scene, and obtain obstacle attributes and obstacle functions; The risk assessment and fence alarm module is used to assess the current risk level based on the obstacle attributes, obstacle functions and the movement trajectory information, determine the fence alarm level and warning method based on the risk level assessment results, and perform weighted fusion of the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold, and generate a virtual fence boundary based on the safety threshold.
[0063] Specifically, the intelligent virtual fence generation device based on a multimodal large model provided in this embodiment of the invention includes: an image acquisition module for acquiring several frames of RGB images in an infant care scenario; a depth calculation and tracking module for performing monocular depth calculation and target tracking on the RGB images to obtain monocular depth images and the infant's motion trajectory information; an obstacle recognition and analysis module for inputting the RGB images and the monocular depth images into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions; and a risk assessment and fence alarm module for assessing the current risk level based on the obstacle attributes, obstacle functions, and motion trajectory information, determining the fence warning level and warning method based on the risk level assessment results, and performing weighted fusion of the infant's real-time position and obstacle position based on the weight coefficients corresponding to the structured parameters output by the visual large model to obtain a safety threshold, and generating a virtual fence boundary based on the safety threshold. This device leverages the multimodal semantic understanding capabilities of a large visual model to upgrade the traditional virtual fence approach, which relies solely on 2D image detection or simple distance thresholds, to dynamic protection based on semantic risk. First, it acquires several frames of RGB images from an infant care scenario and performs monocular depth calculation and target tracking on these images to obtain monocular depth images and the infant's motion trajectory information. This allows the system to simultaneously possess spatial structure information of the scene and the continuous motion characteristics of the infant. Then, the RGB images and monocular depth images are input into a pre-set large visual model. Supported by multimodal input, the large visual model identifies obstacles in the scene and further outputs... By identifying obstacle attributes and functions, the system not only identifies the obstacles but also understands their potential hazards and interaction methods. Finally, the obstacle attributes, functions, and infant movement trajectory information are combined for current risk level assessment. This allows the risk assessment to simultaneously reflect the semantic risks of obstacles and the behavioral trends of infants. Based on this, the system adaptively determines the virtual fence boundary, fence alert level, and warning method, achieving differentiated protection and dynamic adjustment for different risk targets in complex home environments. This effectively solves the problem that existing technologies struggle to integrate monocular depth and large-scale visual model semantic understanding to accurately assess risks and intelligently generate virtual fences.
[0064] Example 3 In addition, combined Figure 1 The intelligent virtual fence generation method based on a multimodal large model described in this embodiment of the invention can be implemented by an electronic device, characterized in that it is used to implement the method described in embodiment 1. Figure 5 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention is shown.
[0065] Electronic devices may include processors and memory storing computer program instructions.
[0066] Specifically, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.
[0067] The memory may include a large-capacity storage device for data or instructions. For example, and not limitingly, the memory may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to a data processing device. In a particular embodiment, the memory is a non-volatile solid-state memory. In a particular embodiment, the memory includes a read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0068] The processor reads and executes computer program instructions stored in memory to implement any of the intelligent virtual fence generation methods based on multimodal large models in the above embodiments.
[0069] In one example, the electronic device may also include a communication interface and a bus. For example, Figure 5 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0070] The communication interface is mainly used to enable communication between various modules, devices, units and / or equipment in the embodiments of the present invention.
[0071] A bus, including hardware, software, or both, couples components of an electronic device together. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, a bus may include one or more buses. While specific buses are described and illustrated in embodiments of the invention, the invention contemplates any suitable bus or interconnect.
[0072] Example 4 Furthermore, in conjunction with the intelligent virtual fence generation method based on a multimodal large model in the above embodiments, this invention can be implemented using a computer-readable storage medium. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the intelligent virtual fence generation methods based on a multimodal large model in the above embodiments.
[0073] In summary, the embodiments of the present invention provide a method, apparatus, and device for generating intelligent virtual fences based on a multimodal large model.
[0074] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0075] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0076] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0077] The above description is merely a specific embodiment of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A method for generating intelligent virtual fences based on a multimodal large model, characterized in that, The method includes: Acquire several frames of RGB images in an infant care scenario; Monocular depth calculation and target tracking are performed on the RGB image to obtain a monocular depth image and the infant's motion trajectory information; The RGB image and the monocular depth image are input into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions. Based on the obstacle attributes, obstacle functions, and movement trajectory information, the current risk level is assessed. The fence warning level and warning method are determined based on the risk level assessment results. The real-time position of the infant and the position of the obstacle are weighted and fused according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
2. The intelligent virtual fence generation method based on a multimodal large model according to claim 1, characterized in that, The step of performing monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the infant's motion trajectory information includes: The RGB image is processed using a preset monocular depth field algorithm to obtain the depth information of each pixel. Based on the depth information, the monocular depth image is obtained; The RGB image is input into a preset target detection model to obtain the target detection box for infants and young children; The target detection box is tracked using a preset target tracking algorithm to obtain the motion trajectory information.
3. The intelligent virtual fence generation method based on a multimodal large model according to claim 2, characterized in that, The step of using a preset target tracking algorithm to track the target detection box and obtain the motion trajectory information includes: In the consecutive frames of the RGB image, the target detection box and the target tracking algorithm are used to perform association matching on the infant target, and the tracking sequence box corresponding to each frame of the RGB image is output; Based on the tracking sequence frame, the real-time position of the infant in each frame is determined, and the real-time position includes the center coordinates; Within a preset time window, the real-time locations are sorted by time to form a sequence of the infant's movement trajectory; Based on the motion trajectory sequence, calculate the displacement difference of the infant in two adjacent RGB images; Based on the displacement difference and the inter-frame time interval, the infant's movement speed and direction information are obtained. The motion trajectory sequence, the movement speed information, and the movement direction information are used as the motion trajectory information.
4. The intelligent virtual fence generation method based on a multimodal large model according to claim 1, characterized in that, The step of inputting the RGB image and the monocular depth image into a preset visual large model to identify obstacles in the scene and obtain obstacle attributes and functions includes: Based on the RGB image and the monocular depth image, multimodal scene data is constructed for the input of the visual large model, wherein the multimodal scene data is used to simultaneously represent the appearance information and spatial depth information of the scene; The multimodal scene data is input into the visual big model to perform overall perception of the care scene, identify obstacle targets in the scene, and obtain the target type of each obstacle; For each identified obstacle target, semantic reasoning is performed on the physical characteristics of the obstacle based on the visual big model to obtain the obstacle attributes, wherein the obstacle attributes are used to characterize the sharpness, softness, stability or easyness to tip over of the obstacle; Based on the obstacle attributes, semantic reasoning is performed on the usage of each obstacle in the care scenario based on the visual big model to obtain the obstacle function, wherein the obstacle function is used to characterize whether the obstacle is climbable, whether it is used to store dangerous items, or whether it constitutes a passage obstruction.
5. The intelligent virtual fence generation method based on a multimodal large model according to claim 1, characterized in that, The process of assessing the current risk level based on the obstacle attributes, obstacle functions, and movement trajectory information; determining the fence alert level and warning method based on the risk level assessment results; and weighting and fusing the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual large model to obtain a safety threshold, and generating a virtual fence boundary based on the safety threshold includes: For each obstacle in the scene, based on the obstacle attributes and the obstacle functions, the visual big model is used to generate semantic risk parameters corresponding to each obstacle; Based on the motion trajectory information, the approach trend information and movement speed pattern information of infants and young children are extracted; Based on the semantic risk parameters, the proximity trend information, and the movement speed pattern information, the current risk level is assessed to obtain the risk level assessment result; The warning method is determined based on the fence alert level corresponding to the risk level assessment result. The warning method includes at least one of sound and light reminders and push notifications to the parent's terminal. Based on the weight coefficients corresponding to the structured parameters output by the visual big model, the real-time position of the infant and the position of the obstacle are weighted and fused to obtain a safety threshold. A virtual fence boundary is generated based on the safety threshold.
6. The intelligent virtual fence generation method based on a multimodal large model according to claim 5, characterized in that, The assessment of the current risk level based on the semantic risk parameters, proximity trend information, and movement speed pattern information yields the following risk level assessment results: Based on the proximity trend information, proximity state parameters of the infant relative to each obstacle are determined, including the proximity direction and the expected degree of proximity. Based on the movement speed pattern information, the speed state parameters of infants and young children are determined, including acceleration tendency and speed stability. Based on the semantic risk parameters, the proximity state parameters, and the velocity state parameters, a current risk score is generated; The current risk score is compared with a preset level threshold. Based on the comparison result, the current risk level is evaluated to obtain the risk level evaluation result.
7. The intelligent virtual fence generation method based on a multimodal large model according to any one of claims 1-6, characterized in that, The method further includes: The care rule information input by the user is input into a preset visual big model for semantic parsing to obtain structured rule parameters, wherein the structured rule parameters include at least the target area, triggering conditions, triggering objects and reminder methods; Based on the structured rule parameters, the virtual fence boundary, the fence alert level, or the early warning method are constrained and updated to ensure that the virtual fence boundary, the fence alert level, or the early warning method meets the care rule information. Within a preset statistical period, the movement trajectory information, risk level assessment results, and early warning trigger records of infants and young children are statistically summarized to obtain the summary results; The summarized results are input into the visual big data model to generate an activity report, wherein the activity report is used to characterize the infant's activity trajectory, exploratory behavior and safety events.
8. A smart virtual fence generation device based on a multimodal large model, characterized in that, The device includes: The image acquisition module is used to acquire several frames of RGB images in infant and toddler care scenarios. The depth calculation and tracking module is used to perform monocular depth calculation and target tracking on the RGB image to obtain a monocular depth image and the movement trajectory information of the infant. The obstacle recognition and analysis module is used to input the RGB image and the monocular depth image into a preset visual large model, identify obstacles in the scene, and obtain obstacle attributes and obstacle functions; The risk assessment and fence alarm module is used to assess the current risk level based on the obstacle attributes, obstacle functions and the movement trajectory information, determine the fence alarm level and warning method based on the risk level assessment results, and perform weighted fusion of the infant's real-time position and obstacle position according to the weight coefficients corresponding to the structured parameters output by the visual big model to obtain a safety threshold, and generate a virtual fence boundary based on the safety threshold.
9. An electronic device, characterized in that, include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method as described in any one of claims 1-7.
10. A storage medium storing computer program instructions thereon, characterized in that, The method as described in any one of claims 1-7 is implemented when the computer program instructions are executed by the processor.