A real-world model training method and system based on a local physical environment

By acquiring a local knowledge base of patterns and multimodal features of the target physical location, classification and model distillation are performed to generate personalized real-world models, solving the problem of low model adaptability in existing technologies and achieving higher task execution accuracy.

CN120874541BActive Publication Date: 2026-05-08BEIJING QIDAISONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING QIDAISONG TECH CO LTD
Filing Date
2025-07-14
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing real-world model training methods rely on learning general patterns from features in a large number of common scenarios, making it difficult to capture personalized modal features in the target scenario. This results in low model adaptability to the target scenario and an inability to accurately understand entity relationships and behavioral patterns.

Method used

By acquiring the local pattern knowledge base and multimodal features of the target physical location, including content features, spatial structure features, entity features and temporal features, classification and model distillation are performed to generate personalized real-world models that dynamically adapt to the personalized patterns and dynamic changes of the target physical location.

Benefits of technology

It improves the representation and adaptability of real-world models to target physical locations, and enhances the accuracy of task execution when performing target application tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874541B_ABST
    Figure CN120874541B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, and more particularly to a real world model training method and system based on a local physical environment, which provides common laws of second modal features of a target physical place in a historical time period through a local law knowledge base, classifies second modal features of the target physical place in a target time period, adds category labels to the second modal features and adds the second modal features to a training set of a cloud world model, and performs model distillation on the cloud world model in the cloud together with the local law knowledge base, first modal features and second modal features, so as to determine corresponding individualized features of the target physical place, so that the real world model can dynamically adapt to individualized laws and dynamic changes of the target physical place, and the representation degree and adaptability of the real world model to the target physical place are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for training real-world models based on local physical environments. Background Technology

[0002] With the development of IoT, computer vision, and sensor technologies, training digital models that can accurately represent physical locations has become crucial for the implementation of smart scenarios.

[0003] Existing real-world model training methods rely on learning general patterns from features across a large number of common scenarios to build models with broad applicability. However, these methods only analyze the learned general patterns and struggle to capture the personalized modal features of the target scenario. This results in a lack of specificity when modeling the target scenario in real time, hindering efficient collaborative training with existing general-world models in the cloud. Consequently, the model fails to fully learn the personalized patterns of the target scenario, leading to low model adaptability and difficulty in accurately understanding entity relationships and behavioral patterns within the target scenario, thus impacting actual task performance.

[0004] Therefore, how to train a real-world model based on the characteristics of the local physical environment in order to improve the adaptability between the real-world model and the target physical location has become an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, the present invention provides a method for training a real-world model based on a local physical environment, which includes the following steps:

[0006] S10, obtain the local pattern knowledge base corresponding to the target physical location before the target time period and the first modal features and second modal features corresponding to the target physical location within the target time period. The first modal features include content feature vector, spatial structure feature vector and temporal feature vector, and the second modal features include entity feature vector and relation feature vector.

[0007] S20, classify the second modality features within the target time period according to the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features. The target classification result includes known feature data and unknown feature data.

[0008] S30, based on the local rule knowledge base, the first modality feature, the second modality feature, and the target classification result corresponding to each feature vector in the second modality feature, perform model distillation on the cloud world model in the cloud to obtain the real world model corresponding to the target physical location. The real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

[0009] The present invention also provides a real-world model training system based on a local physical environment, the real-world model training system based on a local physical environment comprising:

[0010] The data acquisition module is used to acquire the local pattern knowledge base corresponding to the target physical location before the target time period and the first modal features and second modal features corresponding to the target physical location within the target time period. The first modal features include content feature vectors, spatial structure feature vectors and temporal feature vectors, and the second modal features include entity feature vectors and relation feature vectors.

[0011] The feature classification module is used to classify the second modality features within the target time period based on the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features. The target classification result includes known feature data and unknown feature data.

[0012] The model training module is used to perform model distillation on the cloud world model based on the local rule knowledge base, the first modality feature, the second modality feature, and the target classification result corresponding to each feature vector in the second modality feature, to obtain the real world model corresponding to the target physical location. The real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

[0013] This invention has at least the following beneficial effects: By acquiring a local pattern knowledge base corresponding to the target physical location before the target time period, it provides common patterns corresponding to the second modal features of different reference physical locations within historical time periods. It also acquires the content feature vector, spatial structure feature vector, entity feature vector, temporal feature vector, and relational feature vector of the target physical location within the target time period, improving the understanding and representation capabilities of the target physical location and target object under content modality, spatial structure modality, entity modality, temporal modality, and relational modality. Furthermore, based on the local pattern knowledge base, it classifies the second modal features within the target time period, obtaining the target classification result corresponding to each feature vector in the second modal features. This process involves classifying each feature vector in the second modality to clarify the personalized features corresponding to the target physical location. Finally, by adding the local rule knowledge base, the first modality features, the second modality features, and the target classification results corresponding to each feature vector in the second modality features to the training set of the cloud world model, the cloud world model is distilled. This process adjusts the real-world model according to the modal features and personalized features of the target physical location, enabling the real-world model to dynamically adapt to the personalized rules and dynamic changes of the target physical location. This improves the representation and adaptability of the trained real-world model to the target physical location, thereby improving the accuracy of task execution when performing target application tasks. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart of a real-world model training method based on a local physical environment provided in Embodiment 1 of the present invention;

[0016] Figure 2 This is a schematic diagram of a real-world model training system based on a local physical environment, provided in Embodiment 2 of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the terms used to distinguish similar objects can be interchanged so that the invention can also be implemented in other embodiments besides the illustrated or described embodiments. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0019] Example 1

[0020] This first embodiment provides a method for training real-world models based on the local physical environment, such as... Figure 1 As shown, this method for training a real-world model based on the local physical environment includes the following steps:

[0021] S10, obtain the local pattern knowledge base corresponding to the target physical location before the target time period and the first modal features and second modal features corresponding to the target physical location within the target time period. The first modal features include content feature vector, spatial structure feature vector and temporal feature vector, and the second modal features include entity feature vector and relation feature vector.

[0022] Among them, the target physical location is an actual three-dimensional spatial region with a specific spatial structure and functional purpose, which is the object of research and modeling, such as specific spatial regions like shopping malls, campuses, and factories.

[0023] The target time period is a time interval set by the user. In this embodiment, it can be a period of time between two distillation training operations on the world model in the cloud.

[0024] The local pattern knowledge base is a database that stores historical data patterns of the target physical location. It can be trained or extracted based on the second modal feature data of the reference physical location over a relatively long period of time. It serves as the basis for distinguishing between known and unknown feature data of the target physical location in the second modal features corresponding to the target time period. This guides the cloud-based world model to learn the local characteristics of the target physical location during the distillation training process, thereby enhancing the real-world model's ability to represent the target physical location in a more refined way.

[0025] For the first and second modal features of the target physical location within the target time period, the content feature vector corresponding to the content modality can represent the scene semantic features of the image and text corresponding to the target physical location; the spatial structure feature vector corresponding to the spatial structure modality can represent the static layout features such as the spatial layout of the scene and the relative positions between objects in the target physical location; the entity feature vector corresponding to the entity modality can represent the appearance features, postures, and actions of target objects such as people, objects, and robots in the target physical space; the temporal feature vector corresponding to the temporal modality can represent the multi-dimensional dynamic changes of the target physical location over a period of time for the content modality, spatial structure modality, and entity modality; and the relational feature vector can represent the interactions and positions between people, objects, and people and objects in the target physical location.

[0026] In one specific embodiment, S10 includes the following steps:

[0027] S101, obtain the feature vectors of each reference physical location for each second mode within M preset time periods before the target time period, wherein the second mode includes entity mode and relation mode, the feature vector for entity mode is entity feature vector, the feature vector for relation mode is relation feature vector, and M is a positive integer.

[0028] S102, for any second mode, cluster all feature vectors corresponding to all reference physical locations for the current second mode to obtain several feature cluster sets corresponding to the current second mode.

[0029] S103, store all feature cluster sets corresponding to the current second mode into the local pattern knowledge base corresponding to the target physical location.

[0030] Reference physical locations are other physical environments that are similar to or related to the target physical location in terms of physical characteristics, functional attributes, or scene type. By analyzing the feature data of multiple reference physical locations in entity and relational modalities, common patterns (such as crowd flow patterns, object placement rules, and interaction patterns between people and objects) are extracted to provide cross-scene prior knowledge for the target physical location. This allows for the extraction of features beyond the common patterns of the target physical location's deviation from the references, enabling targeted characterization of the target physical location's personalized features. This guides the cloud-based world model to learn the local characteristics of the target physical location during distillation training, thereby enhancing the real-world model's ability to represent the target physical location with greater precision.

[0031] Furthermore, spatial structural features (such as room layout and object positions) typically exhibit strong stability and domain regularity. Their regularity has been pre-modeled by referencing the common patterns of physical locations. If the spatial structure undergoes temporary adjustments (such as rearranging conference room tables and chairs), this change will be indirectly reflected through temporal features, eliminating the need for separate judgment of regularity. In contrast, entity features (such as human appearance and robot models) are highly individualized, with ambiguous definitions of regularity (such as different people wearing different colored clothes being considered normal). Forcibly judging regularity may lead to misjudgments. Temporal features are the trajectories of content modality, spatial structural modality, and entity modality over time. Their regularity needs to be judged based on the patterns of previous modalities. Therefore, the content feature vector, spatial structural feature vector, and temporal feature vector corresponding to the first modality feature are all uploaded to the cloud as the basis for training the cloud world model.

[0032] Relational feature vectors describe the interaction logic between people and objects (such as conversations between people, contact between people and objects, and positional associations between objects), directly reflecting information such as the dynamic patterns and social rules of physical locations. Content feature vectors, including image feature vectors and text feature vectors, can characterize the scene semantic features of images and text corresponding to physical locations, thereby reflecting the type and function of physical locations. Dynamic interactions are prone to deviating from the common patterns due to time, events, or human factors (such as sudden gatherings or abnormal placement of items). The unconventional nature of content features can reflect the special characteristics of the physical location itself (such as the presence of fire drill props in an office). Therefore, it is necessary to identify abnormal patterns under entity modalities and relational modalities through common patterns, and to characterize the personalized features of the target physical location in terms of both static scene semantics and dynamic interaction, thereby enhancing the world model's deep understanding of the target physical location.

[0033] The M preset time periods preceding the target time period to be analyzed are used to extract historical feature data of the reference physical location. Each preset time period can be set by the implementer according to the actual situation.

[0034] Clustering algorithms such as K-means and DBSCAN are used to cluster all feature vectors corresponding to each second mode, resulting in several feature cluster sets corresponding to the current second mode. Each feature cluster set groups together several feature vectors with high similarity. Those skilled in the art will know that any clustering algorithm in the prior art falls within the protection scope of this invention, and will not be described in detail here.

[0035] The above process, through historical data collection, modal clustering, and knowledge storage, transforms the complex second-modal characteristics of the target physical location over a historical period into structured regular knowledge. Reusable prior knowledge is extracted from historical data, providing a foundation for extracting personalized features beyond the common patterns of the target physical location's offset reference.

[0036] In one specific embodiment, S101 includes the following steps:

[0037] S1011, acquire several historical acquisition data corresponding to each reference physical location within M preset time periods, wherein each historical acquisition data can be a historical acquisition image or historical acquisition text, and each historical acquisition data represents the state of several target objects in the corresponding reference physical location within the current time period.

[0038] S1012, for any preset time period corresponding to any reference physical location, encode each historical data collected at the current reference physical location in the current preset time period according to the preset content encoder, and obtain the content feature vector corresponding to each historical data collected.

[0039] S1013, based on the preset spatial structure encoder and the preset human entity encoder, encode each historical acquisition image of the current reference physical location in the current preset time period, and obtain the spatial structure feature vector corresponding to each historical acquisition image, as well as the entity feature vector corresponding to each target object in each historical acquisition image.

[0040] S1014, concatenate the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location within the current preset time period to obtain the scene status feature vector corresponding to the current reference physical location within the current preset time period.

[0041] S1015, concatenate and compress the scene status feature vectors corresponding to the current reference physical location within the current preset time period and N preset time periods before the current preset time period to obtain the temporal feature vector corresponding to the current reference physical location within the current preset time period, where N+1 is the preset number of vector concatenations, and the scene status feature vectors before the first preset time period are preset feature vectors.

[0042] S1016, Input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location in the current preset time period into the first initial world model to obtain several relation feature vectors corresponding to the current reference physical location in the current preset time period.

[0043] Within each preset time period, the target physical location corresponds to several historical data sets. Specifically, each historical data set can be either an image (image mode) or text (text mode). The historical images are visual information of the reference physical location acquired through image acquisition devices such as cameras within the preset time period. They exist in the form of a pixel matrix and can intuitively present information such as the appearance, location, and shape of target objects within the reference physical location, such as the distribution of customers in a shopping mall or the layout of store displays. The historical text can originate from textual information such as signs, displays, and documents within the reference physical location within the preset time period, or from text converted through methods such as speech recognition. It is used to represent the semantic information of the reference physical location and the people, objects, and other entities within it.

[0044] The target object is an entity within the target physical location that has research value, including people, objects, robots, etc. The target object will exhibit different states in different time periods. By setting multiple time periods, the periodic changes, long-term trends, and other characteristics of the reference physical location can be captured.

[0045] The preset content encoder is a model component used to extract and encode features from historical data. The structure and parameters of the preset content encoder are pre-defined and can be built based on deep learning architectures such as convolutional neural networks and Transformers. It can automatically learn and extract effective features from historical data according to the characteristics of historical data, and transform the original historical data into a more representative vector form.

[0046] Correspondingly, the content feature vector is the output of the pre-defined content encoder after encoding the historical data, representing the core content features of the historical data in vector form.

[0047] The spatial structure encoder analyzes the historically acquired images as a whole, extracting structural information such as the spatial layout of the scene and the relative positions of objects, representing the static layout features of the reference physical location, and generating corresponding spatial structure feature vectors. The person entity encoder encodes the appearance features, posture, and actions of target objects in the historically acquired images, generating entity feature vectors for each target object. For example, in a shopping mall scene, the spatial structure encoder analyzes information such as store layout and aisle locations, while the person entity encoder extracts features such as customers' clothing, posture, movement trajectories, and the color, category, and shape of the goods being sold.

[0048] For all feature vectors of the target physical location within the current preset time period, they are concatenated and integrated according to a specific logic, fusing feature information of different modalities and dimensions into a complete unified feature representation, namely the scene status feature vector. Furthermore, for the reference physical location from the current preset time period and the N+1 consecutive preset time periods corresponding to the N preset time periods before the current preset time period, the scene status feature vectors within the N+1 preset time periods are further concatenated and compressed to generate a temporal feature vector, capturing the multi-dimensional dynamic changes of the reference physical location in terms of content modality, spatial structure modality and entity modality over a period of time.

[0049] The first initial world model is a pre-trained model that already has basic multimodal understanding capabilities before model distillation. It is usually trained on general scene data (such as publicly available image, text datasets, and 3D point cloud libraries) and has cross-domain generalization capabilities, but it is not optimized for specific reference physical locations.

[0050] Among them, the feature vector of the current scene status before the first preset time period is the preset feature vector to ensure that the feature vector splicing logic of each preset time period is consistent from the first preset time period, that is, the features of N+1 periods are spliced ​​together to avoid the destruction of the uniformity of time sequence processing due to special boundary conditions.

[0051] In one specific implementation, the preset feature vector can be an all-zero vector, representing "no historical information" or "initial state". This will not introduce false historical feature data and will avoid the temporal feature vector learning incorrect temporal dependencies due to default values ​​such as random values ​​or non-zero constants.

[0052] In another specific implementation, a spatiotemporal database can be configured to store the collection information corresponding to each historical time period, or feature information including content feature vectors, spatial structure feature vectors, and entity feature vectors. Before obtaining the temporal feature vectors corresponding to the first to the Nth preset time period, the collection information of the corresponding historical time period stored in the spatiotemporal database can be extracted, and steps S1011-S1014 can be executed to obtain the scene status feature vectors corresponding to each historical time period. Alternatively, the feature information of the corresponding historical time period stored in the spatiotemporal database can be extracted, and step S1014 can be executed to obtain the scene status feature vectors corresponding to each historical time period.

[0053] As mentioned above, by loading real historical features into the spatiotemporal database, it can be ensured that effective collected data or feature data are used from the first preset time period, thereby enhancing the temporal representation capability of spatiotemporal feature vectors.

[0054] In one specific embodiment, S1011 includes the following steps:

[0055] S10111, acquire several historical data corresponding to the target physical location in each preset time period, wherein the historical data are historically acquired images, historically acquired videos or historically acquired text, and each historical data corresponds to the state of several target objects in the target physical location in the current preset time period.

[0056] S10112, sample each historical video, and use each sampled image frame and each historical image in the historical data as the historical image of the target physical location within the corresponding preset time period.

[0057] Historically acquired images present the visual information of the target physical location as static images. Historically acquired videos are continuous visual information of the target physical location recorded by video capture devices such as cameras within a preset time period. They can dynamically record the behavior and changes of target objects within the location, such as customer walking routes in a shopping mall or the operational dynamics of shops. Acquired text conveys semantic information about the location through text.

[0058] Image frames are the basic building blocks of captured video. Each image frame is a static image that records the scene information of a historically captured video at a specific moment, including the physical location of the target and the state of the target object at that moment. Sampling rules can be based on time intervals (e.g., selecting one frame at fixed intervals) or frame intervals (e.g., selecting one frame at fixed frame intervals) to reduce data volume while preserving key video information and improving the efficiency of subsequent data processing. The fixed time or fixed number of frames for sampling can be set by the implementer based on the actual conditions of the video capture equipment.

[0059] In one specific embodiment, the preset content encoder includes a preset image encoder and a preset text encoder, and S1012 includes the following steps:

[0060] S10121, For any historical data, if the current historical data is a historical image, then the current historical data is encoded according to the preset image encoder to obtain the content feature vector corresponding to the current historical data.

[0061] S10122, If the current historical data is historical text, then the current historical data is encoded according to the preset text encoder to obtain the content feature vector corresponding to the current historical data.

[0062] The pre-defined image encoder is based on a convolutional neural network architecture in deep learning, automatically extracting features from historically acquired images through components such as convolutional layers, pooling layers, and fully connected layers. Specifically, the convolutional kernels in the convolutional layers slide across the historically acquired images, extracting local features such as edges and textures through convolution operations. Pooling layers are used to compress data, reducing computation while retaining key features. Through the stacking of multiple convolutional and pooling layers, higher-level and more abstract image features are gradually extracted. Finally, a fully connected layer maps the extracted image features into a fixed-length content feature vector. For the content feature vector of historically acquired images, each dimension can represent information such as edge feature intensity, texture pattern, and color distribution in the historically acquired images.

[0063] Text encoders can be based on word embedding techniques from natural language processing, such as recurrent neural networks (RNNs) or their variants, including Long Short-Term Memory (LSTM) networks and gated recurrent units (GRUs). First, a word embedding layer captures the semantic and syntactic information of each word in the historically collected text, transforming each word into a low-dimensional word vector representation. Then, an RNN or its variant processes the word vectors sequentially according to the order of the historically collected text, capturing contextual information and long-term dependencies. Finally, the entire text sequence is encoded into a fixed-length content feature vector. For the content feature vector of the historically collected text, each dimension can represent the weights of different semantic words, the distribution of text topics, and semantic and syntactic information.

[0064] As described above, the preset content encoder transforms the original historical images and texts into content feature vectors, realizing the transformation of data from its original form to a feature-based and abstract form. It removes redundant information from the historical data while retaining key features, providing high-quality input data for multimodal data fusion and model training.

[0065] In one specific embodiment, S1013 includes the following steps:

[0066] S10131, for any historical image acquired during the current preset time period, perform dimensional transformation on the current historical image to obtain a three-dimensional image corresponding to the current historical image.

[0067] S10132, Semantic annotation is performed on the 3D image corresponding to the current historical image to obtain the annotated image corresponding to the current historical image.

[0068] S10133, the labeled image corresponding to the current historical acquisition image is encoded according to the preset spatial structure encoder to obtain the spatial structure feature vector corresponding to the current historical acquisition image.

[0069] S10134, Entity extraction is performed on the labeled image corresponding to the current historical acquisition image to obtain the entity image corresponding to each target object in the current historical acquisition image.

[0070] S10135, Encode the entity image corresponding to each target object in the current historical acquisition image according to the preset human entity encoder, and obtain the entity feature vector corresponding to each target object in the current historical acquisition image.

[0071] Specifically, a preset image dimension conversion method is used to convert two-dimensional historical images into three-dimensional images, solving the problem of missing depth information in two-dimensional images and providing richer geometric information for subsequent extraction of spatial structural features. Those skilled in the art will understand that any image dimension conversion method in the prior art falls within the protection scope of this invention. For example, the preset image dimension conversion method could be to use a deep learning model to predict the depth value of each pixel in the historically acquired image, generate a pseudo-depth map, and then combine it with the original RGB image to form an RGB-D three-dimensional representation. Alternatively, it could be to reconstruct a three-dimensional point cloud from multi-angle historically acquired images using structured light motion or multi-view stereo vision technology to obtain a three-dimensional image. Specifically, the intrinsic and extrinsic parameters of the acquisition device corresponding to the acquired image can be obtained, and a conversion matrix from 2D image pixel coordinates to 3D coordinates can be generated based on the intrinsic and extrinsic parameters. The current acquired image is then converted into the corresponding three-dimensional image based on the conversion matrix.

[0072] Pixel-level semantic segmentation is performed on 3D images to label different categories of objects, such as people and objects, providing clear semantic guidance for the spatial structure encoder and the person entity encoder, so as to focus on meaningful person entities and their spatial relationships. In this embodiment, models such as Mask R-CNN can be used to identify the categories of person entities and instance boundaries in 3D images, and each pixel is labeled with the corresponding semantic category. For example, background pixels are labeled as 0, pixels of person A are labeled as 1, pixels of person B are labeled as 2, pixels of object C are labeled as 3, pixels of object D are labeled as 4, pixels of robot E are labeled as 5, and so on.

[0073] By using masks of people and objects in the semantic segmentation results, the regions where people and objects are located are extracted from the labeled images, and a clean entity image corresponding to each target object is generated, providing high-quality input for the subsequent person entity encoder and reducing background noise interference.

[0074] In this embodiment, the preset spatial structure encoder can adopt a graph neural network architecture, mainly including a multi-scale feature extraction layer, a semantic enhancement module, a spatial relationship modeling layer, and a geometric constraint module. The multi-scale feature extraction layer extracts visual features at different levels; for example, low-level features preserve spatial details for accurate region localization, while high-level features capture semantic information to assist in semantic region classification. The semantic enhancement module fuses semantic annotation information, using semantic masks as weights in the attention mechanism to enhance the understanding of functional areas of the scene and focus on meaningful spatial regions. The spatial relationship modeling layer constructs the scene topology and learns spatial relationships between regions. Specifically, it transforms physical space into a graph structure, using regions as nodes and spatial relationships as edges, updating node representations through convolution to capture spatial dependencies. The geometric constraint module introduces 3D geometric prior knowledge and uses a geometric constraint loss function to regulate feature learning and ensure the rationality of spatial relationships. Geometric constraints include distance constraints, angle constraints, and inclusion relationship constraints. Specifically, distance constraints penalize unreasonable inter-region distance predictions, angle constraints ensure that the angular relationships between regions conform to physical common sense, and inclusion relationship constraints force the learning of correct region inclusion levels. Finally, graph pooling is used to integrate all node information to generate the final spatial structure feature vector.

[0075] The pre-defined human entity encoder can include a multimodal input fusion layer, an appearance feature extraction branch, a pose and motion analysis branch, and an attribute classification and behavior prediction branch. The multimodal input fusion layer preprocesses the entity image (such as scaling and normalization) before inputting it into a convolutional layer, where convolutional operations encode the coordinates of human / object joints or corners into heatmaps or vector representations. The appearance feature extraction branch uses convolutional layers to extract visual features of the human region and employs pooling and fully connected layers to generate the human appearance vector. The pose and motion analysis branch calculates geometric features such as joint length and edge length based on human / object joints or corners. The attribute classification and behavior prediction branch predicts and analyzes the basic attributes and behavioral states of humans and objects. Finally, a gating fusion mechanism integrates the features from multiple branches to obtain the entity feature vector.

[0076] As described above, by encoding spatial structure feature vectors, structural information such as the spatial layout of the scene and the relative positions of objects in historically acquired images is extracted to represent the static layout features of the target physical location. By encoding entity feature vectors, the appearance features, posture, and actions of the target object are encoded, which complements the content feature vectors and together constructs a more comprehensive multimodal feature representation.

[0077] In one specific implementation, the preset content encoder, the preset spatial structure encoder, and the preset person entity encoder need to be trained in conjunction with the multimodal base model and the adaptation layer to achieve alignment of different modal data in the latent space, thereby supporting the understanding of entities, positional relationships, and events in the physical space.

[0078] The collaborative training process, relying on cold-start data and data augmentation strategies, is divided into two stages. The cold-start data consists of semantically labeled 3D point cloud datasets (such as ScanNet and ShapeNet), which are combined with multimodal models (such as CLIP and ViT) and specialized models (such as 3D object detection models) to simulate and annotate scenes, forming a multimodal dataset containing content (images / point clouds), spatial structure (geometric coordinates / topological relationships), and human entities (semantic labels / attributes). Further, different scene scripts (such as "office meeting" and "kitchen cooking") are generated using large models, and data augmentation is performed by replacing items / people in the scenes (such as changing table and chair styles, character clothing, etc.), thereby expanding data diversity and serving as the data foundation for collaborative training.

[0079] The first stage involves aligning the latent space of the encoder with that of the multimodal fundamental model, making the output features of the content encoder, spatial structure encoder, and person entity encoder compatible with the latent space of the multimodal fundamental model, and establishing a unified semantic representation. The content encoder extracts visual features (such as edges, textures, object categories, and text content) using convolutional neural networks (CNNs) or Transformers, and aligns the extracted content feature vectors with the latent vectors output by the base model, minimizing the feature distribution differences. This allows the content encoder to map visual information to the latent space of the base model, supporting cross-modal semantic association. The spatial structure encoder uses graph neural networks (GNNs) or geometric deep learning models (such as PointNet) to encode spatial coordinates and structural relationships, generating spatial structure feature vectors. It then matches these spatial structure feature vectors with predefined spatial latent vectors in the base model, using methods such as triplet loss to constrain feature proximity for similar spatial relationships. This allows the spatial structure encoder to represent the positional relationships between people and objects and to unify the spatial semantics with the base model. The person entity encoder transforms discrete labels into continuous entity feature vectors using text encoders (such as BERT) or attribute embedding layers. It compares these entity feature vectors with the entity prototypes stored in the base model, optimizing the feature distribution through cross-entropy loss or metric learning. This allows the person entity encoder to distinguish different identities, actions, and attributes, and to align with the entity semantics of the base model.

[0080] The adaptation layer adds a fully connected layer or attention layer between the encoder and the base module to adjust the feature dimensions and distribution, ensuring cross-module compatibility.

[0081] The second stage involves task-oriented training of the overall model. Based on the aligned encoder, the model is trained to complete physical space understanding tasks, such as entity recognition, text response, relationship reasoning, and event understanding. Specifically, the encoder parameters aligned in the first stage are fixed. Based on cold-start data and data augmentation strategies, combined with more scenario scripts and dynamic interaction data, positive / negative samples (such as reasonable / unreasonable spatial layouts) are constructed. Contrastive loss is used to improve the model's sensitivity to subtle differences, ultimately achieving a unified representation and reasoning of multi-dimensional information in physical space, thus completing the physical space understanding task.

[0082] The above-mentioned method aligns the content encoder, spatial structure encoder, person entity encoder, and the latent space of the multimodal fundamental model to map heterogeneous data such as visual content, spatial structure, and semantics to a unified feature space. This solves the problem of fragmented processing of multimodal data by traditional models, improves the ability to deeply understand physical space and generalize, and meets the needs of practical applications for high accuracy and strong adaptability of the model.

[0083] In one specific embodiment, S1014 includes the following steps:

[0084] S10141, For any target object, concatenate all entity feature vectors corresponding to the current target object to obtain the intermediate feature vector corresponding to the current target object.

[0085] S10142, concatenate the intermediate feature vectors of all target objects corresponding to the target physical location within the current preset time period to obtain the first target vector corresponding to the target physical location within the current preset time period.

[0086] S10143, concatenate all content feature vectors corresponding to the target physical location within the current preset time period to obtain the second target vector corresponding to the target physical location within the current preset time period.

[0087] S10144, concatenate all spatial structure feature vectors corresponding to the target physical location within the current preset time period to obtain the third target vector corresponding to the target physical location within the current preset time period.

[0088] S10145, according to the preset splicing order, splice the first target vector, the second target vector and the third target vector corresponding to the target physical location in the current preset time period to obtain the scene status feature vector corresponding to the target physical location in the current preset time period.

[0089] In this embodiment, all entity feature vectors corresponding to the current target object are concatenated to integrate multi-view or multi-time features of the same target object, forming a more comprehensive individual representation. The concatenation method in this embodiment is serial concatenation.

[0090] The system collects intermediate feature vectors of all target objects within the current time period and concatenates them according to preset rules to obtain the first target vector, integrating the overall human state features of the target physical location. For example, the vectors are concatenated after being sorted by the target object's ID.

[0091] The second target vector integrates all image information and text semantic information corresponding to the target physical location. The third target vector integrates the spatial structure information between all target objects within the target physical location. The scene status feature vector further integrates the overall human status features, all image information and text semantic information, and the spatial structure information between all target objects in the target physical location, forming a unified multimodal feature representation, which provides a complete and rich feature foundation for subsequent model training.

[0092] As described above, by integrating the spatial structure information between all target objects within the target physical location, the scene status feature vector further integrates the overall human state features, all image information and text semantic information, as well as the spatial structure information between all target objects in the target physical location. This unified representation of visual, semantic, and spatial information improves the richness and comprehensiveness of the representation of the target physical location, providing a complete and rich feature foundation for subsequent model training.

[0093] In one specific embodiment, S1015 includes the following steps:

[0094] S10151, according to the time sequence, concatenate the scene status feature vectors of the target physical location in the current preset time period and the N preset time periods before the current preset time period to obtain the reference feature vector of the target physical location in the current preset time period.

[0095] S10152, compress the reference feature vector corresponding to the target physical location within the current preset time period to obtain the temporal feature vector corresponding to the target physical location within the current preset time period.

[0096] Specifically, N+1 scene status feature vectors for preset time periods are arranged in chronological order and concatenated according to their dimensions to form a reference feature vector, thus preserving all feature information within the time window. If the dimension of a single scene status feature vector is K, then the dimension of the reference feature vector is (N+1)×K.

[0097] Then, a dimensionality reduction algorithm is used to compress the reference feature vector to a fixed dimension, thereby reducing the feature dimension and computational complexity, while extracting key patterns from the time series and removing redundant information. In this embodiment, the dimensionality reduction algorithm can be feature mapping through a fully connected layer.

[0098] The specific value of the time window length N can be set by the implementer based on the actual situation. For example, it can be used to stitch features together based on the average time window length in historical experience, or it can be set according to the actual needs of downstream application tasks.

[0099] The above-mentioned method further concatenates and compresses the N+1 scene status feature vectors corresponding to the current preset time period and the N preset time periods before the current preset time period into a temporal feature vector, capturing the multi-dimensional dynamic changes of the target physical location in terms of content modality, spatial structure modality and entity modality over a period of time.

[0100] In one specific embodiment, S10 includes the following steps:

[0101] S110, acquire several target acquisition data corresponding to the target physical location within the target time period, wherein each target acquisition data can be a target acquisition image or target acquisition text, and each target acquisition data corresponds to the state of several target objects in the target physical location within the target time period.

[0102] S120: Encode each target acquisition data corresponding to the target time period according to the preset content encoder to obtain the content feature vector corresponding to each target acquisition data.

[0103] S130: Encode each target acquisition image according to the preset spatial structure encoder and the preset human entity encoder to obtain the spatial structure feature vector corresponding to each target acquisition image, and the entity feature vector corresponding to each target object in each target acquisition image.

[0104] S140: Concatenate the content feature vector, spatial structure feature vector, and entity feature vector of the target physical location within the target time period to obtain the scene status feature vector of the target physical location within the target time period.

[0105] S150, the scene status feature vector corresponding to the target physical location within the target time period and the scene status feature vector corresponding to the N-M+1 to Mth preset time periods before the target time period are concatenated and compressed to obtain the temporal feature vector corresponding to the target physical location within the target time period.

[0106] S160, input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the target physical location within the target time period into the second initial world model to obtain several relation feature vectors corresponding to the target physical location within the target time period.

[0107] Among them, the target physical location corresponds to several target acquisition data within the target time period. Specifically, for each target acquisition data, it can be a target acquisition image in image mode or a target acquisition text in text mode.

[0108] The second initial world model is obtained by training and optimizing based on the first and second modal features corresponding to the reference physical location, based on the first initial model. It has the ability to generalize across domains and various reference physical locations, but it is not optimized for specific target physical locations.

[0109] The methods for obtaining the content feature vector, spatial structure feature vector, entity feature vector, scene status feature vector, temporal feature vector, and relational feature vector of the target physical location within the target time period can refer to the methods for obtaining the content feature vector, spatial structure feature vector, entity feature vector, scene status feature vector, temporal feature vector, and relational feature vector of the target physical location within each preset time period.

[0110] The above-mentioned approach, by comprehensively acquiring target data of the target physical location within the target time period and using preset content encoders, spatial structure encoders, and person entity encoders to specifically encode the target data, can fully extract effective information from the target data and obtain comprehensive and representative content feature vectors, spatial structure feature vectors, entity feature vectors, temporal feature vectors, and relational feature vectors. This improves the ability to understand the target physical location and target object in content modalities, spatial structure modalities, entity modalities, temporal modalities, and relational modalities.

[0111] S20, classify the second modality features within the target time period according to the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features. The target classification result includes known feature data and unknown feature data.

[0112] Specifically, the historical patterns represented by the second modality features and the local regularity knowledge base within the target time period are analyzed to determine whether each feature vector in the second modality features conforms to known patterns, so as to filter out unknown feature data. Thus, each feature vector in the second modality features is labeled, and the corresponding target classification results are added to the training set of the cloud world model to optimize the model parameters and improve the representation and adaptability of the real world model to the target physical location.

[0113] In one specific embodiment, S20 includes the following steps:

[0114] S201, for any second mode, take any feature vector of the target physical location corresponding to the current second mode within the target time period as the target feature vector.

[0115] S202, calculate the first distance between the target feature vector and each feature cluster set corresponding to the current second modality, based on the target feature vector and each feature cluster set corresponding to the current second modality.

[0116] S203, based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality, obtain the target classification result corresponding to the target feature vector.

[0117] S204: Traverse all feature vectors corresponding to all second modalities of the target physical location within the target time period, and obtain the target classification result corresponding to each feature vector in the second modal features of the target physical location within the target time period.

[0118] In one specific embodiment, S203 includes the following steps:

[0119] S2031, Obtain the preset first distance threshold.

[0120] S2032, if the first distance between the target feature vector and each feature cluster set corresponding to the current second modality is less than or equal to the preset first distance threshold, then the target classification result corresponding to the target feature vector is determined to be known feature data.

[0121] S2033, if the first distance between the target feature vector and all feature cluster sets corresponding to the current second modality is greater than the preset first distance threshold, then the target classification result corresponding to the target feature vector is determined to be unknown feature data.

[0122] The first distance is obtained by calculating the Euclidean distance, which is used to measure the similarity between the target feature vector and each feature cluster set. The smaller the first distance, the higher the similarity between the target feature vector and each feature cluster set.

[0123] For any feature cluster set corresponding to the current second modality, in this embodiment, the average distance between the target feature vector and all feature vectors in the current feature cluster set can be used as the first distance, or the distance between the target feature vector and the feature vector corresponding to the cluster center of the current feature cluster set can be used as the first distance.

[0124] Further, a preset first distance threshold is obtained for the current feature cluster set. If the first distance between the target feature vector and the current feature cluster set is less than or equal to the preset first distance threshold, it indicates that the similarity between the target feature vector and the current feature cluster set is high. The target feature vector can then be classified into the current feature cluster set, and the target classification result corresponding to the target feature vector is determined to be known feature data. The preset first distance threshold can be set according to time constraints. For example, the preset first distance threshold can be the maximum or average distance between the feature vector corresponding to the cluster center in the current feature cluster set and other feature vectors in the set.

[0125] If the first distance between the target feature vector and all feature cluster sets is greater than the preset first distance threshold, it means that the similarity between the target feature vector and all feature cluster sets corresponding to the current second modality is low, and the target feature vector cannot be classified into the feature cluster set corresponding to the current second modality. Therefore, the target classification result corresponding to the target feature vector is determined to be unknown feature data, indicating that the local pattern knowledge base cannot understand the target feature vector based on the currently stored local pattern information. Correspondingly, the target feature vector can be regarded as the personalized feature of the target physical location.

[0126] The above describes how target feature vectors are classified based on feature clustering sets to obtain target classification results. These results serve as personalized features of the target physical location. After category labeling, the data is input into the cloud for model training, thereby improving the representation and adaptability of the trained real-world model to the target physical location.

[0127] S30, based on the local rule knowledge base, the first modality feature, the second modality feature, and the target classification result corresponding to each feature vector in the second modality feature, perform model distillation on the cloud world model in the cloud to obtain the real world model corresponding to the target physical location. The real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

[0128] Among them, the cloud world model is a pre-trained general model with broad generalization ability but lacks location-specific knowledge. The feature vectors, first modality features, second modality features, and the target classification results corresponding to each feature vector in the feature cluster set of the local pattern knowledge base are aligned according to the time period and a hybrid training set is constructed according to a preset ratio.

[0129] The cloud-based world model is used as the teacher model, and the real-world model is used as the student model. By constructing a knowledge distillation loss that includes feature matching loss and classification consistency loss, and combining it with task loss for training, the parameters of the student model are updated by minimizing the total loss obtained by task loss + knowledge distillation loss, thus obtaining the real-world model corresponding to the target physical location.

[0130] The real-world model is used to execute the target application task based on the multimodal characteristics corresponding to the target physical location. The task loss is calculated based on the model output and the label corresponding to the target application task. The target application task can be set by the implementer according to the actual situation. For example, the target application task could be a task of predicting pedestrian traffic in a smart retail scenario. Correspondingly, the output of the student model could be the predicted pedestrian traffic value for a future period, and the corresponding label could be the actual pedestrian traffic value for the future period obtained through the counting of deployed cameras. The task loss is then calculated based on the predicted pedestrian traffic value and the corresponding actual pedestrian traffic value, serving as the loss basis for updating the parameters of the student model. Specifically, the task loss can be the mean squared error loss.

[0131] The feature matching loss is used to constrain the student model's extraction results of known features to be close to those of the teacher model. Specifically, the feature matching loss can be the mean squared error loss corresponding to the features in each modality. The classification consistency loss is used to constrain the clustering results of the student model for unknown features to align with the soft labels of the teacher model. Specifically, the temporal consistency loss can be the KL divergence loss.

[0132] In one specific implementation, the number of parameters corresponding to the real-world model is less than the number of parameters corresponding to the cloud-based world model.

[0133] Distillation allows student models to avoid replicating all parameters and computational complexity of teacher models. Instead, it captures only their core features and decision logic, resulting in a significant reduction in the number of parameters. This reduces computational complexity and memory usage, ensuring real-time inference at the edge and meeting low-latency requirements in various scenarios. It also lowers overall computational costs and facilitates large-scale applications.

[0134] As described above, by deploying periodic processing of newly emerging unknown feature data to trigger incremental distillation training, the real-world model can be adjusted according to changes in the target physical location. This allows the real-world model to dynamically adapt to the personalized patterns and dynamic changes of the target physical location, thereby improving the accuracy of task execution when performing target application tasks.

[0135] In one specific embodiment, S10 further includes the following steps:

[0136] S104, analyze several feature vectors in each feature cluster set corresponding to the current second mode to obtain feature change trend data for each feature cluster set of all reference physical sites for the current second mode.

[0137] S105, store all feature cluster sets corresponding to the current second mode and the feature change trend data corresponding to each feature cluster set into the local pattern knowledge base corresponding to the target physical location.

[0138] Specifically, for each feature cluster set, the changes of several feature vectors in each feature cluster set over time are analyzed to obtain the corresponding feature change trend data, so as to characterize the dynamic changes of each second mode of the target physical site. This serves as the basis for predicting the feature vector situation of the corresponding second mode in the future time period, and then the unconventional features are identified and classified through the predicted feature vectors.

[0139] The feature cluster set can serve as the static typical state of the reference physical site, while the feature change trend data can serve as the dynamic evolution law of the reference physical site. The feature cluster set and feature change trend data corresponding to each second mode are stored in the local law knowledge base corresponding to the target physical site. This makes it easier to combine the second mode features within the target time period and judge whether the second mode features of the target physical site within the target time period conform to the historical common laws based on the local law knowledge base. This further improves the accuracy of extracting personalized features of the target physical site in addition to the common laws of the offset reference.

[0140] In one specific embodiment, S104 includes the following steps:

[0141] S1041, for any feature cluster set corresponding to the current second modality, obtain the preset time period corresponding to each feature vector in the current feature cluster set.

[0142] S1042, Based on all feature vectors in the current feature cluster set and the preset time period corresponding to each feature vector, construct a time-series prediction model corresponding to the current feature cluster set. The time-series prediction model is used to characterize the feature change trend of all reference physical sites for the current feature cluster set of the current second mode.

[0143] S1043, traverse all feature cluster sets corresponding to the current second modality, and obtain the time series prediction model corresponding to each feature cluster set corresponding to the current second modality.

[0144] S1044, traverse all second modes and obtain the time series prediction model corresponding to each feature cluster set for each second mode.

[0145] S1045, store the time series prediction model corresponding to each feature cluster set of each second modality into the local regularity knowledge base corresponding to the target physical location.

[0146] The temporal prediction model can be based on a long short-term memory network, using gating mechanisms such as input gates, forget gates, and output gates to capture long-term dependencies in the time series. Furthermore, by fine-tuning all feature vectors in the current feature cluster set and the preset time period corresponding to each feature vector, a temporal prediction model corresponding to the current feature cluster set is constructed. This model is used to characterize the evolution of feature cluster sets corresponding to all reference physical locations over time, and then predicts the feature vectors corresponding to the target physical location for the current second mode in future time periods.

[0147] The above process, through feature clustering, temporal modeling, and knowledge storage, transforms implicit spatiotemporal trends into executable model parameters, improving the interpretability of scene understanding in the reference physical location. This allows for independent modeling for different second modalities and feature cluster sets, accurately adapting to the personalized patterns of the target physical location, and thus enhancing scene adaptability.

[0148] In one specific embodiment, S20 further includes the following steps:

[0149] S205, based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality, obtain the first classification result corresponding to the target feature vector, wherein the first classification result is outlier features and non-outlier features.

[0150] S206, if the first classification result corresponding to the target feature vector is an outlier feature, then the target classification result corresponding to the target feature vector is determined as unknown feature data.

[0151] S207, if the first classification result corresponding to the target feature vector is a non-outlier feature, then the target classification result corresponding to the target feature vector is obtained according to the target time period corresponding to the target feature vector and the time series prediction model corresponding to each feature cluster set corresponding to the current second modality.

[0152] S208, traverse all second modalities and obtain the target classification result corresponding to each feature vector in the multimodal features of the target physical location within the target time period.

[0153] If the first distance between the target feature vector and the current feature cluster set is less than or equal to the preset first distance threshold, it indicates that the target feature vector and the current feature cluster set have a high similarity. The target feature vector can be classified into the current feature cluster set, and the first classification result corresponding to the target feature vector is determined to be a non-outlier feature.

[0154] After the first round of classification of the target feature vector based on the feature cluster set, the time series prediction model corresponding to each feature cluster set corresponding to the current second modality is used to determine whether the target feature vector conforms to the corresponding feature change trend, thereby performing a second classification of the target feature vector to determine whether the target feature vector belongs to unknown feature data. Thus, the accuracy of the target classification result is improved through two rounds of classification operations.

[0155] If the first distance between the target feature vector and all feature cluster sets is greater than the preset first distance threshold, it means that the similarity between the target feature vector and all feature cluster sets corresponding to the current second modality is low, and the target feature vector cannot be classified into the feature cluster set corresponding to the current second modality. Therefore, the first classification result corresponding to the target feature vector is determined to be an outlier feature, and the target classification result corresponding to the target feature vector is further determined to be unknown feature data. This indicates that the local pattern knowledge base cannot understand the target feature vector based on the currently stored local pattern information, thereby clarifying the personalized features corresponding to the target physical location through category labeling.

[0156] In one specific embodiment, S207 includes the following steps:

[0157] S2071, for any feature cluster set corresponding to the current second modality, obtain the predicted feature vector of the target physical location for the current second modality within the target time period.

[0158] S2072, calculate the second distance between the target feature vector and the corresponding predicted feature vector.

[0159] S2073, traverse all feature cluster sets corresponding to the current second modality and obtain all second distances corresponding to the target feature vector.

[0160] S2074: Based on all the second distances corresponding to the target feature vector, obtain the target classification result corresponding to the target feature vector.

[0161] Specifically, the corresponding predicted feature vector is obtained through the time series prediction model, and the second distance is obtained through the Euclidean distance equal distance calculation method to characterize the similarity between the target feature vector and the predicted feature vector. Correspondingly, the smaller the second distance, the higher the similarity between the target feature vector and the predicted feature vector, that is, the higher the degree of agreement between the target feature vector and the feature change trend.

[0162] The preset second distance threshold corresponding to the current second modality is obtained. If any second distance corresponding to the target feature vector is less than or equal to the preset second distance threshold, it indicates that the similarity between the target feature vector and the corresponding feature change trend is high. Then, the target classification result corresponding to the target feature vector is determined to be known feature data.

[0163] If all second distances corresponding to the target feature vector are greater than the preset second distance threshold, it indicates that the similarity between the target feature vector and the corresponding feature change trend is low. In this case, the target classification result corresponding to the target feature vector is determined to be unknown feature data. The local regularity knowledge base cannot understand the target feature vector based on the currently stored local regularity information. Therefore, the personalized features corresponding to the target physical location are clarified through category labeling and transmitted to the cloud to participate in model training, so as to improve the representation degree and adaptability of the real-world model to the target physical location after training.

[0164] The preset second distance threshold can be set according to time conditions. For example, the preset second distance threshold can be the standard deviation of all second distances obtained in historical operations.

[0165] The above method first performs a first-round classification of the target feature vector based on the feature cluster set, quickly eliminating obviously abnormal outliers and reducing subsequent computation. Then, based on the time series prediction model corresponding to each feature cluster set of the current second modality, it judges whether the non-outliers conform to the corresponding feature change trend, thereby obtaining the target classification result and improving the accuracy and reliability of the target classification result.

[0166] The above-mentioned method, by acquiring the local pattern knowledge base corresponding to the target physical location before the target time period, provides the common patterns corresponding to the second modal features of different reference physical locations within historical time periods. It also acquires the content feature vector, spatial structure feature vector, entity feature vector, temporal feature vector, and relational feature vector of the target physical location within the target time period. This improves the understanding and representation ability of the target physical location and target object under content modality, spatial structure modality, entity modality, temporal modality, and relational modality. Furthermore, based on the local pattern knowledge base, it classifies the second modal features within the target time period, obtaining the target classification result corresponding to each feature vector in the second modality features, thereby enabling the classification of the second modality. Each feature vector in the two-modal features is labeled with a category to clarify the personalized features corresponding to the target physical location. Finally, by adding the local rule knowledge base, the first modal features, the second modal features, and the target classification results corresponding to each feature vector in the second modal features to the training set of the cloud world model, the cloud world model is distilled together. This adjusts the real-world model according to the modal features and personalized features of the target physical location, so that the real-world model can dynamically adapt to the personalized rules and dynamic changes of the target physical location. This improves the representation and adaptability of the trained real-world model to the target physical location, thereby improving the accuracy of task execution when performing target application tasks.

[0167] Example 2

[0168] This second embodiment provides a real-world model training system based on the local physical environment, such as... Figure 2 As shown, this real-world model training system based on the local physical environment includes:

[0169] The data acquisition module 21 is used to acquire the local pattern knowledge base corresponding to the target physical location before the target time period and the first modal features and the second modal features corresponding to the target physical location within the target time period. The first modal features include content feature vectors, spatial structure feature vectors and temporal feature vectors, and the second modal features include entity feature vectors and relation feature vectors.

[0170] The feature classification module 22 is used to classify the second modality features within the target time period according to the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features. The target classification result includes known feature data and unknown feature data.

[0171] The model training module 23 is used to perform model distillation on the cloud world model in the cloud based on the local rule knowledge base, the first modality feature, the second modality feature, and the target classification result corresponding to each feature vector in the second modality feature, to obtain the real world model corresponding to the target physical location. The real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

[0172] In one specific embodiment, the data acquisition module 21 includes:

[0173] The feature vector acquisition submodule is used to acquire the feature vectors of each reference physical location for each second mode within M preset time periods before the target time period. The second mode includes entity mode and relation mode. The feature vector for entity mode is the entity feature vector, and the feature vector for relation mode is the relation feature vector. M is a positive integer.

[0174] The feature vector clustering submodule is used to cluster all feature vectors corresponding to all reference physical locations for any second mode, and obtain several feature cluster sets corresponding to the current second mode.

[0175] The first data storage submodule is used to store all feature cluster sets corresponding to the current second modality into the local pattern knowledge base corresponding to the target physical location.

[0176] In one specific implementation, the feature vector acquisition submodule includes:

[0177] The historical data acquisition unit is used to acquire several historical data points corresponding to each reference physical location within M preset time periods. Each historical data point can be a historical image or historical text, and each historical data point represents the state of several target objects in the corresponding reference physical location within the current time period.

[0178] The first feature encoding unit is used to encode each historical data collected at the current reference physical location in the current preset time period according to a preset content encoder for any preset time period corresponding to any reference physical location, and obtain the content feature vector corresponding to each historical data collected.

[0179] The second feature encoding unit is used to encode each historical image of the current reference physical location in the current preset time period according to the preset spatial structure encoder and the preset human entity encoder, so as to obtain the spatial structure feature vector corresponding to each historical image and the entity feature vector corresponding to each target object in each historical image.

[0180] The first feature splicing unit is used to splice the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location within the current preset time period to obtain the scene status feature vector corresponding to the current reference physical location within the current preset time period.

[0181] The second feature splicing unit is used to splice and compress the scene status feature vectors corresponding to the current reference physical location within the current preset time period and N preset time periods before the current preset time period, so as to obtain the temporal feature vector corresponding to the current reference physical location within the current preset time period. Here, N+1 is the preset number of vector splices, and the scene status feature vector before the first preset time period is the preset feature vector.

[0182] The relation feature acquisition unit is used to input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location in the current preset time period into the first initial world model, and obtain several relation feature vectors corresponding to the current reference physical location in the current preset time period.

[0183] In one specific embodiment, the data acquisition module 21 includes:

[0184] The target acquisition data acquisition submodule is used to acquire several target acquisition data corresponding to the target physical location within the target time period. Each target acquisition data can be a target acquisition image or target acquisition text, and each target acquisition data corresponds to the state of several target objects in the target physical location within the target time period.

[0185] The first feature encoding submodule is used to encode each target acquisition data corresponding to the target time period according to the preset content encoder, and obtain the content feature vector corresponding to each target acquisition data.

[0186] The second feature encoding submodule is used to encode each target acquisition image according to the preset spatial structure encoder and the preset human entity encoder, to obtain the spatial structure feature vector corresponding to each target acquisition image, and the entity feature vector corresponding to each target object in each target acquisition image.

[0187] The first feature splicing submodule is used to splice the content feature vector, spatial structure feature vector and entity feature vector of the target physical location within the target time period to obtain the scene status feature vector of the target physical location within the target time period.

[0188] The second feature splicing submodule is used to splice and compress the scene status feature vector corresponding to the target physical location within the target time period and the scene status feature vector corresponding to the N-M+1 to Mth preset time periods before the target time period to obtain the temporal feature vector corresponding to the target physical location within the target time period.

[0189] The relation feature vector acquisition submodule is used to input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the target physical location within the target time period into the second initial world model to obtain several relation feature vectors corresponding to the target physical location within the target time period.

[0190] In one specific embodiment, the feature classification module 82 includes:

[0191] The target feature vector determination submodule is used to take any feature vector of the target physical location corresponding to the current second mode within the target time period as the target feature vector for any second mode.

[0192] The first distance calculation submodule is used to calculate the first distance between the target feature vector and each feature cluster set corresponding to the current second modality, based on the target feature vector and each feature cluster set corresponding to the current second modality.

[0193] The first target classification result acquisition submodule is used to obtain the target classification result corresponding to the target feature vector based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality.

[0194] The second target classification result acquisition submodule is used to obtain the target classification result corresponding to each feature vector in the second modality features of the target physical location within the target time period for all feature vectors corresponding to all second modalities within the target time period.

[0195] In one specific implementation, the first target classification result acquisition submodule includes:

[0196] The first distance threshold acquisition unit is used to acquire a preset first distance threshold.

[0197] The first classification unit is used to determine the target classification result corresponding to the target feature vector as known feature data if the first distance between the target feature vector and each feature cluster set corresponding to the current second modality is less than or equal to a preset first distance threshold.

[0198] The second classification unit is used to determine the target classification result corresponding to the target feature vector as unknown feature data if the first distance between the target feature vector and all feature cluster sets corresponding to the current second modality is greater than a preset first distance threshold.

[0199] In one specific embodiment, the data acquisition module 21 further includes:

[0200] The feature change trend data acquisition submodule is used to analyze several feature vectors in each feature cluster set corresponding to the current second mode, and obtain the feature change trend data corresponding to each feature cluster set of all reference physical sites for the current second mode.

[0201] The second data storage submodule is used to store all feature cluster sets corresponding to the current second modality and the feature change trend data corresponding to each feature cluster set into the local pattern knowledge base corresponding to the target physical location.

[0202] In one specific implementation, the feature change trend data acquisition submodule includes:

[0203] The time period alignment unit is used to obtain the preset time period corresponding to each feature vector in the current feature cluster set for any feature cluster set corresponding to the current second modality.

[0204] The first time-series prediction model construction unit is used to construct a time-series prediction model corresponding to the current feature cluster set based on all feature vectors in the current feature cluster set and the preset time period corresponding to each feature vector. The time-series prediction model is used to characterize the feature change trend of all reference physical sites for the current feature cluster set of the current second mode.

[0205] The second time-series prediction model construction unit is used to traverse all feature cluster sets corresponding to the current second modality and obtain the time-series prediction model corresponding to each feature cluster set corresponding to the current second modality.

[0206] The third time series prediction model construction unit is used to traverse all the second modes and obtain the time series prediction model corresponding to each feature cluster set of each second mode.

[0207] The third data storage unit is used to store the time series prediction model corresponding to each feature cluster set of each second modality into the local regularity knowledge base corresponding to the target physical location.

[0208] In one specific embodiment, the feature classification module 22 further includes:

[0209] The first classification result acquisition submodule is used to obtain the first classification result corresponding to the target feature vector based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality. The first classification result is the outlier feature and the non-outlier feature.

[0210] The second classification result acquisition submodule is used to determine the target classification result corresponding to the target feature vector as unknown feature data if the first classification result corresponding to the target feature vector is an outlier feature.

[0211] The third classification result acquisition submodule is used to obtain the target classification result corresponding to the target feature vector if the first classification result corresponding to the target feature vector is a non-outlier feature, based on the target time period corresponding to the target feature vector and the time series prediction model corresponding to each feature cluster set corresponding to the current second modality.

[0212] The fourth classification result acquisition submodule is used to traverse all the second modalities and obtain the target classification result corresponding to each feature vector in the multimodal features of the target physical location within the target time period.

[0213] In one specific implementation, the third classification result acquisition submodule includes:

[0214] The predictive feature vector acquisition unit is used to obtain the predictive feature vector of the target physical location for the current second mode within the target time period for any feature cluster set corresponding to the current second mode.

[0215] The second distance calculation unit is used to calculate the second distance between the target feature vector and the corresponding predicted feature vector.

[0216] The set traversal unit is used to traverse all feature cluster sets corresponding to the current second modality and obtain all the second distances corresponding to the target feature vector.

[0217] The third classification result acquisition unit is used to obtain the target classification result corresponding to the target feature vector based on all the second distances corresponding to the target feature vector.

[0218] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0219] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for training a real-world model based on a local physical environment, characterized in that, The training method includes the following steps: S10: Obtain the local pattern knowledge base corresponding to the target physical location before the target time period and the first modal feature and the second modal feature corresponding to the target physical location within the target time period. The first modal feature includes a content feature vector, a spatial structure feature vector, and a temporal sequence feature vector. The second modal feature includes an entity feature vector and a relational feature vector. S10 includes the following steps: S101, obtain the feature vectors of each reference physical location for each second mode within M preset time periods before the target time period, wherein the second mode includes entity mode and relation mode, the feature vector for entity mode is entity feature vector, the feature vector for relation mode is relation feature vector, and M is a positive integer; S102, For any second mode, cluster all feature vectors corresponding to all reference physical locations for the current second mode to obtain several feature cluster sets corresponding to the current second mode; S103, store all feature cluster sets corresponding to the current second modality into the local regularity knowledge base corresponding to the target physical location; S20, classify the second modality features within the target time period according to the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features, wherein the target classification result includes known feature data and unknown feature data. S20 includes the following steps: S201, for any second mode, take any feature vector of the target physical location corresponding to the current second mode within the target time period as the target feature vector; S202, based on the target feature vector and each feature cluster set corresponding to the current second modality, calculate the first distance between the target feature vector and each feature cluster set corresponding to the current second modality; S203, based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality, obtain the target classification result corresponding to the target feature vector; S204, traverse all feature vectors corresponding to all second modalities of the target physical location within the target time period, and obtain the target classification result corresponding to each feature vector in the second modal features of the target physical location within the target time period; S30, based on the local rule knowledge base, the first modal feature, the second modal feature, and the target classification result corresponding to each feature vector in the second modal feature, model distillation is performed on the cloud world model in the cloud to obtain the real world model corresponding to the target physical location, wherein the real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

2. The method for training real-world models based on local physical environments according to claim 1, characterized in that, S101 includes the following steps: S1011, Obtain several historical acquisition data corresponding to each reference physical location within M preset time periods, wherein each historical acquisition data is a historical acquisition image or historical acquisition text, and each historical acquisition data represents the state of several target objects in the corresponding reference physical location within the current time period; S1012, For any preset time period corresponding to any reference physical location, encode each historical data collected at the current reference physical location in the current preset time period according to the preset content encoder, and obtain the content feature vector corresponding to each historical data collected. S1013, based on the preset spatial structure encoder and the preset human entity encoder, encode each historical acquisition image of the current reference physical location in the current preset time period, and obtain the spatial structure feature vector corresponding to each historical acquisition image, as well as the entity feature vector corresponding to each target object in each historical acquisition image. S1014, concatenate the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location within the current preset time period to obtain the scene status feature vector corresponding to the current reference physical location within the current preset time period; S1015, the scene status feature vectors corresponding to the current reference physical location in the current preset time period and N preset time periods before the current preset time period are concatenated and compressed to obtain the temporal feature vectors corresponding to the current reference physical location in the current preset time period. Here, N+1 is the preset number of vector concatenations, and the scene status feature vectors before the first preset time period are preset feature vectors. S1016, Input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location in the current preset time period into the first initial world model to obtain several relation feature vectors corresponding to the current reference physical location in the current preset time period.

3. The method for training real-world models based on local physical environments according to claim 2, characterized in that, S10 includes the following steps: S110, acquire several target acquisition data corresponding to the target physical location within the target time period, wherein each target acquisition data is a target acquisition image or target acquisition text, and each target acquisition data corresponds to characterize the state of several target objects in the target physical location within the target time period; S120, Encode each target acquisition data corresponding to the target time period according to the preset content encoder to obtain the content feature vector corresponding to each target acquisition data; S130, each target acquisition image is encoded according to the preset spatial structure encoder and the preset human entity encoder to obtain the spatial structure feature vector corresponding to each target acquisition image and the entity feature vector corresponding to each target object in each target acquisition image. S140, the content feature vector, spatial structure feature vector and entity feature vector corresponding to the target physical location within the target time period are concatenated to obtain the scene status feature vector corresponding to the target physical location within the target time period; S150, the scene status feature vector corresponding to the target physical location within the target time period and the scene status feature vector corresponding to the N-M+1 to Mth preset time periods before the target time period are spliced ​​and compressed to obtain the temporal feature vector corresponding to the target physical location within the target time period. S160, the content feature vector, spatial structure feature vector and entity feature vector corresponding to the target physical location within the target time period are input into the second initial world model to obtain several relation feature vectors corresponding to the target physical location within the target time period.

4. The method for training real-world models based on local physical environments according to claim 1, characterized in that, S203 includes the following steps: S2031, Obtain the preset first distance threshold; S2032, if the first distance between the target feature vector and each feature cluster set corresponding to the current second modality is less than or equal to the preset first distance threshold, then the target classification result corresponding to the target feature vector is determined to be known feature data; S2033, if the first distance between the target feature vector and all feature cluster sets corresponding to the current second modality is greater than the preset first distance threshold, then the target classification result corresponding to the target feature vector is determined to be unknown feature data.

5. The method for training real-world models based on local physical environments according to claim 1, characterized in that, The number of parameters corresponding to the real-world model is less than the number of parameters corresponding to the cloud-based world model.

6. A real-world model training system based on a local physical environment, characterized in that, The training system includes: A data acquisition module is used to acquire a local pattern knowledge base corresponding to the target physical location before the target time period and a first modal feature and a second modal feature corresponding to the target physical location within the target time period. The first modal feature includes a content feature vector, a spatial structure feature vector, and a temporal sequence feature vector; the second modal feature includes an entity feature vector and a relational feature vector. The data acquisition module includes: The feature vector acquisition submodule is used to acquire the feature vectors of each reference physical location for each second mode within M preset time periods before the target time period. The second mode includes entity mode and relation mode. The feature vector for entity mode is the entity feature vector, and the feature vector for relation mode is the relation feature vector. M is a positive integer. The feature vector clustering submodule is used to cluster all feature vectors corresponding to all reference physical locations for any second mode, and obtain several feature cluster sets corresponding to the current second mode. The data storage submodule is used to store all feature cluster sets corresponding to the current second modality into the local pattern knowledge base corresponding to the target physical location; A feature classification module is used to classify the second modality features within the target time period according to the local pattern knowledge base, and obtain the target classification result corresponding to each feature vector in the second modality features, wherein the target classification result includes known feature data and unknown feature data. The feature classification module includes: The target feature vector determination submodule is used to take any feature vector of the target physical location corresponding to the current second mode within the target time period as the target feature vector for any second mode; The first distance calculation submodule is used to calculate the first distance between the target feature vector and each feature cluster set corresponding to the current second modality based on the target feature vector and each feature cluster set corresponding to the current second modality; The first target classification result acquisition submodule is used to obtain the target classification result corresponding to the target feature vector based on the first distance between the target feature vector and each feature cluster set corresponding to the current second modality; The second target classification result acquisition submodule is used to traverse all feature vectors corresponding to all second modes of the target physical location within the target time period, and obtain the target classification result corresponding to each feature vector in the second mode features of the target physical location within the target time period; The model training module is used to perform model distillation on the cloud world model in the cloud based on the local rule knowledge base, the first modal feature, the second modal feature, and the target classification result corresponding to each feature vector in the second modal feature, to obtain the real world model corresponding to the target physical location. The real world model is used to execute the target application task based on the multimodal features corresponding to the target physical location.

7. The real-world model training system based on local physical environment according to claim 6, characterized in that, The feature vector acquisition submodule includes: The historical data acquisition unit is used to acquire several historical data points corresponding to each reference physical location within M preset time periods. Each historical data point is either a historical image or historical text, and each historical data point represents the state of several target objects in the corresponding reference physical location within the current time period. The first feature encoding unit is used to encode each historical data collected at the current reference physical location in the current preset time period according to the preset content encoder for any preset time period corresponding to any reference physical location, and obtain the content feature vector corresponding to each historical data collected. The second feature encoding unit is used to encode each historical acquisition image of the current reference physical location in the current preset time period according to the preset spatial structure encoder and the preset human entity encoder, and to obtain the spatial structure feature vector corresponding to each historical acquisition image, as well as the entity feature vector corresponding to each target object in each historical acquisition image. The first feature splicing unit is used to splice the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location within the current preset time period to obtain the scene status feature vector corresponding to the current reference physical location within the current preset time period. The second feature splicing unit is used to splice and compress the scene status feature vector corresponding to the current reference physical location in the current preset time period and N preset time periods before the current preset time period, so as to obtain the temporal feature vector corresponding to the current reference physical location in the current preset time period. Here, N+1 is the preset number of vector splicing, and the scene status feature vector before the first preset time period is the preset feature vector. The relation feature acquisition unit is used to input the content feature vector, spatial structure feature vector and entity feature vector corresponding to the current reference physical location in the current preset time period into the first initial world model, and obtain several relation feature vectors corresponding to the current reference physical location in the current preset time period.

Citation Information

Patent Citations

  • Unmanned system control method and device based on multistage world model, and medium

    CN119270885A

  • Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search

    US20240386015A1