Extensible smart community knowledge graph construction method and system

By constructing a spatial topology map in the intelligent space and combining graph neural networks and large language models, the problems of dynamics and contextual relevance of multimodal perception data are solved, enabling accurate modeling and reasoning of dynamic events, behavioral chains and user intentions, thereby improving the responsiveness and service efficiency of the intelligent space.

CN121351973BActive Publication Date: 2026-03-20CHINA NAT POSTAL & TELECOMM APPLIANCES CORP +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511915475.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-20
Estimated Expiration
2045-12-18

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle the dynamic nature of multimodal sensing data, spatial contextual relationships, and the semantic structure of complex tasks in intelligent spaces, resulting in limited capabilities for accurate perception and response to user behavior and environmental conditions.

Method used

By constructing a spatial topology graph, multimodal perception data is associated with spatial nodes. Combining the contextual semantic enhancement of graph neural networks and the semantic reasoning capabilities of large language models, accurate modeling and reasoning of dynamic events, behavioral chains, and user intentions in intelligent spaces can be achieved.

Benefits of technology

It significantly improves the ability of smart spaces to understand and respond to complex interactive tasks, enhances the intelligence level and service efficiency of smart communities, and can flexibly adapt to the diverse needs of residents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121351973B_ABST
    Figure CN121351973B_ABST
Patent Text Reader

Abstract

The application provides an extensible smart community knowledge graph construction method and system, relating to the technical field of artificial intelligence, the method comprising: preprocessing multi-modal sensing data collected by multiple spatial regions in a smart community; constructing a spatial topology graph with spatial regions as nodes and physical adjacency relationships as edges according to multiple spatial regions and their physical adjacency relationships in the smart community; processing the preprocessed multi-modal sensing data to generate fusion features representing dynamic events in the region, and taking the fusion features as the attributes of the corresponding nodes in the spatial topology graph to construct a fusion graph for the knowledge graph, the method uses the spatial topology structure to enhance the correlation of multi-modal data, improve the interpretability and reasoning ability of the data, and thus significantly improve the understanding and response ability of the intelligent space to complex interaction tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to an extensible smart community knowledge graph construction method and system. BACKGROUND

[0002] With the popularity of Internet of Things and artificial intelligence technology, intelligent spaces (such as smart homes and smart communities) collect multi-modal data such as video, audio, and environmental parameters through a large number of sensors, aiming to achieve intelligent perception and response to user behavior and environmental state. However, the current mainstream technology faces bottlenecks in understanding complex, context-related human activities. On the one hand, traditional multi-modal fusion methods usually rely on strict spatio-temporal alignment, making it difficult to handle the problem of asynchronous and missing data in actual scenarios. On the other hand, they mostly focus on low-level feature splicing, lacking effective modeling of high-level dynamic semantics such as user behavior chains and task processes.

[0003] Some advanced existing solutions attempt to use large models for end-to-end scene understanding. For example, some solutions input visual and audio information into a large language model (LLM) for reasoning, but such methods generally ignore the inherent physical topology and task context of intelligent spaces, resulting in isolated and shallow understanding of events, making it difficult to support complex multi-step, cross-space tasks. Meanwhile, although knowledge graphs can represent entity relationships, existing construction methods mainly rely on static knowledge, which cannot effectively express dynamic causal chains centered on events and supported by multi-modal data.

[0004] In summary, there is an urgent need for a new technical solution in the current technical field that can bridge the gap between unstructured multi-modal perception data and structured knowledge representation. Specifically, there is an urgent need for a knowledge representation method that can organically integrate spatial topology, temporal dynamics, and multi-modal evidence, and combine it with the powerful semantic reasoning capabilities of large language models, thereby achieving accurate, dynamic, and structured understanding of user high-level intentions in intelligent spaces. SUMMARY

[0005] The present application provides an extensible smart community knowledge graph construction method and system to address the deficiencies of static knowledge representation, lack of spatial context, insufficient multi-modal fusion, and limited understanding of complex tasks in existing technologies, enabling accurate modeling and semantic reasoning of dynamic events, behavior chains, and user intentions in intelligent spaces.

[0006] The present application provides an extensible smart community knowledge graph construction method, comprising:

[0007] preprocessing multi-modal perception data collected from multiple spatial regions within the smart community;

[0008] According to the plurality of space regions and the physical adjacency relationship thereof pre-set in the smart community, a space topology graph is constructed, taking the space regions as nodes and the physical adjacency relationship as edges.

[0009] The pre-processed multi-modal perception data is processed to generate fusion features representing dynamic events in the region, and the fusion features are taken as attributes of corresponding nodes in the space topology graph to construct a fusion graph of the knowledge graph.

[0010] According to the expandable smart community knowledge graph construction method provided by the application, the multi-modal perception data includes at least one of video data, audio data and environmental sensor data.

[0011] The multi-modal perception data collected in the plurality of space regions in the smart community is pre-processed, specifically including: when the multi-modal perception data includes video data, uniformly sampling and normalizing the video stream; when the multi-modal perception data includes audio data, performing noise reduction processing and normalization on the audio data; and when the multi-modal perception data includes environmental sensor data, filtering and denoising the environmental sensor data using a low-pass filter and normalizing.

[0012] According to the expandable smart community knowledge graph construction method provided by the application, according to the plurality of space regions and the physical adjacency relationship thereof pre-set in the smart community, a space topology graph is constructed, taking the space regions as nodes and the physical adjacency relationship as edges, specifically including:

[0013] The space regions are divided according to the physical layout of the smart community, as nodes of the space topology graph; the sensors deployed in each space region are mapped to the nodes according to their physical positions; the physical adjacency relationship between nodes is established according to whether there is a physical connection relationship between each space region, to jointly construct the nodes and the physical adjacency relationship into the space topology graph.

[0014] According to the expandable smart community knowledge graph construction method provided by the application, the pre-processed multi-modal perception data is processed to generate fusion features representing dynamic events in the region, specifically including:

[0015] The multi-modal perception data is time-series encoded in a pre-set time window to generate time-series representations of each modality; the time-series representations of each modality are compressed to generate compressed features of each modality; the compressed features of each modality are cross-modally fused to generate the fusion features; the cross-modal fusion is to use a set of learnable query vectors to generate the fusion features through cross-attention interaction aggregation with the compressed features of each modality.

[0016] According to the scalable smart community knowledge graph construction method provided by the application, after the fusion graph forming the knowledge graph is constructed, the method further comprises: converting information in the fusion graph into natural language description; inputting the natural language description into a pre-trained large language model for semantic reasoning to generate a structured recognition result of user intention or behavior prediction.

[0017] According to the scalable smart community knowledge graph construction method provided by the application, the information in the fusion graph is converted into natural language description, specifically comprising:

[0018] A graph neural network is applied to the fusion graph, and the context semantic enhancement of the fusion features of each node is performed through an adjacency propagation mechanism to obtain enhanced fusion features.

[0019] The spatial topological relationship of the target node and its neighbor nodes and the enhanced fusion features are jointly converted into the natural language description.

[0020] According to the scalable smart community knowledge graph construction method provided by the application, the method further comprises: feeding back the semantic reasoning result generated by the large language model to the fusion graph as new knowledge to dynamically update the attributes of the corresponding nodes.

[0021] The application also provides a scalable smart community knowledge graph construction system, comprising:

[0022] A preprocessing module is configured to preprocess multi-modal perception data collected from multiple spatial regions in the smart community.

[0023] A topological graph construction module is configured to construct a spatial topological graph taking the spatial regions as nodes and the physical adjacency relationship as edges based on the pre-set multiple spatial regions and their physical adjacency relationship in the smart community.

[0024] A knowledge graph construction module is configured to process the preprocessed multi-modal perception data to generate fusion features representing dynamic events in the region, and take the fusion features as the attributes of the corresponding nodes in the spatial topological graph to construct a fusion graph forming the knowledge graph.

[0025] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the scalable smart community knowledge graph construction method according to any of the above when executing the computer program.

[0026] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the scalable smart community knowledge graph construction method according to any of the above.

[0027] The application further provides a computer program product comprising a computer program which, when executed by a processor, implements the scalable smart community knowledge graph construction method according to any one of the above.

[0028] The scalable smart community knowledge graph construction method and system provided by the application organizes multi-modal perception data of multiple spatial regions in a smart community in a structured manner through construction of a spatial topology graph, realizes accurate representation of dynamic events, behavior chains and user intentions, and enhances the correlation of multi-modal data by using the spatial topology structure, improves the data interpretability and reasoning ability, thereby significantly improving the understanding and response ability of the intelligent space to complex interaction tasks, enabling more flexible and effective adaptation to the diversified needs of residents, and enhancing the intelligent level and service efficiency of the smart community. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0030] Figure 1 is one of the flowcharts of the scalable smart community knowledge graph construction method provided by the application.

[0031] Figure 2 is the flowchart of constructing a spatial topology graph provided by the application.

[0032] Figure 3 is the generation flowchart of the fusion features provided by the application.

[0033] Figure 4 is the generation flowchart of the fusion features provided by the application.

[0034] Figure 5 is the flowchart of semantic reasoning by a large language model provided by the application.

[0035] Figure 6 is the second flowchart of the scalable smart community knowledge graph construction method provided by the application.

[0036] Figure 7 is the structural diagram of the scalable smart community knowledge graph construction system provided by the application.

[0037] Figure 8 is the structural diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0038] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0039] In the scenario of intelligent space (such as smart community, smart home, etc.), the collection and processing of multi-modal perception data are the basis for realizing intelligent services. These data include video, audio, environmental parameters (such as temperature, humidity, light intensity), etc., which are collected by cameras, microphone arrays, and sensors, etc. However, there are many characteristics in the actual collection process of these data, such as sampling asynchrony (different modal data collection frequencies), data incompleteness (some modal data may be missing due to device failure or environmental interference), and semantic shift (different modal data have differences in semantic expression), etc. These characteristics make it difficult for traditional multi-modal fusion methods to effectively process, especially when complex interactive tasks need to be modeled and understood.

[0040] In addition, user behaviors and events in intelligent space often have dynamicity and context correlation. For example, the activities of users between different space areas (such as living room, kitchen, bedroom) may involve multiple steps and tasks, and these behavior chains and task flows need to be modeled and reasoned in the time and space dimensions. However, existing knowledge graph construction methods mainly focus on static entity relationships, lacking effective representation of dynamic events and behavior chains. At the same time, although large language models have strong semantic understanding ability, their input is usually mainly natural language, which is difficult to directly process graph structure or multi-modal perception data, resulting in limitations in their application in intelligent space.

[0041] In summary, the existing technology fails to fully consider the dynamicity of data, spatial context correlation, and semantic structure of complex tasks when processing multi-modal perception data in intelligent space, which limits the precise perception and response ability of intelligent space systems to user behavior and environmental state.

[0042] In order to solve the above problems, the present application proposes an extensible smart community knowledge graph construction method and system. This method models the spatial topology structure, associates multi-modal perception data with spatial nodes, combines the context semantic enhancement of graph neural networks and the semantic reasoning ability of large language models, and realizes the precise modeling and reasoning of dynamic events, behavior chains and user intentions in intelligent space.

[0043] Before the technical solutions of the embodiments of the present application are described in detail, the terms involved in the embodiments of the present application are first explained.

[0044] Smart space: refers to a space environment that is perceived, analyzed, decided and automatically executed in real time in a physical environment through technologies such as the Internet of Things, sensors and artificial intelligence.

[0045] Multi-modal perception data: refers to data types from multiple heterogeneous perception channels, including video, voice, text, environmental parameters, etc., which have asynchronicity in the time and space dimensions.

[0046] Space topology graph: refers to the physical connection relationship or functional linkage relationship between entities such as devices, areas and personnel in a smart space, which is usually modeled in a graph structure.

[0047] Knowledge graph: a structured knowledge representation method that represents semantic relationships between entities through triples (entity, relationship, entity).

[0048] Transformer: a deep learning model architecture based on self-attention mechanism.

[0049] Large language model (LLM): a pre-trained large-scale language model for natural language processing tasks.

[0050] The scalable smart community knowledge graph construction method and system of the present application is implemented through a set of collaborative hardware architecture, mainly including a sensor network, a data processing unit, a computing server and a user interaction device. The sensor network is deployed in various areas of the smart community, covering cameras, microphone arrays, environmental sensors and other types, for real-time collection of multi-modal perception data such as video, audio, temperature, humidity, light intensity, etc., and transmission of the data to the data processing unit through wired or wireless networks. The data processing unit is responsible for receiving, storing and preprocessing these data, including normalization, noise reduction, filtering and other operations to eliminate noise and outliers, and indexing and managing the data to support subsequent analysis.

[0051] The computing server, as the core executive body of the system, undertakes the construction and reasoning tasks of the knowledge graph. It first constructs a space topology graph according to the physical layout and sensor deployment information of the smart community, then performs time series encoding and cross-modal fusion on the preprocessed multi-modal data, generates fusion features and embeds them into the space topology graph to form a fusion graph. On this basis, the computing server enhances the context semantic of the fusion graph through a graph neural network, and converts the information into natural language description with the help of a graph abstraction module, and inputs it into a large language model for semantic reasoning. The reasoning result is not only used to generate structured recognition of user intent or behavior prediction, but also fed back to the fusion graph to dynamically update the graph state, support multi-round interaction and complex task reasoning.

[0052] User interaction devices are used to realize real-time interaction with users, including voice assistants, touchscreens, mobile devices, and other input devices, as well as display screens, speakers, and other output devices. Users send instructions or requests through input devices, and the system-generated response results are displayed to users through output devices. User interaction devices are connected to computing servers through networks to ensure information flow and interactive response between users and the system.

[0053] In actual operation, the data collected by the sensor network is transmitted to the computing server after preprocessing by the data processing unit, the computing server completes the knowledge graph construction and reasoning task, and the result is fed back to the user through the user interaction device. This hardware architecture not only supports efficient processing and reasoning of large-scale multi-modal data, but also has good scalability and flexibility, which can adapt to different smart community scenarios and resource constraints, providing a solid hardware foundation for precise modeling and reasoning of dynamic events, behavior chains, and user intentions in intelligent spaces.

[0054] The following will be described in detail Figures 1-7 The scalable smart community knowledge graph construction method and system of the embodiment of the application.

[0055] Figure 1 is a flowchart of the scalable smart community knowledge graph construction method provided by the application, which comprises:

[0056] Step 101, preprocessing the multi-modal perception data collected in the multiple space areas in the smart community.

[0057] In a smart community, multiple space areas (such as living rooms, kitchens, bedrooms, public corridors, activity rooms, etc.) are equipped with multiple sensors for real-time collection of multi-modal perception data. These data include but are not limited to video data, audio data, and environmental sensor data. In order to ensure the accuracy and efficiency of subsequent processing, it is necessary to preprocess these multi-modal perception data first.

[0058] The main purpose of preprocessing is to clean, standardize and format the collected raw data, so as to eliminate noise, unify data format and extract useful information. The specific steps of preprocessing are as follows:

[0059] The multi-modal data collected by the sensor network in the smart community is first transmitted to the data processing unit. These data are attached with accurate time stamps when collected, so as to facilitate subsequent time series analysis and synchronous processing. After receiving these data, the data processing unit performs preliminary formatting processing to ensure the integrity and consistency of the data.

[0060] The collected multi-modal data is normalized to convert the data to a unified numerical range for subsequent analysis and processing. At the same time, the data is denoised to remove possible environmental noise and outliers. This process can improve the quality of the data and reduce the interference of noise on subsequent analysis.

[0061] The preprocessed data is stored in a local or cloud database and indexed. Through indexing, data within a specific time range or specific spatial area can be quickly retrieved and extracted, providing efficient data support for subsequent analysis tasks.

[0062] Through the above preprocessing steps, the multi-modal sensing data collected in the smart community is cleaned, standardized and formatted, providing high-quality and consistent input data for subsequent spatial topology graph construction, fusion graph generation and semantic reasoning. The preprocessed data not only reduces noise interference, but also unifies data format, improves data usability and processing efficiency, and lays a solid foundation for precise modeling and reasoning of dynamic events, behavior chains and user intentions in intelligent space.

[0063] Step 102, according to the pre-set multiple spatial regions and their physical adjacency relationships in the smart community, a spatial topology graph is constructed with the spatial regions as nodes and the physical adjacency relationships as edges.

[0064] In the smart community, the construction of spatial topology graph is the basis for multi-modal data fusion and dynamic event modeling. In this embodiment, the smart community refers to a community environment that realizes intelligent management through technologies such as Internet of Things, sensor network and artificial intelligence, which includes residential units, public facilities and activity areas, etc. The spatial region refers to a region within the community with a clear function or physical boundary, such as the living room, kitchen, bedroom in a residential unit, and public areas such as corridors, elevator rooms and activity rooms, etc. These regions are the basic constituent units of the smart community, and there is a "physical adjacency relationship" between them, that is, whether there is a direct physical connection or passage path between two spatial regions, such as the connection between the living room and the kitchen through a door, or the connection between two residential units through a corridor.

[0065] In constructing the spatial topology graph, the node represents a spatial region in the smart community, and each node corresponds to a specific region, such as the living room or the kitchen. The edge represents the physical adjacency relationship between the spatial regions, and if there is a direct physical connection between two spatial regions, an edge is drawn between the two nodes. For example, in a residential unit, there is a door between the living room and the kitchen, which means that there is a physical adjacency relationship between the living room and the kitchen, so in the topology graph, there will be an edge connecting the living room node and the kitchen node.

[0066] In this step, a spatial topology graph is constructed based on the pre-set spatial regions and their physical adjacency relationships within the smart community. For example, taking a residential unit as an example, the living room is a node, the kitchen is a node, and the door between the living room and the kitchen is an edge, thus forming a simple spatial topology graph. In this way, the spatial topology graph can clearly represent the layout and connection relationship of each spatial region in the smart community, providing a basic framework for multi-modal data fusion and dynamic event modeling of spatial perception.

[0067] Step 103, processing the pre-processed multi-modal perception data to generate fusion features representing dynamic events in the region, and taking the fusion features as attributes of corresponding nodes in the spatial topology graph, constructing a fusion graph of the knowledge graph.

[0068] The core task of step 103 is to convert the pre-processed multi-modal data into fusion features that can represent dynamic events in the region, and assign these fusion features to corresponding nodes in the spatial topology graph, thereby constructing a fusion graph of the knowledge graph. This process aims to effectively integrate data from different modalities to form a structured knowledge representation that reflects dynamic events.

[0069] First, after pre-processing, multi-modal perception data contains various types of information such as video, audio, environmental parameters, etc. Although these data have been preliminarily cleaned and standardized, they still need further processing to extract features that can represent dynamic events. For example, video data can extract features such as human activity trajectories and object movements; audio data can extract features such as speech content and background noise types; environmental parameter data can extract features such as temperature changes and light intensity changes. These features reflect specific events occurring in the region, such as personnel access, device usage, and environmental changes.

[0070] Next, through specific algorithms or models, these features from different modalities are fused. The purpose of fusion is to integrate information from different perception channels together to form a comprehensive feature representation, thereby more comprehensively reflecting dynamic events in the region. For example, by combining the activity trajectory of a person in the video and the speech content in the audio, it can be more accurately determined whether the event occurring in the region is a person talking to others or engaging in some activity alone. This fusion feature not only contains single modality information, but also reflects the association and interaction between different modalities.

[0071] Finally, the generated fusion features are used as attributes of the corresponding nodes in the spatial topology graph. In the spatial topology graph, each node represents a spatial region, and the fusion features endow these nodes with semantic information about dynamic events. For example, the fusion features of the living room node may include information such as the type of human activity and changes in environmental parameters, reflecting specific events occurring in the living room. By combining fusion features with the spatial topology graph, the constructed knowledge graph's fusion map not only possesses a spatial structure but also reflects the occurrence and development of dynamic events within the region.

[0072] Through step 103, the knowledge graph of the smart community can more comprehensively and accurately represent the dynamic events of various spatial areas within the community in the form of a fused graph. This fused graph provides rich semantic information for multimodal data fusion for spatial perception and dynamic event modeling, laying a solid foundation for subsequent intelligent analysis and decision-making.

[0073] The scalable smart community knowledge graph construction method provided in this invention constructs a spatial topology graph to structurally organize multimodal perception data from multiple spatial areas within a smart community. This enables accurate representation of dynamic events, behavioral chains, and user intentions. The method utilizes spatial topology to enhance the correlation of multimodal data, improve the interpretability and reasonability of the data, and thus significantly improve the ability of the smart space to understand and respond to complex interactive tasks. This allows the system to adapt to the diverse needs of residents more flexibly and effectively, enhancing the intelligence level and service efficiency of the smart community.

[0074] Furthermore, in smart communities, the preprocessing of multimodal sensing data is a crucial step in realizing subsequent knowledge graph construction and dynamic event modeling. This data includes video data, audio data, and environmental sensor data, collected through a sensor network deployed in various spatial areas of the smart community. To ensure data quality and consistency, targeted preprocessing of this multimodal sensing data is necessary.

[0075] For video data, the preprocessing process includes uniform sampling and normalization of the video stream. Specifically, the video data is captured at a frame rate of 30 frames per second using a fixed-mount camera. The total number of frames after sampling is N, and the time interval between video frames is... The video frames, after being uniformly sampled, are divided into days to obtain a daily frame sequence. Where T is the total number of frames for the day. To reduce data volume and improve processing efficiency, the system samples the video stream uniformly, for example, extracting one frame per second. The sampled video frames are then resized to a uniform image size, such as 224×224 pixels, to suit the input requirements of subsequent deep learning models.

[0076] Subsequently, the video frames are normalized, and the pixel values of each frame are normalized using the following formula:

[0077]

[0078] where, and are the mean and standard deviation of the pixel sequence of the video frame, represent the original video frame data at time t. This process can standardize the pixel values of the video data within a unified range, reduce the differences between different video sources, and provide high-quality input for subsequent feature extraction and analysis.

[0079] For audio data, the preprocessing process includes noise reduction and normalization. Audio data is collected by a microphone array with a sampling rate of 44.1kHz. Since the audio data may contain environmental noise and echo interference, the system first uses the Wav2Vec2 model to perform noise reduction on the audio data to remove these interferences and improve the quality of the audio signal.

[0080] The noise-reduced audio signal is further normalized using the following formula to normalize the amplitude value of the audio signal:

[0081]

[0082] where, and are the mean and standard deviation of the audio signal; is the amplitude value of the audio signal at time t. Normalized audio data can better adapt to subsequent speech recognition and semantic analysis tasks, ensuring the comparability and consistency of audio data in different environments.

[0083] For environmental sensor data, the preprocessing process includes filtering and noise reduction and normalization. The environmental sensor data collected includes temperature (unit: ℃), humidity (unit: %), and light intensity (unit: lux). These data may be affected by environmental interference or sensor accuracy during collection, so filtering is needed. The system uses a low-pass filter to filter and denoise the environmental sensor data to smooth short-term fluctuations in the data and remove high-frequency noise.

[0084] The filtered environmental data is further normalized using the following formula to normalize the original measurement value of each modality:

[0085]

[0086] where, and are the mean and standard deviation of the data of this modality; The original measurement value at time t. The normalized environmental data can better reflect the changes in the environmental state, providing standardized input for subsequent environmental state analysis and behavior association.

[0087] At the end of every 24 hours, all video, audio and environmental data of the day are uniformly stored and managed according to the timestamp, and daily data sets are provided for subsequent analysis tasks. Real-time data processing and real-time decision-making will be based on these daily data sets.

[0088] Through the above preprocessing steps, the multi-modal perception data collected in the smart community is cleaned, standardized and formatted, providing high-quality and consistent input data for subsequent spatial topology graph construction, fusion graph generation and semantic reasoning. The preprocessed data not only reduces noise interference, but also unifies the data format, improves data usability and processing efficiency, and lays a solid foundation for precise modeling and reasoning of dynamic events, behavior chains and user intentions in intelligent space.

[0089] Further, in the smart community, constructing a spatial topology graph is the basis for multi-modal data fusion and dynamic event modeling. The spatial topology graph takes spatial regions as nodes and physical adjacency relationships between spatial regions as edges, which can intuitively reflect the layout and connection relationship of each region in the community. Referring to Figure 2 , the specific implementation steps of constructing a spatial topology graph include:

[0090] 201. Divide the spatial regions according to the physical layout of the smart community as nodes of the spatial topology graph.

[0091] First, according to the physical layout information of the smart community, the community is divided into multiple preset spatial regions. These regions can be functional regions within a residential unit, such as living room, kitchen, bedroom, bathroom, etc., or public areas, such as corridors, elevator rooms, activity rooms, parking lots, etc. Each spatial region node R i defined by its three-dimensional boundary coordinates, represented as:

[0092] .

[0093] wherein, represents the minimum coordinate point of the region, represents the maximum coordinate point of the region.

[0094] For example, a residential unit can be divided into the following spatial regions: Living Room, Kitchen, Bedroom, Bathroom, and Corridor. These spatial regions serve as nodes in a spatial topology graph, with each spatial region node R i representing a specific spatial region and being assigned a unique identifier. For example, the Living Room node can be represented as R 1 , the Kitchen node can be represented as R 2 .

[0095] 202. The sensors deployed within each spatial region are homologously mapped to the nodes according to their physical locations.

[0096] In a smart community, various sensors (such as cameras, microphone arrays, environmental sensors, etc.) are deployed in different spatial regions. In order to associate sensors with spatial regions, it is necessary to homologously assign sensors to the corresponding spatial region nodes according to their physical locations. The specific steps are as follows:

[0097] Determine the physical location coordinates (x j , y j , z j ) of each sensor S j .

[0098] According to the location of the sensor, it is mapped to the nearest spatial region node R i . The homologous function is defined as: .

[0099] In this way, each sensor is mapped to the spatial region node to which it belongs, forming a set of sensors within each region : .

[0100] For example, if a camera is installed on the ceiling of the living room, its location coordinates are within the three-dimensional boundary of the living room, then the camera is homologously assigned to the Living Room node R 1.

[0101] 203. According to whether there is a physical connection relationship between each spatial region, the physical adjacency relationship between nodes is established to construct the spatial topology graph together with the nodes and the physical adjacency relationship.

[0102] The physical adjacency relationship between the space regions refers to whether there is a direct physical connection or a passable path between two regions. For example, the living room is connected to the kitchen through a door, or two residential units are connected through a corridor. These physical connection relationships will be edges in the topological graph, used to represent the connectivity between nodes. The specific steps are as follows:

[0103] Determine the physical connectivity relationship between each space region. If two space regions R p and R q have a passable region at their boundary, and the boundary overlap area width w pq is greater than or equal to the minimum passable threshold , then it is considered that there is a physical adjacency relationship between the two regions: 。

[0104] In the case of meeting the above conditions, an edge is established between the space regions R p and R q e pq .

[0105] For example, if the width of the door between the living room and the kitchen is greater than the passable threshold , then an edge is established between the living room node R 1 and the kitchen node R 2 e 12 .

[0106] Through the above steps, the divided space regions are taken as nodes, and the physical adjacency relationship is taken as edges, and finally a complete space topological graph G ( V , E A) is constructed, where:

[0107] V={v1,v2,…,v n} represents the set of space region nodes;

[0108] E represents the set of adjacency edges between nodes;

[0109] A i represents the attribute set of each node v i , used to represent the sensing ability of the region, including: sensor type distribution , sensor number |A i |, unit volume sensing density , and sensing update frequency set F i ​the functional area tag to which the region belongs.

[0110] For example, the living room node R 1 The attribute set of the living room node may include:

[0111] Sensor types: camera, microphone;

[0112] Number of sensors: 2;

[0113] Per-unit volume perception density: 0.5 (assuming the living room volume is 100 cubic meters, with 2 sensors);

[0114] Perception update frequency: update once per second;

[0115] Functional area tag: residential area.

[0116] Through the above steps, the constructed spatial topology graph not only clearly represents the layout and connection relationship of each space region in the smart community, but also provides an important structured basis for multi-modal data fusion and dynamic event modeling of space perception.

[0117] Further, in order to realize accurate modeling and understanding of dynamic events, it is necessary to further process the preprocessed multi-modal perception data to generate fusion features that can represent dynamic events in the region. This process includes three key steps: time series encoding, compression processing and cross-modal fusion, as shown in Figure 3 .

[0118] 301、In a preset time window, the multi-modal perception data is time series encoded to generate time series representations of each modality.

[0119] In a preset time window, the preprocessed multi-modal perception data is time series encoded to generate time series representations of each modality. This process aims to capture the dynamic changes of data in the time dimension, providing a basis for subsequent feature compression and fusion.

[0120] Time series encoding of video data: using a sliding window time series Transformer to process the video frame sequence. The sliding window mechanism can effectively capture the temporal relationship between video frames, and the Transformer architecture can model the inter-frame dependency relationship using self-attention mechanism. In this way, the time series representation r t f .

[0121] Temporal encoding of audio data: A bidirectional GRU (Gated Recurrent Unit) network is used to process the audio feature sequence. The bidirectional GRU can consider both forward and backward dependencies of the audio signal, thus better capturing the temporal features of the audio signal. In this way, the temporal representation r t a .

[0122] Temporal encoding of environmental sensor data: A one-dimensional convolutional network is used to process environmental sensor data (such as temperature, humidity, and light intensity). The one-dimensional convolutional network can effectively capture the local features and periodic changes of environmental data in the time series. In this way, the temporal representation r t e .

[0123] 302, compress the temporal representation of each modality to generate compressed features of each modality.

[0124] After generating the temporal representation of each modality, in order to reduce the feature dimension and improve the calculation efficiency, it is necessary to compress these temporal representations to generate more compact feature representations. This process is implemented through a specific compression function , which is implemented as follows:

[0125] Compression of video modality: The sliding window temporal Transformer is used to compress the temporal representation of the video modality { r t f ∣ t ∈ T} to generate compressed features .

[0126] Compression of audio modality: The bidirectional GRU network is used to compress the temporal representation of the audio modality { r t a ∣ t ∈ T} to generate compressed features .

[0127] Compression of environmental modality: The one-dimensional convolutional network is used to compress the temporal representation of the environmental modality { r t e ∣ t ∈ T} to generate compressed features .

[0128] The specific formula is as follows:

[0129]

[0130] wherein, denotes the compression function of modality k, and T denotes a preset time window; denotes the feature representation of modality k at time t.

[0131] 303、 The compressed features of each modality are cross-modally fused to generate the fusion features.

[0132] After generating the compressed features of each modality, cross-modal fusion is needed to generate fusion features that can represent dynamic events in the region. The purpose of cross-modal fusion is to integrate information from different modalities to form a comprehensive feature representation, thereby more comprehensively reflecting dynamic events occurring in the region.

[0133] Cross-modal fusion method: The present application adopts a cross-modal fusion method based on query vectors. Specifically, a set of learnable query vectors (Query Vectors) is introduced, which interacts with the compressed features of each modality through a multi-head attention mechanism. This method can dynamically aggregate semantic features of different modalities and capture the correlation and complementarity between modalities.

[0134] Fusion process: Let be the compressed features of the video, audio and environmental modalities of the spatial region node R i Through the Q-Former module (a lightweight cross-modal fusion module), the compressed features are fused through a multi-head attention mechanism to generate the final fusion features z i :

[0135]

[0136] wherein, the Q-Former module dynamically aggregates semantic features through the interaction of query vectors and modality compressed features to generate fusion features z i that can represent dynamic events in the region.

[0137] Finally, the fusion features z are embedded into the graph structure as the unified representation of the spatial region node , and the node feature set is obtained, and the spatial topology graph is expanded into a fusion graph .

[0138] Taking a living room node in a smart community as an example, suppose that within a preset time window T, the node has collected the following multi-modal data:

[0139] Video data: captures the activity trajectory of the personnel in the living room.

[0140] Audio data: record the content of the conversation to the person.

[0141] Environmental data: monitor changes in temperature and light intensity.

[0142] See Figure 4 , the specific steps include:

[0143] (1) Time sequence coding.

[0144] Video data is encoded by sliding window time sequence Transformer to generate time sequence representation r t f of the video modality.

[0145] Audio data is encoded by bidirectional GRU to generate time sequence representation r t a of the audio modality.

[0146] Environmental data is encoded by one-dimensional convolutional network to generate time sequence representation r t e of the environment modality.

[0147] (2) Compression processing.

[0148] The time sequence representation {r t f | r t ∈ T} of the video modality is compressed by sliding window time sequence Transformer to generate compressed feature .

[0149] The time sequence representation {r t a | r t ∈ T} of the audio modality is compressed by bidirectional GRU to generate compressed feature .

[0150] The time sequence representation {r t e | r t ∈ T} of the environment modality is compressed by one-dimensional convolutional network to generate compressed feature .

[0151] (3) Cross-modal fusion.

[0152] Use Q-Former module to perform cross-modal fusion on the compressed features to generate the final fusion feature z i .

[0153] Through the above steps, the generated fusion feature z i As the attributes of the corresponding nodes in the spatial topology graph, it provides rich semantic information for subsequent knowledge graph construction and dynamic event modeling. This process not only captures the temporal dynamics of multi-modal data, but also integrates information from different modalities through cross-modal fusion, thereby more comprehensively reflecting the dynamic events occurring in the region.

[0154] After constructing the fusion graph that forms the knowledge graph, in order to achieve accurate recognition of user intent and behavior prediction, the information in the fusion graph needs to be further processed and input into a pre-trained large language model for semantic reasoning. Referring to Figure 5 This process includes two key steps: natural language description conversion of fusion graph information and semantic reasoning based on large language models.

[0155] Step 501, converting the information in the fusion graph into a natural language description.

[0156] After the construction of the fusion graph is completed, the semantic information it contains needs to be converted into a natural language description in order to be input into a large language model for further processing.

[0157] This process specifically includes the following two sub-steps:

[0158] 1) Apply a graph neural network to the fusion graph to enhance the context semantic information of the fusion features of each node through an adjacency propagation mechanism to obtain enhanced fusion features.

[0159] A graph neural network (GNN) is applied to the fusion graph to enhance the context semantic information of the fusion features of each node. The graph neural network interacts the fusion features of a node with the features of its neighbor nodes through an adjacency propagation mechanism, thereby capturing the structured relationships and context information between nodes. Specifically, the update process of each layer of the graph neural network is defined as:

[0160]

[0161] Where:

[0162] represents the feature representation of node i at the l layer;

[0163] represents the set of neighbor nodes of node i;

[0164] is the learnable weight matrix of the l

[0165] is an activation function, such as ReLU;

[0166] AGG represents an aggregation function, such as a multi-head attention mechanism.

[0167] Through the propagation of the multi-layer graph neural network, the fused features of the nodes are enhanced, which can better reflect their contextual semantic information in the graph structure.

[0168] 2) Convert the spatial topological relationships of the target node and its neighbor nodes and the enhanced fused features into the natural language description.

[0169] After processing by the graph neural network, the enhanced features of each node not only contain its own semantic information, but also fuse the contextual information of its neighbor nodes. Next, these enhanced features and their corresponding spatial topological relationships need to be converted into natural language descriptions. This process is achieved through a graph structure abstraction module, with the following specific steps:

[0170] Extract the features of the target node and its neighbor nodes: For the target node i, extract its enhanced fused features , and the feature set of its neighbor nodes .

[0171] Integrate spatial topological relationships: Integrate the spatial topological relationships (such as node types, adjacency relationships, etc.) of the target node and its neighbor nodes with the enhanced fused features.

[0172] Generate natural language description: Through the graph structure abstraction module, convert the above information into a natural language description. For example, for the living room node R 1 , its natural language description may be as follows: In the living room, personnel activity is detected, and the environmental temperature is 25℃, with a light intensity of 500 lux. The kitchen area has related activities, connected to the living room through a door.

[0173] Step 502, input the natural language description into a pre-trained large language model for semantic reasoning to generate a structured recognition result or behavior prediction of the user's intent.

[0174] Input the generated natural language description into a pre-trained large language model (such as GPT or BERT), and use its powerful semantic understanding ability for reasoning. The large language model can analyze the semantic information in the natural language description and generate a structured recognition result or behavior prediction of the user's intent. For example:

[0175] User intent recognition: If the natural language description mentions "someone in the living room is looking for a remote control", the large language model can recognize the user's intent as "looking for a remote control".

[0176] Behavior prediction: If the description mentions "cooking activity in the kitchen, someone waiting in the living room", the large language model can predict the user's possible behavior as "preparing for a meal".

[0177] Since the large language model cannot directly process graph structure information, the graph structure summary module S is introduced to encode the structural state of the graph node and its neighbors into a natural language description (Prompt):

[0178]

[0179] wherein, is the natural language description of node i, is the fusion feature enhanced by the graph neural network, L is the maximum number of layers of the neural network, Meta i contains the meta information of the node.

[0180] To enhance the adaptability of the system to complex events and delayed behaviors, a multi-round interaction update mechanism of graph-LLM is designed. The semantic reasoning result generated by the large language model is fed back as new knowledge into the fusion graph to dynamically update the attributes of the corresponding nodes.

[0181] Specifically, each round of LLM response can be injected into the graph node as new knowledge and update the graph state for the next round of reasoning. The update process can be expressed as:

[0182]

[0183] wherein, denotes the graph update method, which is realized by a gated recurrent unit (GRU);

[0184] Response i is the response generated by the model.

[0185] The graph node continuously obtains semantic enhancement and state update in multiple rounds of language interaction, and finally builds an integrated intelligent framework that supports context awareness, structure memory and behavior guidance, supporting human-machine collaborative interaction in intelligent space.

[0186] In order to further understand the scheme of the embodiments of the present application, Figure 6 a flowchart of the scalable smart community knowledge graph construction method of the present embodiment is shown, which fuses multi-modal data and uses graph neural networks to enhance the perception and reasoning ability of intelligent space. It includes:

[0187] S1: Multi-modal data acquisition and preprocessing.

[0188] Data is collected from various sensors in the smart space, including video, audio, and environmental sensor data.

[0189] The collected multi-modal data is pre-processed, including data cleaning, format unification, and preliminary feature extraction.

[0190] S2: Physical space topology construction.

[0191] According to the physical layout of the smart community, the spatial topology structure is constructed, and each spatial region (node) and its physical adjacency relationship (edge) are defined.

[0192] The sensors are mapped to the corresponding spatial region nodes according to their deployment locations, and node attributes are added.

[0193] S3: Pre-trained encoder feature extraction.

[0194] The pre-processed multi-modal data is extracted using a pre-trained encoder.

[0195] The video data is compressed and represented at a single time, and the same operation is performed on audio and environmental data.

[0196] S4: Graph neural network structure enhancement.

[0197] The graph neural network is applied to the context semantic enhancement of the fused features of the nodes, and the node features are updated through the adjacency propagation mechanism. The output of the graph neural network is used to generate natural language descriptions that integrate the modal state of the nodes and the context events of the adjacent nodes.

[0198] The enhanced node features and natural language descriptions of the graph neural network are input into a pre-trained large language model (LLM) for semantic reasoning. The LLM automatically reasons the current interaction state based on the input description, identifies user intent, and generates structured response instructions or prediction results.

[0199] The semantic reasoning results generated by the LLM are fed back as new knowledge into the graph nodes, and the graph state is updated for the next round of reasoning. Through methods such as gated recurrent unit (GRU), the graph state is dynamically updated to adapt to new information and environmental changes.

[0200] Through the above steps, the knowledge graph of the intelligent space is continuously updated and improved to reflect the dynamic events and user behavior occurring in the community. This method not only improves the understanding and response ability of the intelligent space to user intent, but also enhances the adaptability and reasoning ability of the system to complex events.

[0201] The extensible smart community knowledge graph construction system provided by the embodiment of the present application is described below, and the extensible smart community knowledge graph construction system described below can be correspondingly referred to the extensible smart community knowledge graph construction method described above.

[0202] The embodiment of the present application provides an extensible smart community knowledge graph construction system, referring to Figure 7 , comprising:

[0203] The preprocessing module 710 is configured to preprocess the multi-modal perception data collected by the plurality of spatial regions in the smart community.

[0204] The topology graph construction module 720 is configured to construct a spatial topology graph with the spatial regions as nodes and the physical adjacency relationship as edges according to the plurality of spatial regions and the physical adjacency relationship thereof in the smart community.

[0205] The knowledge graph construction module 730 is configured to process the preprocessed multi-modal perception data to generate fusion features representing dynamic events in the region, and construct a fusion graph of the knowledge graph by taking the fusion features as attributes of corresponding nodes in the spatial topology graph.

[0206] Figure 8 An example of an entity structure diagram of an electronic device is shown in Figure 8 The electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 can communicate with each other through the communications bus 840. The processor 810 can invoke the logical instructions in the memory 830 to execute the extensible smart community knowledge graph construction method, which includes preprocessing the multi-modal perception data collected by the plurality of spatial regions in the smart community, constructing a spatial topology graph with the spatial regions as nodes and the physical adjacency relationship as edges according to the plurality of spatial regions and the physical adjacency relationship thereof in the smart community, and processing the preprocessed multi-modal perception data to generate fusion features representing dynamic events in the region, and constructing a fusion graph of the knowledge graph by taking the fusion features as attributes of corresponding nodes in the spatial topology graph.

[0207] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or parts of the present application that essentially contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0208] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the scalable smart community knowledge graph construction method provided by the above-mentioned methods. The method comprises: pre-processing multi-modal perception data collected by a plurality of spatial regions in the smart community; constructing a spatial topology graph taking the spatial regions as nodes and the physical adjacency relationship as edges according to a plurality of spatial regions and their physical adjacency relationship pre-set in the smart community; processing the pre-processed multi-modal perception data to generate fusion features representing dynamic events in the region, and taking the fusion features as attributes of corresponding nodes in the spatial topology graph to construct a fusion graph forming the knowledge graph.

[0209] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the scalable smart community knowledge graph construction method provided by the above-mentioned methods. The method comprises: pre-processing multi-modal perception data collected by a plurality of spatial regions in the smart community; constructing a spatial topology graph taking the spatial regions as nodes and the physical adjacency relationship as edges according to a plurality of spatial regions and their physical adjacency relationship pre-set in the smart community; processing the pre-processed multi-modal perception data to generate fusion features representing dynamic events in the region, and taking the fusion features as attributes of corresponding nodes in the spatial topology graph to construct a fusion graph forming the knowledge graph.

[0210] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0211] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0212] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A scalable method for constructing a knowledge graph for smart communities, characterized in that, include: Preprocessing of multimodal sensing data collected from multiple spatial areas within the smart community; Based on the multiple spatial regions and their physical adjacency relationships within the smart community, a spatial topology graph is constructed with the spatial regions as nodes and the physical adjacency relationships as edges. The preprocessed multimodal sensing data is processed to generate fusion features that characterize dynamic events in the region, and the fusion features are used as attributes of corresponding nodes in the spatial topology graph to construct the fusion graph of the knowledge graph. After constructing the fusion graph that forms the knowledge graph, the method further includes: Convert the information in the fused graph into a natural language description; The natural language description is input into a pre-trained large language model for semantic reasoning to generate a structured recognition result or behavior prediction of the user's intent. Converting the information in the fused graph into a natural language description specifically includes: A graph neural network is applied to the fused graph, and the contextual semantics of the fused features of each node are enhanced through the adjacency propagation mechanism to obtain the enhanced fused features; The spatial topological relationship between the target node and its neighboring nodes, along with the enhanced fusion features, are combined and converted into the natural language description. The method further includes: injecting the semantic reasoning results generated by the large language model as new knowledge feedback into the fusion graph to dynamically update the attributes of the corresponding nodes; Specifically, the conversion of information in the fused graph into natural language descriptions is achieved through the graph structure summarization module S, and the process is represented as follows: ; in, It is the natural language description of node i. These are fused features enhanced by a graph neural network, where L is the maximum number of layers in the neural network, and Meta... i Includes the node's metadata; Specifically, the semantic reasoning results generated by the large language model are injected into the fusion graph as new knowledge feedback to dynamically update the attributes of the corresponding nodes. The specific update process is as follows: ; Among them, h i (L) The features of node i before the update, Response i U is the semantic reasoning result generated by the large language model for node i, U is the graph update function implemented using a gated recurrent unit, and h is the semantic reasoning result generated by the large language model for node i. i (L+1) It is the updated feature of node i.

2. The method according to claim 1, characterized in that, The multimodal sensing data includes at least one of video data, audio data, and environmental sensor data; Preprocessing of multimodal sensing data collected from multiple spatial areas within the smart community specifically includes: When the multimodal sensing data includes video data, the video stream is uniformly sampled and normalized; When the multimodal sensing data includes audio data, the audio data is denoised and normalized. When the multimodal sensing data includes environmental sensor data, a low-pass filter is used to filter, denoise, and normalize the environmental sensor data.

3. The method according to claim 1, characterized in that, Based on multiple pre-defined spatial regions within the smart community and their physical adjacency relationships, a spatial topology graph is constructed, with the spatial regions as nodes and the physical adjacency relationships as edges. Specifically, this includes: The spatial areas are divided according to the physical layout of the smart community, serving as nodes in the spatial topology map; The sensors deployed in each spatial area are assigned to the nodes based on their physical locations; Physical adjacency relationships between nodes are established based on whether there are physical connections between different spatial regions, so that the nodes and the physical adjacency relationships together construct the spatial topology graph.

4. The method according to claim 1, characterized in that, The preprocessed multimodal sensing data is then processed to generate fusion features characterizing dynamic events within the region, specifically including: The multimodal sensing data is temporally encoded within a preset time window to generate temporal representations of each modality. The temporal representations of each mode are compressed to generate compressed features for each mode; The compressed features of each modality are fused across modalities to generate the fused features; the cross-modal fusion is to use a set of learnable query vectors to generate the fused features through cross-attention interaction aggregation with the compressed features of each modality.

5. A scalable smart community knowledge graph construction system, characterized in that, include: The preprocessing module is used to preprocess the multimodal sensing data collected from multiple spatial areas within the smart community. The topology graph construction module is used to construct a spatial topology graph with the spatial regions as nodes and the physical adjacency relationships as edges, based on multiple preset spatial regions within the smart community and their physical adjacency relationships. The knowledge graph construction module is used to process the preprocessed multimodal perception data to generate fusion features that characterize dynamic events in the region, and use the fusion features as attributes of corresponding nodes in the spatial topology graph to construct the fusion graph of the knowledge graph. After constructing the fusion graph that forms the knowledge graph, the following is also included: Convert the information in the fused graph into a natural language description; The natural language description is input into a pre-trained large language model for semantic reasoning to generate a structured recognition result or behavior prediction of the user's intent. Converting the information in the fused graph into a natural language description specifically includes: A graph neural network is applied to the fused graph, and the contextual semantics of the fused features of each node are enhanced through the adjacency propagation mechanism to obtain the enhanced fused features; The spatial topological relationship between the target node and its neighboring nodes, along with the enhanced fusion features, are combined and converted into the natural language description. It also includes: injecting the semantic reasoning results generated by the large language model as new knowledge feedback into the fusion graph to dynamically update the attributes of the corresponding nodes; Specifically, the conversion of information in the fused graph into natural language descriptions is achieved through the graph structure summarization module S, and the process is represented as follows: ; in, It is the natural language description of node i. These are fused features enhanced by a graph neural network, where L is the maximum number of layers in the neural network, and Meta... i Includes the node's metadata; Specifically, the semantic reasoning results generated by the large language model are injected into the fusion graph as new knowledge feedback to dynamically update the attributes of the corresponding nodes. The specific update process is as follows: ; Among them, h i (L) The features of node i before the update, Response i U is the semantic reasoning result generated by the large language model for node i, U is the graph update function implemented using a gated recurrent unit, and h is the semantic reasoning result generated by the large language model for node i. i (L+1) It is the updated feature of node i.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the scalable smart community knowledge graph construction method as described in any one of claims 1 to 4.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the scalable smart community knowledge graph construction method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-modal fusion power grid dispatching knowledge graph generation method and device

    CN119721208A

  • Community intelligent monitoring and emergency linkage method and system fusing BIM spatial semantics

    CN120561322A