Cross-domain multi-modal dynamic reasoning and generating method and system for space sensing equipment linkage, terminal and storage medium
By weighted fusion and dynamic partitioning of real-time sensing data streams from heterogeneous devices, a cross-domain spatial sensing data stream with a unified spatiotemporal benchmark is generated. This solves the problems of response latency and resource waste in highly dynamic and heterogeneous scenarios, and achieves efficient collaborative control and adaptive capabilities, thus meeting the flexible production needs of modern manufacturing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies suffer from high response latency, poor system flexibility, serious resource waste, and insufficient scenario adaptability when facing highly dynamic, strongly coupled, and heterogeneous industrial collaborative scenarios, thus failing to meet the high real-time and flexible production requirements of modern manufacturing.
By receiving real-time sensing data streams from multiple heterogeneous devices, weighted fusion and spatial alignment are performed to dynamically divide semantic cooperation regions, generate cross-domain spatial sensing data streams with unified spatiotemporal benchmarks, perform multimodal fusion and semantic bridging, generate collaborative control strategies, update parallel rendering visualization output in real time, achieve quality assessment and local regeneration, and adapt to target device protocols.
It achieves millisecond-level real-time response, cross-protocol non-stop self-adaptation, and precise localized anomaly repair, improving the system's response speed and resource utilization, and adapting to dynamic adjustments in complex scenarios.
Smart Images

Figure CN121842230A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent device network technology in industrial scenarios, and in particular to a cross-domain multimodal dynamic reasoning and generation method, system, terminal, and computer-readable storage medium for spatial sensing device linkage. Background Technology
[0002] Driven by the current wave of Industry 4.0 and intelligentization, traditional industrial control, logistics scheduling, and infrastructure operation and maintenance are undergoing a profound transformation from automation to autonomy, and from local optimization to global collaboration. Typical scenarios, such as industrial manufacturing, port logistics, and power systems, increasingly rely on real-time collaboration and intelligent decision-making based on multi-device and multi-modal data to significantly improve production efficiency, response speed, and resource utilization. However, existing technologies reveal numerous systemic bottlenecks when facing complex collaborative scenarios characterized by high dynamism, strong coupling, and heterogeneity, severely restricting the in-depth deployment and effectiveness release of intelligent applications.
[0003] Existing technologies face the following problems when dealing with large-scale, high-real-time, and heterogeneous intelligent collaboration scenarios: (1) Response latency is insufficient to meet the requirements of high real-time collaboration: In high-precision control scenarios such as industrial manufacturing, flexible production lines such as automotive welding and 3C electronic assembly often require millisecond-level collaboration of dozens or even hundreds of devices (such as robotic arms, AGVs, and sensors). Traditional edge computing or centralized control solutions often adopt fixed task partitioning and resource allocation strategies (for example, pre-dividing the workshop into four fixed control areas). When order changes, process adjustments, or sudden disturbances occur, this static partitioning can easily lead to severe uneven load distribution among computing nodes. Critical control commands (such as robot trajectory compensation commands) may generate transmission delays exceeding safety thresholds due to queue congestion or lengthy paths. For example, in automotive spot welding processes, the control latency is required to be no more than 100 milliseconds, but in traditional solutions with more than 50 devices collaborating, the latency often exceeds 300 milliseconds, which not only affects product quality but may also lead to production safety risks.
[0004] (2) High cost of cross-domain heterogeneous device protocol interoperability and insufficient system flexibility: In industrial settings, information technology (IT) and operational technology (OT) systems coexist for a long time with fragmented protocols. Devices from different manufacturers and at different times may use multiple industrial communication protocols such as Modbus, OPC UA, and PROFINET, with different data semantics and formats. To achieve device interconnection, it is usually necessary to develop customized conversion layers or adapters for each protocol combination, resulting in high costs for single device access. More importantly, whenever a new device is introduced to the production line or a protocol version is upgraded, system-level configuration modifications and testing verification are required, which may take several days and seriously affect the availability and rapid reconfiguration capabilities of the production line. This rigid integration method cannot meet the flexible production needs of modern manufacturing industries for small batches and multiple varieties.
[0005] (3) Global recalculation leads to severe waste of computing resources: In IoT-based monitoring and diagnostic systems (such as power grid fault diagnosis and predictive maintenance of equipment), traditional architectures typically employ data processing logic where a change in one part affects the whole system. That is, any abnormal data reported by any sensor or device in the system will trigger a full recalculation of the entire system's state and a refresh of the panoramic view. For example, if a transformer monitoring point reports abnormal data, the traditional solution will send back and process the data from all related sensors to regenerate a panoramic diagnostic map of the entire substation. According to statistics, this global recalculation can consume more than 70% of computing resources on edge nodes, most of which is consumed in processing unchanged or irrelevant data. In resource-constrained edge environments, this results in a huge waste of computing power and also increases the delay in the system's response to critical events.
[0006] (4) Lack of dynamic adjustment and precise processing capabilities for scenario adaptation: Existing technologies lack the ability to dynamically and intelligently adjust system behavior based on real-time scenarios. Fixed partitions cannot cope with load fluctuations, rigid protocol adaptation cannot support rapid expansion, and extensive full updates cannot achieve resource efficiency. Especially in scenarios such as port scheduling and disaster emergency response, environmental conditions and task priorities change rapidly. The system needs to be able to understand scenario semantics (such as avoidance and priority processing) and dynamically adjust resource allocation, data processing granularity, and communication strategies. However, most existing methods rely on preset rules and thresholds, have limited intelligence, and cannot achieve closed-loop optimization based on data and model feedback.
[0007] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0008] The main objective of this invention is to provide a cross-domain multimodal dynamic reasoning and generation method, system, terminal, and computer-readable storage medium for the linkage of spatial sensing devices. This invention aims to solve the problems faced by existing technologies in dealing with large-scale, high-real-time, and heterogeneous intelligent collaborative scenarios, such as excessive response latency, poor system flexibility, low resource efficiency, and insufficient scene adaptability.
[0009] To achieve the above objectives, the present invention provides a cross-domain multimodal dynamic reasoning and generation method for linkage of spatial sensing devices, the method comprising the following steps: Through the device access layer, real-time sensing data streams from multiple heterogeneous devices are received, and the multiple real-time sensing data streams are weighted, fused, and spatially aligned to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference. Based on the cross-domain spatial perception data stream, according to the real-time spatial topology relationship of the device cluster and the preset business constraints, the physical space is dynamically divided into one or more semantic collaboration regions, and a mapping relationship is established between each semantic collaboration region and the corresponding visualization generation region. For each semantic collaboration region, the parsed spatial semantics, text instructions, and visual information are integrated to perform contextual enhancement and semantic bridging, generating a collaborative control strategy for the regional device cluster. Based on the mapping relationship, the visualization content corresponding to the collaborative control strategy is rendered in parallel spatial blocks, and the status information of each associated device is updated synchronously to generate preliminary visualization output. The preliminary visualization output is subjected to regional-level quality reflection and evaluation to determine whether the output results of each semantic cooperation region meet the preset quality standards. If all areas meet the standards, the final generated collaborative control strategy and visualization content will be adapted and converted into control commands or status mapping information that conform to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display. If there are non-compliant areas, a local regeneration protocol is initiated for the identified non-compliant areas to recalibrate the perception data associated with the non-compliant areas, and to update the collaborative control strategy generation and visualization rendering for the non-compliant areas.
[0010] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein the step of receiving real-time sensing data streams from multiple heterogeneous devices through a device access layer, and performing weighted fusion and spatial alignment on the multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference, specifically includes: Acquire real-time sensing data streams reported by multiple heterogeneous devices in at least one scenario in industrial manufacturing, port logistics, or power grid systems, wherein the heterogeneous devices include industrial robots, AGVs, sensors, and cameras; Based on preset equipment weighting factors and data quality evaluation indicators, multiple real-time sensing data streams are weighted and fused to obtain weighted fused data. The equipment weighting factors are dynamically adjusted based on equipment accuracy, historical reliability and current operating conditions. Based on a unified global coordinate system and time server, the weighted and fused data is transformed into spatial coordinates and synchronized with timestamps to generate a cross-domain spatial perception data stream with a unified spatiotemporal reference that can be directly accessed.
[0011] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein the step of dynamically dividing the physical space into one or more semantic cooperation regions based on the cross-domain spatial sensing data stream, according to the real-time spatial topology relationship of the device cluster and preset business constraints, and establishing a mapping relationship between each semantic cooperation region and the corresponding visualization generation region, specifically includes: The cross-domain spatial sensing data stream is analyzed to construct and update the dynamic spatial topology diagram of the device cluster in real time. Combining preset business constraint rules, including process steps, safety radius, or task coupling degree, the current optimal number and boundary of semantic cooperation regions are determined through a dynamic K-value partitioning algorithm; Establish a mapping relationship between the physical spatial extent of each semantic collaboration region and the corresponding logical generation region in the visualization rendering engine, wherein the mapping relationship is implemented through a coordinate transformation matrix or a grid index.
[0012] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein the step of fusing parsed spatial semantics, text commands, and visual information for each semantic cooperation region to perform contextual enhancement and semantic bridging, and generating a collaborative control strategy for a cluster of regional devices, specifically includes: For each semantic collaboration region, the spatial layout features, device state set, and environmental context information of each semantic collaboration region are extracted to form a region spatial semantic vector. Receive and parse natural language text instructions or predefined task scripts, and associate and enrich the natural language text instructions or predefined task scripts with the regional spatial semantic vector; By using a multimodal semantic bridging network, enhanced text instructions, regional spatial semantic vectors, and image features from visual sensors are fused together to infer and generate a collaborative control strategy for all devices within the semantic cooperation area.
[0013] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein the step of performing parallel spatial block rendering of the visualization content corresponding to the collaborative control strategy according to the mapping relationship, and synchronously updating the status information of each associated device to generate preliminary visualization output, specifically includes: Based on the mapping relationship, the collaborative control strategy for each region is assigned to different parallel rendering calculators; Each of the rendering calculators independently renders the visual content corresponding to the semantic collaboration area it is responsible for, wherein the visual content includes a quality heatmap, a device status panel, or a dynamic trajectory line. During parallel rendering, a distributed state synchronization mechanism is used to update and broadcast the latest status information of each associated device in real time, generating preliminary visualization output.
[0014] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial perception device linkage, wherein the step of performing regional-level quality reflection and evaluation on the preliminary visualization output to determine whether the output results of each semantic cooperation region meet the preset quality standards, specifically includes: Set a region-based quality assessment function, which sets differentiated thresholds for the accuracy, real-time performance, and completeness of the visualization output for different business scenarios; The visualization output is segmented according to semantic cooperation regions, and the corresponding quality assessment function is called to perform region-by-region quantitative scoring. The scoring results of each region are compared with the preset quality standards, and the regional quality compliance judgment results are output. The identification of the regions that fail to meet the standards and the specific non-compliance items are recorded.
[0015] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein if all regions meet the criteria, the final generated collaborative control strategy and visualization content are adapted and converted into control commands or state mapping information conforming to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display, specifically includes: Once all regions have been determined to meet the standards, the final collaborative control strategies and visualization content determined for all semantic collaborative regions are integrated. Through the protocol adapter, the integrated policies and content are converted into control command streams or status mapping data packets that conform to a specific communication protocol format, based on the brand and model of the target device. The converted instructions or data packets are distributed to various physical devices through the corresponding industrial network or IoT link to drive and execute corresponding actions or to be visualized on the terminal interface.
[0016] Optionally, the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage, wherein if there are substandard areas, a local regeneration protocol is initiated for the identified substandard areas to recalibrate the sensing data associated with the substandard areas, and the collaborative control strategy generation and visualization rendering for the substandard areas are updated, specifically including: Based on the recorded identity of the non-compliant area, a local regeneration protocol is initiated, suspending the policy and rendering update process for other compliant areas outside of that area; For the non-compliant areas, the associated sensing devices are triggered to re-acquire data or to perform filtering, completion, and calibration operations on the existing sensing data. Based on the recalibrated perception data, the collaborative control strategy generation and visualization rendering process for the non-compliant areas is re-executed to generate updated local results, and the updated local results are integrated into the overall output.
[0017] Optionally, the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: When a new device connects to the network, it actively broadcasts the device type, spatial awareness capabilities, supported communication protocols, and computing resource margin of the new device. After receiving broadcast information through the central scheduling node or adjacent devices, the capability information is registered to a unified device capability map, and an initial logical cooperation group is assigned to the new device based on spatial topology. Based on the capability map of the new equipment, dynamically adjust the data fusion weights, allocate rendering tasks, or select protocol adaptation paths.
[0018] Optionally, the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: Record the boundaries of the semantic collaboration region for each dynamic division, the typical collaborative control strategy patterns generated, and various cases of substandard quality and corresponding solutions; Based on historical records, the optimal partitioning strategies and inference patterns under different spatial topologies, business loads, and external events are summarized using machine learning methods and stored in a knowledge base; When similar scene characteristics are detected again, the best historical strategy is retrieved from the knowledge base and recommended as the initial parameter to accelerate the strategy generation process.
[0019] Furthermore, to achieve the above objectives, the present invention also provides a cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices, wherein the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices includes: The device access and spatiotemporal alignment module is used to receive real-time sensing data streams from multiple heterogeneous devices through the device access layer, and to perform weighted fusion and spatial alignment on the multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference. The semantic region dynamic partitioning module is used to dynamically divide the physical space into one or more semantic collaboration regions based on the cross-domain spatial perception data stream, according to the real-time spatial topology relationship of the device cluster and the preset business constraints, and to establish a mapping relationship between each semantic collaboration region and the corresponding visualization generation region. The multimodal strategy generation module is used to integrate parsed spatial semantics, text instructions and visual information for each semantic cooperation region, perform contextual enhancement and semantic bridging, and generate a collaborative control strategy for the regional device cluster. The visualization parallel rendering module is used to perform spatial block parallel rendering of the visualization content corresponding to the collaborative control strategy according to the mapping relationship, and synchronously update the status information of each associated device to generate preliminary visualization output. The quality reflection and evaluation module is used to perform regional-level quality reflection and evaluation on the preliminary visualization output, and to determine whether the output results of each semantic cooperation region meet the preset quality standards. The protocol adaptation and linkage output module is used to adapt and convert the final generated collaborative control strategy and visualization content into control commands or status mapping information that conform to the target device communication protocol if all areas meet the standards, and distribute them to the corresponding physical devices for execution or display. The local regeneration protocol execution module is used to, if there are non-compliant areas, initiate the local regeneration protocol for the identified non-compliant areas, recalibrate the perception data associated with the non-compliant areas, and update the collaborative control strategy generation and visualization rendering for the non-compliant areas.
[0020] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for spatial sensing device linkage, the device access and spatiotemporal alignment module includes: A sensing data stream acquisition unit is used to acquire real-time sensing data streams reported by multiple heterogeneous devices in at least one scenario in industrial manufacturing, port logistics or power grid systems, wherein the heterogeneous devices include industrial robots, AGVs, sensors and cameras. The sensing data stream weighted fusion unit is used to perform weighted fusion on multiple real-time sensing data streams based on preset equipment weight factors and data quality evaluation indicators to obtain weighted fused data. The equipment weight factors are dynamically adjusted based on equipment accuracy, historical reliability and current operating conditions. The spatial transformation and time synchronization unit is used to perform spatial coordinate transformation and timestamp synchronization on the weighted and fused data based on a unified global coordinate system and time server, so as to generate a cross-domain spatial perception data stream with a unified spatiotemporal reference and direct access.
[0021] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for spatial perception device linkage, the semantic region dynamic partitioning module includes: The dynamic spatial topology graph construction unit is used to parse the cross-domain spatial sensing data stream and construct and update the dynamic spatial topology graph of the device cluster in real time. The semantic cooperation region information confirmation unit is used to combine preset business constraint rules, wherein the business constraint rules include process steps, safety radius or task coupling degree, and determine the current optimal number and boundary of semantic cooperation regions through a dynamic K-value partitioning algorithm; The mapping relationship establishment unit is used to establish the mapping relationship between the physical spatial range of each semantic cooperation region and the corresponding logical generation region in the visualization rendering engine, wherein the mapping relationship is implemented through a coordinate transformation matrix or a grid index.
[0022] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices, the multimodal policy generation module includes: The regional spatial semantic vector generation unit is used to extract the spatial layout features, device state set and environmental context information of each semantic cooperation region for each semantic cooperation region, and to form a regional spatial semantic vector. The regional spatial semantic vector association unit is used to receive and parse natural language text instructions or predefined task scripts, and associate and enrich the natural language text instructions or predefined task scripts with the regional spatial semantic vectors; The collaborative control strategy generation unit is used to infer and generate a collaborative control strategy for all devices within the semantic collaboration area by fusing enhanced text instructions, regional spatial semantic vectors, and image features from visual sensors through a multimodal semantic bridging network.
[0023] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for spatial perception device linkage, the visualization parallel rendering module includes: The collaborative control strategy allocation unit is used to allocate the collaborative control strategy of each region to different parallel rendering calculators according to the mapping relationship. A visualization content rendering unit is used for each of the rendering calculators to independently render the visualization content corresponding to the semantic collaboration area they are responsible for, wherein the visualization content includes a quality heatmap, a device status panel, or a dynamic trajectory line. The preliminary visualization output generation unit is used to generate preliminary visualization output by updating and broadcasting the latest status information of each associated device in real time through a distributed state synchronization mechanism during parallel rendering.
[0024] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices, the quality reflection and evaluation module includes: The quality assessment function setting unit is used to set a region-based quality assessment function, which sets differentiated thresholds for the accuracy, real-time performance and completeness of the visualization output for different business scenarios. The region segmentation and quantization scoring unit is used to segment the visualization output according to the semantic cooperation region and call the corresponding quality evaluation function to perform region-by-region quantization scoring. The results comparison and recording unit is used to compare the scoring results of each region with the preset quality standards, output the regional quality compliance judgment results, and record the identity of the non-compliant regions and the specific non-compliant items.
[0025] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices, the protocol adaptation and linkage output module includes: The strategy and content integration unit is used to integrate the final collaborative control strategy and visualization content determined by all semantic collaboration areas after it is determined that all areas have met the standards. The policy and content conversion unit is used to convert the integrated policy and content into control command streams or status mapping data packets that conform to a specific communication protocol format, based on the brand and model of the target device, through a protocol adapter. The instruction or data packet distribution unit is used to distribute the converted instructions or data packets to various physical devices through the corresponding industrial network or IoT link to drive and execute corresponding actions or to display them visually on the terminal interface.
[0026] Optionally, in the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices, the local regeneration protocol execution module includes: The local regeneration protocol initiation unit is used to initiate the local regeneration protocol based on the recorded identity identifier of the non-compliant area, and to suspend the policy and rendering update process for other compliant areas outside of that area. The data processing unit for non-compliant areas is used to trigger the sensing devices associated with the non-compliant areas to re-acquire data or to perform filtering, completion and calibration operations on the existing sensing data for the non-compliant areas. The local result generation and integration unit is used to re-execute the collaborative control strategy generation and visualization rendering process for the non-compliant areas based on the recalibrated perception data, generate updated local results, and integrate the updated local results into the overall output.
[0027] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a cross-domain multimodal dynamic reasoning and generation program for spatial sensing device linkage stored in the memory and executable on the processor, wherein when the cross-domain multimodal dynamic reasoning and generation program for spatial sensing device linkage is executed by the processor, it implements the steps of the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage as described above.
[0028] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a cross-domain multimodal dynamic reasoning and generation program for linkage with spatial sensing devices, and when the cross-domain multimodal dynamic reasoning and generation program for linkage with spatial sensing devices is executed by a processor, it implements the steps of the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices as described above.
[0029] In this invention, a device access layer receives real-time sensing data streams from multiple heterogeneous devices, and performs weighted fusion and spatial alignment on these data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference. Based on this cross-domain spatial sensing data stream, and according to the real-time spatial topology of the device cluster and preset business constraints, the physical space is dynamically divided into one or more semantic collaboration regions, and a mapping relationship is established between each semantic collaboration region and its corresponding visualization generation region. For each semantic collaboration region, the parsed spatial semantics, text commands, and visual information are fused, and contextual enhancement and semantic bridging are performed to generate a collaborative control strategy for the regional device cluster. Based on the mapping relationship, the collaborative control strategy is... The corresponding visualization content is rendered in parallel spatial blocks, and the status information of each associated device is updated synchronously to generate preliminary visualization output. The preliminary visualization output is then subjected to regional-level quality reflection and evaluation to determine whether the output results of each semantically collaborative region meet preset quality standards. If all regions meet the standards, the final generated collaborative control strategy and visualization content are adapted and converted into control commands or status mapping information that conform to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display. If there are substandard regions, a local regeneration protocol is initiated for the identified substandard regions to recalibrate the perception data associated with the substandard regions and update the collaborative control strategy generation and visualization rendering for the substandard regions. This invention achieves millisecond-level real-time response for multi-device collaboration, cross-protocol uninterrupted self-adaptation, and precise localized repair of anomalies through a closed-loop mechanism of dynamic partitioning, multimodal fusion, and quality reflection. Attached Figure Description
[0030] Figure 1 This is a flowchart of a preferred embodiment of the cross-domain multimodal dynamic reasoning and generation method for the linkage of spatial sensing devices according to the present invention; Figure 2 This is a flowchart illustrating the entire process in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 3 This is a flowchart illustrating the specific implementation process of step S10 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 4 This is a flowchart illustrating the specific implementation process of step S20 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices; Figure 5 This is a flowchart illustrating the specific implementation process of step S30 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 6This is a flowchart illustrating the specific implementation process of step S40 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 7 This is a flowchart illustrating the specific implementation process of step S50 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 8 This is a flowchart illustrating the specific implementation process of step S60 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 9 This is a flowchart illustrating the specific implementation process of step S70 in a preferred embodiment of the invention of a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices. Figure 10 This is a flowchart illustrating the specific implementation process of dynamically adjusting parameters in a preferred embodiment of the invention's cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices; Figure 11 This is a flowchart illustrating the specific implementation process of accelerating the strategy generation process in a preferred embodiment of the invention's cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices; Figure 12 This is a schematic diagram of a cross-domain multimodal dynamic reasoning and generation system for the linkage of spatial sensing devices according to the present invention; Figure 13 This is a schematic diagram illustrating the cooperation of various modules in the cross-domain multimodal dynamic reasoning and generation system for spatial sensing device linkage, which is a schematic diagram of the present invention to complete the entire method. Figure 14 This is another structural schematic diagram of the cross-domain multimodal dynamic reasoning and generation system for linkage with spatial sensing devices according to the present invention; Figure 15 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0032] The preferred embodiment of the present invention describes a cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices, such as... Figure 1 and Figure 2 As shown, the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices includes the following steps: Step S10: Receive real-time sensing data streams from multiple heterogeneous devices through the device access layer, and perform weighted fusion and spatial alignment on the multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference.
[0033] Specifically, such as Figure 3 As shown, step S10 specifically includes: Step S11: Obtain real-time sensing data streams reported by multiple heterogeneous devices in at least one scenario of industrial manufacturing, port logistics or power grid system, wherein the heterogeneous devices include industrial robots, AGVs, sensors and cameras; Step S12: Based on the preset equipment weight factor and data quality evaluation index, the multiple real-time sensing data streams are weighted and fused to obtain weighted fused data. The equipment weight factor is dynamically adjusted based on equipment accuracy, historical reliability and current operating conditions. Step S13: Based on a unified global coordinate system and time server, the weighted and fused data is transformed into spatial coordinates and synchronized with timestamps to generate a cross-domain spatial perception data stream with a unified spatiotemporal reference that can be directly accessed.
[0034] For example, in an automotive welding workshop, the system simultaneously acquires real-time data streams from laser sensors (high-precision coordinates), welding torch controllers (process parameters), industrial cameras (weld seam vision), and AGV scheduling systems (material location). The system dynamically assigns weights based on the characteristics of each device: laser sensors, due to their high precision and reliability, have the highest weight in coordinate fusion; industrial cameras have an increased weight when identifying weld seam appearance; if a sensor experiences recent drift, its weight is dynamically reduced. All data is converted to the workshop's global coordinate system and synchronized with the central clock, ultimately generating a standardized data stream that integrates location, process, vision, and material flow information, reflecting the complete status of the welding station in real time.
[0035] By dynamically weighted fusion and strict spatiotemporal alignment, heterogeneous device data from different manufacturers and using different protocols are transformed in real time into a standardized data stream with a unified spatiotemporal benchmark. This mechanism can automatically adjust data weights dynamically based on device accuracy and operating conditions, effectively suppressing outlier interference and fusing multi-dimensional information, thereby eliminating data silos and spatiotemporal deviations at the source. This not only provides a highly reliable and consistent input foundation for subsequent dynamic partitioning and collaborative decision-making, but also significantly improves the system's scalability through standardized interfaces, enabling new devices to be quickly connected in a plug-and-play manner, laying a solid data foundation for the entire system's millisecond-level real-time response and precise collaboration.
[0036] Step S20: Based on the cross-domain spatial perception data stream, according to the real-time spatial topology relationship of the device cluster and the preset business constraints, the physical space is dynamically divided into one or more semantic collaboration regions, and a mapping relationship is established between each semantic collaboration region and the corresponding visualization generation region.
[0037] Specifically, such as Figure 4 As shown, step S20 specifically includes: Step S21: Analyze the cross-domain spatial sensing data stream and construct and update the dynamic spatial topology relationship diagram of the device cluster in real time; Step S22: Combining preset business constraint rules, wherein the business constraint rules include process steps, safety radius or task coupling degree, the current optimal number and boundary of semantic cooperation regions are determined by dynamic K-value partitioning algorithm; Step S23: Establish a mapping relationship between the physical spatial extent of each semantic collaboration region and the corresponding logical generation region in the visualization rendering engine, wherein the mapping relationship is implemented through a coordinate transformation matrix or a grid index.
[0038] For example, in an automotive welding production line, the system analyzes data streams from robots, AGVs, and sensors in real time to construct a dynamic positional relationship diagram between the equipment. When switching to the door assembly process, the system automatically divides the production line into three collaborative areas based on the process constraints and the robot's safety radius using a dynamic partitioning algorithm. Each area corresponds to an independent rendering block in the visualization system and is precisely mapped using a coordinate matrix.
[0039] The adaptive division of collaboration boundaries based on real-time operating conditions and business logic replaces the traditional fixed partitioning mode. This fundamentally solves the problem of uneven equipment load caused by production line changes or task changes, raising the flexibility and efficiency of collaborative scheduling to a new level and laying the core foundation for subsequent precise collaboration and parallel processing.
[0040] Step S30: For each semantic collaboration region, the parsed spatial semantics, text instructions and visual information are integrated to perform contextual enhancement and semantic bridging, and a collaborative control strategy for the regional device cluster is generated.
[0041] Specifically, such as Figure 5 As shown, step S30 specifically includes: Step S31: For each semantic cooperation region, extract the spatial layout features, device state set and environmental context information of each semantic cooperation region to form a regional spatial semantic vector; Step S32: Receive and parse natural language text instructions or predefined task scripts, and associate and enrich the natural language text instructions or predefined task scripts with the regional spatial semantic vector; Step S33: Through a multimodal semantic bridging network, the enhanced text instructions, regional spatial semantic vectors, and image features from the visual sensor are fused to infer and generate a collaborative control strategy for all devices within the semantic cooperation area.
[0042] For example, in a port scheduling scenario, the system first extracts the equipment layout, real-time AGV location, and container stacking information within a terminal's loading and unloading area to form a semantic vector for that area. Upon receiving a text instruction to "prioritize unloading for the cargo ship at berth A03," the system associates the instruction with the semantics of that area. Then, through a multimodal network and combined with real-time monitoring footage, it ultimately generates specific collaborative strategies such as "schedule AGV #07 to accelerate to 8 m / s and plan an avoidance path."
[0043] By transforming high-level abstract instructions into executable action sequences that fit specific scenarios and real-time situations through a multimodal fusion reasoning mechanism, a key leap has been achieved from "human understanding of tasks" to "system autonomous collaboration," significantly improving the intelligence level and decision-making accuracy of cross-device collaboration in complex environments.
[0044] Step S40: Based on the mapping relationship, perform parallel spatial block rendering on the visualization content corresponding to the collaborative control strategy, and synchronously update the status information of each associated device to generate preliminary visualization output.
[0045] Specifically, such as Figure 6 As shown, step S40 specifically includes: Step S41: Based on the mapping relationship, the collaborative control strategy of each region is assigned to different parallel rendering calculators; Step S42: Each rendering calculator independently renders the visual content corresponding to the semantic collaboration area it is responsible for, wherein the visual content includes a quality heatmap, a device status panel, or a dynamic trajectory line. Step S43: During the parallel rendering process, the latest status information of each associated device is updated and broadcast in real time through a distributed state synchronization mechanism to generate preliminary visualization output.
[0046] For example, in power grid fault diagnosis, the system assigns diagnostic strategies for different areas such as transformers and transmission lines to multiple rendering calculators based on mapping relationships. Each rendering calculator generates temperature heatmaps and equipment status panels for its assigned area in parallel. At the same time, the system updates the current and voltage data of all equipment in real time through a synchronization mechanism, and finally integrates them into a panoramic diagnostic view.
[0047] This parallel rendering and synchronization mechanism liberates the visualization generation of large-scale complex scenes from the serial bottleneck by performing block-based processing and distributed computing on the space. While ensuring global state consistency, it achieves an order-of-magnitude improvement in rendering efficiency, providing a high refresh rate and low latency visual interaction foundation for real-time monitoring and decision-making.
[0048] Step S50: Perform regional-level quality reflection and evaluation on the preliminary visualization output to determine whether the output results of each semantic collaboration region meet the preset quality standards.
[0049] Specifically, such as Figure 7 As shown, step S50 specifically includes: Step S51: Set a region-based quality assessment function. The quality assessment function sets differentiated thresholds for the accuracy, real-time performance and completeness of the visualization output for different business scenarios. Step S52: The visualization output is segmented according to semantic cooperation regions, and the corresponding quality assessment function is called to perform region-by-region quantitative scoring. Step S53: Compare the scoring results of each region with the preset quality standards, output the regional quality compliance judgment results, and record the identity of the non-compliant regions and the specific non-compliant items.
[0050] For example, in automotive welding scenarios, the system sets differentiated quality standards for heat maps of different workstations such as car doors and side panels: the accuracy threshold for door welds is ±0.1mm, while the allowable accuracy for side panel areas is ±0.2mm. The system independently scores the output of each area. When it detects that the accuracy of the heat map at a certain car door workstation exceeds the standard due to robot vibration, it marks that area as substandard and records the "positioning deviation" item.
[0051] This mechanism enables refined measurement of collaborative output results and rapid anomaly localization through differentiated regional-level quality reflection. It transforms the traditional coarse evaluation based on overall or fixed thresholds into precise closed-loop control oriented towards multiple objectives and constraints, providing key decision-making basis for subsequent targeted repair and continuous system optimization.
[0052] Step S60: If all areas meet the standards, the final generated collaborative control strategy and visualization content will be adapted and converted into control commands or status mapping information that conform to the target device communication protocol, and distributed to the corresponding physical devices for execution or display.
[0053] Specifically, such as Figure 8 As shown, step S60 specifically includes: Step S61: After determining that all regions meet the standards, integrate the final collaborative control strategies and visualization content determined for all semantic collaborative regions; Step S62: Through the protocol adapter, the integrated policies and content are converted into control command streams or status mapping data packets that conform to a specific communication protocol format according to the brand and model of the target device. Step S63: Distribute the converted instructions or data packets to each physical device through the corresponding industrial network or IoT link to drive and execute corresponding actions or to display them visually on the terminal interface.
[0054] For example, in collaborative production line operations, once the control strategies (such as the robotic arm welding trajectory) and visualization content (such as the quality monitoring panel) of all areas have passed the quality assessment, the system automatically converts these unified instructions into the PROFINET protocol that conforms to the PLC, the specific interface instructions of the specified robot, and the WebSocket data stream of the monitoring screen through the protocol adapter, and then sends them out for execution simultaneously.
[0055] The protocol's adaptive conversion and linkage output mechanism builds a seamless bridge from unified intelligent strategies to execution by heterogeneous physical devices, enabling "one-time generation and precise distribution" across brands and protocols. This completely eliminates the pain points of high system integration costs and large collaboration delays caused by protocol fragmentation, ensuring the final implementation of the intelligent decision-making closed loop.
[0056] Step S70: If there are substandard areas, start the local regeneration protocol for the identified substandard areas, recalibrate the perception data associated with the substandard areas, and update the collaborative control strategy generation and visualization rendering for the substandard areas.
[0057] Specifically, such as Figure 9 As shown, step S70 specifically includes: Step S71: Based on the recorded identity identifier of the non-compliant area, initiate the local regeneration protocol and suspend the strategy and rendering update process for other compliant areas outside of that area; Step S72: For the non-compliant area, trigger the sensing device associated with the non-compliant area to re-acquire data or perform filtering, completion and calibration operations on the existing sensing data; Step S73: Based on the recalibrated perception data, only re-execute the collaborative control strategy generation and visualization rendering process for the non-compliant areas to generate updated local results, and integrate the updated local results into the overall output.
[0058] For example, in power grid monitoring, when the system detects an abnormal peak in the temperature heat map of a certain substation area, it initiates a local regeneration protocol: only the sensor data of that station is reacquired and calibrated, and the risk assessment strategy and heat map of that area are regenerated. Then, it is seamlessly updated to the panoramic monitoring interface, while the data processing and display of other normal areas are completely unaffected.
[0059] This targeted regeneration mechanism breaks through the traditional extensive mode of "one point of anomaly, global recalculation", and realizes the precise location of the problem and closed-loop repair. While ensuring the continuous availability of the overall system output, it greatly reduces the computing power and communication overhead caused by anomaly handling, and greatly improves resource utilization efficiency and system robustness.
[0060] In addition, such as Figure 10 As shown, the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: Step S81: When a new device connects to the network, actively broadcast the new device's device type, spatial awareness capabilities, supported communication protocols, and computing resource margin. Step S82: After receiving broadcast information through the central scheduling node or adjacent devices, register the capability information to the unified device capability map, and assign an initial logical cooperation group to the new device based on the spatial topology relationship. Step S83: Dynamically adjust data fusion weights, allocate rendering tasks, or select protocol adaptation paths based on the new device's capability map.
[0061] For example, when a new high-precision 3D vision sensor is connected to the smart production line, it actively broadcasts its own capability information. After the system registers it in the equipment capability map, it is assigned to the "body inspection" collaboration group according to its spatial location, and its weight is automatically increased in the subsequent data fusion. At the same time, high-precision point cloud rendering tasks are assigned to it.
[0062] This dynamic discovery and registration mechanism enables "plug-and-play" integration of new devices. The system can automatically identify and utilize their new capabilities without manual downtime configuration, significantly improving the expansion flexibility and overall resource utilization of the device cluster. This allows the system to adaptively absorb technological advancements and continuously optimize collaborative performance.
[0063] In addition, such as Figure 11 As shown, the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: Step S91: Record the boundaries of the semantic collaboration region for each dynamic division, the typical collaborative control strategy patterns generated, and various cases of substandard quality and corresponding solutions. Step S92: Based on historical records, summarize the optimal partitioning strategies and inference patterns under different spatial topologies, business loads and external events using machine learning methods, and store them in the knowledge base; Step S93: When similar scene features are detected again, the best historical strategy is retrieved from the knowledge base and recommended as the initial parameter to accelerate the strategy generation process.
[0064] For example, the system records the special zoning strategy and priority inspection plan for the area of tilted power towers in the power grid during each typhoon. Through learning, the system summarizes the correspondence between the strategy of "strong wind weather + specific geographical area" and "strengthening local monitoring and reducing the detection frequency of surrounding areas". When similar meteorological and terrain features are detected again, the system automatically recommends and applies the historical optimization strategy and quickly generates inspection instructions.
[0065] This continuous learning and knowledge recommendation mechanism enables the system to accumulate and reuse historical best experiences, transforming the decision-making for complex scenarios from "recalculating each time" to "rapid initialization and fine-tuning based on experience". This significantly improves the system's response speed and decision quality in repetitive or predictable scenarios, and achieves the autonomous evolution of the system's intelligence level.
[0066] This invention achieves millisecond-level real-time response, cross-protocol uninterrupted self-adaptation, and precise localized anomaly repair through a closed-loop mechanism of dynamic partitioning, multimodal fusion, and quality reflection. Thus, while ensuring the safety threshold, it systematically solves the three core pain points of existing technologies: excessive response latency, insufficient system flexibility, and waste of computing resources.
[0067] The technical effects that this invention can bring are as follows: (1) Achieve millisecond-level real-time accurate collaboration: Through dynamic partitioning and multimodal fusion reasoning, the end-to-end delay of key control commands is reduced to within the process safety threshold (e.g., <100ms), solving the response delay bottleneck in large-scale equipment collaboration.
[0068] (2) Construct a flexible and scalable protocol adaptive system: With the help of protocol adaptation and dynamic discovery mechanisms, support uninterrupted access and command coordination of cross-brand and cross-protocol devices, and shorten the time for system integration and protocol expansion from days in traditional solutions to minutes.
[0069] (3) Break through the resource bottleneck of "one point of abnormality, global recalculation": Based on quality reflection to trigger targeted regeneration, only the abnormal area is repaired in a closed loop, avoiding the waste of global computing power and greatly reducing the computing power and communication overhead caused by abnormality handling.
[0070] (4) Improve the level of intelligent decision-making in complex environments: Through multimodal semantic fusion and scenario-based reasoning, high-level instructions are automatically transformed into executable action sequences that fit the real-time situation, which significantly reduces the reliance on human experience and enhances the system's autonomous decision-making ability in dynamic scenarios.
[0071] (5) Forming a system capability of continuous evolution and autonomous optimization: With the help of knowledge base accumulation and machine learning induction, the system can reuse historical best strategies and achieve rapid initialization and decision recommendation in similar scenarios, promoting the intelligent evolution of the system from "repeated execution" to "experience learning".
[0072] Furthermore, such as Figure 12 and Figure 13 As shown, based on the above-mentioned cross-domain multimodal dynamic reasoning and generation method for linkage of spatial sensing devices, the present invention also provides a corresponding cross-domain multimodal dynamic reasoning and generation system for linkage of spatial sensing devices, wherein the cross-domain multimodal dynamic reasoning and generation system for linkage of spatial sensing devices includes: The device access and spatiotemporal alignment module 50 is used to receive real-time sensing data streams from multiple heterogeneous devices through the device access layer, and to perform weighted fusion and spatial alignment on the multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference. The semantic region dynamic partitioning module 60 is used to dynamically partition the physical space into one or more semantic collaboration regions based on the cross-domain spatial perception data stream, according to the real-time spatial topology relationship of the device cluster and the preset business constraints, and to establish a mapping relationship between each semantic collaboration region and the corresponding visualization generation region. The multimodal strategy generation module 70 is used to fuse parsed spatial semantics, text instructions and visual information for each semantic cooperation region, perform contextual enhancement and semantic bridging, and generate a collaborative control strategy for the regional device cluster. The visualization parallel rendering module 80 is used to perform spatial block parallel rendering of the visualization content corresponding to the collaborative control strategy according to the mapping relationship, and synchronously update the status information of each associated device to generate preliminary visualization output. The quality reflection and evaluation module 90 is used to perform regional-level quality reflection and evaluation on the preliminary visualization output, and to determine whether the output results of each semantic cooperation region meet the preset quality standards. The protocol adaptation and linkage output module 100 is used to adapt and convert the final generated collaborative control strategy and visualization content into control instructions or status mapping information that conform to the target device communication protocol if all areas meet the standards, and distribute them to the corresponding physical devices for execution or display. The local regeneration protocol execution module 110 is used to, if there are non-compliant areas, initiate the local regeneration protocol for the identified non-compliant areas, recalibrate the perception data associated with the non-compliant areas, and update the collaborative control strategy generation and visualization rendering for the non-compliant areas.
[0073] like Figure 14As shown in this embodiment of the invention, another embodiment of the cross-domain multimodal dynamic reasoning and generation system for spatial sensing device linkage is described. In this embodiment, the device access and spatiotemporal alignment module 50 includes: The sensing data stream acquisition unit 501 is used to acquire real-time sensing data streams reported by multiple heterogeneous devices in at least one scenario in industrial manufacturing, port logistics or power grid systems, wherein the heterogeneous devices include industrial robots, AGVs, sensors and cameras. The sensing data stream weighted fusion unit 502 is used to perform weighted fusion on multiple real-time sensing data streams based on preset equipment weight factors and data quality evaluation indicators to obtain weighted fused data. The equipment weight factors are dynamically adjusted based on equipment accuracy, historical reliability and current operating conditions. The spatial transformation and time synchronization unit 503 is used to perform spatial coordinate transformation and timestamp synchronization on the weighted and fused data based on a unified global coordinate system and time server, so as to generate a cross-domain spatial perception data stream with a unified spatiotemporal reference and direct access.
[0074] In this embodiment, the semantic region dynamic partitioning module 60 includes: The dynamic spatial topology graph construction unit 601 is used to parse the cross-domain spatial sensing data stream and construct and update the dynamic spatial topology graph of the device cluster in real time. The semantic cooperation region information confirmation unit 602 is used to combine preset business constraint rules, wherein the business constraint rules include process steps, safety radius or task coupling degree, and determine the current optimal number and boundary of semantic cooperation regions through a dynamic K-value partitioning algorithm; The mapping relationship establishment unit 603 is used to establish a mapping relationship between the physical spatial range of each semantic cooperation region and the corresponding logical generation region in the visualization rendering engine, wherein the mapping relationship is implemented through a coordinate transformation matrix or a grid index.
[0075] In this embodiment, the multimodal strategy generation module 70 includes: The regional spatial semantic vector generation unit 701 is used to extract the spatial layout features, device state set and environmental context information of each semantic cooperation region for each semantic cooperation region, and to form a regional spatial semantic vector. The regional spatial semantic vector association unit 702 is used to receive and parse natural language text instructions or predefined task scripts, and associate and enrich the natural language text instructions or predefined task scripts with the regional spatial semantic vectors. The collaborative control strategy generation unit 703 is used to infer and generate a collaborative control strategy for all devices in the semantic cooperation area by fusing enhanced text instructions, regional spatial semantic vectors and image features from the visual sensor through a multimodal semantic bridging network.
[0076] In this embodiment, the visualization parallel rendering module 80 includes: The collaborative control strategy allocation unit 801 is used to allocate the collaborative control strategy of each region to different parallel rendering calculators according to the mapping relationship. The visualization content rendering unit 802 is used for each of the rendering calculators to independently perform graphic rendering on the visualization content corresponding to the semantic collaboration area they are responsible for, wherein the visualization content includes a quality heatmap, a device status panel, or a dynamic trajectory line. The preliminary visualization output generation unit 803 is used to generate preliminary visualization output by updating and broadcasting the latest status information of each associated device in real time through a distributed state synchronization mechanism during parallel rendering.
[0077] In this embodiment, the quality reflection and evaluation module 90 includes: The quality assessment function setting unit 901 is used to set a region-based quality assessment function, which sets differentiated thresholds for the accuracy, real-time performance and completeness of the visualization output for different business scenarios. The region segmentation and quantification scoring unit 902 is used to segment the visualization output according to the semantic cooperation region and call the corresponding quality evaluation function to perform region-by-region quantification scoring. The result comparison and recording unit 903 is used to compare the scoring results of each region with the preset quality standards, output the regional quality compliance judgment results, and record the identity of the non-compliant regions and the specific non-compliant items.
[0078] In this embodiment, the protocol adaptation and linkage output module 100 includes: The strategy and content integration unit 1001 is used to integrate the final determined collaborative control strategy and visualization content of all semantic collaborative areas after it is determined that all areas have met the standards. The policy and content conversion unit 1002 is used to convert the integrated policy and content into a control command stream or status mapping data packet that conforms to a specific communication protocol format, according to the brand and model of the target device, through a protocol adapter. The instruction or data packet distribution unit 1003 is used to distribute the converted instructions or data packets to various physical devices through the corresponding industrial network or Internet of Things link to drive and execute corresponding actions or to display them visually on the terminal interface.
[0079] In this embodiment, the local regeneration protocol execution module 110 includes: The local regeneration protocol initiation unit 1101 is used to initiate the local regeneration protocol based on the recorded identity identifier of the non-compliant area, and to suspend the policy and rendering update process for other compliant areas outside the non-compliant area. The data processing unit 1102 for non-compliant areas is used to trigger the sensing devices associated with the non-compliant areas to re-acquire data or to perform filtering, completion and calibration operations on the existing sensing data for the non-compliant areas. The local result generation and integration unit 1103 is used to re-execute the collaborative control strategy generation and visualization rendering process for the non-compliant area based on the recalibrated perception data, generate updated local results, and integrate the updated local results into the overall output.
[0080] Furthermore, such as Figure 15 As shown, based on the above-mentioned cross-domain multimodal dynamic reasoning and generation method and system for linkage of spatial perception devices, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 15 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0081] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a cross-domain multimodal dynamic reasoning and generation program 40 for spatial sensing device linkage. This cross-domain multimodal dynamic reasoning and generation program 40 for spatial sensing device linkage can be executed by the processor 10, thereby implementing the cross-domain multimodal dynamic reasoning and generation method for spatial sensing device linkage in this application.
[0082] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices.
[0083] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminal's processor 10, memory 20, and display 30 communicate with each other via a system bus.
[0084] In one embodiment, when the processor 10 executes the cross-domain multimodal dynamic reasoning and generation program 40 for spatial sensing device linkage stored in the memory 20, the following steps are performed: Through the device access layer, real-time sensing data streams from multiple heterogeneous devices are received, and the multiple real-time sensing data streams are weighted, fused, and spatially aligned to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference. Based on the cross-domain spatial perception data stream, according to the real-time spatial topology relationship of the device cluster and the preset business constraints, the physical space is dynamically divided into one or more semantic collaboration regions, and a mapping relationship is established between each semantic collaboration region and the corresponding visualization generation region. For each semantic collaboration region, the parsed spatial semantics, text instructions, and visual information are integrated to perform contextual enhancement and semantic bridging, generating a collaborative control strategy for the regional device cluster. Based on the mapping relationship, the visualization content corresponding to the collaborative control strategy is rendered in parallel spatial blocks, and the status information of each associated device is updated synchronously to generate preliminary visualization output. The preliminary visualization output is subjected to regional-level quality reflection and evaluation to determine whether the output results of each semantic cooperation region meet the preset quality standards. If all areas meet the standards, the final generated collaborative control strategy and visualization content will be adapted and converted into control commands or status mapping information that conform to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display. If there are non-compliant areas, a local regeneration protocol is initiated for the identified non-compliant areas to recalibrate the perception data associated with the non-compliant areas, and to update the collaborative control strategy generation and visualization rendering for the non-compliant areas.
[0085] Specifically, the process of receiving real-time sensing data streams from multiple heterogeneous devices through the device access layer, and performing weighted fusion and spatial alignment on these multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference, includes: Acquire real-time sensing data streams reported by multiple heterogeneous devices in at least one scenario in industrial manufacturing, port logistics, or power grid systems, wherein the heterogeneous devices include industrial robots, AGVs, sensors, and cameras; Based on preset equipment weighting factors and data quality evaluation indicators, multiple real-time sensing data streams are weighted and fused to obtain weighted fused data. The equipment weighting factors are dynamically adjusted based on equipment accuracy, historical reliability and current operating conditions. Based on a unified global coordinate system and time server, the weighted and fused data is transformed into spatial coordinates and synchronized with timestamps to generate a cross-domain spatial perception data stream with a unified spatiotemporal reference that can be directly accessed.
[0086] Specifically, based on the cross-domain spatial perception data stream, and according to the real-time spatial topology of the device cluster and preset business constraints, the physical space is dynamically divided into one or more semantic collaboration regions, and a mapping relationship is established between each semantic collaboration region and its corresponding visualization generation region. The cross-domain spatial sensing data stream is analyzed to construct and update the dynamic spatial topology diagram of the device cluster in real time. Combining preset business constraint rules, including process steps, safety radius, or task coupling degree, the current optimal number and boundary of semantic cooperation regions are determined through a dynamic K-value partitioning algorithm; Establish a mapping relationship between the physical spatial extent of each semantic collaboration region and the corresponding logical generation region in the visualization rendering engine, wherein the mapping relationship is implemented through a coordinate transformation matrix or a grid index.
[0087] Specifically, for each semantic cooperation region, the parsed spatial semantics, text commands, and visual information are fused to perform contextual enhancement and semantic bridging, generating a collaborative control strategy for the regional device cluster. This includes: For each semantic collaboration region, the spatial layout features, device state set, and environmental context information of each semantic collaboration region are extracted to form a region spatial semantic vector. Receive and parse natural language text instructions or predefined task scripts, and associate and enrich the natural language text instructions or predefined task scripts with the regional spatial semantic vector; By using a multimodal semantic bridging network, enhanced text instructions, regional spatial semantic vectors, and image features from visual sensors are fused together to infer and generate a collaborative control strategy for all devices within the semantic cooperation area.
[0088] Specifically, the step of performing parallel spatial block rendering of the visualization content corresponding to the collaborative control strategy based on the mapping relationship, and synchronously updating the status information of each associated device to generate preliminary visualization output includes: Based on the mapping relationship, the collaborative control strategy for each region is assigned to different parallel rendering calculators; Each of the rendering calculators independently renders the visual content corresponding to the semantic collaboration area it is responsible for, wherein the visual content includes a quality heatmap, a device status panel, or a dynamic trajectory line. During parallel rendering, a distributed state synchronization mechanism is used to update and broadcast the latest status information of each associated device in real time, generating preliminary visualization output.
[0089] Specifically, the step of performing regional-level quality reflection and evaluation on the preliminary visualization output to determine whether the output results of each semantic cooperation region meet the preset quality standards includes: Set a region-based quality assessment function, which sets differentiated thresholds for the accuracy, real-time performance, and completeness of the visualization output for different business scenarios; The visualization output is segmented according to semantic cooperation regions, and the corresponding quality assessment function is called to perform region-by-region quantitative scoring. The scoring results of each region are compared with the preset quality standards, and the regional quality compliance judgment results are output. The identification of the regions that fail to meet the standards and the specific non-compliance items are recorded.
[0090] If all regions meet the standards, the final generated collaborative control strategy and visualization content will be adapted and converted into control commands or status mapping information that conform to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display. Specifically, this includes: Once all regions have been determined to meet the standards, the final collaborative control strategies and visualization content determined for all semantic collaborative regions are integrated. Through the protocol adapter, the integrated policies and content are converted into control command streams or status mapping data packets that conform to a specific communication protocol format, based on the brand and model of the target device. The converted instructions or data packets are distributed to various physical devices through the corresponding industrial network or IoT link to drive and execute corresponding actions or to be visualized on the terminal interface.
[0091] Specifically, if there are substandard areas, a local regeneration protocol is initiated for the identified substandard areas to recalibrate the sensing data associated with the substandard areas, and the collaborative control strategy generation and visualization rendering for the substandard areas are updated, including: Based on the recorded identity of the non-compliant area, a local regeneration protocol is initiated, suspending the policy and rendering update process for other compliant areas outside of that area; For the non-compliant areas, the associated sensing devices are triggered to re-acquire data or to perform filtering, completion, and calibration operations on the existing sensing data. Based on the recalibrated perception data, the collaborative control strategy generation and visualization rendering process for the non-compliant areas is re-executed to generate updated local results, and the updated local results are integrated into the overall output.
[0092] The cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: When a new device connects to the network, it actively broadcasts the device type, spatial awareness capabilities, supported communication protocols, and computing resource margin of the new device. After receiving broadcast information through the central scheduling node or adjacent devices, the capability information is registered to a unified device capability map, and an initial logical cooperation group is assigned to the new device based on spatial topology. Based on the capability map of the new equipment, dynamically adjust the data fusion weights, allocate rendering tasks, or select protocol adaptation paths.
[0093] The cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices further includes: Record the boundaries of the semantic collaboration region for each dynamic division, the typical collaborative control strategy patterns generated, and various cases of substandard quality and corresponding solutions; Based on historical records, the optimal partitioning strategies and inference patterns under different spatial topologies, business loads, and external events are summarized using machine learning methods and stored in a knowledge base; When similar scene characteristics are detected again, the best historical strategy is retrieved from the knowledge base and recommended as the initial parameter to accelerate the strategy generation process.
[0094] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a cross-domain multimodal dynamic reasoning and generation program for linkage with spatial sensing devices, and the cross-domain multimodal dynamic reasoning and generation program for linkage with spatial sensing devices, when executed by a processor, implements the steps of the cross-domain multimodal dynamic reasoning and generation method for linkage with spatial sensing devices as described above.
[0095] In summary, this invention provides a method, system, terminal, and computer-readable storage medium for cross-domain multimodal dynamic reasoning and generation of spatial sensing devices. The method includes: receiving real-time sensing data streams from multiple heterogeneous devices through a device access layer, and performing weighted fusion and spatial alignment on the multiple real-time sensing data streams to generate a cross-domain spatial sensing data stream with a unified spatiotemporal reference; based on the cross-domain spatial sensing data stream, dynamically dividing the physical space into one or more semantic collaboration regions according to the real-time spatial topology of the device cluster and preset business constraints, and establishing a mapping relationship between each semantic collaboration region and its corresponding visualization generation region; for each semantic collaboration region, fusing the parsed spatial semantics, text instructions, and visual information, performing contextual enhancement and semantic bridging, and generating a region-oriented spatial sensing data stream. The invention establishes a collaborative control strategy for the cluster. Based on the mapping relationship, it performs parallel spatial block rendering of the visualization content corresponding to the collaborative control strategy and synchronously updates the status information of each associated device to generate preliminary visualization output. It then performs regional-level quality reflection and evaluation on the preliminary visualization output to determine whether the output results of each semantic collaboration region meet preset quality standards. If all regions meet the standards, the final generated collaborative control strategy and visualization content are adapted and converted into control commands or status mapping information conforming to the target device's communication protocol, and distributed to the corresponding physical devices for execution or display. If there are non-compliant regions, a local regeneration protocol is initiated for the identified non-compliant regions to recalibrate the perception data associated with the non-compliant regions and update the collaborative control strategy generation and visualization rendering for the non-compliant regions. This invention achieves millisecond-level real-time response, cross-protocol downtime self-adaptation, and precise localized anomaly repair through a closed-loop mechanism of dynamic partitioning, multimodal fusion, and quality reflection.
[0096] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0097] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.
[0098] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A cross-domain multi-modal dynamic inference and generation method oriented to space-aware device linkage, characterized in that, The method comprises the following steps: Through the device access layer, real-time perception data streams of multiple heterogeneous devices are received, and multiple real-time perception data streams are weightedly fused and spatially aligned to generate cross-domain spatial perception data streams with unified space-time reference; Based on the cross-domain spatial perception data streams, according to the real-time spatial topological relationship of the device cluster and the preset business constraint, the physical space is dynamically divided into one or more semantic cooperation areas, and the mapping relationship between each semantic cooperation area and the corresponding visual generation area is established; For each semantic cooperation area, the spatial semantics, text instructions and visual information after fusion analysis are subjected to situational reinforcement and semantic bridging to generate a cooperative control strategy for the regional device cluster; According to the mapping relationship, the visual content corresponding to the cooperative control strategy is rendered in parallel in the spatial block, and the state information of each associated device is updated synchronously to generate a preliminary visual output; The preliminary visual output is subjected to regional quality reflection and evaluation to determine whether the output results of each semantic cooperation area meet the preset quality standard; If all areas meet the standard, the finally generated cooperative control strategy and visual content are adapted and converted into control instructions or state mapping information conforming to the target device communication protocol, and are distributed to the corresponding physical devices for execution or display; If there are non-compliant areas, a local regeneration protocol is started for the identified non-compliant areas, the perception data associated with the non-compliant areas is recalibrated, and the cooperative control strategy generation and visual rendering for the non-compliant areas are updated.
2. The space-aware device collaboration oriented cross-domain multi-modal dynamic inference and generation method according to claim 1, characterized in that, The method comprises the following steps: Obtain real-time perception data streams reported by multiple heterogeneous devices in at least one of industrial manufacturing, port logistics or power grid systems, wherein the heterogeneous devices include industrial robots, AGVs, sensors and cameras; According to the preset device weight factor and data quality evaluation index, multiple real-time perception data streams are weightedly fused to obtain weightedly fused data, wherein the device weight factor is dynamically adjusted based on device accuracy, historical reliability and current working conditions; Based on a unified global coordinate system and a time server, the weightedly fused data is subjected to spatial coordinate transformation and time stamp synchronization to generate cross-domain spatial perception data streams with unified space-time reference and direct calling.
3. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, wherein, The method comprises the following steps: The cross-domain spatial perception data streams are analyzed to construct and update the dynamic spatial topological relationship diagram of the device cluster in real time; In combination with a preset business constraint rule, the current optimal semantic collaboration region quantity and boundary are determined by a dynamic K-value partition algorithm, wherein the business constraint rule includes a process procedure, a safety radius, or a task coupling degree; A mapping relationship between a physical space range of each semantic collaboration region and a corresponding logical generated region in a visual rendering engine is established, wherein the mapping relationship is realized by a coordinate conversion matrix or a grid index.
4. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, wherein, For each semantic collaboration region, the parsed spatial semantics, text instructions, and visual information are fused, contextualized reinforcement and semantic bridging are performed, and a collaborative control strategy for a regional device cluster is generated, specifically including: For each semantic collaboration region, spatial layout features, device state sets, and environmental context information of each semantic collaboration region are extracted to form a regional spatial semantic vector; Natural language text instructions or predefined task scripts are received and parsed, and the natural language text instructions or the predefined task scripts are associated and enriched with the regional spatial semantic vector; Through a multi-modal semantic bridging network, the reinforced text instructions, the regional spatial semantic vector, and image features from a visual sensor are fused to infer and generate a collaborative control strategy for all devices in the semantic collaboration region.
5. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, wherein, According to the mapping relationship, the visual content corresponding to the collaborative control strategy is spatially block-parallel rendered, and the state information of each associated device is synchronously updated to generate a preliminary visual output, specifically including: According to the mapping relationship, the collaborative control strategies of each region are distributed to different parallel rendering calculators; Each rendering calculator independently performs graphical rendering on the visual content corresponding to the semantic collaboration region it is responsible for, wherein the visual content includes a quality heat map, a device state panel, or a dynamic trajectory line; During parallel rendering, the latest state information of each associated device is updated and broadcast in real time through a distributed state synchronization mechanism to generate a preliminary visual output.
6. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, wherein, The preliminary visual output is regionally quality-reflected and evaluated to determine whether the output results of each semantic collaboration region meet the preset quality standards, specifically including: A regional quality evaluation function is set, which sets differential thresholds for the accuracy, real-time performance, and completeness of the visual output for different business scenarios; The visual output is segmented by semantic collaboration region, and the corresponding quality evaluation function is called for regional quantitative scoring; The scoring results of each region are compared with the preset quality standards, and a regional quality compliance determination result is output, and the identity of the non-compliant region and the specific non-compliant items are recorded.
7. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 6, characterized in that, If all regions meet the standards, the final generated collaborative control strategy and visual content are adapted and converted into control instructions or state mapping information that meet the target device communication protocol, and are distributed to the corresponding physical devices for execution or display, specifically including: After determining that all regions meet the standards, all semantic collaboration regions finally determine the collaborative control strategy and visual content; The integrated strategy and content are converted into control instruction streams or state mapping data packets conforming to specific communication protocol formats through a protocol adapter according to the brand and model of the target device; The converted instructions or data packets are distributed to physical devices through corresponding industrial networks or Internet of Things links to drive and execute corresponding actions or perform visual display on terminal interfaces.
8. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 6, characterized in that, If there are substandard areas, a local regeneration protocol is started for the identified substandard areas, the perception data associated with the substandard areas are recalibrated, and the collaborative control strategy generation and visual rendering for the substandard areas are updated, specifically including: According to the recorded identity of the substandard area, a local regeneration protocol is started, and the strategy and rendering update process for other standard areas outside the area are suspended; For the substandard area, the sensing devices associated with the substandard area are triggered to perform data reacquisition or filtering, completion and calibration operations on existing perception data; Based on the recalibrated perception data, only the collaborative control strategy generation and visual rendering process for the substandard area is re-executed, an updated local result is generated, and the updated local result is integrated into the overall output.
9. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, characterized in that, The spatial perception device-oriented cross-domain multi-modal dynamic reasoning and generation method further includes: When a new device accesses the network, the device type, spatial perception capability, supported communication protocol and computing resource margin of the new device are actively broadcasted; After receiving the broadcast information through the central scheduling node or adjacent devices, the capability information is registered to a unified device capability map, and an initial logical collaboration group is allocated for the new device based on the spatial topology relationship; The data fusion weight is dynamically adjusted, the rendering task is allocated or the protocol adaptation path is selected according to the capability map of the new device.
10. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation method according to claim 1, characterized in that, The spatial perception device-oriented cross-domain multi-modal dynamic reasoning and generation method further includes: The semantic collaboration area boundary of each dynamic division, the generated typical collaborative control strategy mode, and various quality substandard cases and corresponding solutions are recorded; Based on the historical records, the optimal partitioning strategy and reasoning mode under different spatial topologies, business loads and external events are induced by machine learning methods and stored in a knowledge base; When similar scenario features are detected again, the historical optimal strategy is retrieved and recommended from the knowledge base as initial parameters to accelerate the strategy generation process.
11. A cross-domain multi-modal dynamic inference and generation system oriented to space-aware device collaboration, characterized in that, The spatial perception device-oriented cross-domain multi-modal dynamic reasoning and generation system includes: A device access and space-time alignment module for receiving real-time perception data streams of multiple heterogeneous devices through a device access layer and performing weighted fusion and space-time alignment on the multiple real-time perception data streams to generate a cross-domain spatial perception data stream with a unified space-time reference; A semantic area dynamic division module for dynamically dividing a physical space into one or more semantic collaboration areas based on the cross-domain spatial perception data stream according to the real-time spatial topology relationship of the device cluster and the preset business constraints, and establishing a mapping relationship between each semantic collaboration area and the corresponding visual generation area; A multi-modal strategy generation module is configured to, for each semantic cooperation region, fuse the parsed spatial semantics, text instructions and visual information, perform contextual reinforcement and semantic bridging, and generate a cooperative control strategy for a regional device cluster; A visual parallel rendering module is configured to, according to the mapping relationship, perform spatial block parallel rendering on visual content corresponding to the cooperative control strategy, and synchronously update state information of each associated device to generate preliminary visual output; A quality reflection and evaluation module is configured to perform regional-level quality reflection and evaluation on the preliminary visual output, and determine whether the output result of each semantic cooperation region meets preset quality standards; A protocol adaptation and linkage output module is configured to, if all regions meet the standards, adapt and convert the finally generated cooperative control strategy and visual content into control instructions or state mapping information conforming to a target device communication protocol, and distribute the control instructions or state mapping information to corresponding physical devices for execution or display; A local regeneration protocol execution module is configured to, if there is a non-compliant region, start a local regeneration protocol for the identified non-compliant region, recalibrate perception data associated with the non-compliant region, and update the cooperative control strategy generation and visual rendering for the non-compliant region.
12. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system according to claim 11, wherein, The device access and space-time alignment module includes: A perception data stream acquisition unit is configured to acquire real-time perception data streams reported by multiple heterogeneous devices in at least one of an industrial manufacturing, port logistics or power grid system, wherein the heterogeneous devices include industrial robots, AGVs, sensors and cameras; A perception data stream weighted fusion unit is configured to perform weighted fusion on multiple real-time perception data streams according to preset device weight factors and data quality evaluation indexes, to obtain weighted fusion data, wherein the device weight factors are dynamically adjusted based on device accuracy, historical reliability and current working conditions; A space transformation and time synchronization unit is configured to perform space coordinate transformation and time stamp synchronization on the weighted fusion data based on a unified global coordinate system and a time server, to generate cross-domain spatial perception data streams with a unified space-time reference and direct invocation.
13. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 11, wherein, The semantic region dynamic division module includes: A dynamic space topology relationship graph construction unit is configured to analyze the cross-domain spatial perception data streams, and construct and update a dynamic space topology relationship graph of a device cluster in real time; A semantic cooperation region information confirmation unit is configured to combine preset business constraint rules, wherein the business constraint rules include process procedures, safety radii or task coupling degrees, and determine the current optimal number and boundaries of semantic cooperation regions through a dynamic K-value partitioning algorithm; A mapping relationship establishment unit is configured to establish a mapping relationship between a physical space range of each semantic cooperation region and a corresponding logical generation region in a visual rendering engine, wherein the mapping relationship is realized through a coordinate conversion matrix or a grid index.
14. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 11, wherein, The multi-modal strategy generation module includes: A regional space semantic vector generation unit is configured to, for each semantic cooperation region, extract spatial layout features, device state sets and environmental context information of each semantic cooperation region to form a regional space semantic vector; The regional space semantic vector association unit is configured to receive and parse a natural language text instruction or a predefined task script, and associate and enrich the natural language text instruction or the predefined task script with the regional space semantic vector; The collaborative control strategy generation unit is configured to infer and generate a collaborative control strategy for all devices in the semantic collaboration region by fusing the reinforced text instruction, the regional space semantic vector, and image features from a visual sensor through a multi-modal semantic bridging network.
15. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 11, wherein, The visualization parallel rendering module includes: The collaborative control strategy distribution unit is configured to distribute the collaborative control strategies of the regions to different parallel rendering calculators according to the mapping relationship; The visualization content rendering unit is configured to independently perform graphical rendering on the visualization content corresponding to the responsible semantic collaboration region by each rendering calculator, where the visualization content includes a quality heat map, a device state panel, or a dynamic trajectory line; The preliminary visualization output generation unit is configured to update and broadcast the latest state information of each associated device in real time through a distributed state synchronization mechanism during the parallel rendering process, and generate a preliminary visualization output.
16. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 11, wherein, The quality reflection and evaluation module includes: The quality evaluation function setting unit is configured to set a regional-based quality evaluation function, which sets differential thresholds for the accuracy, real-time performance, and completeness of the visualization output for different business scenarios; The regional segmentation and quantitative scoring unit is configured to segment the visualization output according to the semantic collaboration region, and call the corresponding quality evaluation function to perform regional quantitative scoring; The result comparison and recording unit is configured to compare the scoring results of the regions with preset quality standards, output a regional-level quality compliance determination result, and record the identity of the non-compliant region and specific non-compliant items.
17. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 16, wherein, The protocol adaptation and linkage output module includes: The strategy and content integration unit is configured to integrate the final determined collaborative control strategies and visualization content of all semantic collaboration regions when it is determined that all regions are compliant; The strategy and content conversion unit is configured to convert the integrated strategies and content into control instruction streams or state mapping data packets in a format conforming to a specific communication protocol through a protocol adapter according to the brand and model of the target device; The instruction or data packet distribution unit is configured to distribute the converted instructions or data packets to each physical device through a corresponding industrial network or Internet of Things link to drive and execute corresponding actions or perform visualization display on a terminal interface.
18. The space-aware device-orchestrated cross-domain multi-modal dynamic inference and generation system of claim 16, wherein, The local regeneration protocol execution module includes: The local regeneration protocol start unit is configured to start a local regeneration protocol according to the recorded non-compliant region identity, and suspend the strategy and rendering update process for other compliant regions outside the region; The non-compliant region data processing unit is configured to trigger the associated sensing device of the non-compliant region to perform data reacquisition or filtering, completion, and calibration operations on existing perception data for the non-compliant region. A local result generation and integration unit is configured to, based on the re-calibrated perception data, only re-execute the collaborative control strategy generation and visualization rendering procedure for the under-performing region, generate updated local results, and integrate the updated local results into the overall output.
19. A terminal, characterized by The terminal comprises a memory, a processor, and a spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation program stored on the memory and executable on the processor, and the spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation program, when executed by the processor, implements the steps of the spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation method according to any one of claims 1-10.
20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation program, and the spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation program, when executed by the processor, implements the steps of the spatial perception device linkage oriented cross-domain multi-modal dynamic inference and generation method according to any one of claims 1-10.