Visual prompting method and system for interaction between image and real scene
By using multimodal data fusion and augmented reality technology, high-precision identification and real-time feedback of invisible obstacles in complex environments are achieved, solving the problems of identification accuracy and feedback lag in existing technologies and providing personalized intelligent assistance effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-01
AI Technical Summary
Existing image recognition systems struggle to accurately identify deep or hidden obstacles in complex real-world environments, lack real-time feedback, and have insufficient adaptive capabilities, making it difficult to meet the demands for high-precision and robust real-time visualization and interactive decision-making.
By collecting multi-source sensing data, preprocessing and fusing multimodal image features, a spatially structured expression system is constructed for obstacle recognition and feedback. Combined with augmented reality technology, visual prompts are provided, and the feedback content is dynamically optimized.
It achieves high-precision identification of invisible obstacles such as tiny blood vessels and deep tissues, provides real-time feedback and personalized intelligent assistance, and improves operational safety and efficiency in complex scenarios.
Smart Images

Figure CN121959097A_ABST
Abstract
Description
Visual prompting methods and systems for interaction between images and real-world scenes Technical Field
[0001] This invention relates to the field of image interaction technology, and in particular to a visual prompting method and system for interaction between images and real-world scenes. Background Technology
[0002] Currently, image recognition and data fusion technologies are widely used in medical diagnosis, industrial inspection, and security patrol. Traditional systems rely on multiple types of sensors to collect images and environmental information, combining computer vision and intelligent analysis algorithms to achieve target detection and feature recognition in complex scenes. With the continuous advancement of artificial intelligence and augmented reality technologies, more and more application scenarios are beginning to focus on the fusion processing of multi-source sensor data, the dynamic management of historical task data, and the impact of environmental factors on recognition and interaction effects. Through methods such as partitioning analysis, feature extraction, and visual prompts, related systems are continuously improving the accuracy of obstacle detection and the efficiency of human-computer interaction, providing strong support for intelligent assistance and safe operation in complex environments.
[0003] For example, the invention patent with publication number CN114201048A discloses an intelligent human-computer interaction system for digital team building, including a team building machine and a group of mini-programs that can be run on the team building machine. The team building machine is characterized by including an image recording module, an audio recording module, a basic component module, a character input module, and a communication module. It can interact with the processor and memory in multiple ways through a touch screen, camera, microphone, speaker, 5G or 4G communication module, Bluetooth communication module, and WIFI communication module, and supports input methods such as external keyboards and virtual keyboards. The mini-program group includes task mini-programs and theme mini-programs, enabling the collection, display, editing, and cloud storage of team building image and audio data, meeting diverse data management and display needs during team building activities. Furthermore, the system supports remote collaboration, multi-platform push, and subsequent analysis and management of team building activity data, which helps improve the collaborative efficiency and data utilization value of corporate team building.
[0004] For example, the invention patent with publication number CN107422862A discloses a method for virtual image interaction in a virtual reality scene. This method includes selecting a song, a singing scene, a virtual image of a user avatar, and virtual images of chorus partners. The system changes the interactive behavior of the virtual images in real time according to preset control information or the user's live singing information. By using voice analysis, semantic recognition, facial expression capture, and motion tracking, the system adjusts the facial expressions, lip movements, and actions of the user avatar and chorus partners, achieving multimodal real-time interaction between the user and the virtual objects. At the same time, the system also supports collecting the user's facial expressions and actions through various interactive tools such as VR headsets and wearable motion-sensing devices. Combined with 3D modeling and panoramic video, it achieves high-fidelity generation of virtual scenes and virtual images, and can share the singing process data to social media, enhancing the human-computer interaction experience in a virtual reality environment.
[0005] While existing interaction methods can achieve multimodal data acquisition, virtual scene rendering, and basic interactive operations, they still have shortcomings in practical applications. The system has limited ability to deeply integrate and intelligently evaluate multi-source perception data in complex real-world environments, making it difficult to accurately identify deep or hidden obstacles. It also has weak adaptive processing capabilities for environmental changes and task diversity. Existing solutions have limitations in dynamic risk warnings, intelligent spatial partitioning analysis, and real-time information overlay and personalized feedback in augmented reality scenarios. They are unable to achieve high-precision, high-robust real-time visualization and interactive decision-making, and cannot meet the higher demands of operators for safety, intelligence, and efficiency in complex scenarios.
[0006] Therefore, in response to the above problems, there is an urgent need for a visual prompting method and system that enables interaction between images and real-world scenes. Summary of the Invention
[0007] To address the technical challenges of existing image recognition systems in effectively identifying hidden obstacles and providing accurate real-time feedback in deep structures and complex environments, this invention provides a visual prompting method and system for image-real-world scene interaction. The technical solution is as follows:
[0008] On the one hand, a visualization prompting method for interaction between images and real-world scenes is provided. This method includes: S1, collecting real-time perception data and acquiring historical task interaction data, and preprocessing the real-time perception data and historical task interaction data; S2, constructing a spatially structured expression system, dividing and layering complex image information, and assigning numbers; S3, performing multimodal image feature fusion extraction and analysis on the real-time perception data, improving image quality based on the results of multimodal image feature fusion extraction and analysis, and entering the obstacle recognition and feedback process; S4, performing hierarchical evaluation on historical task interaction data and real-time perception data based on priority optimization and feedback adjustment assessment, identifying and marking obstacles and providing intensity feedback based on the hierarchical evaluation results; S5, performing spatial accessibility and risk detection on historical task interaction data and real-time perception data, and performing visualization linkage based on the spatial accessibility and risk detection results.
[0009] Optionally, real-time sensing data is collected, and task interaction history data is obtained. The specific process for preprocessing the real-time sensing data and task interaction history data is as follows: Real-time sensing data is collected, including: ultrasound image data, scan image data, digital image data, sensor calibration data, location spatial coordinates, real-time environmental data, real-time task data, and obstacle spatial coordinates; task interaction history data is obtained, including: task execution records, task priority data, user operation log data, historical abnormal alarm counts, and historical task completion status data; through multimodal noise separation and adaptive residual correction algorithms, the ultrasound image data, scan image data, and digital image data are preprocessed. Preliminary denoising and signal correction are performed; data filtering and anomaly detection methods based on spatial coordinates and task labels are used to verify the temporal consistency of location spatial coordinates, obstacle spatial coordinates, and real-time task data, and to identify anomalies; multi-modal feature fusion and adaptive filtering algorithms are used to suppress noise and enhance details in ultrasonic image data, scanned image data, and digital image data; dynamic priority reconstruction and sliding window statistical algorithms are used to identify abnormal fluctuations and smooth trends in task execution records, task priority data, and historical anomaly alarm frequency time-series information; and distribution standardization and linear normalization algorithms are used to standardize and normalize real-time sensing data and historical task interaction data.
[0010] Optionally, the specific process of constructing a spatially structured representation system and partitioning and numbering complex image information is as follows: multimodal registration is completed using SIFT features, and ultrasound image data, scan image data, and digital image data are fused using Laplacian pyramid weights to obtain fused image data; for the fused image data, superpixel segmentation algorithm and deep learning region segmentation method are used to divide it into several image partitions with similar spatial structure and texture features, and each partition is assigned a unique number. The partition number will serve as the core index for spatial mapping in subsequent calculations and feedback, and the spatial coordinates, pixel set, and boundary information of the partition will be recorded simultaneously to perform image partitioning and spatial attribute structuring.
[0011] Optionally, the specific process of multimodal image feature fusion extraction and analysis of real-time perceived data is as follows: acquiring ultrasound image data, scan image data, digital image data, and fused image data; calculating the local contrast of the fused image data using Gaussian smoothing-based local difference calculation; extracting image sharpness from the fused image data using the Laplacian operator; obtaining the image noise level from the fused image data using the Wiener filtering algorithm and image residual method; obtaining the fused image enhancement value from the fused image data, ultrasound image data, scan image data, and digital image data using a weighted fusion algorithm; obtaining the image gradient value from the fused image data using the Canny edge detection algorithm; and extracting depth information from the fused image data using a convolutional neural network. The image depth information value is obtained from the feature; the image occlusion degree is obtained from the fused image data through image segmentation and occlusion detection algorithms; the product of local image contrast and image sharpness is calculated, divided by the image noise level, and the quotient is double-integrated over the entire image range to obtain a double integral term; the sum of the enhanced values of the fused image is calculated, and the sum is multiplied by the data fusion coefficient to obtain the fusion enhancement term; the double integral term and the fusion enhancement term are added to obtain the manifestation numerator term; the product of the image gradient value and image depth information value at each location is calculated and divided by the image occlusion degree at that location, the quotients over all locations are summed, multiplied by the deep structure enhancement coefficient, and the product is added to the deviation protection constant to obtain the manifestation denominator term; finally, the manifestation numerator term is divided by the manifestation denominator term to obtain the obstacle manifestation value.
[0012] Optionally, the specific process of improving image quality based on the multimodal image feature fusion extraction and analysis results and entering the obstacle recognition and feedback process is as follows: Real-time comparison of obstacle display value and obstacle display threshold, which includes a primary display threshold and a secondary display threshold: When the obstacle display value is less than the secondary display threshold, high-frequency details of the image are restored through super-resolution reconstruction technology, and missing areas in the image are repaired using a deep convolutional neural network to restore the obstacles in the image, thus entering the obstacle recognition and feedback process; When the obstacle display value is greater than or equal to the secondary display threshold and less than the primary display threshold, image sharpening, contrast stretching, and edge sharpening are performed using the Laplacian operator, and artifacts are avoided using wavelet transform and image denoising deep networks, thus entering the obstacle recognition and feedback process; When the obstacle display value is greater than or equal to the primary display threshold, the obstacle is clearly displayed, requiring no further intervention, and the process enters the obstacle recognition and feedback process.
[0013] Optionally, the specific process for hierarchical evaluation of historical task interaction data and real-time perception data based on priority optimization and feedback adjustment is as follows: Acquire sensor calibration data, task execution records, fused image data, task priority data, historical task completion status data, user operation log data, real-time task data, and real-time environmental data; calculate task priority and response sensitivity based on sensor calibration data and task execution records, and obtain feedback response adjustment coefficients using dynamic time warping algorithm and weighted average method; classify obstacles in fused image data using K-means clustering algorithm, and obtain the importance value of obstacle regions using weighted average method; obtain task priority values from task priority data using weighted algorithm and Bayesian optimization; obtain feedback type enhancement coefficients from historical task completion status data using machine learning algorithm; obtain feedback type priority from user operation log data and real-time task data using multi-objective optimization algorithm and adaptive algorithm; obtain feedback adjustment refinement coefficients from user operation log data and real-time task data using deep learning and reinforcement learning algorithms; and obtain feedback adjustment refinement coefficients from user operation log data and historical task completion status data using KNN algorithm and self-reinforcement learning algorithm. The adaptive adjustment algorithm obtains the user feedback adjustment coefficient; the Kalman filter is used to analyze environmental changes in real-time environmental data to obtain the environmental change adjustment coefficient; the time-series prediction algorithm is used to predict the task change trend in real-time task data to obtain the real-time task change factor; the importance values of all obstacle regions are summed, and the priority values of all tasks are summed, and the ratio of the two is calculated to obtain the task priority ratio term; the task priority ratio term is multiplied by the feedback response adjustment coefficient and the negative number is taken to obtain the feedback response adjustment term; the feedback response adjustment term is exponentially operated to obtain the exponential activation term, and added to form a unity gain superposition term; the obstacle manifestation value is multiplied by the unity gain superposition term to obtain the manifestation coupling term; the feedback type enhancement coefficient is multiplied by the sum of the feedback type priority and the feedback adjustment refinement coefficient, and accumulated over the entire feedback type range to obtain the feedback type synthesis term; the user feedback adjustment coefficient is multiplied by the sum of the environmental change adjustment coefficient and the real-time task change factor, and accumulated over the entire user range to obtain the user environment synthesis term; the feedback type synthesis term is added to the user environment synthesis term as the denominator assembly term; the manifestation coupling term is divided by the denominator assembly term to obtain the obstacle feedback intensity value.
[0014] Optionally, the specific process of identifying and marking obstacles and providing feedback on their intensity based on the graded assessment results is as follows: The obstacle feedback intensity value and the obstacle feedback intensity threshold are compared in real time. The obstacle feedback intensity threshold includes a primary intensity threshold and a secondary intensity threshold. The obstacle identification and feedback process is executed as follows: When the obstacle feedback intensity value is less than the secondary obstacle feedback intensity threshold, a circular light gray mark is used to mark the obstacle, and a warning is only issued when the obstacle approaches; when the obstacle feedback intensity value is greater than or equal to the secondary intensity threshold but less than the primary intensity threshold, a triangle and yellow mark are used to mark the obstacle, triggering a warning sound, and Laplacian sharpening and contrast enhancement processing are applied to the image area; when the obstacle feedback intensity value is greater than or equal to the primary intensity threshold, a star and red mark are used to mark the obstacle, triggering a warning sound and a vibration sound, and super-resolution reconstruction and edge sharpening are used to reveal obstacle details.
[0015] Optionally, the specific process for spatial accessibility and risk detection of historical task interaction data and real-time perception data is as follows: Acquire location spatial coordinates, obstacle spatial coordinates, user operation log data, historical abnormal alarm counts, and historical task completion status data; use an IMU (Inertial Measurement Unit) to collect location spatial coordinates and obstacle spatial coordinates, and calculate the obstacle-user spatial distance using the Euclidean distance formula; obtain scene environment values by using principal component analysis algorithm on local image contrast, image sharpness, image noise level, image depth information value, and image occlusion degree; obtain historical obstacle risk feedback values by using exponential moving average algorithm and risk memory decay function on user operation log data, historical abnormal alarm counts, and historical task completion status data; obtain the variance of the environment score by using standard variance algorithm on the scene environment values; and obtain the obstacle feedback intensity value and scene environment value by using Pearson correlation coefficient algorithm. Obtain the intensity-environment correlation; calculate the obstacle feedback intensity value and spatial distance attenuation term, multiply the obstacle feedback intensity value by the negative exponent of the spatial distance attenuation coefficient to obtain the spatial distance attenuation modulation term; calculate the historical feedback adjustment term, multiply the historical feedback weighting coefficient by the historical obstacle risk feedback value, and then multiply by the hyperbolic tangent of the product of the environmental gain adjustment coefficient and the scene environment value to obtain the historical feedback adjustment term; add the spatial distance attenuation modulation term and the historical feedback adjustment term to obtain the control numerator; calculate the variance of the environment score and the intensity-environment correlation, respectively calculate the absolute values of the variance of the environment score and the intensity-environment correlation, add the variance of the environment score and the correlation adjustment coefficient multiplied by the absolute value of the intensity-environment correlation, multiply by the global fluctuation suppression coefficient, and add the result to obtain the control denominator; divide the control numerator by the control denominator and take the natural logarithm to obtain the actual linkage control value.
[0016] Optionally, the specific process of visual linkage based on spatial accessibility and risk detection results is as follows: real-time comparison of the real linkage control value and the real linkage control threshold: when the real linkage control value is less than the real linkage control threshold, it is marked as a non-visualized interaction state; the obstacle recognition and feedback process is invoked to re-evaluate the quality of the fused image; when the real linkage control value is greater than or equal to the real linkage control threshold, it is marked as a visualized interaction state. Through augmented reality technology, the information and decision suggestions of invisible obstacles are directly superimposed on the real scene. The hierarchical markings output by the obstacle recognition and feedback process are displayed on the AR interface. All feedback is superimposed synchronously, and dynamic path projection, voice, and interactive commands are performed.
[0017] On the other hand, a visualization prompting system for image-real-scene interaction is provided. This system is applied to visualization prompting methods for image-real-scene interaction and includes: an acquisition and preprocessing module for acquiring real-time perception data, obtaining historical task interaction data, and preprocessing the real-time perception data and historical task interaction data; a pixel segmentation and image partitioning module for constructing a spatially structured expression system, partitioning and numbering complex image information; a data fusion and image enhancement module for performing multimodal image feature fusion extraction and analysis on real-time perception data, improving image quality based on the results of multimodal image feature fusion extraction and analysis, and entering the obstacle recognition and feedback process; an obstacle recognition and comprehensive feedback module for performing hierarchical evaluation of historical task interaction data and real-time perception data based on priority optimization and feedback adjustment assessment, identifying and marking obstacles and providing intensity feedback based on the hierarchical evaluation results; and an augmented reality interaction and visualization prompting module for performing spatial accessibility and risk detection on historical task interaction data and real-time perception data, and performing visualization linkage based on the spatial accessibility and risk detection results.
[0018] The beneficial effects of the technical solution provided by the embodiments of the present invention include at least the following: (1) It integrates augmented reality technology and multimodal interaction methods, and directly overlays obstacle recognition results and decision suggestions onto the real scene, thereby achieving the effect of intuitive information visualization and human-computer collaborative interaction, effectively solving the problems of scattered information and unintuitive prompts in the prior art.
[0019] (2) By using multimodal perception data fusion and high-resolution image enhancement, the detection sensitivity of invisible obstacles such as microvascular vessels, deep tissues and structural changes can be effectively improved, thereby achieving high-precision recognition of invisible obstacles in complex scenes and effectively solving the problem that conventional imaging technology is difficult to capture subtle obstacles.
[0020] (3) By combining real-time task information and environmental adaptive algorithms, the obstacle recognition results can be dynamically optimized according to the actual needs of surgery or industrial operation, thereby realizing the real-time recognition and feedback effect of invisible obstacles, effectively solving the problems of feedback lag and difficulty in dynamic control of operational risks in the existing technology.
[0021] (4) By adopting a user behavior analysis and feedback control mechanism, the feedback content and interaction method can be dynamically adjusted according to the operator's real-time needs and historical habits, thereby achieving the effect of personalized intelligent assistance and effectively solving the problem of fixed feedback methods and lack of personalized adaptation in the existing technology. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 is a flowchart of a visualization prompting method for interaction between images and real-world scenes provided in an embodiment of the present invention; Figure 2 is a structural schematic diagram of a visualization prompting system for interaction between images and real-world scenes provided in an embodiment of the present invention; Figure 3 is a three-dimensional bubble diagram of the spatial distribution of feedback intensity parameters provided in an embodiment of the present invention; Figure 4 is a flowchart of the three-level processing of feedback intensity provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0025] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0026] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0027] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0028] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0029] This invention provides a visualization prompting method for interaction between images and real-world scenes. As shown in Figure 1, the flowchart of this method includes the following steps: S1, collecting real-time perception data and acquiring historical task interaction data, and preprocessing the real-time perception data and historical task interaction data; S2, constructing a spatially structured expression system, partitioning and numbering complex image information; S3, performing multimodal image feature fusion extraction and analysis on the real-time perception data, improving image quality based on the multimodal image feature fusion extraction and analysis results, and proceeding to obstacle recognition and feedback; S4, performing hierarchical evaluation of historical task interaction data and real-time perception data based on priority optimization and feedback adjustment assessment, identifying and marking obstacles and providing intensity feedback based on the hierarchical evaluation results; S5, performing spatial accessibility and risk detection on historical task interaction data and real-time perception data, and performing visualization linkage based on the spatial accessibility and risk detection results.
[0030] Optionally, real-time sensing data is collected, and historical task interaction data is obtained. The specific process of preprocessing the real-time sensing data and historical task interaction data is as follows: Real-time sensing data is collected, including: ultrasound image data, scan image data, digital image data, sensor calibration data, position spatial coordinates, real-time environmental data, real-time task data, and obstacle spatial coordinates. Ultrasonic image data is collected using an ultrasound imager; scan image data is collected using a CT scanner; digital image data is collected using an infrared sensor; fused image data is obtained by fusing ultrasound image data, scan image data, and digital image data through scale-invariant feature transformation; position spatial coordinates and obstacle spatial coordinates are collected using an IMU inertial measurement unit; and real-time environmental data includes temperature, humidity, and background noise information.
[0031] Acquire historical task interaction data, which includes: task execution records, task priority data, user operation log data, historical abnormal alarm counts, and historical task completion status data. For task execution records and user operation log data, extract key interaction nodes and abnormal operation behaviors through log clustering analysis and sequence slicing technology.
[0032] This paper employs a multimodal noise separation and adaptive residual correction algorithm to perform preliminary denoising and signal correction on ultrasound image data, scan image data, and digital image data. In the noise separation stage, a combined approach of time-frequency domain filtering and spatial domain residual modeling is used to achieve hierarchical suppression of structural noise and background noise. A data filtering and anomaly detection method based on spatial coordinates and task labels is used to verify the temporal consistency of location spatial coordinates, obstacle spatial coordinates, and real-time task data, and to identify outliers. A dual criterion based on cluster radius constraints and temporal jump detection is used to effectively identify outliers and label misalignment events, and to correct position drift and synchronization deviations between task data in real time. Finally, a multimodal feature fusion and adaptive filtering algorithm is used to further refine the ultrasound image data, scan image data, and digital image data. Multi-channel noise suppression and detail enhancement are implemented, dynamically allocating fusion weights for each modality feature based on local image contrast, texture gradient, and spatial correlation. Adaptive filtering adjusts the smoothing intensity according to image content. Dynamic priority reconstruction and sliding window statistical algorithms are used to identify abnormal fluctuations and smooth trends in task execution records, task priority data, and historical anomaly alarm frequency time-series information. This allows for rapid identification of abnormally high-fluctuation areas in the task flow, while sliding window statistics ensure simultaneous tracking of short-term bursts and long-term trends. Distribution standardization and linear normalization algorithms are used to standardize and normalize real-time sensing data and historical task interaction data. Different data sources are normalized to standard intervals based on their statistical characteristics, effectively eliminating the influence of units and the interference of outliers on subsequent analysis.
[0033] In this implementation scheme, the data quality and robustness of the obstacle recognition and integrated feedback system are significantly improved by multi-source acquisition and fusion preprocessing of real-time perception data and historical task interaction data. By combining multiple types of sensors, including ultrasound imaging, CT scanning, infrared sensing, IMU positioning, and environmental monitoring, comprehensive acquisition of obstacle spatial information, environmental status, and task execution dynamics is achieved. Multimodal noise separation, feature fusion, and adaptive filtering algorithms are used to perform deep denoising and detail enhancement on the original multi-source data, ensuring the accuracy and consistency of the fused image and spatial coordinate information. At the same time, through temporal consistency verification, anomaly detection, and dynamic priority reconstruction, outliers and fluctuating data are effectively eliminated, improving the stability and continuity of spatiotemporal data. Finally, through standardization and normalization, different data sources are uniformly mapped to a standard range, significantly enhancing the comparability of multi-source data and the adaptability of model processing. This preprocessing process provides a high-quality and reliable data foundation and algorithmic guarantee for subsequent obstacle feedback intensity determination, multi-dimensional grading, and personalized interaction strategies.
[0034] Optionally, the specific process of constructing a spatially structured representation system and partitioning, layering, and numbering complex image information is as follows: multimodal registration is completed using SIFT features; ultrasound image data, scan image data, and digital image data are fused using Laplacian pyramid weights to obtain fused image data; during multimodal registration, the spatial distribution density of SIFT feature points is used for matching verification to further improve the spatial consistency and detail restoration capability of the fused image; the Laplacian pyramid weight parameters can be adaptively adjusted according to the edge intensity of different modalities to achieve effective integration of multi-scale information; for the fused image data, superpixel segmentation algorithm and deep learning region segmentation method are used to divide it into several image partitions with similar spatial structure and texture features, and a minimum partition size is set. The number of pixels ranges from 30 to 100 pixels, the compactness parameter is in the range of 0.1 to 0.4, and the desired number of partitions is selected from 10 to 60 according to the image size to ensure that the partitions can cover fine-grained structures without over-segmentation. Each partition is assigned a unique number, which will serve as the core index for spatial mapping in subsequent calculations and feedback. The spatial coordinates, pixel set, and boundary information of the partition are recorded simultaneously to perform image partitioning and spatial attribute structuring. The numbering is assigned incrementally according to the spatial scanning order of the partition centroid to ensure the reproducibility of segmentation results in different rounds of the same image. Each partition synchronously records its spatial centroid, pixel set, boundary point set, and bounding box coordinates and other structured attributes to facilitate subsequent partition-level obstacle attribute quantification and feedback intensity parameter extraction.
[0035] In this implementation scheme, by constructing a spatially structured representation system, efficient partitioning, layering, and standardized numbering of complex multimodal image data are achieved. Based on multimodal registration and weighted fusion, the fused image not only improves spatial consistency and detail reproduction but also lays a solid foundation for subsequent superpixel and depth segmentation. By adopting reasonable partitioning parameters and numbering rules, the geometric structure, pixel set, and boundary information of each spatial partition are uniformly and structurally managed. Ultimately, the entire image data achieves a spatially clear, numbered, and attribute-complete partitioned representation, which significantly improves the accuracy and efficiency of subsequent obstacle attribute quantification, partition feedback, and parameter extraction, providing solid data support for refined analysis and intelligent decision-making.
[0036] Optionally, the specific process of multimodal image feature fusion extraction and analysis of real-time perceived data is as follows: Ultrasonic image data, scanned image data, digital image data, and fused image data are acquired. All three types of data use a unified sampling raster to ensure a one-to-one correspondence between the multimodal data in space and resolution. The fused image data serves as the basic input for subsequent processing, and spatial coordinate information is recorded synchronously for each type of data. The local contrast of the fused image data is calculated using Gaussian smoothing-based local difference calculations, performed within a window region of each spatial pixel. Image sharpness is extracted from the fused image data using the Laplacian operator. The Laplacian operator is applied to the entire effective pixel region of the fused image, and symmetrical filling is used for boundary processing to ensure alignment between the operator output and subsequent features. The image noise level is obtained from the fused image data using the Wiener filtering algorithm and image residual method. The Wiener filtering output and image residual analysis are processed synchronously using a block window. The noise estimation results of each region are fused using a spatial weighting method. The fused image data, ultrasound image data, scan image data, and digital image data are weighted and fused using a weighted fusion algorithm to obtain the enhanced value of the fused image. All modal inputs are aligned with a uniform grid to avoid the influence of spatial misalignment on the fusion result. The Canny edge detection algorithm is used to obtain image gradient values from the fused image data. The Canny algorithm's edge detection output is a pixel-level gradient magnitude matrix, with the integration domain completely overlapping with the effective area of the fused image, and the spatial coordinates of each edge point are recorded synchronously. A convolutional neural network is used to extract depth features from the fused image data to obtain image depth information values. A pre-trained convolutional neural network is used for depth feature extraction, and the output image depth information matrix corresponds pixel-by-pixel to the input image. Image occlusion is obtained from the fused image data using image segmentation and occlusion detection algorithms. The calculation area for image occlusion is all effective pixels, and multi-layer segmentation and occlusion detection are used in conjunction to generate an occlusion matrix of the same dimension as the original image.
[0037] The product of local image contrast and image sharpness is calculated, divided by the image noise level, and then double-integrated over the entire image range to obtain a double integral term. The integration domain of the double integral term is the coordinate region Ω of all effective pixels in the fused image, i.e., the sum is accumulated over all pixels within Ω. The actual calculation uses full-image scanning and numerical integration. The sum of the enhanced values of the fused image is calculated, and the sum is multiplied by the data fusion coefficient to obtain the fusion enhancement term. The image enhancement values are added point-by-point according to spatial location, and the data fusion coefficient is adjusted according to the actual application scenario and image type. The double integral term and the fusion enhancement term are added to obtain the display numerator. The product of the image gradient value and the image depth information value at each location is calculated and divided by the image occlusion degree at that location. The quotients over all locations are summed and multiplied by the deep structure enhancement coefficient. The product is then added to the deviation protection constant to obtain the display denominator. The set of all locations is defined as the effective pixel set Λ of the fused image. The calculation is performed for each point within Λ to ensure the numerical stability of the denominator and robustness to special scenes. Finally, the display numerator is divided by the display denominator to obtain the obstacle display value. The specific calculation formula is as follows: In the formula, The obstacle visibility value is used to comprehensively measure the structural saliency of each region in the fused image and is the core basis for subsequent obstacle risk classification and feedback intensity R calculation. Indicates local contrast of an image; It indicates image sharpness, reflects the visual separation between obstacles and the background, and is exceptionally sensitive to low-contrast areas; It indicates the level of image noise, has a strong ability to distinguish imaging defects such as focus and motion blur, and is good at highlighting boundary changes; It represents the image enhancement value after fusion, which can quantitatively describe the intensity of random noise and device interference in the image and has guiding significance for signal-to-noise ratio optimization; Represents the image gradient value; the enhancement result directly affects the overall visual performance and detail reproduction of the image. It represents the image depth information value, which is used to characterize the edge intensity and structural inflection points in the image, and is an important basis for detecting obstacle boundaries and spatial abrupt changes; It indicates the degree of image occlusion, reducing the interference of occluded areas on the determination of display values; The data fusion coefficient is determined by an image quality assessment algorithm, which evaluates the contribution of the sensor image to the fused data from ultrasound image data, scan image data, and digital image data. The data fusion coefficient is obtained by a weighted fusion algorithm and its value ranges from 0.3 to 0.7. The deep structure enhancement coefficient is derived by extracting image depth information, image gradient values, and image occlusion enhancement characteristics based on deep learning algorithms. The value ranges from 1.2 to 3.0. This represents the deviation protection constant, which is obtained by calculating the fixed deviation using error analysis methods. Its value ranges from 0.005 to 0.01.
[0038] In this implementation scheme, multi-modal image feature fusion and hierarchical extraction are used to achieve multi-dimensional quantification and accurate spatial distribution representation of obstacle visibility values. Data is fused in an orderly manner under a unified sampling grid and spatial coordinates. Through comprehensive analysis of multiple features, the structural details and visual heterogeneity of different modalities and spatial levels are fully explored. After processing multiple features, the sensitivity and saliency recognition ability of obstacle areas are effectively improved by integral and weighted combination methods. Noise interference and occlusion errors are significantly suppressed, and highly reliable restoration of obstacle structural information is achieved in complex image scenes. This method provides a solid data foundation and feature support for subsequent obstacle classification assessment and dynamic feedback adjustment, and improves adaptability and fine processing capabilities in complex environments.
[0039] Optionally, the specific process of improving image quality based on the multimodal image feature fusion extraction and analysis results and entering the obstacle recognition and feedback process is as follows: Real-time comparison of obstacle display value and obstacle display threshold. The obstacle display threshold includes a primary display threshold and a secondary display threshold, which correspond to the discrimination situation under high sensitivity and basic perception, respectively. When the obstacle display value is less than the secondary display threshold, high-frequency details of the image are restored through super-resolution reconstruction technology. Both the super-resolution and missing region repair networks use the original partition numbers of the fused image for batch processing to ensure that the repair of each obstacle region is strictly aligned with the partition spatial information. The missing regions in the image are repaired using a deep convolutional neural network to restore the obstacles in the image and enter the obstacle recognition and feedback process.
[0040] When the obstacle visibility value is greater than or equal to the secondary visibility threshold and less than the primary visibility threshold, the Laplacian operator is used for image sharpening, contrast stretching, and edge sharpening. Wavelet transform and image denoising depth network are used to avoid artifacts. Image processing operations are performed sequentially. All enhancement and denoising modules are synchronized with spatial pixel labels. The processing order is sharpening, contrast adjustment, and denoising to ensure consistent output features before entering the obstacle recognition and feedback process.
[0041] When the obstacle display value is greater than or equal to the first-level display threshold, the obstacle is clearly displayed and no further intervention is required. Only the obstacle area partition number and related metadata are synchronized, and the obstacle recognition and feedback process begins.
[0042] In this implementation scheme, dynamic comparison of display value thresholds and multi-strategy image enhancement effectively improve the adaptability of processing obstacle regions with different display levels. For low display value regions, the collaborative operation of super-resolution reconstruction and deep inpainting network achieves comprehensive completion of spatial information and structural details of obstacle regions. For medium display value regions, a multi-step image enhancement and denoising process ensures balanced optimization of various obstacle boundaries, contrast, and noise levels. For high display value regions, image enhancement is skipped, and the process directly enters the subsequent recognition and feedback stage. This process achieves a high degree of synergy between obstacle region spatial partitioning, image features, and image quality enhancement strategies, providing a more accurate and high-quality data foundation for subsequent intelligent recognition and personalized feedback output.
[0043] Optionally, based on priority optimization and feedback adjustment evaluation, the specific process for hierarchical evaluation of historical task interaction data and real-time sensing data is as follows: Sensor calibration data, task execution records, fused image data, task priority data, historical task completion status data, user operation log data, real-time task data, and real-time environmental data are acquired synchronously using unified timestamps and spatial labels to ensure that all features achieve sample alignment and consistent partitioning during the algorithm input stage; the task priority and response sensitivity are calculated based on the sensor calibration data and task execution records; feedback response adjustment coefficients are obtained using dynamic time warping algorithms and weighted average methods; and the sensor calibration data is used to correct the task priority and response sensitivity in real time. To improve the reliability of task priority and response sensitivity calculations, the system addresses trajectory errors during task execution. K-means clustering is used to classify obstacles in fused image data, and a weighted average method is employed to obtain the importance value of each obstacle region. After obstacle classification, the importance value is automatically linked to the spatial centroid and boundary features, facilitating subsequent partition-level feedback and strategy discrimination. Task priority data is processed using a weighted algorithm and Bayesian optimization to obtain task priority values. Bayesian optimization helps balance priority allocation in multi-objective tasks. Historical task completion status data is used to obtain feedback type enhancement coefficients through machine learning algorithms. These machine learning algorithms employ a multilayer perceptron to minimize the deviation between actual and ideal feedback. To enhance feedback differentiation, user operation log data and real-time task data are used to obtain feedback type priorities through multi-objective optimization and adaptive algorithms. Multi-objective optimization aims to maximize user satisfaction and minimize task conflict, with inputs being user behavior sequences and task temporal features, and employing Pareto optimal solution set constraints for the optimization process. Feedback adjustment refinement coefficients are obtained from user operation log data and real-time task data through deep learning and reinforcement learning algorithms. The deep learning part takes user operation sequences and task state features as input, using prediction error minimization as the loss function. The reinforcement learning part aims to maximize cumulative reward, with the action space representing the feedback adjustment magnitude and the environmental state representing the interaction history trajectory. Training is performed using policy gradients; user feedback adjustment coefficients are obtained from user operation log data and historical task completion status data through KNN and adaptive adjustment algorithms. These coefficients are correlated with specific operating habits and historical behavioral characteristics, supporting real-time fine-tuning and dynamic adaptation of personalized feedback strategies; environmental change adjustment coefficients are obtained from real-time environmental data through Kalman filtering analysis, improving the robustness and response speed of the adjustment coefficients in abrupt environments; and real-time task data are used to predict task change trends through time-series prediction algorithms, obtaining real-time task change factors. These algorithms support multi-timescale inputs, ensuring that task change factors can be sensitively reflected under both sudden and periodic tasks.
[0044] The importance values of all obstacle regions are summed, and the priority values of all tasks are summed. The ratio of the two is calculated to obtain the task priority ratio. All obstacle regions correspond one-to-one with task partition numbers. Data deduplication and spatial aggregation are performed before summing and ratio calculation. The task priority ratio is multiplied by the feedback response adjustment coefficient and the negative is taken to obtain the feedback response adjustment term. The result of the adjustment term is dynamically limited to the allowable feedback range to prevent extreme inputs from causing uncontrolled response. An exponential operation is performed on the feedback response adjustment term to obtain an exponential activation term, which is then added to form a unity gain superposition term. The exponential activation term ensures that the feedback mechanism has a gradual change in sensitivity, and the superposition of the gain term ensures that even low priority ratio regions can obtain minimum response. The obstacle display value is superimposed with the unity gain. Multiplying the terms yields the explicit coupling term, which matches the feedback flow for each obstacle spatial partition. Multiplying the feedback type enhancement coefficient by the sum of the feedback type priority and feedback adjustment refinement coefficients, and accumulating this sum across all feedback types, yields the feedback type composite term, summarizing all category features. This facilitates unified output and risk stratification for subsequent multi-type feedback suggestions. Multiplying the user feedback adjustment coefficient by the sum of the environmental change adjustment coefficient and the real-time task change factor, and accumulating this sum across all users, yields the user environment composite term, achieving cross-user policy unification and personalized compatibility. Adding the feedback type composite term to the user environment composite term serves as the denominator assembly term. Dividing the explicit coupling term by the denominator assembly term yields the obstacle feedback intensity value, calculated using the following formula: In the formula, This represents the obstacle feedback intensity value, used to comprehensively measure the risk of obstacle visibility at the zone level; This represents the obstacle visibility value, reflecting the structural saliency within the fused image partition; This represents the feedback response adjustment coefficient, supporting personalized feedback adjustment for different task types. This indicates the importance value of the obstacle area, reflecting the relative degree of impact of different obstacle zones on task execution; This indicates the task priority value and supports dynamic adjustment of priority in a multi-task environment. This represents the feedback type enhancement coefficient, which strengthens the output weight of a specific type of feedback. This indicates the priority of feedback types, providing a reference for sorting feedback types and allocating resources; This indicates that feedback adjustments can be made to refine the coefficients, adapting to different user behaviors and real-time changes; This represents the user feedback adjustment coefficient, which aggregates user history and interaction preferences to achieve adaptive adjustment of personalized feedback at the zone level. It represents the environmental change adjustment coefficient, which corrects the sensitivity of the feedback parameters to external disturbances in real time; This represents the real-time task change factor, reflecting the immediate impact of dynamic task changes on the feedback strategy.
[0045] In this embodiment, Table 1 is a table of obstacle feedback intensity parameters and classification data. It records in detail the obstacle display value, feedback response adjustment coefficient, obstacle area importance value, task priority value, denominator assembly term, and finally calculated obstacle feedback intensity corresponding to different serial numbers in typical obstacle recognition and integrated feedback scenarios. This table is used to quantify the differences in obstacle feedback intensity in each partition under multi-factor collaboration. Among them: Sample No. 1 has an obstacle visibility value of 0.32, a feedback response modifier of 0.7, an obstacle region importance value of 2.2, a task priority value of 4.4, a denominator assembly term of 0.46, and a final feedback intensity of 2.37, which belongs to the high intensity level; Sample No. 2 has an obstacle visibility value of 0.25, a feedback response modifier of 0.7, an obstacle region importance value of 2.2, a task priority value of 4.4, a denominator assembly term of 0.46, and a feedback intensity of 1.85, which is at a relatively high feedback level; Sample No. 3 has an obstacle region importance value increased to 4.0, a task priority value of 4.0, and a feedback intensity of 1.70; Sample No. 4 has an increased feedback response modifier of 1.0 and a feedback intensity of 1.79. Furthermore, as the obstacle visibility value and the denominator assembly term decrease, the feedback intensity of serial numbers 6, 7, and 8 decreases to 0.81, 0.74, and 0.34 respectively, reflecting the medium-low and extremely low level scenarios. The obstacle feedback intensity parameters and the level data table provide basic data support for the level control and dynamic adjustment of the feedback strategy.
[0046] Table 1 Obstacle Feedback Intensity Parameters and Grading Data
[0047] Figure 3 shows a three-dimensional bubble diagram of the spatial distribution of feedback intensity parameters. The diagram visually illustrates the distribution characteristics of feedback intensity in the parameter space under different obstacle display values and denominator assembly terms. Each bubble represents a sequence number in Table 1. The spatial position of the bubble is jointly determined by the obstacle display value, the denominator assembly term, and the feedback intensity. The size and color of the bubble further reflect the level of feedback intensity. It can be seen that as the obstacle display value increases and the denominator assembly term decreases, the feedback intensity significantly increases, and the bubbles exhibit a clear stratified and gradient distribution. High-intensity samples such as sequences 1, 2, and 3 cluster in the high and low denominator regions, reflecting a sensitive response to high-risk obstacle regions; while low-intensity samples such as sequences 7 and 8 are distributed in the lower and higher denominator regions, demonstrating the suppression and graded control capabilities at lower risks. This three-dimensional bubble diagram not only effectively reveals the influence of multi-parameter coupling on feedback intensity but also provides a theoretical basis and intuitive reference for dynamic grading judgment and subsequent visual interactive optimization.
[0048] In this implementation plan, through multi-level hierarchical analysis based on priority optimization and feedback adjustment evaluation, deep integration and hierarchical judgment of historical task interaction data and real-time perception data are achieved. Under a unified timestamp and spatial label standard, various data features are collected and aligned, laying a solid foundation for subsequent parameter calculation and partition identification. Combined with sensor calibration correction and dynamic time warping during task execution, the accuracy of priority and response sensitivity judgment is effectively improved, ensuring the timeliness and accuracy of feedback strategies. Multimodal image data are spatially clustered and weighted fused, and the importance and spatial distribution characteristics of obstacle areas are synchronously linked, significantly enhancing the pertinence of partition-level risk assessment. The combined use of multi-objective optimization, Bayesian optimization, and machine learning algorithms ensures that diverse parameters such as task priority, feedback type weight, and user behavior can dynamically adapt to complex task scenarios. Through standardized aggregation and fine-tuning, parameters of various partitions and types are ultimately integrated into a feedback intensity parameter system, realizing real-time hierarchical output of obstacle feedback intensity. The accompanying hierarchical data table and 3D bubble chart intuitively reflect the parameter distribution and spatial hierarchy, providing high-quality data support and theoretical foundation for system partitioned hierarchical control, interactive feedback, and dynamic strategy optimization, and comprehensively improving the adaptability and intelligence level in complex environments.
[0049] Optionally, the specific process of identifying and marking obstacles and providing feedback on their intensity based on the graded assessment results is as follows: As shown in Figure 4, the three-level processing flowchart for feedback intensity is as follows: The obstacle feedback intensity value and the obstacle feedback intensity threshold are compared in real time. The obstacle feedback intensity threshold includes a first-level intensity threshold and a second-level intensity threshold. The obstacle identification and feedback process is executed, and the threshold is mapped using the Lab standard color space. The first-level and second-level intensity thresholds correspond one-to-one with the UI prompt color parameters. When the obstacle feedback intensity value is less than the second-level obstacle feedback intensity threshold, a circular light gray mark is used to mark the obstacle. A reminder is only given when the obstacle is close. The size of the circular mark is adaptively scaled according to the actual pixel area of the obstacle. The mark position is precisely aligned with the spatial centroid of the obstacle. The transparency parameter is superimposed to keep the background content visible. The reminder only pops up when the distance between the obstacle and the user is less than the distance threshold.
[0050] When the obstacle feedback intensity value is greater than or equal to the secondary intensity threshold and less than the primary intensity threshold, the obstacle is marked with a triangle and yellow, triggering a warning sound. The image area is then processed with Laplacian sharpening and contrast enhancement. The triangle marker adopts the standard shape with the vertex facing upwards. The yellow color value in the HSV space is H=55~60, S>0.7, V>0.8. The marker can appropriately cover the boundary area of the obstacle. The volume and frequency of the warning sound can be customized. Image sharpening and contrast enhancement adopt a local ROI overlay mode.
[0051] When the obstacle feedback intensity value is greater than or equal to the first-level intensity threshold, the obstacle is marked with a star and red, triggering a warning sound and vibration. Super-resolution reconstruction and edge sharpening are used to reveal obstacle details. The star mark adopts the shape of a pentagram and a multi-pointed star. The red color value takes a>60 and b<20 in the Lab space. The mark border is thickened to highlight the high risk. The warning sound and vibration are triggered synchronously. The UI supports layered display and interactive highlighting when marking multiple obstacles. The reconstruction and sharpening processing results are rendered in real time in the obstacle partition area.
[0052] In this implementation scheme, real-time comparison and threshold mapping based on feedback intensity grading achieve precise linkage between obstacle recognition, visual marking, and multimodal feedback. The Lab standard color space is used to map the grading thresholds one-to-one with the UI prompt color parameters. Obstacles in different intensity ranges are accurately marked with different shapes such as circles, triangles, and stars. The size and position of the markers are adjusted according to the actual pixel area and spatial centroid. The multi-level feedback mechanism supports the adjustment of multi-dimensional parameters such as transparency, volume, and vibration, taking into account both user experience and risk warning. Image enhancement, sharpening, and super-resolution processing all adopt local area dynamic rendering to further improve the structural salience and detail visualization level of obstacles. This process achieves a close integration of grading evaluation results and visual interactive markings, providing a comprehensive perception and feedback foundation for obstacle detection, intelligent reminders, and dynamic interaction strategies, effectively improving the intelligence and practicality in complex environments.
[0053] Optionally, the specific process for spatial accessibility and risk detection of historical task interaction data and real-time perception data is as follows: Location spatial coordinates, obstacle spatial coordinates, user operation log data, historical abnormal alarm counts, and historical task completion status data are acquired. All spatial coordinates, user data, and obstacle data are calibrated using a unified world coordinate system and collected synchronously via timestamps and event tags. Location spatial coordinates and obstacle spatial coordinates are acquired using an IMU (Inertial Measurement Unit). The spatial distance between the obstacle and the user is calculated using the Euclidean distance formula. The raw IMU measurements are then registered in real-time with the image spatial coordinates using a coordinate transformation matrix to ensure that the physical dimensions of the spatial distance are consistent with those of the image analysis module. Principal component analysis is used to calculate local image contrast, image sharpness, image noise level, image depth information values, and image occlusion. The method obtains scene environment values. Principal component analysis inputs multidimensional feature data within a specified sampling window. The sampling window size can be dynamically adjusted according to the real-time environment to improve the timeliness and representativeness of the environment assessment. Historical obstacle risk feedback values are obtained from user operation log data, historical abnormal alarm counts, and historical task completion status data using an exponential moving average algorithm and a risk memory decay function to reflect the different impact weights of short-term high-risk and long-term accumulated events. The variance of the environment score is obtained from the scene environment values using a standard variance algorithm. The variance statistics of the environment score use the scene environment values of each partition within the sliding window as input to improve the sensitivity to environmental fluctuations. The intensity-environment correlation between obstacle feedback intensity values and scene environment values is obtained using a Pearson correlation coefficient algorithm. Random associations are eliminated through significance testing to improve the effectiveness of correlation analysis.
[0054] The obstacle feedback intensity value and spatial distance attenuation term are calculated. The spatial distance attenuation modulation term is obtained by multiplying the obstacle feedback intensity value by the negative exponent of the spatial distance attenuation coefficient. This negative exponent attenuation ensures priority response to near-range obstacle risks. The historical feedback adjustment term is calculated by multiplying the historical feedback weighting coefficient by the historical obstacle risk feedback value, then multiplying this by the hyperbolic tangent of the product of the environmental gain adjustment coefficient and the scene environment value. This ensures a balance between historical risk feedback and real-time environmental gain. The spatial distance attenuation modulation term is added to the historical feedback adjustment term to obtain the control numerator. The intermediate term calculation result is synchronized with the spatial partition number for subsequent partition-level control value statistics. The variance of the environmental score and the intensity-environment correlation are calculated. The absolute values of the variance of the environmental score and the intensity-environment correlation are calculated separately. The variance of the environmental score is multiplied by the correlation adjustment coefficient, then multiplied by the absolute value of the intensity-environment correlation, and finally multiplied by the global fluctuation suppression coefficient. The result is added to the control denominator. The control numerator is divided by the control denominator and the natural logarithm is taken to obtain the actual linkage control value. The specific calculation formula is as follows: In the formula, This represents the actual linkage control value, used to dynamically adjust the risk feedback output of obstacle space partitioning; This represents the obstacle feedback intensity value, indicating the risk weight of obstacles within the current partition and the system response priority; It indicates the spatial distance to the user from obstacles, supports units of meters and pixels, and can output in sections and frames for easy dynamic prompts of high-risk situations at close range; It represents the scene environment value, reflecting the overall trend of environmental variables such as lighting, noise, and occlusion within the partition; This indicates the historical risk feedback value of the obstacle; The variance of the environmental score is used to highlight zones with long-term high risk and frequent risk. It indicates the degree of environmental relevance and measures sensitivity to environmental fluctuations and local disturbances. This represents the spatial distance attenuation coefficient, which is obtained by analyzing and fitting the spatial distance between the obstacle and the user through a nonlinear regression algorithm. The value ranges from 0.005 to 0.05. This represents the environmental gain adjustment coefficient, which is obtained by applying a Bayesian optimization algorithm to the scene environment values. The value ranges from 0.1 to 1.5. This represents the historical feedback weighting coefficient, which is obtained by using a multinomial regression algorithm on the historical risk feedback values of obstacles. The value ranges from 0.01 to 0.2. This represents the global fluctuation suppression coefficient, which is obtained by using the least squares method with respect to the variance of the environmental score. The value ranges from 0.5 to 1.5. The correlation moderating coefficient is obtained by performing Pearson correlation coefficient analysis on the correlation between intensity and environment, and its value ranges from 0.5 to 1.5.
[0055] This implementation scheme achieves accessibility analysis and refined risk detection for obstacle spatial partitioning through multi-dimensional spatial fusion and dynamic risk assessment of historical task interaction data and real-time perception data. A unified world coordinate system is used to calibrate and synchronize user, obstacle, and environmental variables, ensuring spatiotemporal consistency of all spatial and behavioral data. Real-time registration of IMU and image coordinates effectively improves the accuracy of spatial distance and risk assessment. After principal component analysis, the scene environment values and their variance reflect environmental fluctuations and uncertainties within the partition. Historical risk feedback, combined with operation logs and event memory decay, achieves dynamic compatibility for short-term and long-term risks. Significance tests of correlation and variance statistics further enhance the scientific rigor and accuracy of risk linkage analysis. Finally, a linkage control value is constructed based on multiple modulation terms such as distance decay, historical feedback, and environmental fluctuations, achieving adaptive adjustment of obstacle risk output and dynamic control at the partition level. This provides solid technical support and data foundation for intelligent partitioning, proactive risk control, and multi-modal linkage feedback, significantly enhancing the dynamic response and intelligent adaptation capabilities to complex spatial scenarios.
[0056] Optionally, the specific process of visual linkage based on spatial accessibility and risk detection results is as follows: real-time comparison of the real linkage control value and the real linkage control threshold: when the real linkage control value is less than the real linkage control threshold, it is marked as a non-visualized interaction state; the obstacle recognition and feedback process is invoked to re-evaluate the quality of the fused image. In the non-visualized interaction state, the AR interface obstacle highlighting and path projection will be paused, only the background status monitoring will be retained, and the image quality re-evaluation and obstacle risk re-judgment will be triggered.
[0057] When the real-world linkage control value is greater than or equal to the real-world linkage control threshold, it is marked as a visual interactive state. Through augmented reality technology, information and decision suggestions of invisible obstacles are directly superimposed onto the real-world scene. The AR superimposition adopts multi-level, priority sorting and spatial avoidance rules. First, the visualization level is allocated according to the obstacle feedback intensity and spatial priority to ensure the prominence of high-risk areas. Low-priority feedback is displayed in a semi-transparent and small size to prevent multiple obstacle information from obscuring each other. The hierarchical markers of the obstacle recognition and feedback process output are displayed on the AR interface. All feedback is superimposed synchronously and dynamically projected, voice and interactive commands are used. The hierarchical markers adopt a spatial non-overlapping layout first. The markers and path projections avoid the main area of the current user's line of sight. All real-time feedback is displayed in order of risk level. When conflicts occur, they are dynamically switched through carousel, flashing and transparency adjustment. The voice and command module can customize the output frequency and content to achieve effective management and information deduplication of multi-channel high-concurrency feedback.
[0058] In this implementation plan, by real-time determination of linkage control values and flexible switching of AR visualization interaction states, a deep integration of obstacle risk perception and augmented reality interface is achieved. When in non-visual interaction state, highlighted markers and path projection are paused, retaining only background monitoring and periodic risk reassessment, effectively reducing interface interference. When entering visualization interaction state, all obstacle information and decision suggestions are superimposed on the real scene according to multi-level, spatial avoidance, and priority ranking principles. Through multi-channel output of hierarchical markers, path projection, voice, and interactive commands, high-risk information is always prioritized, while low-priority feedback is displayed semi-transparently or in a small size to avoid information congestion and obstruction. Dynamically adjusting the marker layout and display order, combined with carousel, flashing, and transparency adjustment mechanisms, enables accurate presentation and effective management of multi-obstacle zones, complex paths, and high-concurrency feedback. Overall, it significantly improves the intelligence, interactivity, and practicality of obstacle visualization linkage in complex scenarios, providing high-quality visual and decision support for applications such as safe operation, intelligent navigation, and human-machine collaboration.
[0059] On the other hand, a visualization prompting system for image-real-scene interaction is provided. Figure 2 is a schematic diagram of the structure of a visualization prompting system for image-real-scene interaction according to an exemplary embodiment. This system is used for a visualization prompting method for image-real-scene interaction. Referring to Figure 2, the system includes an acquisition and preprocessing module, a pixel segmentation and image partitioning module, a data fusion and image enhancement module, an obstacle recognition and comprehensive feedback module, and an augmented reality interaction and visualization prompting module. The system comprises the following modules: Acquisition and Preprocessing Module: Acquires real-time sensing data and historical task interaction data, and preprocesses both. Pixel Segmentation and Image Partitioning Module: Constructs a spatially structured representation system, partitioning and numbering complex image information. Data Fusion and Image Enhancement Module: Performs multimodal image feature fusion extraction and analysis on real-time sensing data, improves image quality based on the results, and initiates obstacle recognition and feedback. Obstacle Recognition and Comprehensive Feedback Module: Based on priority optimization and feedback adjustment evaluation, performs hierarchical evaluation of historical task interaction data and real-time sensing data, identifies and marks obstacles based on the evaluation results, and provides feedback on their intensity. Augmented Reality Interaction and Visualization Prompt Module: Performs spatial accessibility and risk detection on historical task interaction data and real-time sensing data, and implements visual linkage based on the results.
[0060] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0061] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0062] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0063] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0064] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0065] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0066] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0067] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0068] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0069] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0070] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A visual prompting method for interaction between images and real-world scenes, characterized in that, The method includes: S1, collecting real-time perception data and acquiring historical task interaction data, and preprocessing the real-time perception data and historical task interaction data; S2, constructing a spatial structured expression system, dividing and layering complex image information, and assigning numbers; S3, performing multimodal image feature fusion extraction and analysis on the real-time perception data, improving image quality based on the results of multimodal image feature fusion extraction and analysis, and entering the obstacle recognition and feedback process; S4, performing hierarchical evaluation on the historical task interaction data and real-time perception data based on priority optimization and feedback adjustment evaluation, identifying and marking obstacles and providing feedback on their intensity based on the hierarchical evaluation results; S5, performing spatial accessibility and risk detection on the historical task interaction data and real-time perception data, and performing visual linkage based on the spatial accessibility and risk detection results.
2. The visualization prompting method for image-real-scene interaction according to claim 1, characterized in that, The specific process of collecting real-time sensing data, acquiring task interaction history data, and preprocessing the real-time sensing data and task interaction history data is as follows: Real-time sensing data is collected, including: ultrasound image data, scan image data, digital image data, sensor calibration data, location spatial coordinates, real-time environmental data, real-time task data, and obstacle spatial coordinates; task interaction history data is acquired, including: task execution records, task priority data, user operation log data, historical abnormal alarm counts, and historical task completion status data; multimodal noise separation and adaptive residual correction algorithms are used to preprocess the ultrasound image data, scan image data, and digital image data. Preliminary denoising and signal correction are performed; data filtering and anomaly detection methods based on spatial coordinates and task labels are used to verify the temporal consistency of location spatial coordinates, obstacle spatial coordinates, and real-time task data, and to identify anomalies; multi-modal feature fusion and adaptive filtering algorithms are used to suppress noise and enhance details in ultrasonic image data, scanned image data, and digital image data; dynamic priority reconstruction and sliding window statistical algorithms are used to identify anomaly fluctuations and smooth trends in task execution records, task priority data, and historical anomaly alarm frequency time-series information; and distribution standardization and linear normalization algorithms are used to standardize and normalize real-time sensing data and historical task interaction data.
3. The visualization prompting method for interaction between images and real-world scenes according to claim 1, characterized in that, The specific process of constructing a spatially structured representation system and partitioning and numbering complex image information is as follows: multimodal registration is completed using SIFT features; ultrasound image data, scan image data, and digital image data are fused using Laplacian pyramid weights to obtain fused image data; for the fused image data, superpixel segmentation algorithm and deep learning region segmentation method are used to divide it into several image partitions with similar spatial structure and texture features. Each partition is assigned a unique number, which will serve as the core index for spatial mapping in subsequent calculations and feedback. The spatial coordinates, pixel set, and boundary information of the partition are recorded simultaneously to perform image partitioning and spatial attribute structuring.
4. The visualization prompting method for image-real-scene interaction according to claim 1, characterized in that, The specific process of multimodal image feature fusion extraction and analysis of real-time perceived data is as follows: acquiring ultrasound image data, scan image data, digital image data, and fused image data; calculating the local contrast of the fused image data using Gaussian smoothing-based local difference calculation; extracting image sharpness from the fused image data using the Laplacian operator; obtaining the image noise level from the fused image data using the Wiener filtering algorithm and image residual method; obtaining the fused image enhancement value from the fused image data, ultrasound image data, scan image data, and digital image data using a weighted fusion algorithm; obtaining the image gradient value from the fused image data using the Canny edge detection algorithm; and extracting depth features from the image using a convolutional neural network to obtain image depth information values. The occlusion degree of the fused image data is obtained by using image segmentation and occlusion detection algorithms; the product of local image contrast and image sharpness is calculated, divided by the image noise level, and the quotient is double-integrated over the entire image range to obtain the double integral term; Calculate the sum of the enhanced values of the fused image, and multiply the sum by the data fusion coefficient to obtain the fusion enhancement term; The double integral term is added to the fusion enhancement term to form the manifest numerator term; Calculate the product of the image gradient value and the image depth information value at each location and divide it by the image occlusion degree at that location. Sum the quotients for all locations and multiply them by the deep structure enhancement coefficient. Then add the deviation protection constant to the product as the denominator of the display term. Finally, divide the display numerator by the display denominator to obtain the obstacle display value.
5. The visualization prompting method for image-real-scene interaction according to claim 1, characterized in that, The specific process of improving image quality based on the multimodal image feature fusion extraction and analysis results and entering the obstacle recognition and feedback process is as follows: Real-time comparison of obstacle display value and obstacle display threshold, the obstacle display threshold includes a first-level display threshold and a second-level display threshold: When the obstacle display value is less than the second-level display threshold, high-frequency details of the image are restored by super-resolution reconstruction technology, and the missing areas in the image are repaired by deep convolutional neural network to restore the obstacles in the image and enter the obstacle recognition and feedback process; When the obstacle display value is greater than or equal to the secondary display threshold and less than the primary display threshold, the Laplacian operator is used to sharpen the image, stretch the contrast and sharpen the edges. Wavelet transform and image denoising depth network are used to avoid artifacts and enter the obstacle recognition and feedback process. When the obstacle visibility value is greater than or equal to the first-level visibility threshold, the obstacle is clearly visible and no further intervention is required; the process then proceeds to the obstacle recognition and feedback process.
6. The visualization prompting method for interaction between images and real-world scenes according to claim 1, characterized in that, The specific process of hierarchically evaluating historical task interaction data and real-time perception data based on priority optimization and feedback adjustment is as follows: Acquire sensor calibration data, task execution records, fused image data, task priority data, historical task completion status data, user operation log data, real-time task data, and real-time environmental data; calculate task priority and response sensitivity based on sensor calibration data and task execution records, and obtain feedback response adjustment coefficients using dynamic time warping algorithm and weighted average method; classify obstacles in fused image data using K-means clustering algorithm, and obtain the importance value of obstacle regions using weighted average method; obtain task priority values from task priority data using weighted algorithm and Bayesian optimization; obtain feedback type enhancement coefficients from historical task completion status data using machine learning algorithm; obtain feedback type priority from user operation log data and real-time task data using multi-objective optimization algorithm and adaptive algorithm; obtain feedback adjustment refinement coefficients from user operation log data and real-time task data using deep learning and reinforcement learning algorithms; obtain user feedback adjustment coefficients from user operation log data and historical task completion status data using KNN algorithm and adaptive adjustment algorithm; and obtain environmental change adjustment coefficients from real-time environmental data by analyzing environmental changes using Kalman filtering. The real-time task change trend is predicted using a time-series prediction algorithm to obtain the real-time task change factor. The importance values of all obstacle regions and the priority values of all tasks are summed, and their ratio is calculated to obtain the task priority ratio term. The task priority ratio term is multiplied by the feedback response adjustment coefficient and the negative number is taken to obtain the feedback response adjustment term. An exponential operation is performed on the feedback response adjustment term to obtain the exponential activation term, which is then added to form a unity gain superposition term. The obstacle visibility value is multiplied by the unity gain superposition term to obtain the visibility coupling term. The feedback type enhancement coefficient is multiplied by the sum of the feedback type priority and the feedback adjustment refinement coefficient, and accumulated over all feedback types to obtain the feedback type composite term. The user feedback adjustment coefficient is multiplied by the sum of the environmental change adjustment coefficient and the real-time task change factor, and accumulated over all users to obtain the user environment composite term. The feedback type composite term is added to the user environment composite term as the denominator assembly term. The visibility coupling term is divided by the denominator assembly term to obtain the obstacle feedback intensity value.
7. The visualization prompting method for interaction between images and real-world scenes according to claim 1, characterized in that, The specific process of identifying and marking obstacles and providing feedback on their intensity based on the graded assessment results is as follows: Real-time comparison of the obstacle feedback intensity value and the obstacle feedback intensity threshold, which includes a first-level intensity threshold and a second-level intensity threshold. The obstacle identification and feedback process is executed: When the obstacle feedback intensity value is less than the second-level obstacle feedback intensity threshold, the obstacle is marked with a circular light gray mark, and a warning is only given when the obstacle is approaching. When the obstacle feedback intensity value is greater than or equal to the secondary intensity threshold and less than the primary intensity threshold, the obstacle is marked with a triangle and yellow, a warning sound is triggered, and the image area is processed with Laplacian sharpening and contrast enhancement. When the obstacle feedback intensity value is greater than or equal to the first-level intensity threshold, the obstacle is marked with a star and red, triggering a warning sound and a vibration sound, and using super-resolution reconstruction and edge sharpening to reveal obstacle details.
8. The visualization prompting method for interaction between images and real-world scenes according to claim 1, characterized in that, The specific process of spatial accessibility and risk detection based on historical task interaction data and real-time perception data is as follows: Location spatial coordinates, obstacle spatial coordinates, user operation log data, historical abnormal alarm counts, and historical task completion status data are acquired; location spatial coordinates and obstacle spatial coordinates are collected using an IMU (Inertial Measurement Unit); the obstacle-user spatial distance is calculated using the Euclidean distance formula; the obstacle-user spatial distance is analyzed and fitted using a nonlinear regression algorithm to obtain a spatial distance attenuation coefficient; scene environment values are obtained using principal component analysis algorithms for image local contrast, image sharpness, image noise level, image depth information value, and image occlusion degree; historical obstacle risk feedback values are obtained using exponential moving average algorithms and risk memory decay functions based on user operation log data, historical abnormal alarm counts, and historical task completion status data. The variance of the environment score is obtained by applying the standard variance algorithm to the scene environment values; The intensity-environment correlation between obstacle feedback intensity values and scene environment values is obtained using the Pearson correlation coefficient algorithm. The obstacle feedback intensity value and spatial distance attenuation term are calculated. The spatial distance attenuation modulation term is obtained by multiplying the obstacle feedback intensity value by the negative power of the spatial distance attenuation coefficient. The historical feedback adjustment term is calculated by multiplying the historical feedback weighting coefficient by the historical obstacle risk feedback value and then multiplying it by the hyperbolic tangent of the product of the environmental gain adjustment coefficient and the scene environment value. The spatial distance attenuation modulation term and the historical feedback adjustment term are added together to obtain the control numerator. The variance of the environment score and the intensity environment correlation are calculated. The absolute values of the variance of the environment score and the intensity environment correlation are calculated respectively. The variance of the environment score and the correlation adjustment coefficient are multiplied by the absolute value of the intensity environment correlation and then multiplied by the global fluctuation suppression coefficient. The result is added together to obtain the control denominator. The control numerator is divided by the control denominator and the natural logarithm is taken to obtain the actual linkage control value.
9. The visualization prompting method for interaction between images and real-world scenes according to claim 1, characterized in that, The specific process of visual linkage based on spatial accessibility and risk detection results is as follows: real-time comparison of the actual linkage control value and the actual linkage control threshold; when the actual linkage control value is less than the actual linkage control threshold, it is marked as a non-visualized interaction state; the obstacle recognition and feedback process is invoked to re-evaluate the quality of the fused image; When the real-world linkage control value is greater than or equal to the real-world linkage control threshold, it is marked as a visual interactive state. Through augmented reality technology, the information and decision suggestions of invisible obstacles are directly superimposed on the real-world scene. The hierarchical markings of the obstacle recognition and feedback process output are displayed on the AR interface. All feedback is superimposed synchronously, and dynamic path projection, voice and interactive commands are performed.
10. A visualization prompting system for image-real-scene interaction, wherein the visualization prompting system for image-real-scene interaction is used to implement the visualization prompting method for image-real-scene interaction as described in any one of claims 1-9, characterized in that, The system includes: an acquisition and preprocessing module for acquiring real-time sensing data, obtaining historical task interaction data, and preprocessing the real-time sensing data and historical task interaction data; a pixel segmentation and image partitioning module for constructing a spatially structured expression system, partitioning and numbering complex image information; a data fusion and image enhancement module for performing multimodal image feature fusion extraction and analysis on real-time sensing data, improving image quality based on the results of multimodal image feature fusion extraction and analysis, and entering the obstacle recognition and feedback process; an obstacle recognition and comprehensive feedback module for performing hierarchical evaluation of historical task interaction data and real-time sensing data based on priority optimization and feedback adjustment assessment, identifying and marking obstacles and providing intensity feedback based on the hierarchical evaluation results; and an augmented reality interaction and visualization prompt module for performing spatial accessibility and risk detection on historical task interaction data and real-time sensing data, and performing visualization linkage based on the spatial accessibility and risk detection results.
Citation Information
Patent Citations
Virtual image interaction method in virtual-reality scene
CN107422862A
Intelligent man-machine interaction system for digital group building
CN114201048A