Interaction method and device based on image recognition
By real-time monitoring of dynamic factors in training scenarios, adjusting the confidence threshold and visual feedback type of target recognition, the problem of information overload and low recognition efficiency caused by traditional visual feedback methods in complex environments is solved, and efficient target recognition and decision-making for soldiers in dynamic environments is achieved.
Patent Information
- Application Number
- CN202510490010.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In a complex and dynamic military training environment, traditional visual feedback methods can easily lead to overloading of soldiers' information or disconnection from the actual environment, affecting the efficiency of target identification and decision-making.
By monitoring the movement speed and frequency of target and non-target elements in the training scenario in real time, calculate the environment dynamic factor, dynamically adjust the target recognition confidence threshold and visual feedback type, set the visual feedback intensity parameters, and superimpose the feedback in the field of view of the technical object.
Reduce soldiers' cognitive load, improve target rapid identification ability, and improve the effectiveness of training systems and decision-making efficiency.
Smart Images

Figure CN120356128A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of human-computer interaction. Specifically, it relates to an interaction method and device based on image recognition. Background Art
[0002] In modern military training, especially in virtual-real interaction training scenarios such as simulated urban warfare, soldiers usually need to perform target recognition and decision-making in complex and dynamic environments. Such training environments ingeniously integrate the physical structure of the real scene with virtual-generated elements, aiming to simulate the complexity and variability of the real battlefield environment to the greatest extent. In such an environment, soldiers rely on multi-modal image recognition technology to quickly and accurately identify various targets including real objects and virtual targets. To effectively assist soldiers in making quick decisions in a dynamic environment, the training system usually needs to provide real-time visual feedback.
[0003] However, in a scenario like urban warfare where the environment changes rapidly, traditional visual feedback methods have obvious limitations. For example, continuously providing static feedback methods such as detailed target contour highlighting or indicating arrows is likely to cause information overload for soldiers, which instead interferes with their rapid capture and response to key information and reduces decision-making efficiency. On the other hand, if the visual feedback is too simplified, such as only providing simple dot prompts, it may not be sufficient to provide effective target guidance for soldiers in a complex environment. Especially when the environmental dynamics are relatively high and target and non-target elements move and change frequently, the confidence level of image recognition will be significantly affected, and fixed visual feedback strategies are difficult to adapt to this uncertainty of recognition results, which may lead to the disconnection between the feedback and the actual environment and reduce the effectiveness of training.
[0004] In view of the above problems, the existing technology urgently needs to be improved. Summary of the Invention
[0005] The purpose of the present application is to provide an interaction method and device based on image recognition, which can reduce the cognitive load of the technical object (soldier) and improve the technical object's (soldier's) ability to quickly recognize targets in complex and uncertain battlefield environments.
[0006] In the first aspect, the present application provides an interaction method based on image recognition for a virtual-real interaction system, and the technical solution is as follows: The steps of this method include: Monitor the training scenario in real time, analyze the moving speeds and change frequencies of target and non-target elements in the training scenario, and calculate the environmental dynamic factor, where the environmental dynamic factor quantifies the degree of environmental change; According to the environmental dynamic factor, adjust the high threshold and low threshold of the target recognition confidence level, and the greater the environmental dynamic factor, the higher the confidence level threshold; Compare the object recognition confidence of the multi-modal image recognition output with the adjusted high threshold and low threshold to determine the visual feedback type. If the object recognition confidence is higher than the high threshold, use concise visual feedback; if the object recognition confidence is between the high threshold and the low threshold, use moderate visual feedback; if the object recognition confidence is lower than the low threshold, use no feedback or weakened visual feedback; Set the visual feedback intensity parameter according to the determined visual feedback type and object type, with the virtual object feedback intensity higher than that of the real object; Overlay the set visual feedback on the field of view of the technical object to assist the technical object in object recognition and decision-making.
[0007] Further, in the present application, the step of adjusting the high threshold and low threshold of the object recognition confidence according to the environmental dynamic factor includes: Obtain a preset confidence adjustment function, which is a piecewise function, and different environmental dynamic factor intervals correspond to different confidence adjustment strategies; Obtain the value of the environmental dynamic factor in real time; According to the interval where the value of the environmental dynamic factor is located, select the corresponding confidence adjustment strategy from the confidence adjustment function, and the confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold; According to the selected confidence adjustment strategy, calculate the adjusted high threshold and low threshold of the object recognition confidence respectively. The greater the environmental dynamic factor, the greater the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold, and the adjustment amplitude of the high threshold is greater than the adjustment amplitude of the low threshold.
[0008] Further, in the present application, the step of selecting the corresponding confidence adjustment strategy from the confidence adjustment function according to the interval where the value of the environmental dynamic factor is located, and the confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold includes: Obtain the value of the environmental dynamic factor at the current moment; Use a smoothing factor to perform a weighted average on the value of the current environmental dynamic factor and the smoothed value of the environmental dynamic factor at the previous moment to calculate the smoothed value of the environmental dynamic factor at the current moment. Among them, the value range of the smoothing factor is from 0 to 1, and the greater the smoothing factor, the greater the influence of the smoothed value of the environmental dynamic factor by the previous moment; According to the interval where the smoothed value of the environmental dynamic factor is located, select the corresponding confidence adjustment strategy from the confidence adjustment function, and the confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold.
[0009] Further, in the present application, after the step of setting the visual feedback intensity parameter according to the determined visual feedback type and object type, and the virtual object feedback intensity is higher than that of the real object, it also includes: Obtain the angular information of the observation target of the technical object and the area ratio of the target being occluded. The angular information is provided by the head tracking device, and the occlusion ratio is calculated by the image recognition module. Calculate the target occlusion factor O according to the formula O = α * A + β * (1 - cosθ), where A is the occlusion ratio, θ is the angle between the observation angle and the front of the target, and α and β are weight coefficients; Obtain the physiological state data of the technical object. The physiological state data includes eye movement data and heart rate data. Calculate the technical object state factor S according to the formula S = γ * E + ε * H, where E is the eye movement data, H is the heart rate data, and γ, ε are weight coefficients. Then calculate the fatigue degree factor F and the attention concentration factor C of the technical object according to the technical object state factor S; According to the target occlusion factor O, the fatigue degree factor F, and the attention concentration factor C, use the function I' = I * (1 + k1 * O + k2 * F + k3 * (1 - C)) to adjust the visual feedback intensity parameter I to obtain the adjusted visual feedback intensity parameter I', where I is the initial visual feedback intensity parameter, and k1, k2, and k3 are adjustment coefficients.
[0010] Further, in this application, the step of calculating the fatigue degree factor F and the attention concentration factor C of the technical object according to the technical object state factor S includes: Obtain the historical physiological data of the technical object. The historical physiological data includes historical eye movement data and historical heart rate data, and establish an individual physiological baseline model; According to the current technical object state factor S, calculate the deviation value between S and the individual physiological baseline model to obtain the deviated technical object state factor S'; Decompose the deviated technical object state factor S' into the fatigue degree factor F and the attention concentration factor C. When decomposing, dynamically adjust the weights of the eye movement data and the heart rate data in the calculation of the fatigue degree factor F and the attention concentration factor C according to the correlation between the eye movement data and the heart rate data in the historical data and the fatigue degree and the attention concentration.
[0011] Further, in this application, the step of comparing the target recognition confidence of the multi-modal image recognition output with the adjusted high threshold and low threshold to determine the visual feedback type includes: Obtain the current frame target recognition result output by the multi-modal image recognition module. The target recognition result includes the visible light image recognition confidence, the infrared image recognition confidence, and the millimeter wave radar recognition confidence; Calculate the multi-modal fusion confidence level W according to the formula W = (w1 * C + w2 * I + w3 * R) / (w1 + w2 + w3), where C is the visible light image recognition confidence level, I is the infrared image recognition confidence level, R is the millimeter wave radar recognition confidence level, w1, w2, and w3 are the weight coefficients of the three modalities respectively, and the weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero; If the multi-modal fusion confidence level W is greater than or equal to the high threshold, determine the visual feedback type as concise visual feedback. If the multi-modal fusion confidence level W is less than the low threshold, determine the visual feedback type as no feedback or weakened visual feedback. Otherwise, determine the visual feedback type as moderate visual feedback.
[0012] Furthermore, in the present application, the step of calculating the multi-modal fusion confidence level W according to the formula W = (w1 * C + w2 * I + w3 * R) / (w1 + w2 + w3), where C is the visible light image recognition confidence level, I is the infrared image recognition confidence level, R is the millimeter wave radar recognition confidence level, w1, w2, and w3 are the weight coefficients of the three modalities respectively, and the weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero includes: Obtain the acquisition timestamps of each modality image, calculate the time difference between each modality image and the current moment, and perform time decay on the confidence levels of each modality image according to the formula Ct = C * exp(-λ * Δt) to obtain the time-decayed visible light image recognition confidence level Ct, the time-decayed infrared image recognition confidence level It, and the time-decayed millimeter wave radar recognition confidence level Rt, where C is the visible light image recognition confidence level, I is the infrared image recognition confidence level, R is the millimeter wave radar recognition confidence level, Δt is the time difference, and λ is the time decay coefficient; Construct a modality interference detector to detect whether each modality image is interfered. If it is detected that a certain modality image is interfered, then according to the interference degree, use the function wi' = wi * (1 - Di) to adjust the weight coefficient of this modality to obtain the interference-adjusted weight coefficients w1', w2', and w3', where wi is the original weight coefficient and Di is the interference degree factor with a value range of 0 to 1; Obtain the technical object or target motion speed provided by the inertial measurement unit. If the visible light image recognition confidence level is lower than the preset threshold and the motion speed exceeds the preset speed threshold, then perform motion blur compensation on the visible light image to improve the visible light image recognition confidence level; Calculate the multi-modal fusion confidence level W according to the formula W=(w1'*Ct+w2'*It+w3'*Rt) / (w1'+w2'+w3'), where Ct is the recognition confidence level of the visible light image after time decay, It is the recognition confidence level of the infrared image after time decay, Rt is the recognition confidence level of the millimeter wave radar after time decay, w1', w2', and w3' are the weight coefficients of the three modalities after interference adjustment, and the weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero.
[0013] Further, in the present application, the steps of constructing the modality interference detector, based on the image quality evaluation algorithm and the sensor data anomaly detection algorithm, detecting whether each modality image is interfered, and if it is detected that a certain modality image is interfered, then according to the degree of interference, using the function wi'=wi*(1-Di) to adjust the weight coefficient of this modality, and obtaining the weight coefficients w1', w2' and w3' after interference adjustment, where wi is the original weight coefficient and Di is the interference degree factor with a value range from 0 to 1, include: Construct an image quality evaluation model, adopt a no-reference image quality evaluation algorithm, extract the clarity, contrast, and color saturation feature values of the visible light image, and use a support vector regression machine to fit the mapping relationship between the image quality and the feature values, and output the quality score of the visible light image; Construct a sensor data anomaly detection model, obtain the historical data of the infrared sensor and the millimeter wave radar, calculate the mean and variance, establish a Gaussian distribution model, input the current-time infrared sensor and millimeter wave radar data into the Gaussian distribution model, calculate the probability density value, and if the probability density value is lower than the preset threshold, it is determined that the data is abnormal; If the quality score of the visible light image is lower than the preset threshold, it is determined that the visible light image is interfered. If the infrared sensor or millimeter wave radar data is determined to be abnormal, it is determined that the corresponding modality is interfered; If it is detected that a certain modality image is interfered, then according to the output quality score or the output probability density value, use the function wi'=wi*(1-Di) to adjust the weight coefficient of this modality, and obtain the weight coefficients w1', w2' and w3' after interference adjustment, where wi is the original weight coefficient and Di is the interference degree factor, and Di is calculated according to the image quality score or the probability density value, Di=1-(Qi / Qmax) or Di=1-(Pi / Pmax), Qi is the quality score of the visible light image, Qmax is the maximum quality score of the visible light image, Pi is the probability density value, and Pmax is the maximum probability density value.
[0014] Further, in the present application, the steps of the real-time monitoring of the training scenario, analyzing the moving speeds and change frequencies of the target and non-target elements in the training scenario, and calculating the environmental dynamic factor include: Obtain the initial position information of all target and non-target elements in the training scenario; Calculate the displacement vectors of each element in two adjacent frames of images, calculate the moving speeds of each element according to the displacement vectors, and count the number of elements whose speed changes exceed a preset threshold within a unit time as the speed change frequency; Calculate the environmental dynamic factor according to the moving speeds and speed change frequencies of each element. The weights of the moving speeds and speed change frequencies of target elements are higher than those of non-target elements. The environmental dynamic factor calculation formula is D = Σ(αi * Vi + βi * Fi), where Vi is the moving speed of the i-th element, Fi is the speed change frequency of the i-th element, αi is the weight coefficient of the moving speed, βi is the weight coefficient of the speed change frequency, and the weight coefficients of target elements are higher than those of non-target elements.
[0015] Furthermore, the present application also proposes an image recognition-based interaction device for a virtual-real interaction system. The device includes: A calculation module that monitors the training scenario in real time, analyzes the moving speeds and change frequencies of target and non-target elements in the training scenario, and calculates the environmental dynamic factor, where the environmental dynamic factor quantifies the severity of environmental changes; A setting module that adjusts the high threshold and low threshold of the target recognition confidence according to the environmental dynamic factor. The greater the environmental dynamic factor, the higher the confidence threshold; A first adjustment module that compares the target recognition confidence output by multi-modal image recognition with the adjusted high threshold and low threshold to determine the visual feedback type. If the target recognition confidence is higher than the high threshold, simple visual feedback is adopted. If the target recognition confidence is between the high threshold and the low threshold, moderate visual feedback is adopted. If the target recognition confidence is lower than the low threshold, no feedback or weakened visual feedback is adopted; A second adjustment module that sets the visual feedback intensity parameter according to the determined visual feedback type and the target type. The feedback intensity of virtual targets is higher than that of real targets; A display module that superimposes and displays the set visual feedback in the field of view of the technical object to assist the technical object in target recognition and decision-making.
[0016] As can be seen from the above, an image recognition-based interaction method and device provided by the present application adjust the target recognition confidence threshold and visual feedback type through the environmental dynamic factor, and set the visual feedback intensity according to the target type, which has the advantages of reducing the cognitive load of the technical object and improving the target rapid recognition ability of the technical object in complex and uncertain battlefield environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of an image recognition-based interaction method provided by the present application.
[0018] Figure 2 Schematic diagram of a structure of an interaction device based on image recognition provided by this application.
[0019] In the figure: 210, computing module; 220, setting module; 230, first adjustment module; 240, second adjustment module; 250, display module. Specific implementation manners
[0020] Next, the technical solutions in this application will be clearly and completely described in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of this application.
[0021] Please refer to Figure 1 , this application proposes an interaction method based on image recognition for a virtual-real interaction system. The steps of this method include: S110. Monitor the training scenario in real time, analyze the moving speeds and change frequencies of the target and non-target elements in the training scenario, calculate the environmental dynamic factor, and the environmental dynamic factor quantifies the severity of environmental changes; S120. Adjust the high threshold and low threshold of the target recognition confidence according to the environmental dynamic factor. The greater the environmental dynamic factor, the higher the confidence threshold; S130. Compare the target recognition confidence output by multi-modal image recognition with the adjusted high threshold and low threshold to determine the visual feedback type. If the target recognition confidence is higher than the high threshold, use simple visual feedback. If the target recognition confidence is between the high threshold and the low threshold, use medium visual feedback. If the target recognition confidence is lower than the low threshold, use no feedback or weakened visual feedback; S140. Set the visual feedback intensity parameter according to the determined visual feedback type and the target type. The virtual target feedback intensity is higher than that of the real target; S150. Superimpose the set visual feedback on the field of view of the technical object to assist the technical object in target recognition and decision-making.
[0022] Among them, the environmental dynamic factor refers to a parameter that quantifies the degree of environmental change calculated by analyzing the moving speeds and change frequencies of the target and non-target elements in the training scenario.
[0023] Among them, the high threshold and low threshold of the target recognition confidence refer to the boundary values of the confidence interval dynamically adjusted according to the environmental dynamic factor.
[0024] Among them, the target recognition confidence output by multi-modal image recognition refers to the comprehensive confidence evaluation value that fuses the recognition results of multi-modal such as visible light, infrared, and millimeter-wave radar.
[0025] Among them, the visual feedback type refers to a differentiated feedback mode determined based on the comparison result of the confidence level and the threshold value.
[0026] Among them, the visual feedback intensity parameter refers to an intensity index for displaying feedback information set according to the target type and the feedback type.
[0027] Among them, the visual feedback superimposed display refers to an augmented reality technology that projects feedback information onto the field of view of a technical object in real time.
[0028] The core innovation of this application lies in constructing an adaptive visual feedback control system driven by environmental dynamic factors, achieving real-time matching of feedback strategies and scene complexity by dynamically adjusting the confidence threshold, combining multi-modal fusion recognition and setting of differentiated feedback intensities, balancing the feedback effectiveness and the risk of information overload in a dynamically changing urban combat training environment, and significantly improving the rapid decision-making ability of technical objects in a virtual-real fusion scenario.
[0029] As a preferred implementation manner, the specific embodiments of this application can be described as follows: A multi-modal sensor array is used to monitor the environment, including a high-resolution visible light camera, an infrared thermal imager, and a millimeter wave radar. These sensors collect data to form a fused scene image.
[0030] The following method is used to calculate the environmental dynamic factor: Calculate the motion vector of each pixel point in the scene, and then divide the scene into grids. For each grid, calculate its average motion speed and direction change frequency. Set a speed threshold and a direction change threshold. Count the number of grids exceeding these thresholds, and divide it by the total number of grids to obtain a value between 0 and 1 as the environmental dynamic factor D.
[0031] The following functional relationship is used for the dynamic adjustment of the target recognition confidence threshold: High threshold = 0.8 + 0.15 * D; Low threshold = 0.5 + 0.2 * D; Where D is the environmental dynamic factor.
[0032] Deep learning models such as YOLOv5 can be used for multi-modal image recognition to output the category and confidence level of the target. The system compares this confidence level with the dynamically adjusted threshold to determine the visual feedback type: Among them, for high confidence (higher than the high threshold), concise visual feedback is adopted: At this time, the system is very certain that what it has recognized is the target. For a technical object, what is needed is a quick and definite confirmation message, rather than a complex or overly conspicuous interference. Since the system is already very certain, there is no need to emphasize the uncertainty with fancy or strong feedback. A simple mark (such as a simple box or dot) is sufficient to tell the technical object the recognized target, and then the technical object can quickly shift its attention to the next action (such as aiming, evading, or further observing the target's behavior). Overly strong feedback may instead obscure important details of the target itself or the surrounding environment.
[0033] Among them, for low confidence (below the low threshold), no feedback or weakened visual feedback is adopted: At this time, the system is very uncertain, or believes that the probability that the recognized object is not the target is very high. In this case, giving an obvious hint to the technical object is actually harmful. If the system is uncertain but gives a strong feedback, it may cause the technical object to be distracted to focus on a threat that is actually unimportant or does not exist, wasting precious time and energy, and even making wrong judgments. To avoid false alarms and unnecessary interferences, the best strategy is not to display feedback, or only give a very inconspicuous and almost negligible hint (weakened feedback), so that the technical object mainly relies on its own observations and judgments.
[0034] Among them, for medium confidence (between the high and low thresholds), moderate visual feedback is adopted: At this time, the system has a certain degree of confidence that it is the target, but it has not reached a very certain level. There is a certain degree of uncertainty. For the technical object, this is the situation where the system's auxiliary judgment information is most needed. The technical object needs to be reminded that there is a potential target here, and at the same time, it also needs to know that the system is not completely certain. Therefore, moderate feedback is appropriate, which can attract the attention of the technical object, prompt them that they need to further confirm or pay special attention to this area, but will not "confirm" directly like high confidence, nor completely ignore it like low confidence. Balance the reminder and uncertainty.
[0035] Then, set the visual feedback intensity parameter. Specifically, the target type can be considered. For example, for a virtual target, the contour brightness is set to 255, and the information text size is 24-point font. For a real target: the contour brightness is set to 200, and the information text size is 18-point font. In a mixed reality (MR) or augmented reality (AR) environment, there are both real objects and virtual information in the user's field of view. It is very important to clearly distinguish between the two. By making the feedback of the virtual target stronger, it can help the technical object see at a glance which are the information added by the system and which is the real environment itself, avoiding confusion.
[0036] These visual feedbacks can be superimposed on the actual field of view of the technical object through a head-mounted augmented reality display.
[0037] Through the above solution, the present application realizes the adaptive adjustment of the visual feedback strategy to the dynamic changes of the environment. The environmental dynamic factor quantifies the environmental complexity and drives the dynamic adjustment of the confidence threshold, enabling the feedback strategy to adapt to environmental changes. The setting of multiple levels of visual feedback types avoids the problem of information overload and improves the feedback effectiveness. The differential setting of the feedback intensities of virtual and real targets balances the requirements of training guidance and actual combat simulation. The overall solution improves the target recognition and decision-making efficiency of the technical object in a complex dynamic environment and enhances the environmental adaptability and training effect of the virtual-real interaction training system.
[0038] The present application further proposes to obtain a preset confidence adjustment function, which is a piecewise function, and different environmental dynamic factor intervals correspond to different confidence adjustment strategies; obtain the value of the environmental dynamic factor in real time; according to the interval where the value of the environmental dynamic factor is located, select the corresponding confidence adjustment strategy from the confidence adjustment function, and the confidence adjustment strategy includes a high threshold adjustment amplitude and a low threshold adjustment amplitude; according to the selected confidence adjustment strategy, calculate the adjusted high threshold and low threshold of the target recognition confidence respectively. The larger the environmental dynamic factor, the larger the high threshold adjustment amplitude and the low threshold adjustment amplitude, and the high threshold adjustment amplitude is greater than the low threshold adjustment amplitude.
[0039] Among them, the confidence adjustment function can adopt an interval division method. For example, the environment is divided into three intervals, and each interval corresponds to a different linear adjustment coefficient. The adjustment amplitude calculation can be realized based on an exponential function. For example, the high threshold adjustment amplitude ΔH = K1 * D^2, and the low threshold adjustment amplitude ΔL = K2 * D, where D is the value of the environmental dynamic factor, and K1 and K2 are preset coefficients and satisfy K1 > K2. A hysteresis mechanism can be adopted when switching intervals to avoid frequent jumps.
[0040] Specifically, when the system detects that the value of the environmental dynamic factor is in the corresponding interval, it selects the adjustment strategy corresponding to that interval. This non-linear adjustment method ensures the strict screening of reliable targets by the high threshold when the environment changes violently, and at the same time avoids the over-elevation of the low threshold resulting in missed detections. Through the combination of segmented intervals and differential adjustment amplitudes, the system maintains a high detection sensitivity in a low-dynamic environment and gives priority to ensuring recognition accuracy in a high-dynamic environment, thus effectively balancing the accuracy of target recognition and the real-time requirement of feedback in the virtual-real interaction scenario.
[0041] The present application further proposes a method for optimizing the selection of the confidence adjustment strategy by introducing a smoothing processing mechanism, which specifically includes the following steps: obtaining the value of the environmental dynamic factor at the current moment; using a smoothing factor to perform a weighted average on the current value of the environmental dynamic factor and the smoothed value of the environmental dynamic factor at the previous moment to calculate the smoothed value of the environmental dynamic factor at the current moment; and selecting the corresponding confidence adjustment strategy according to the interval where the smoothed value of the environmental dynamic factor is located.
[0042] Specifically, after the environmental dynamic factor value at the current moment is collected in real time, a new smoothed value is generated through weighted calculation with the historical smoothed value. This smoothed value is then input into a piecewise function for policy matching. Through continuous multi-frame smoothed iteration, the fluctuations of the environmental dynamic factor are effectively suppressed, avoiding policy mis-switching caused by abnormal single-frame data.
[0043] For example, in a scenario where smoke suddenly spreads, the system gradually increases the confidence threshold through the smoothed dynamic factor value, which not only avoids feedback flicker caused by threshold jumps but also ensures timely enhancement of the recognition standard in a continuously harsh environment.
[0044] This application further proposes to obtain the angle information of the observation target of the technical object and the area ratio of the target being occluded. The angle information is provided by a head tracking device, and the occlusion ratio is calculated by an image recognition module. The target occlusion factor O is calculated according to the formula O = α * A + β * (1 - cosθ); obtain the physiological state data of the technical object, where the physiological state data includes eye movement data and heart rate data. The state factor S of the technical object is calculated according to the formula S = γ * E + ε * H, and then the fatigue degree factor F and the attention concentration factor C of the technical object are calculated according to the state factor S of the technical object; according to the target occlusion factor O, the fatigue degree factor F, and the attention concentration factor C, the visual feedback intensity parameter I is adjusted using the function I' = I * (1 + k1 * O + k2 * F + k3 * (1 - C)) to obtain the adjusted visual feedback intensity parameter I'.
[0045] Among them, the acquisition of the angle information is realized through a head tracking device, and the calculation of θ can be based on a preset azimuth reference on the front of the target. The occluded area ratio A is calculated by the image recognition module through a pixel-level segmentation algorithm to calculate the ratio of the occluded area to the total target area. The weight coefficients α and β in the formula can be set according to different scenario requirements. For example, in an urban combat environment, α = 0.6 and β = 0.4 to strengthen the influence of the occluded area on target recognition.
[0046] The eye movement data of the physiological state data is used to collect the blink frequency and the fixation point dispersion degree through an eye tracker, and the heart rate data is collected through a wrist-worn sensor. The γ and ε in the formula can be adjusted according to individual differences. The fatigue degree factor F and the attention concentration factor C are obtained by decomposing the state factor S into a pre-trained neural network model, and this model establishes an individual cognitive state mapping relationship based on historical physiological data. k1, k2, and k3 in the visual feedback intensity adjustment function can be set to 0.2, 0.15, and 0.25 respectively. The combination method of the above parameters can be dynamically optimized according to the real-time training performance of the technical object. For example, when the system detects three consecutive recognition errors, the weight coefficient of k1 is automatically increased.
[0047] Specifically, when the technical object encounters dynamic occlusion of the target in an urban combat environment, the head tracking device continuously collects the observation angle θ, and the image recognition module can update the occlusion ratio A every 200 milliseconds. The occlusion factor O is calculated in real time through a linear combination formula. For example, when θ = 60 degrees and A = 0.3, O = 0.6 * 0.3 + 0.4 * (1 - cos60°) = 0.18 + 0.4 * (1 - 0.5) = 0.38. At the same time, the blink frequency E = 0.8Hz and heart rate H = 95bpm collected by the physiological monitoring device are input into the state factor calculation formula, and S = 0.7 * 0.8 + 0.3 * 95 = 0.56 + 28.5 = 29.06 is obtained. This state factor is decomposed into F = 0.65 and C = 0.4 through a trained LSTM network, reflecting that the technical object is in a state of moderate fatigue and distracted attention.
[0048] Taking I = 50 as an example, the initial feedback intensity visual parameter is calculated by the adjustment function as follows: I' = 50 * (1 + 0.2 * 0.38 + 0.15 * 0.65 + 0.25 * (1 - 0.4)) = 50 * (1 + 0.076 + 0.0975 + 0.15) = 50 * 1.3235 ≈ 66.2.
[0049] On the basis of maintaining the original environmental dynamic factor adjustment strategy, this dynamic adjustment mechanism additionally integrates the target visibility and human body state parameters, so that the visual salience of virtual and real targets always matches the current battlefield situation and the cognitive load of the technical object.
[0050] This application further proposes to obtain the historical physiological data of the technical object, where the historical physiological data includes historical eye movement data and historical heart rate data, and establish an individual physiological baseline model; according to the current state factor S of the technical object, calculate the deviation value between S and the individual physiological baseline model to obtain the deviated state factor S' of the technical object; decompose the deviated state factor S' of the technical object into a fatigue degree factor F and an attention concentration factor C. When decomposing, according to the correlation between the eye movement data and heart rate data in the historical data and the fatigue degree and attention concentration degree, dynamically adjust the weights of the eye movement data and heart rate data in the calculation of the fatigue degree factor F and the attention concentration factor C.
[0051] Among them, to establish an individual physiological baseline model, historical eye movement data and historical heart rate data of the technical object in a normal state can be collected, and baseline parameters can be obtained through statistical analysis, such as calculating the mean value, standard deviation of historical data or establishing a probability distribution model. The deviation value can be calculated by the difference method or the ratio method. For example, S' = S - B, where B is the reference value corresponding to the baseline model. When dynamically adjusting the weight, the weight can be allocated based on the correlation coefficient in historical data. For example, when the correlation coefficient between the historical eye movement data of a certain technical object and the fatigue level is 0.8, the weight of the eye movement data in the calculation of the fatigue level factor is set to 0.6, and when the correlation coefficient of the heart rate data is 0.5, the corresponding weight is 0.4. The correlation analysis can be realized by linear regression or machine learning methods, and the weight adjustment period can be set to be updated in real time or periodically. Combined with the calculation of the technical object state factor S in the previous step, through the decomposed fatigue level factor F and attention concentration factor C, the visual feedback intensity parameter can be adjusted more accurately. For example, when the fatigue level factor exceeds the threshold, the feedback intensity is automatically reduced to avoid cognitive overload.
[0052] Specifically, the establishment of an individual physiological baseline model requires continuous collection of eye movement frequency, pupil diameter, and heart rate variability data of the technical object in a non-fatigued and attention-concentrated state. When calculating the deviation value between the current technical object state factor S and the baseline, normalization processing is used to eliminate the dimension difference. For example, S' = (S - μ) / σ, where μ is the baseline mean value and σ is the baseline standard deviation. During the decomposition process, through historical data analysis, it is found that the correlation between the eye movement data of a certain technical object and the fatigue level is as high as 0.85, while the correlation between the heart rate data and the attention concentration level is 0.7. Then, when calculating the fatigue level factor F, the weight of the eye movement data is dynamically set to 0.7, and the weight of the heart rate data is 0.3; when calculating the attention concentration factor C, the weight of the heart rate data is increased to 0.6, and the weight of the eye movement data is reduced to 0.4. This dynamic weight adjustment mechanism forms a synergy with the feedback intensity calculation based on the target occlusion factor and physiological state data in the previous step. For example, when the technical object is in a high-fatigue state, even if the target occlusion factor is low, the system will still reduce the feedback intensity to avoid interference. By eliminating the influence of individual differences on the analysis of physiological data, the decomposition accuracy of the technical object state factor is improved, making the adjustment of the visual feedback intensity parameter more in line with the actual training needs.
[0053] The present application further proposes to obtain the current frame target recognition result output by the multi-modal image recognition module. The target recognition result includes the visible light image recognition confidence, the infrared image recognition confidence, and the millimeter wave radar recognition confidence. Calculate the multi-modal fusion confidence W according to the formula W=(w1*C+w2*I+w3*R) / (w1+w2+w3), where C is the visible light image recognition confidence, I is the infrared image recognition confidence, R is the millimeter wave radar recognition confidence, and w1, w2, and w3 are the weight coefficients of the three modalities respectively. The weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero. When the multi-modal fusion confidence W is greater than or equal to the high threshold, determine the visual feedback type as concise visual feedback. When the multi-modal fusion confidence W is less than the low threshold, determine the visual feedback type as no feedback or weakened visual feedback. Otherwise, determine the visual feedback type as moderate visual feedback.
[0054] Among them, the steps of obtaining the visible light, infrared, and millimeter wave radar recognition confidences enhance the coverage of target recognition through multi-modal data fusion. In the step of dynamically adjusting the weights, the weight coefficients are updated according to the historical recognition accuracy; when a certain modality data is missing, its weight is set to zero, and the weights of the remaining modalities are re-allocated proportionally. The fusion confidence calculation adopts a weighted average method to avoid the abnormal situation of a single modality having too much impact on the overall result.
[0055] Specifically, after the multi-modal data is synchronously collected, the historical accuracy of each modality is statistically analyzed through a sliding window. When the environmental dynamic factor is relatively high, the judgment threshold is increased through a threshold adjustment strategy. At this time, the fusion confidence needs to meet both the dynamic weight calculation and the adjusted threshold conditions. In the scenario of sensor failure, for example, when the millimeter wave radar has a hardware failure, its weight is immediately set to zero, and the fusion results of the remaining modalities can still support the visual feedback decision. When the fusion confidence is between the high and low thresholds, a moderate visual feedback is used to display the target contour without annotating details, balancing the amount of information and the interference risk. This solution realizes maintaining a reliable feedback type decision in the case of sensor performance fluctuations or partial failures through a dynamic weight allocation and missing data processing mechanism.
[0056] This application further proposes to obtain the acquisition timestamps of each modality image, calculate the time differences between each modality image and the current moment, and perform time attenuation on the confidence levels of each modality image according to the formula Ct = C * exp(-λ * Δt) to obtain the visible light image recognition confidence level Ct after time attenuation, the infrared image recognition confidence level It after time attenuation, and the millimeter-wave radar recognition confidence level Rt after time attenuation; construct a modality interference detector to detect whether each modality image is interfered. If it is detected that a certain modality image is interfered, then according to the degree of interference, use the function wi' = wi * (1 - Di) to adjust the weight coefficient of this modality to obtain the weight coefficients w1', w2', and w3' after interference adjustment; obtain the technical object or target motion speed provided by the inertial measurement unit. If the visible light image recognition confidence level is lower than the preset threshold and the motion speed exceeds the preset speed threshold, then perform motion blur compensation on the visible light image to improve the visible light image recognition confidence level; calculate the multi-modal fusion confidence level W according to the formula W = (w1' * Ct + w2' * It + w3' * Rt) / (w1' + w2' + w3').
[0057] Among them, the value range of the time attenuation coefficient λ can be from 0.1 to 0.5. The confidence level decays with the time difference Δt through an exponential function. The larger the time difference Δt, the greater the confidence level decay amplitude.
[0058] The modality interference detector can adopt a no-reference image quality assessment algorithm to extract the clarity, contrast, and color saturation feature values of the visible light image, and combine a support vector regression machine to calculate the quality score. When the quality score is lower than the preset threshold, it is determined that the visible light image is interfered. The interference degree factor Di is calculated according to the quality score. For example, Di = 1 - (Qi / Qmax), where Qi is the current quality score and Qmax is the maximum quality score.
[0059] For the infrared and millimeter-wave radar modalities, a Gaussian distribution model is established to detect data anomalies. When the probability density value is lower than the preset threshold, it is determined to be interfered. The interference degree factor Di = 1 - (Pi / Pmax), where Pi is the current probability density value and Pmax is the maximum probability density value. The preset speed threshold in the motion blur compensation trigger condition can be 5 m / s. When the visible light image recognition confidence level is lower than 0.7 and the motion speed exceeds this threshold, compensation is performed. The weight coefficients are dynamically adjusted according to the historical recognition accuracy. The weight coefficients are positively correlated with the accuracy. When the data of a certain modality is missing, its weight coefficient is automatically set to zero.
[0060] Specifically, when the rapid movement of the technical object causes the time asynchronization of multimodal data, by obtaining the acquisition timestamps of each modal image, calculating the time difference from the current moment, and using the exponential decay formula to dynamically decay the confidence level, the impact of stale data on the fusion result is effectively reduced. When the sensor is interfered by smoke, heat source or metal, the modal interference detector identifies the interfered modality through image quality assessment and anomaly detection, and linearly reduces the weight of that modality according to the degree of interference, avoiding incorrect data from dominating the decision-making. When the visible light image is blurred due to the high-speed movement of the technical object, combining the motion speed data provided by the inertial measurement unit, when the recognition confidence level is lower than the threshold, the image restoration algorithm is automatically triggered to improve the effectiveness of the visible light modality. Finally, by weighted-fusing the multimodal confidence levels after time decay, interference adjustment and blur compensation, it is ensured that the fusion result accurately reflects the current environmental state. Compared with the traditional fixed-weight fusion method, this solution effectively solves the problem of confidence distortion caused by multimodal data asynchrony, interference and motion blur through the triple cooperation of the time decay coefficient, the interference degree factor and the dynamic weight adjustment mechanism.
[0061] This application further proposes to construct an image quality assessment model, adopt a no-reference image quality evaluation algorithm, extract the clarity, contrast, and color saturation feature values of the visible light image, use a support vector regression machine to fit the mapping relationship between the image quality and the feature values, and output the quality score of the visible light image; construct a sensor data anomaly detection model, obtain the historical data of the infrared sensor and the millimeter-wave radar, calculate the mean and variance, establish a Gaussian distribution model, input the current moment's infrared sensor and millimeter-wave radar data into the Gaussian distribution model, calculate the probability density value, if the probability density value is lower than the preset threshold, it is determined that the data is abnormal; if the quality score of the visible light image is lower than the preset threshold, it is determined that the visible light image is interfered, if the infrared sensor or millimeter-wave radar data is determined to be abnormal, it is determined that the corresponding modality is interfered; if it is detected that a certain modal image is interfered, then according to the output quality score or the output probability density value, use the function wi' = wi*(1 - Di) to adjust the weight coefficient of that modality, and obtain the weight coefficients w1', w2' and w3' after interference adjustment, where wi is the original weight coefficient, Di is the interference degree factor, Di is calculated according to the image quality score or the probability density value, Di = 1 - (Qi / Qmax) or Di = 1 - (Pi / Pmax), Qi is the quality score of the visible light image, Qmax is the maximum quality score of the visible light image, Pi is the probability density value, and Pmax is the maximum probability density value.
[0062] Among them, the no-reference image quality assessment algorithm can adopt the clarity calculation method based on the gradient magnitude; the contrast eigenvalue can be obtained by calculating the variance or dynamic range of the image grayscale histogram; the color saturation eigenvalue can be obtained by extracting the mean value of the S channel after converting the image to the HSV color space. The kernel function of the support vector regression machine can select the radial basis function, and its hyperparameters are determined by cross-validation. In the sensor data anomaly detection model, the historical data of the infrared sensor can include the statistical features of the temperature distribution matrix, the historical data of the millimeter-wave radar can include the point cloud density or reflection intensity, and the mean and variance of the Gaussian distribution model are updated by the sliding window method. In the calculation of the interference degree factor Di, Qmax can be set to the highest score on a 100-point scale, and Pmax can be set to the 99% quantile of the historical probability density value.
[0063] Specifically, the visible light image quality score extracts three feature dimensions of clarity, contrast, and color saturation through the no-reference algorithm, and uses the support vector regression machine to establish a nonlinear mapping model. For example, when the input feature vector is [clarity = 85.3, contrast = 72.1, color saturation = 68.5], the output quality score Qi = 78.2 points. When Qi is lower than the preset threshold of 60 points, it is determined that the visible light mode is interfered. At this time, the interference degree factor Di = 1 - (78.2 / 100) = 0.218, and the corresponding weight coefficient wi' = wi * (1 - 0.218). For the infrared sensor data, the mean value of the historical temperature data is 35.6 °C, the variance is 2.3, the current frame temperature data is 42.1 °C, and its Gaussian distribution probability density value is 0.05, which is lower than the threshold of 0.1, so it is determined to be abnormal. The interference degree factor Di = 1 - (0.05 / 0.15) = 0.667, and the corresponding weight coefficient wi' = wi * (1 - 0.667). By inputting the adjusted weight coefficient into the multi-modal fusion confidence calculation, the influence weight of the interfered mode is effectively reduced, thereby improving the reliability of the target recognition result.
[0064] The present application further proposes to monitor the training scenario in real time, analyze the moving speeds and change frequencies of the target and non-target elements in the training scenario, calculate the environmental dynamic factor, and the environmental dynamic factor quantifies the severity of environmental changes. The specific steps include: obtaining the initial position information of all target and non-target elements in the training scenario; calculating the displacement vectors of each element in two adjacent frames of images, calculating the moving speed of each element according to the displacement vectors, and counting the number of elements whose speed changes exceed a preset threshold within a unit time as the speed change frequency; calculating the environmental dynamic factor according to the moving speeds and speed change frequencies of each element, and the weights of the moving speed and speed change frequency of the target element are higher than those of the non-target element. The formula for calculating the environmental dynamic factor is D = Σ(αi * Vi + βi * Fi), where Vi is the moving speed of the i-th element, Fi is the speed change frequency of the i-th element, αi is the weight coefficient of the moving speed, and βi is the weight coefficient of the speed change frequency, and the weight coefficients of the target elements are higher than those of the non-target elements.
[0065] Among them, the initial position information is obtained through the multi-modal image recognition module and includes the coordinate information of the target and non-target elements. The displacement vector is calculated by the coordinate difference of the same element in adjacent frames of images, and the moving speed is obtained by dividing the displacement vector by the time interval. The speed change frequency is determined by counting the number of times the speed change amount exceeds the set threshold within a unit time. When setting the weight coefficients, the value ranges of αi and βi for the target elements can be 0.5 - 0.7, and for the non-target elements can be 0.1 - 0.3. For example, when the target element is a virtual enemy unit, αi is set to 0.6 and βi is set to 0.65, and the speed change threshold is set to 0.5 - 1.2 meters per second according to the scene complexity, and preferably 0.8 meters in the urban combat environment.
[0066] Specifically, the initial position information provides a reference coordinate system for subsequent dynamic parameter calculations to ensure the continuity of the motion trajectory analysis of each element. The accurate calculation of the displacement vector depends on the image sampling frequency. The moving speed is calculated using a sliding window algorithm to avoid interference from instantaneous jitters. The statistics of the speed change frequency are realized through a circular buffer, and the number of speed mutations is recorded within the time window. When performing weighted calculations, the weight coefficients of the target elements are dynamically adjusted by a neural network model. When the target type is a high-threat target, αi and βi are automatically increased by 10% - 15%. The environmental dynamic factor is finally input into the confidence adjustment module, and when the D value exceeds the threshold range, an exponential adjustment of the visual feedback intensity parameter is triggered.
[0067] Second, referring to Figure 2 ,the present application also proposes an interaction device based on image recognition for a virtual-real interaction system, and the device includes: The calculation module 210 monitors the training scenario in real time, analyzes the moving speeds and change frequencies of the target and non-target elements in the training scenario, and calculates the environmental dynamic factor, where the environmental dynamic factor quantifies the severity of environmental changes; The setting module 220 adjusts the high threshold and low threshold of the target recognition confidence according to the environmental dynamic factor. The greater the environmental dynamic factor, the higher the confidence threshold; The first adjustment module 230 compares the target recognition confidence output by the multi-modal image recognition with the adjusted high threshold and low threshold to determine the type of visual feedback. If the target recognition confidence is higher than the high threshold, simple visual feedback is adopted. If the target recognition confidence is between the high threshold and the low threshold, medium visual feedback is adopted. If the target recognition confidence is lower than the low threshold, no feedback or weakened visual feedback is adopted; The second adjustment module 240 sets the visual feedback intensity parameter according to the determined type of visual feedback and the type of target. The virtual target feedback intensity is higher than that of the real target; The display module 250 superimposes and displays the set visual feedback in the field of view of the technical object to assist the technical object in target recognition and decision-making.
[0068] Adjusting the target recognition confidence threshold and the type of visual feedback through the environmental dynamic factor, and setting the visual feedback intensity according to the type of target, which can reduce the cognitive load of the technical object and improve the target rapid recognition ability of the technical object in complex and uncertain battlefield environments.
[0069] In addition, in some preferred embodiments, an interaction device based on image recognition proposed in this application can execute any one of the steps in the above method.
[0070] The above are only the embodiments of this application and are not used to limit the protection scope of this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. An interaction method based on image recognition, for a virtual-real interaction system, characterized in that, The steps of the method include: Monitor the training scenario in real time, analyze the moving speeds and change frequencies of the target and non-target elements in the training scenario, calculate the environmental dynamic factor, and the environmental dynamic factor quantifies the severity of environmental changes; Adjust the high threshold and low threshold of the target recognition confidence according to the environmental dynamic factor. The larger the environmental dynamic factor, the higher the confidence threshold; Compare the target recognition confidence output by the multi-modal image recognition with the adjusted high threshold and low threshold to determine the type of visual feedback. If the target recognition confidence is higher than the high threshold, adopt concise visual feedback. If the target recognition confidence is between the high threshold and the low threshold, adopt moderate visual feedback. If the target recognition confidence is lower than the low threshold, adopt no feedback or weakened visual feedback; Set the visual feedback intensity parameter according to the determined type of visual feedback and the type of target. The virtual target feedback intensity is higher than that of the real target; Overlay the set visual feedback on the field of view of the technical object to assist the technical object in target recognition and decision-making.
2. The interactive method based on image recognition according to claim 1, wherein The step of adjusting the high threshold and low threshold of the target recognition confidence according to the environmental dynamic factor includes: Obtain a preset confidence adjustment function. The confidence adjustment function is a piecewise function, and different environmental dynamic factor intervals correspond to different confidence adjustment strategies; Obtain the value of the environmental dynamic factor in real time; Select the corresponding confidence adjustment strategy from the confidence adjustment function according to the interval where the value of the environmental dynamic factor is located. The confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold; Calculate the adjusted high threshold and low threshold of the target recognition confidence respectively according to the selected confidence adjustment strategy. The larger the environmental dynamic factor, the larger the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold, and the adjustment amplitude of the high threshold is greater than the adjustment amplitude of the low threshold.
3. An image recognition-based interaction method according to claim 2, characterized in that, The step of selecting the corresponding confidence adjustment strategy from the confidence adjustment function according to the interval where the value of the environmental dynamic factor is located, and the confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold includes: Obtain the value of the environmental dynamic factor at the current moment; Use a smoothing factor to perform weighted averaging on the value of the current environmental dynamic factor and the smoothed value of the environmental dynamic factor at the previous moment to calculate the smoothed value of the environmental dynamic factor at the current moment. The smoothing factor ranges from 0 to 1. The larger the smoothing factor, the greater the influence of the smoothed value of the environmental dynamic factor by the previous moment; Select the corresponding confidence adjustment strategy from the confidence adjustment function according to the interval where the smoothed value of the environmental dynamic factor is located. The confidence adjustment strategy includes the adjustment amplitude of the high threshold and the adjustment amplitude of the low threshold.
4. An interactive method based on image recognition according to claim 1, characterized in that, After the step of setting the visual feedback intensity parameter according to the determined type of visual feedback and the type of target, and the virtual target feedback intensity is higher than that of the real target, it further includes: Obtain the angle information of the technical object observing the target and the area ratio of the target being occluded. The angle information is provided by the head tracking device, and the occlusion ratio is calculated by the image recognition module. Calculate the target occlusion factor O according to the formula O = α * A + β * (1 - cosθ), where A is the occlusion ratio, θ is the included angle between the observation angle and the front of the target, and α and β are weight coefficients; Obtain the physiological state data of the technical object, where the physiological state data includes eye movement data and heart rate data. Calculate the state factor S of the technical object according to the formula S = γ * E + ε * H, where E is the eye movement data, H is the heart rate data, and γ and ε are weight coefficients. Then calculate the fatigue degree factor F and the attention concentration factor C of the technical object according to the state factor S of the technical object; According to the target occlusion factor O, the fatigue degree factor F, and the attention concentration factor C, use the function I' = I * (1 + k1 * O + k2 * F + k3 * (1 - C)) to adjust the visual feedback intensity parameter I to obtain the adjusted visual feedback intensity parameter I', where I is the initial visual feedback intensity parameter, and k1, k2, and k3 are adjustment coefficients.
5. An interactive method based on image recognition according to claim 4, characterized in that, The step of calculating the fatigue degree factor F and the attention concentration factor C of the technical object according to the state factor S of the technical object includes: Obtain the historical physiological data of the technical object, where the historical physiological data includes historical eye movement data and historical heart rate data, and establish an individual physiological baseline model; According to the current state factor S of the technical object, calculate the deviation value between S and the individual physiological baseline model to obtain the deviated state factor S' of the technical object; Decompose the deviated state factor S' of the technical object into the fatigue degree factor F and the attention concentration factor C. When decomposing, dynamically adjust the weights of the eye movement data and the heart rate data in the calculation of the fatigue degree factor F and the attention concentration factor C according to the correlation between the eye movement data and the heart rate data in the historical data and the fatigue degree and the attention concentration.
6. The interactive method based on image recognition according to claim 1, wherein The step of comparing the target recognition confidence output by the multimodal image recognition with the adjusted high threshold and low threshold to determine the visual feedback type includes: Obtain the current frame target recognition result output by the multimodal image recognition module, where the target recognition result includes the visible light image recognition confidence, the infrared image recognition confidence, and the millimeter wave radar recognition confidence; Calculate the multimodal fusion confidence W according to the formula W = (w1 * C + w2 * I + w3 * R) / (w1 + w2 + w3), where C is the visible light image recognition confidence, I is the infrared image recognition confidence, R is the millimeter wave radar recognition confidence, and w1, w2, and w3 are the weight coefficients of the three modalities respectively. The weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero; If the multimodal fusion confidence W is greater than or equal to the high threshold, determine the visual feedback type as concise visual feedback. If the multimodal fusion confidence W is less than the low threshold, determine the visual feedback type as no feedback or weakened visual feedback. Otherwise, determine the visual feedback type as moderate visual feedback.
7. The interactive method based on image recognition according to claim 6, wherein The step of calculating the multimodal fusion confidence W according to the formula W = (w1 * C + w2 * I + w3 * R) / (w1 + w2 + w3), where C is the visible light image recognition confidence, I is the infrared image recognition confidence, R is the millimeter wave radar recognition confidence, and w1, w2, and w3 are the weight coefficients of the three modalities respectively. The weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero includes: Obtain the acquisition timestamps of each modality image, calculate the time difference between each modality image and the current moment, and perform time decay on the confidence levels of each modality image according to the formula Ct = C * exp(-λ * Δt) to obtain the confidence level Ct of the visible light image recognition after time decay, the confidence level It of the infrared image recognition after time decay, and the confidence level Rt of the millimeter-wave radar recognition after time decay, where C is the confidence level of the visible light image recognition, I is the confidence level of the infrared image recognition, R is the confidence level of the millimeter-wave radar recognition, Δt is the time difference, and λ is the time decay coefficient; Construct a modality interference detector to detect whether each modality image is interfered. If it is detected that a certain modality image is interfered, then according to the degree of interference, use the function wi' = wi * (1 - Di) to adjust the weight coefficient of this modality to obtain the weight coefficients w1', w2', and w3' after interference adjustment, where wi is the original weight coefficient and Di is the interference degree factor with a value range of 0 to 1; Obtain the motion speed of the technical object or target provided by the inertial measurement unit. If the confidence level of the visible light image recognition is lower than the preset threshold and the motion speed exceeds the preset speed threshold, then perform motion blur compensation on the visible light image to improve the confidence level of the visible light image recognition; Calculate the multi-modal fusion confidence level W according to the formula W = (w1' * Ct + w2' * It + w3' * Rt) / (w1' + w2' + w3'), where Ct is the confidence level of the visible light image recognition after time decay, It is the confidence level of the infrared image recognition after time decay, Rt is the confidence level of the millimeter-wave radar recognition after time decay, and w1', w2', and w3' are the weight coefficients of the three modalities after interference adjustment respectively. The weight coefficients are dynamically adjusted according to the historical recognition accuracy. When the recognition result of any modality is missing, the corresponding weight coefficient is set to zero.
8. An image recognition-based interaction method according to claim 7, characterized in that, The step of constructing the modality interference detector, based on the image quality assessment algorithm and the sensor data anomaly detection algorithm, to detect whether each modality image is interfered. If it is detected that a certain modality image is interfered, then according to the degree of interference, use the function wi' = wi * (1 - Di) to adjust the weight coefficient of this modality to obtain the weight coefficients w1', w2', and w3' after interference adjustment, where wi is the original weight coefficient and Di is the interference degree factor with a value range of 0 to 1 includes: Construct an image quality assessment model, adopt a no-reference image quality evaluation algorithm, extract the clarity, contrast, and color saturation feature values of the visible light image, and use a support vector regression machine to fit the mapping relationship between the image quality and the feature values, and output the quality score of the visible light image; Construct a sensor data anomaly detection model, obtain the historical data of the infrared sensor and the millimeter-wave radar, calculate the mean and variance, establish a Gaussian distribution model, input the current moment's infrared sensor and millimeter-wave radar data into the Gaussian distribution model, calculate the probability density value, and if the probability density value is lower than the preset threshold, then determine it as data anomaly; If the quality score of the visible light image is lower than the preset threshold, the visible light image is determined to be interfered with; if the infrared sensor or millimeter wave radar data is determined to be abnormal, the corresponding modality is determined to be interfered with; If it is detected that a certain modality image is disturbed, the weight coefficient of the modality is adjusted according to the output quality score or the output probability density value using the function wi'=wi*(1-Di) to obtain the interference-adjusted weight coefficients w1', w2' and w3', where wi is the original weight coefficient, Di is the interference degree factor, Di is calculated according to the image quality score or probability density value, Di=1-(Qi / Qmax) or Di=1-(Pi / Pmax), Qi is the quality score of the visible light image, Qmax is the maximum quality score of the visible light image, Pi is the probability density value, and Pmax is the maximum probability density value.
9. The interactive method based on image recognition according to claim 1, wherein The steps of real-time monitoring the training scene, analyzing the moving speed and change frequency of the target and non-target elements in the training scene, and calculating the environmental dynamic factors include: Obtain the initial position information of all target and non-target elements in the training scene; Calculate the displacement vector of each element in two adjacent frames of images, calculate the moving speed of each element according to the displacement vector, and count the number of elements whose speed change exceeds a preset threshold per unit time as the speed change frequency; The environmental dynamic factor is calculated according to the moving speed and speed change frequency of each element. The weights of the moving speed and speed change frequency of the target element are higher than those of the non-target elements. The calculation formula of the environmental dynamic factor is D=Σ(αi*Vi+βi*Fi), where Vi is the moving speed of the i-th element, Fi is the speed change frequency of the i-th element, αi is the weight coefficient of the moving speed, and βi is the weight coefficient of the speed change frequency. The weight coefficient of the target element is higher than that of the non-target element.
10. An interaction device based on image recognition, for a virtual-real interaction system, characterized in that, The device includes: The computing module monitors the training scene in real time, analyzes the movement speed and change frequency of the target and non-target elements in the training scene, and calculates the environmental dynamic factor, which quantifies the severity of environmental changes; The setting module adjusts the high and low thresholds of target recognition confidence according to the dynamic factors of the environment. The larger the dynamic factors of the environment, the higher the confidence threshold. The first adjustment module compares the target recognition confidence output by the multimodal image recognition with the adjusted high threshold and low threshold, and determines the type of visual feedback. If the target recognition confidence is higher than the high threshold, simple visual feedback is adopted; if the target recognition confidence is between the high threshold and the low threshold, moderate visual feedback is adopted; if the target recognition confidence is lower than the low threshold, no feedback or weakened visual feedback is adopted; The second adjustment module sets the visual feedback intensity parameter according to the determined visual feedback type and target type, and the feedback intensity of the virtual target is higher than that of the real target; The display module superimposes and displays the set visual feedback in the field of view of the technical object to assist the technical object in target recognition and decision-making.
Citation Information
Cited By
Unmanned aerial vehicle target detection method and device based on multi-modal detection information fusion
CN120891493A
Unmanned aerial vehicle target detection method and device based on multi-modal detection information fusion
CN120891493B
Container operation behavior identification and analysis method and system based on AI
CN121148018A
Methane leakage laser detection method and system based on industrial vision
CN121558650A
Video picture optimization method and device, equipment and medium
CN121644889A