Visual detection method and device based on artificial intelligence, equipment and medium
By constructing a multimodal physical field and visual perception technology, the expected behavior spectrum of the detected target in the logistics and warehousing scenario is generated, which solves the problem of lack of correlation between safety and efficiency analysis, and realizes unified quantitative assessment and risk warning optimization of safety risks and collaborative efficiency.
Patent Information
- Application Number
- CN202610192593.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-19
AI Technical Summary
The lack of inherent connection between safety inspection and efficiency analysis in logistics and warehousing scenarios leads to congestion in workflows and increased safety risks.
By constructing a multimodal physical field, including an attention field, a task potential gradient field, a path probability field, and a robot action field, and combining it with visual perception to obtain the real-time behavior of the detected target, a expected behavior spectrum is generated and a comprehensive risk entropy value is calculated, thereby achieving a unified assessment of safety and efficiency.
It enables a unified quantitative assessment of the safety risks and collaboration efficiency risks of personnel behavior in logistics and warehousing scenarios, providing a more comprehensive and refined basis for risk warning and decision optimization.
Smart Images

Figure CN122067196A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of visual inspection, and in particular relates to a visual inspection method, device, equipment and medium based on artificial intelligence. Background Technology
[0002] With the popularization of smart office and mobile robot applications, logistics warehousing has become a dynamic and complex system that tightly couples people, robots, task flows and physical space.
[0003] Currently, in logistics and warehousing scenarios, security detection (such as intrusion detection) and efficiency analysis (such as workstation utilization) are usually performed by two independent systems, lacking a unified understanding of the intrinsic relationship between "safety" and "efficiency." For example, a layout that leads to congested workflows may also increase the safety risk of collisions between personnel and robots. Therefore, there is an urgent need for an artificial intelligence-based visual inspection method, device, equipment, and medium. Summary of the Invention
[0004] This application provides an artificial intelligence-based visual inspection method, device, equipment, and medium that can intrinsically link the traditionally separate evaluation dimensions of safety and efficiency through a unified model of "multimodal physical field." A single human action can simultaneously produce effects in both the safety and efficiency fields, thereby achieving coupled evaluation, which is more in line with the nature of complex systems.
[0005] On the one hand, this application provides a visual detection method based on artificial intelligence, the method comprising: Acquire field energy data in the target scenario, wherein the field energy data is used to characterize task behavior data in the target scenario; Based on the field energy data, a multimodal physical field is constructed, which includes at least two of the following: attention field, task potential gradient field, path probability field, and robot action field. For the target to be detected in the target scene, based on the target's identity information, current task status and multimodal physical field parameters of the target's location, an expected behavior spectrum of the target at the current moment is generated. The expected behavior spectrum is used to describe the probability distribution of the target's behavior. The real-time behavior vector of the detected target is obtained based on visual perception, and the real-time behavior vector includes a sequence of position, velocity, orientation, and posture. The original real-time behavior vector is substituted into the current multimodal physical field to generate a real-time behavior field effect vector. Based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum, the comprehensive risk entropy value of the detected target is calculated and output. The comprehensive risk entropy value is used to describe the quantitative assessment of security risk and collaboration efficiency risk.
[0006] Optionally, constructing a multimodal physical field based on the field energy data includes: Acquire visual images of the target scene; The visual image is used to estimate the action state using a preset neural network model to obtain a set of suspected gaze point coordinates. The set of suspected gaze point coordinates is associated and matched with the location of the preset information source entity to obtain the dominant information source at each moment; Using the spatial location of the dominant information source as the field source, different intensity weights are assigned according to the semantic type of the information source entity, and the attenuated field strengths from any point in the space to all effective field sources are superimposed using a radial basis function network to generate the attention field.
[0007] Optionally, the field energy data includes a task title, a list of participants, pre-deadline task dependencies, and task tags. Constructing a multimodal physical field based on the field energy data includes: Based on the task title, the list of participants, the dependencies of the tasks before the deadline, and the task tags, a dynamic task relationship graph is constructed, where nodes represent tasks and edges represent overlapping relationships between personnel for task keys. The dynamic task relationship graph is analyzed using graph attention networks to calculate the global urgency score and collaboration heat score for each task node. The real-time location of the detected target is used as the anchor point of the task's potential energy in physical space. The potential energy intensity of the anchor point is calculated based on the global urgency score and the cooperation heat score. A potential energy diffusion algorithm is used to diffuse the potential energy intensity of each anchor point to the entire target scene according to spatial accessibility, forming a continuous task potential energy gradient field.
[0008] Optionally, constructing a multimodal physical field based on the field energy data includes: An anisotropic Gaussian risk field is constructed with the robot's real-time planned path as the central axis and the robot's physical outline and safe braking distance as parameters. A collaborative attraction field is constructed centered on the robot's preset collaboration points; The Gaussian risk field and the cooperative attraction field are dynamically modulated to generate the robot action field.
[0009] Optionally, generating the expected behavior spectrum of the detected target at the current moment based on the target's identity information, current task status, and multimodal physical field parameters of the target's location includes: Using the identity information and current task status of the detected target as the query entry point, relevant nodes and subgraphs are activated in the preset behavioral knowledge graph; The multimodal sociophysical field strength of the detected target's location is transformed into the contextual conditions of the preset behavioral knowledge graph. Conditional probability propagation and reasoning are then performed in the preset behavioral knowledge graph to obtain the expected behavioral spectrum.
[0010] Optionally, calculating and outputting the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum includes: Calculate the distance between the real-time behavior field effect vector and the expected behavior spectrum in the feature space to obtain the initial deviation. Based on the current multimodal physical field state, several preset scenarios are enumerated, and the hypothetical behavioral field effect vector that the detected target should produce under each scenario is inferred. Calculate the deviation of each hypothetical behavior field effect vector from the expected behavior spectrum, and take the minimum value as the reasonable deviation baseline; The initial deviation is compared with the reasonable deviation baseline. If the former is significantly greater than the latter, a high safety risk is determined, and the safety risk entropy is calculated based on the degree of deviation. Otherwise, the initial deviation is reduced before calculating the safety risk entropy.
[0011] Optionally, after calculating and outputting the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum, the method further includes: Based on the multimodal physical field state of the previous moment, the field perturbation caused by the real-time behavior of the detected target is calculated. The field perturbation includes the redistribution of the local probability density of the movement probability field or the redirection of high-intensity regions in the attention field. Using the multimodal physical field after the field perturbation as input, the expected behavior spectrum of the affected entity is recalculated, and the degree of spectrum change is evaluated. The risk entropy of cooperation efficiency is obtained by weighted aggregation of the spectral changes of all affected entities. The comprehensive risk entropy value is corrected based on the collaborative efficiency risk entropy to obtain the final risk entropy value.
[0012] On the other hand, this application provides a visual inspection device based on artificial intelligence, the device comprising: The first acquisition module allows the user to acquire field energy data in a target scenario, where the field energy data is used to characterize task behavior data in the target scenario. A construction module is used to construct a multimodal physical field based on the field energy data. The multimodal physical field includes at least two of the following: attention field, task potential gradient field, path probability field, and robot action field. The first generation module is used to generate the expected behavior spectrum of the detected target at the current moment based on the identity information, current task status and multimodal physical field parameters of the detected target's location in the target scene. The expected behavior spectrum is used to describe the behavior probability distribution of the detected target. The second acquisition module is further configured to acquire the real-time behavior raw vector of the detected target based on visual perception, wherein the real-time behavior raw vector includes a sequence of position, velocity, orientation, and posture. The second generation module is also used to substitute the original real-time behavior vector into the current multimodal physical field to generate a real-time behavior field effect vector. The calculation module is used to calculate and output the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum. The comprehensive risk entropy value is used to describe the quantitative assessment of security risk and collaboration efficiency risk.
[0013] In another aspect, this application provides an electronic device, the device comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the artificial intelligence-based visual detection method as described in the first aspect.
[0014] In another aspect, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the artificial intelligence-based visual detection method as described in the first aspect.
[0015] In another aspect, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the artificial intelligence-based visual detection method as described in the first aspect.
[0016] This application discloses an AI-based visual inspection method, device, equipment, and medium. By acquiring field energy data in a target scene, it determines the task data of the target at that moment and the required behavior to perform that task. Then, it constructs a multimodal physical field using the field energy data, analyzing from multiple perspectives whether the target exhibits dangerous or abnormal behavior. By using the target's identity information, current task status, and multimodal physical field parameters of its location, it generates the target's expected behavior spectrum at the current moment—the probability of the target's behavior within a preset future timeframe. Finally, it calculates the target's comprehensive risk entropy value using the real-time behavior field effect vector and the expected behavior spectrum, ensuring user safety and generating and evaluating the deviation between expected and actual behavior. This achieves a unified quantitative assessment of the safety and collaborative efficiency risks of personnel behavior in logistics and warehousing scenarios, overcoming the limitations of traditional solutions where safety and efficiency analyses are independent and lack correlation. This provides a more comprehensive and refined basis for risk warning and decision optimization in dynamic and complex systems. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic flowchart of an artificial intelligence-based visual inspection method provided in one embodiment of this application; Figure 2 This is a schematic diagram of the structure of an artificial intelligence-based visual inspection device provided in another embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0019] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0020] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0021] To address the problems of existing technologies, this application provides an artificial intelligence-based visual inspection method, apparatus, device, and medium. In this application, by acquiring field energy data in a target scene, the task data of the target being detected at that moment and the required behavior for performing that task can be determined. Then, a multimodal physical field is constructed using the field energy data to analyze from multiple perspectives whether the detected target exhibits dangerous or abnormal behavior. By using the target's identity information, current task status, and multimodal physical field parameters of the target's location, the expected behavior spectrum of the detected target at the current moment is generated, representing the probability of the target's behavior within a preset future timeframe. Finally, the comprehensive risk entropy value of the detected target is calculated using the real-time behavior field effect vector and the expected behavior spectrum, ensuring user safety and generating and evaluating the deviation between expected and actual behavior. This achieves a unified quantitative assessment of the safety risks and collaborative efficiency risks of personnel behavior in logistics and warehousing scenarios, overcoming the limitations of traditional solutions where safety and efficiency analyses are independent and lack correlation. This provides a more comprehensive and refined basis for risk warning and decision optimization in dynamic and complex systems.
[0022] The following section first introduces the AI-based visual detection method provided in the embodiments of this application.
[0023] Figure 1 A flowchart illustrating an artificial intelligence-based visual inspection method according to an embodiment of this application is shown. Figure 1 As shown, the AI-based visual inspection method may include S101-S106: S101, acquire field energy data under the target scenario.
[0024] In this embodiment of the application, field energy data is used to characterize task behavior data in the target scenario. Field energy data may include information related to task execution, personnel activities and environmental status, such as task priority, personnel skills, equipment operating status, etc., to provide basic input for subsequent construction of physical field.
[0025] S102, based on field energy data, constructs a multimodal physical field.
[0026] In this embodiment, the multimodal physical field includes at least two of the following: attention field, task potential gradient field, movement probability field, and robot action field. It is an abstract spatial model composed of the superposition or combination of multiple different modal physical fields. This physical field can comprehensively reflect the interaction and potential influence between elements such as personnel, robots, and tasks in the target scene.
[0027] As an example, an attention field describes the degree to which different areas in a target scene attract the attention of a person or robot. The field strength is usually associated with the location and importance of the information source entity (such as a display screen or control panel) and the person's gaze behavior.
[0028] The task potential gradient field is used to characterize the relationship between different locations in the target scene and the current task completion or urgency. The field strength usually indicates the tendency of a person or robot to complete the task, and its gradient direction can guide the actor to move towards the task target.
[0029] The path probability field is used to describe the probability distribution of the movement path of people or robots in a target scene. The field strength reflects the probability of a specific area being traversed or stopped, and can be used to predict potential congestion or conflict areas.
[0030] The robot action field describes the impact of a robot on its surrounding environment and people. This field can include the robot's risk area and cooperative attraction area, guiding people to interact with the robot safely and effectively.
[0031] S103, for the target object in the target scene, based on the target object's identity information, current task status and multimodal physical field parameters of the target object's location, generate the expected behavior spectrum of the target object at the current moment.
[0032] In this embodiment of the application, the expected behavior spectrum is used to describe the probability distribution of the behavior of the detected target; the spectrum is a set of predictions based on the target's identity information, current task state, and multimodal physical field parameters of its location, predicting the various behaviors that the detected target may take at the current moment and their probabilities.
[0033] S104, real-time behavior vector of the detected target is obtained based on visual perception.
[0034] In this embodiment, the real-time behavior raw vector includes a sequence of position, velocity, orientation, and attitude; used to describe the actual behavior state of the detected target at a certain moment.
[0035] S105 substitutes the original real-time behavior vector into the current multimodal physical field to generate the real-time behavior field effect vector.
[0036] The real-time behavior field effect vector represents the impact or response of the detected target in the current multimodal physical field after its original real-time behavior vector is substituted into the physical field. This vector reflects the interaction between the actual behavior and the environmental field.
[0037] S106, based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum, calculate and output the comprehensive risk entropy value of the detected target.
[0038] In this application embodiment, the comprehensive risk entropy value is used to describe the quantitative assessment of security risks and collaboration efficiency risks. This entropy value provides a unified indicator to measure potential dangers and collaboration efficiency losses in the target scenario by calculating the degree of deviation between real-time behavior and expected behavior and the impact of behavior on collaboration efficiency.
[0039] In this embodiment, information such as task details, personnel schedules, and equipment status can be obtained manually. Alternatively, various data related to task execution can be passively collected by deploying sensors, such as RFID tag readers or barcode scanners. Furthermore, static task descriptions and personnel role assignments can be directly retrieved from a pre-set database or task management system.
[0040] Secondly, different types of fields, such as attention fields and task potential fields, can be linearly superimposed to form a composite field. Alternatively, fixed weights can be pre-set to weighted average the fields of different modes, thereby constructing the multimodal physical field. In another implementation, based on empirical rules, two or more physical fields can be selectively activated and combined according to specific keywords or numerical thresholds in the field energy data.
[0041] Furthermore, a rule-based expert system can be pre-defined. Based on the target's identity information and current task state, it searches for corresponding behavioral patterns in a fixed rule base and makes simple probability adjustments based on the field strength at the target's location. Alternatively, a statistical model-based approach can be used. A simple classifier is trained using historical behavioral data to predict the behavior category based on input information and assign an average probability. Additionally, a simple behavioral state machine can be maintained, searching within a predefined state transition graph based on identity and task state, and making limited probability adjustments based on the field strength.
[0042] Subsequently, a monocular camera can be used for target detection and tracking. Image processing algorithms can be used to estimate the target's position, velocity, and orientation, while pose information is roughly estimated using a simple skeleton model to obtain the original vector of the target's behavior. In another implementation, multiple cameras with fixed viewing angles can be deployed. The target's position can be obtained using simple triangulation, and its velocity and orientation can be calculated from the position changes in consecutive frames. Pose information is obtained by matching a preset pose template. Alternatively, traditional vision algorithms based on background subtraction or optical flow can be used to identify moving targets and extract their center point position. Velocity and orientation are obtained through inter-frame subtraction, while pose information is not extracted in detail.
[0043] Next, the components of the original real-time behavior vector, such as position information, can be used as query points to directly read the multimodal physical field parameters at that location as the field effect vector. Alternatively, basic vector operations such as dot product or cross product can be performed between the original real-time behavior vector and the multimodal physical field to obtain a scalar or vector as the field effect. In another implementation, a simple mapping function can be predefined to compare a specific dimension of the original real-time behavior vector, such as the velocity direction, with a specific mode of the multimodal physical field, such as the task potential gradient direction, to generate a field effect vector representing the degree of matching.
[0044] Finally, the Euclidean distance between the real-time behavior field-effect vector and the behavior with the highest probability in the expected behavior spectrum can be simply calculated as a direct measure of risk. Alternatively, a threshold can be preset; when the matching degree between the real-time behavior field-effect vector and a certain behavior in the expected behavior spectrum is lower than this threshold, a fixed high-risk value can be directly output. Furthermore, a simple similarity calculation, such as cosine similarity, can be performed between the real-time behavior field-effect vector and the expected behavior spectrum, and then mapped to a risk entropy value using a linear function.
[0045] In this embodiment, by acquiring field energy data in the target scenario, the task data of the target at that moment and the behavior required to perform the task can be determined. Then, a multimodal physical field is constructed using the field energy data to analyze whether the target exhibits dangerous or abnormal behavior from multiple perspectives. By using the target's identity information, current task status, and multimodal physical field parameters of its location, the expected behavior spectrum of the target at the current moment is generated, representing the probability of the target's behavior within a preset future timeframe. Finally, the comprehensive risk entropy value of the target is calculated using the real-time behavior field effect vector and the expected behavior spectrum, ensuring user safety and generating and evaluating the deviation between expected and actual behavior. This achieves a unified quantitative assessment of the safety and collaboration efficiency risks of personnel behavior in logistics and warehousing scenarios, overcoming the limitations of traditional solutions where safety and efficiency analyses are independent and lack correlation. This provides a more comprehensive and refined basis for risk warning and decision optimization in dynamic and complex systems.
[0046] In some other embodiments, S102 may include: Acquire visual images of the target scene; By using a pre-defined neural network model to estimate the action state of visual images, a set of possible gaze point coordinates is obtained. The set of suspected gaze point coordinates is associated and matched with the location of the preset information source entity to obtain the dominant information source at each moment; Using the spatial location of the dominant information source as the field source, different intensity weights are assigned according to the semantic type of the information source entity. The radial basis function network is used to calculate the superposition of the attenuated field strengths from any point in the space to all effective field sources to generate the attention field.
[0047] In this embodiment, when constructing the attention field, it is first necessary to acquire visual images of the target scene. This is typically achieved by deploying one or more visual sensors (e.g., RGB cameras, depth cameras, or infrared cameras) in the target scene. These sensors can capture visual information in the scene in real time, including human activity, objects in the environment, and lighting conditions. The acquired visual images serve as the foundational data for subsequent analysis.
[0048] Subsequently, a pre-defined neural network model is used to estimate the action state of the visual images, resulting in a set of potential gaze points. This pre-defined neural network model is a deep learning model trained on a large amount of data, capable of recognizing information such as a person's body posture, head orientation, and eye features (if the image resolution allows) from the input visual image. By comprehensively analyzing this information, the model can infer the person's visual focus or attention direction, thereby outputting the spatial coordinates of one or more potential gaze points in the scene. For example, if a person's head and torso are both facing a specific area, that area may be identified as a potential gaze point.
[0049] Next, the set of suspected gaze points is correlated and matched with the locations of predefined information source entities to obtain the dominant information source at each moment. In the target scene, some entities with specific meanings are usually predefined, such as workbenches, tools, monitors, other collaborators, or potentially hazardous areas. These entities are called information source entities, and their precise spatial locations are known. By comparing the spatial relationships between suspected gaze points and these information source entities (e.g., calculating distance, determining whether they are within entity boundaries), it can be determined which information source entity a person's attention is most likely to focus on. At a given moment, the information source entity with the highest or closest match to the suspected gaze point is identified as the dominant information source.
[0050] Finally, using the spatial location of the dominant information source as the field source, different intensity weights are assigned based on the semantic type of the information source entity. A radial basis function network is then used to calculate the superposition of the attenuated field strengths from any point in space to all valid field sources, generating the attention field. Specifically, the spatial location of each identified dominant information source is considered the "field source" of the attention field. A specific intensity weight is assigned to the information source entity based on its intrinsic attributes or importance (i.e., "semantic type," for example, the semantic type of a safety warning sign may be more important than that of a common tool). For example, an information source entity representing a high-risk area may be assigned a higher intensity weight. Subsequently, a radial basis function network (e.g., using Gaussian radial basis functions) is used to simulate the attenuation process of attention intensity in space. For any spatial point in the scene, its attention field strength is the superposition of the attenuated field strengths generated by all currently valid dominant information sources at that point. This superposition mechanism reflects the combined influence of multiple information sources on human attention, forming a continuous attention field with gradient changes, where regions with higher field strengths indicate a greater likelihood of concentrated human attention.
[0051] When constructing an attention field, this method can accurately identify and quantify a person's attention focus based on visual images and transform it into a spatially continuous attention field. This approach overcomes the limitation of traditional behavior analysis in capturing a person's internal cognitive state, enabling the multimodal physical field to more accurately reflect a person's actual focus and potential intentions. Given that attention is a key precursor to behavioral decisions, a precisely constructed attention field can significantly improve the prediction accuracy of expected behavior spectra, thereby making the calculation of comprehensive risk entropy values more reliable and instructive, and effectively reducing the risks of safety accidents and low collaboration efficiency caused by attention bias.
[0052] In other embodiments, the field energy data includes a task title, a list of participants, deadline-preceding task dependencies, and task tags. S102 may also include: Based on the task title, list of participants, deadline and pre-task dependencies, and task tags, a dynamic task relationship graph is constructed, where nodes represent tasks and edges represent overlapping relationships between personnel for task keys. Graph attention network analysis is used to analyze dynamic task relationship graphs and calculate global urgency score and collaboration heat score for each task node. The real-time location of the detected target is used as the anchor point of the task's potential energy in physical space. The potential energy intensity of the anchor point is calculated based on the global urgency score and the cooperation heat score. A potential energy diffusion algorithm is used to diffuse the potential energy intensity of each anchor point to the entire target scene according to spatial accessibility, forming a continuous task potential energy gradient field.
[0053] In this embodiment, the dynamic task relationship graph is a graph structure used to represent the complex relationships between tasks in a target scenario. Nodes represent specific tasks, such as "material handling" or "equipment debugging," while edges represent overlapping personnel relationships between tasks; that is, two tasks may involve the same personnel or require the collaboration of the same personnel. This graph can be constructed by parsing the task titles, participant lists, deadline dependencies, and task tags from the field energy data. For example, if the participant lists of two tasks overlap, or if one task is a prerequisite for another, an edge can be established between them. The graph is "dynamic," meaning its structure and attributes can be updated in real time with task status, personnel allocation, or the passage of time.
[0054] Subsequently, a graph attention network (GAT) is used to analyze the dynamic task relationship graph, calculating a global urgency score and a collaboration heat score for each task node. A graph attention network (GAT) is a deep learning model particularly suitable for processing graph-structured data. It learns the importance weights between a node and its neighbors by applying an attention mechanism to nodes in the graph, thereby aggregating neighbor information to update the node representation. Here, GAT is used to analyze the dynamic task relationship graph to capture the deep connections between tasks. The global urgency score reflects the importance and time pressure of a task within the overall task hierarchy; for example, tasks nearing deadlines or preceding multiple key tasks will have higher urgency scores. The collaboration heat score quantifies the degree of collaboration required for a task; for example, tasks involving multiple departments or individuals will have higher collaboration heat scores. These scores are calculated based on the results of GAT's learning and aggregation of features of task nodes and their neighboring nodes (i.e., related tasks).
[0055] Next, when a detected target (e.g., a worker) is performing or associated with a task, the real-time location of that target is considered the anchor point of the task's potential energy. Potential energy intensity is an indicator of the task's importance and attractiveness at that anchor point. It comprehensively considers the task's global urgency score and collaboration heat score. For example, a task with high urgency and high collaboration heat will have a higher potential energy intensity at its anchor point, indicating that the area has a stronger attraction or repulsion for human behavior. The calculation method can be a weighted sum of the two scores or a combination using other functions. As an example, the anchor point is a specific location mapping of the task's potential energy in physical space.
[0056] Finally, the potential energy diffusion algorithm is a method that extends discrete potential energy anchor point information to the entire continuous space, simulating the propagation process of potential energy in a physical field. Spatial accessibility refers to the ease with which one can reach another from one location, taking into account factors such as physical obstacles, distance, and pathways. This algorithm uses the potential energy intensity of each anchor point as a source and diffuses the potential energy value into the surrounding space according to a preset attenuation function and spatial accessibility rules. For example, the farther away from the anchor point and the worse the accessibility, the faster the potential energy value attenuates. Ultimately, every point in the entire target scene is assigned a potential energy value, thus forming a continuous and smooth task potential energy gradient field. This field can intuitively represent the intensity of the "energy" or "attraction" associated with the task in different regions.
[0057] By introducing richer field energy data, such as task titles, participant lists, deadline dependencies, and task tags, this application can construct a dynamic task relationship graph, thus more comprehensively depicting the complex relationships between tasks. In-depth analysis of this graph using a graph attention network allows for the calculation of a global urgency score and a collaboration heat score for each task node, quantifying the importance of tasks and their need for collaboration. Furthermore, the real-time location of the detected target is used as the anchor point for the task potential energy in the physical space, and these scores are combined to calculate the potential energy intensity, tightly linking abstract task attributes with human behavior in the physical space. Finally, a potential energy diffusion algorithm smoothly diffuses these discrete potential energy intensities throughout the target scene, forming a continuous task potential energy gradient field that accurately reflects the distribution of tasks in the physical space and their impact on human behavior. This effectively solves the problem of traditional methods struggling to accurately model complex task relationships and their potential energy distribution in the physical space, providing more refined and accurate physical field support for subsequent generation of expected behavior spectra and risk assessment of human behavior, thereby improving the accuracy and reliability of visual detection.
[0058] In some other embodiments, S102 may further include: An anisotropic Gaussian risk field is constructed with the robot's real-time planned path as the central axis and the robot's physical outline and safe braking distance as parameters. A collaborative attraction field is constructed centered on the robot's preset collaboration points; The robot's action field is generated by dynamically modulating the Gaussian risk field and the cooperative attraction field.
[0059] In this embodiment, in complex human-robot collaboration scenarios involving robots, relying solely on these physical fields may not be sufficient to capture the comprehensive impact of robots on human behavior, nor may it be easy to accurately assess the potential risks or collaboration opportunities caused by robot movement and tasks, which may lead to inaccurate expectations of human behavior or incomplete risk assessments.
[0060] To improve the accuracy of predicting human behavior, a robot action field can be constructed to ensure robot-human collaboration. Specifically, when constructing an anisotropic Gaussian risk field, the aim is to quantify the potential safety threats posed by robot movement to surrounding personnel. Traditional risk fields may use isotropic models, but the danger of robot movement is often related to its direction of travel, speed, and physical outline, thus requiring anisotropy. This risk field can be constructed with the robot's real-time planned path as the central axis, and the robot's physical outline and safe braking distance as parameters. The central axis defines the main direction of risk distribution, while the physical outline and safe braking distance determine the range and intensity of the risk field. For example, a Gaussian function or other decay functions can be used to describe the attenuation of risk intensity with distance from the central axis, and anisotropy can be achieved by adjusting the function parameters or covariance matrix, resulting in different distribution characteristics of risk along the robot's movement direction.
[0061] When constructing a collaborative attraction field, this field aims to guide personnel towards the robot's intended collaborative location, or to represent the robot's intention to collaborate with personnel. It reflects the area where the robot hopes to interact with personnel or jointly complete tasks. This attraction field can be constructed centered on the robot's pre-defined collaboration points. These collaboration points can be specific spatial locations where the robot needs to deliver or receive items, jointly operate tools, or exchange information. The intensity of the attraction can be set according to the importance and urgency of the collaborative task and the distance to the collaboration point, typically exhibiting a characteristic of stronger attraction closer to the collaboration point and gradually decreasing outwards.
[0062] When dynamically modulating the Gaussian risk field and the cooperative attraction field to generate the robot's action field, the robot action field is a comprehensive reflection of these two fields. It fully reflects the robot's combined impact on the surrounding environment and personnel at a specific moment, including potential hazards and cooperative opportunities. Dynamic modulation means that these two fields are not simply superimposed, but adjusted according to the real-time context. The Gaussian risk field and the cooperative attraction field can be dynamically modulated based on various factors such as the robot's current task state (e.g., performing a hazardous transport task or a cooperative delivery task), the density of surrounding personnel, the relative speed and distance between personnel and the robot, and the task priority. Modulation methods can include weighted summation, nonlinear fusion, or switching based on decision logic. The final generated robot action field is a comprehensive spatial field whose intensity and direction indicate the behavioral tendencies of personnel in the presence of the robot, thus providing more comprehensive environmental information for subsequent behavior prediction and risk assessment.
[0063] In this embodiment, by constructing a robot action field, the potential safety threats posed by the robot's real-time path planning, physical outline, and safe braking distance can be accurately quantified. This transforms risk assessment from a simple distance judgment into a refined assessment considering the robot's motion characteristics and the direction of danger. Simultaneously, by constructing a cooperative attraction field, the robot's preset cooperation points are clearly defined, providing clear cooperation signals and behavioral guidance for personnel. Furthermore, by dynamically modulating these two fields, safety risks and cooperation opportunities can be flexibly balanced according to the robot's real-time state and task requirements, generating a comprehensive robot action field. This robot action field can more comprehensively and accurately reflect the robot's overall impact on human behavior in a shared space. This allows the system to fully consider robot factors when generating the expected human behavior spectrum, significantly improving the accuracy of predicting human behavior in human-robot collaboration scenarios. It also provides a more refined and comprehensive input for calculating the comprehensive risk entropy value, effectively reducing safety risks in human-robot collaboration and optimizing collaboration efficiency.
[0064] In some other embodiments, S103 may include: Using the target's identity information and current task status as the query entry point, relevant nodes and subgraphs are activated in the preset behavioral knowledge graph; The multimodal sociophysical field strength of the detected target's location is transformed into contextual conditions of a pre-defined behavioral knowledge graph. Conditional probability propagation and reasoning are then performed within the pre-defined behavioral knowledge graph to obtain the expected behavioral spectrum.
[0065] Specifically, the target's identity information can include its unique employee ID, preset role type (e.g., "operator," "inspector"), skill level (e.g., "beginner," "advanced"), and historical behavior records. The current task status can include the task ID the target is currently executing, the task type, the task's stage (e.g., "started," "in progress," "completed"), and the task's priority. This information serves as a query entry point for retrieval and matching within a preset behavioral knowledge graph. The preset behavioral knowledge graph is a structured knowledge base storing a wealth of knowledge about personnel behavior patterns, environmental contexts, task processes, and the relationships between them. Upon receiving identity information and task status, the system uses graph traversal or graph matching algorithms to search for nodes directly related to this information in the knowledge graph and expands along predefined relationships (e.g., "possible behaviors," "in...context," "requires...skills"), thereby activating a behavioral subgraph highly relevant to the current context. For example, if the target's identity is "new employee" and the current task is "operating device A," the activated subgraph might contain knowledge such as "common erroneous behavior patterns of new employees operating device A" and "standard procedures for operating device A."
[0066] Building upon this, if the target is located in a region with high attention field strength (e.g., near an important information source), this high field strength can be used as a condition to increase the activation weight of nodes in the knowledge graph related to behaviors such as "information acquisition" and "observation." Similarly, if the target is located in a region with high task potential gradient (e.g., near the next task point), this field strength can be used as a condition to increase the activation weight of nodes in the knowledge graph related to behaviors such as "moving to a task point" and "performing a task." This transformation can be achieved through pre-defined mapping rules, fuzzy logic systems, or small neural networks, mapping continuous field strength values to weight factors or probability correction values on specific relationships or nodes in the knowledge graph. Subsequently, on the activated subgraph, conditional probability propagation and inference are performed using probabilistic graphical models (such as Bayesian networks) or rule-based inference engines. During this process, the transformed contextual conditions serve as evidence or prior probabilities, influencing the calculation of the posterior probability of each behavioral node in the knowledge graph under the current contextual conditions. Through message passing algorithms or logical reasoning rules, the system comprehensively considers the causal, temporal, and mutually exclusive relationships between behaviors, and finally calculates the probability of various behaviors occurring, thus forming a distribution describing the various behaviors that the detected target may take at the current moment and their probability of occurrence, i.e., the expected behavior spectrum.
[0067] By using the target's identity information and current task status as query entry points, and activating relevant nodes and subgraphs in a pre-defined behavioral knowledge graph, the system can quickly focus on behavioral knowledge most relevant to the current context. Furthermore, the multimodal sociophysical field strength of the target's location is transformed into contextual conditions for the knowledge graph, and conditional probability propagation and reasoning are performed. This ensures that the generation of the expected behavioral spectrum fully considers the dynamic influence of the environment and the inherent logic of individual behavior. This not only improves the accuracy of the expected behavioral spectrum generation, enabling it to more accurately describe the behavioral probability distribution of the target, but also provides a more reliable benchmark for the subsequent calculation of the comprehensive risk entropy value, thereby enhancing the overall accuracy and robustness of the visual detection method.
[0068] In some other embodiments, S106 may include: Calculate the distance between the real-time behavior field effect vector and the expected behavior spectrum in the feature space to obtain the initial deviation. Based on the current multimodal physical field state, several preset scenarios are enumerated, and the hypothetical behavioral field effect vector that the detected target should produce under each scenario is inferred. Calculate the deviation between the hypothetical behavior field effect vector and the expected behavior spectrum for each hypothesis, and take the minimum value as the reasonable deviation baseline; The initial deviation is compared with the reasonable deviation from the baseline. If the former is significantly greater than the latter, a high safety risk is determined, and the safety risk entropy is calculated based on the degree of deviation. Otherwise, the initial deviation is reduced before calculating the safety risk entropy.
[0069] In this embodiment, the real-time behavior field effect vector and the expected behavior spectrum are high-dimensional data. Their distance in the feature space can quantify the difference between the actual behavior and the expected behavior. This distance can be calculated using various metrics, such as Euclidean distance, cosine similarity, Mahalanobis distance, Kullback-Leibler divergence, or Jensen-Shannon divergence. The specific choice depends on the representation of the vector and spectrum, as well as the definition of "distance." For example, if the expected behavior spectrum is represented as a probability distribution, then Kullback or Jensen-Shannon divergence is more suitable. This distance value is the initial deviation, directly reflecting the difference between the current actual behavior and the ideal expected behavior.
[0070] Building upon this, considering other reasonable or acceptable behavioral patterns besides the expected behavior under the current physical state, the "preset scenario" can be a series of predefined alternative behavioral paths or states related to the current task and environment. For example, in a robot collaboration scenario, when a robot changes its path, a person may need to take different actions such as avoidance, following, or waiting. Inferring the "hypothetical behavioral field effect vector" can be accomplished through a pre-trained behavior prediction model, a rule-based expert system, or a simulator. This model or system takes the current multimodal physical state and the preset scenario as input and outputs a behavioral field effect vector that the detected target may produce under that scenario. For example, conditional reasoning can be performed using a behavioral knowledge graph to simulate behavioral responses under different scenarios.
[0071] Subsequently, for each enumerated preset scenario, the deviation between its corresponding hypothetical behavior field effect vector and the expected behavior spectrum is calculated in the same way as the initial deviation. By calculating all these hypothetical deviations and selecting the minimum value, a minimum deviation that the detected target may produce within a reasonable behavior range under the current multimodal physical field state can be obtained. This minimum value serves as the "reasonable deviation baseline," representing the minimum behavioral difference that the system can tolerate after considering multiple reasonable behavioral paths.
[0072] Finally, the initial deviation is compared with the reasonable deviation baseline. If the former is significantly greater than the latter, a high safety risk is identified, and the safety risk entropy is calculated based on the degree of deviation. Otherwise, the initial deviation is reduced before calculating the safety risk entropy. When comparing the initial deviation with the reasonable deviation baseline, a threshold or scaling factor can be set to define "significantly greater than." For example, if the initial deviation exceeds the reasonable deviation baseline by a certain percentage (such as 1.5 or 2 times), a high safety risk is considered to exist. The calculation of the safety risk entropy can be based on the degree of deviation; for example, the greater the deviation, the higher the safety risk entropy. Linear, exponential, or logarithmic functions can be used for mapping. If the initial deviation is not significantly greater than the reasonable deviation baseline, it indicates that although the current behavior differs from the expected behavior, it is still within the acceptable range of reasonable behavior. In this case, the initial deviation needs to be reduced to avoid oversensitivity or false alarms. For example, it can be multiplied by a reduction factor less than 1, or it can be set to a low fixed value before calculating the safety risk entropy.
[0073] In this embodiment, not only is the direct difference between actual and expected behavior considered, but also the various reasonable behavioral paths that the detected target may take under the current complex and ever-changing multimodal physical field conditions. This multi-scenario consideration avoids misjudgments caused by the limitations of a single expected behavioral path, making the risk assessment results more robust and accurate. When there is a difference between actual and expected behavior, but the difference is still within the range of reasonable behavior, the safety risk entropy is calculated by subtracting the initial deviation, effectively reducing the false alarm rate and avoiding excessive intervention in normal behavior. Conversely, when actual behavior deviates significantly from all reasonable behavioral paths, the system can promptly and accurately identify high safety risks and quantify the risk based on the degree of deviation, thereby providing a reliable basis for subsequent risk warnings and interventions, significantly improving the safety and practicality of the visual detection method.
[0074] In some other embodiments, after S106, the method may further include: Based on the previous moment's multimodal physical field state, the field perturbation caused by the real-time behavior of the detected target is calculated. The field perturbation includes the redistribution of the local probability density of the movement probability field or the redirection of high-intensity regions in the attention field. Using the multimodal physical field after field perturbation as input, the expected behavior spectrum of the affected entity is recalculated, and the degree of spectrum change is evaluated. The risk entropy of cooperation efficiency is obtained by weighted aggregation of the spectral changes of all affected entities. The comprehensive risk entropy value is corrected based on the risk entropy of collaboration efficiency to obtain the final risk entropy value.
[0075] Specifically, to accurately capture the dynamic impact of the detected target's behavior on the environment, the system uses the previous moment's multimodal physical field state as a baseline. This means the system stores and references the instantaneous state of the multimodal physical field (e.g., attention field, task potential gradient field, path probability field, and robot interaction field) before the current detected target's real-time behavior occurs. This baseline state provides a reference point for subsequent field perturbation calculations, enabling the system to quantify the field state changes caused by the detected target's real-time behavior.
[0076] Based on this, the system calculates the field perturbations caused by the real-time behavior of the detected target. Field perturbations refer to local or global changes in the multimodal physical field caused by the specific behavior of the detected target. These perturbations are key to assessing the risk to collaborative efficiency. For example, when a detected target suddenly changes its movement path, it may cause a redistribution of the local probability density in its surrounding area in the movement probability field, meaning that the expected movement probability of other entities in that area will change accordingly. Similarly, when a detected target performs an unusual or attention-grabbing action, it may cause a redirection of high-intensity areas in the attention field, meaning that the attention focus of other entities may be attracted to the detected target. Calculating field perturbations can be achieved by comparing the current field state with the reference field state from the previous time step, for example, through field intensity differences, gradient changes, or detection using specific pattern recognition algorithms.
[0077] Subsequently, using the perturbed multimodal physical field as input, the system recalculates the expected behavior spectrum of the affected entities. Once the field perturbation is identified, the system needs to assess how these perturbations affect the expected behavior of other entities in the scene. For each entity affected by the field perturbation, the system regenerates the entity's expected behavior spectrum using its identity information, current task state, and the updated multimodal physical field parameters of its location. This process is similar to generating the expected behavior spectrum of the detected target itself, but the input is the perturbed field state, thus reflecting the behavioral tendencies of other entities under the new environmental conditions.
[0078] Simultaneously, the system assesses the degree of spectral change in the affected entity. The degree of spectral change refers to the magnitude of the alteration in the expected behavioral spectrum of the affected entity before and after the perturbation. This can be quantified by calculating the distance or similarity between two expected behavioral spectra (before and after the perturbation) in the feature space, using metrics such as KL divergence, JS divergence, or Euclidean distance. A greater degree of spectral change indicates a greater impact on the entity's expected behavior.
[0079] Next, the spectral changes of all affected entities are weighted and aggregated to obtain the collaboration efficiency risk entropy. To obtain a comprehensive assessment of collaboration efficiency risk, the system integrates the spectral changes of all affected entities. Weighted aggregation means that behavioral changes of different entities may have different importance; for example, behavioral changes of key task performers may be given higher weight. These weighted changes are aggregated to form a single quantitative indicator: the collaboration efficiency risk entropy. This entropy value reflects the field disturbances caused by the detected target behavior, and consequently, the potential risks to the entire collaboration system in terms of behavioral coordination and efficiency.
[0080] Finally, the comprehensive risk entropy value is corrected based on the collaboration efficiency risk entropy to obtain the final risk entropy value. By incorporating the collaboration efficiency risk entropy into the calculation, the system can correct the comprehensive risk entropy value of the detected target, thus obtaining a more comprehensive and accurate final risk entropy value. This correction mechanism ensures that risk assessment not only focuses on the safe deviation of individual behaviors but also considers the potential negative impact of such behaviors on the overall smoothness and efficiency of collaboration.
[0081] This application's embodiments can identify and quantify the impact of real-time behavior of a single detected target on a multimodal physical field. Specifically, by monitoring the redistribution of local probability density in the movement probability field or the redirection of high-intensity regions in the attention field, the system can accurately capture the chain reaction of individual behavior on the environment. Based on this, the expected behavior spectrum of the affected entities is recalculated and its degree of change is assessed, enabling the system to understand the behavioral adjustment needs of other entities due to environmental changes. Furthermore, by weighted aggregation of these spectrum changes, this application can generate a collaborative efficiency risk entropy, thereby quantifying the overall collaborative efficiency loss caused by deviations in individual behavior. Finally, this collaborative efficiency risk entropy is corrected with the original comprehensive risk entropy value, so that risk assessment is no longer limited to individual safety but extends to the efficiency and stability of the entire collaborative system, providing a more comprehensive, dynamic, and forward-looking risk assessment mechanism. This helps to promptly identify and mitigate potential collaborative obstacles, optimizing the smoothness and security of multi-entity collaborative operations.
[0082] Based on the AI-based visual inspection method provided in the above embodiments, this application also provides specific implementations of an AI-based visual inspection device. Please refer to the following embodiments.
[0083] First see Figure 2 The artificial intelligence-based visual inspection device 200 provided in this application embodiment may include: The first acquisition module 201 allows the user to acquire field energy data in the target scenario. The field energy data is used to characterize task behavior data in the target scenario. Module 202 is used to construct a multimodal physical field based on field energy data. The multimodal physical field includes at least two of the following: attention field, task potential gradient field, path probability field, and robot action field. The first generation module 203 is used to generate the expected behavior spectrum of the detected target at the current moment based on the identity information of the detected target, the current task status, and the multimodal physical field parameters of the detected target's location in the target scene. The expected behavior spectrum is used to describe the behavior probability distribution of the detected target. The second acquisition module 204 is also used to acquire the real-time behavior raw vector of the detected target based on visual perception. The real-time behavior raw vector includes a sequence of position, velocity, orientation and attitude. The second generation module 205 is also used to substitute the original real-time behavior vector into the current multimodal physical field to generate a real-time behavior field effect vector. The calculation module 206 is used to calculate and output the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum. The comprehensive risk entropy value is used to describe the quantitative assessment of security risk and collaboration efficiency risk.
[0084] As an optional implementation, the building module 202 can be used for: Acquire visual images of the target scene; By using a pre-defined neural network model to estimate the action state of visual images, a set of possible gaze point coordinates is obtained. The set of suspected gaze point coordinates is associated and matched with the location of the preset information source entity to obtain the dominant information source at each moment; Using the spatial location of the dominant information source as the field source, different intensity weights are assigned according to the semantic type of the information source entity. The radial basis function network is used to calculate the superposition of the attenuated field strengths from any point in the space to all effective field sources to generate the attention field.
[0085] As an optional implementation, the field energy data includes task titles, a list of participants, deadline-preceding task dependencies, and task tags. The construction module 202 can be used for: Based on the task title, list of participants, deadline and pre-task dependencies, and task tags, a dynamic task relationship graph is constructed, where nodes represent tasks and edges represent overlapping relationships between personnel for task keys. Graph attention network analysis is used to analyze dynamic task relationship graphs and calculate global urgency score and collaboration heat score for each task node. The real-time location of the detected target is used as the anchor point of the task's potential energy in physical space. The potential energy intensity of the anchor point is calculated based on the global urgency score and the cooperation heat score. A potential energy diffusion algorithm is used to diffuse the potential energy intensity of each anchor point to the entire target scene according to spatial accessibility, forming a continuous task potential energy gradient field.
[0086] As an optional implementation, the building module 202 can be used for: An anisotropic Gaussian risk field is constructed with the robot's real-time planned path as the central axis and the robot's physical outline and safe braking distance as parameters. A collaborative attraction field is constructed centered on the robot's preset collaboration points; The robot's action field is generated by dynamically modulating the Gaussian risk field and the cooperative attraction field.
[0087] As an optional implementation, the first generation module 203 can also be used for: Using the target's identity information and current task status as the query entry point, relevant nodes and subgraphs are activated in the preset behavioral knowledge graph; The multimodal sociophysical field strength of the detected target's location is transformed into contextual conditions of a pre-defined behavioral knowledge graph. Conditional probability propagation and reasoning are then performed within the pre-defined behavioral knowledge graph to obtain the expected behavioral spectrum.
[0088] As an optional implementation, the computing module 206 can also be used for: Calculate the distance between the real-time behavior field effect vector and the expected behavior spectrum in the feature space to obtain the initial deviation. Based on the current multimodal physical field state, several preset scenarios are enumerated, and the hypothetical behavioral field effect vector that the detected target should produce under each scenario is inferred. Calculate the deviation between the hypothetical behavior field effect vector and the expected behavior spectrum for each hypothesis, and take the minimum value as the reasonable deviation baseline; The initial deviation is compared with the reasonable deviation from the baseline. If the former is significantly greater than the latter, a high safety risk is determined, and the safety risk entropy is calculated based on the degree of deviation. Otherwise, the initial deviation is reduced before calculating the safety risk entropy.
[0089] As an optional implementation, the computing module 206 can also be used for: Based on the previous moment's multimodal physical field state, the field perturbation caused by the real-time behavior of the detected target is calculated. The field perturbation includes the redistribution of the local probability density of the movement probability field or the redirection of high-intensity regions in the attention field. Using the multimodal physical field after field perturbation as input, the expected behavior spectrum of the affected entity is recalculated, and the degree of spectrum change is evaluated. The risk entropy of cooperation efficiency is obtained by weighted aggregation of the spectral changes of all affected entities. The comprehensive risk entropy value is corrected based on the risk entropy of collaboration efficiency to obtain the final risk entropy value.
[0090] This application also provides an electronic device. (See reference...) Figure 3 As shown, the electronic device includes a processor 301 and a memory 302 storing computer program instructions. When the processor 301 executes the computer program instructions, it implements the aforementioned artificial intelligence-based visual detection method.
[0091] Specifically, the processor 301 described above may include a central processing unit (CPU) or an application-specific integrated circuit (ASIC) that can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application.
[0092] The memory 302 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, which is used to store application code that executes the scheme of this application and is controlled to execute by the processor 301.
[0093] In this embodiment, the electronic device may further include a communication interface 303 and a bus 304. The processor 301, memory 302, and communication interface 303 are connected via the bus 304 and communicate with each other.
[0094] Specifically, the communication interface 303 is mainly used to realize communication between various modules, devices, units, and / or equipment in the embodiments of this application. The bus 304 includes hardware, software, or both, coupling the components of the electronic device together. The bus 304 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0095] Furthermore, in conjunction with the AI-based visual inspection method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement the AI-based visual inspection method in the above embodiments.
[0096] This application also provides a computer program product in which instructions, when executed by the processor of an electronic device, cause the electronic device to perform the artificial intelligence-based visual detection method as described in the above embodiments.
[0097] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0098] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A visual inspection method based on artificial intelligence, characterized in that, include: Acquire field energy data in the target scenario, wherein the field energy data is used to characterize task behavior data in the target scenario; Based on the field energy data, a multimodal physical field is constructed, which includes at least two of the following: attention field, task potential gradient field, path probability field, and robot action field. For the target to be detected in the target scene, based on the target's identity information, current task status and multimodal physical field parameters of the target's location, an expected behavior spectrum of the target at the current moment is generated. The expected behavior spectrum is used to describe the probability distribution of the target's behavior. The real-time behavior vector of the detected target is obtained based on visual perception, and the real-time behavior vector includes a sequence of position, velocity, orientation, and posture. The original real-time behavior vector is substituted into the current multimodal physical field to generate a real-time behavior field effect vector. Based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum, the comprehensive risk entropy value of the detected target is calculated and output. The comprehensive risk entropy value is used to describe the quantitative assessment of security risk and collaboration efficiency risk.
2. The method according to claim 1, characterized in that, The construction of a multimodal physical field based on the field energy data includes: Acquire visual images of the target scene; The visual image is used to estimate the action state using a preset neural network model to obtain a set of suspected gaze point coordinates. The set of suspected gaze point coordinates is associated and matched with the location of the preset information source entity to obtain the dominant information source at each moment; Using the spatial location of the dominant information source as the field source, different intensity weights are assigned according to the semantic type of the information source entity, and the attenuated field strengths from any point in the space to all effective field sources are superimposed using a radial basis function network to generate the attention field.
3. The method according to claim 1, characterized in that, The field energy data includes a task title, a list of participants, pre-deadline task dependencies, and task tags. The construction of a multimodal physical field based on the field energy data includes: Based on the task title, the list of participants, the dependencies of the tasks before the deadline, and the task tags, a dynamic task relationship graph is constructed, where nodes represent tasks and edges represent overlapping relationships between personnel for task keys. The dynamic task relationship graph is analyzed using graph attention networks to calculate the global urgency score and collaboration heat score for each task node. The real-time location of the detected target is used as the anchor point of the task's potential energy in physical space. The potential energy intensity of the anchor point is calculated based on the global urgency score and the cooperation heat score. A potential energy diffusion algorithm is used to diffuse the potential energy intensity of each anchor point to the entire target scene according to spatial accessibility, forming a continuous task potential energy gradient field.
4. The method according to any one of claims 1-3, characterized in that, The construction of a multimodal physical field based on the field energy data includes: An anisotropic Gaussian risk field is constructed with the robot's real-time planned path as the central axis and the robot's physical outline and safe braking distance as parameters. A collaborative attraction field is constructed centered on the robot's preset collaboration points; The Gaussian risk field and the cooperative attraction field are dynamically modulated to generate the robot action field.
5. The method according to claim 1, characterized in that, The step of generating the expected behavior spectrum of the detected target at the current moment based on the target's identity information, current task status, and multimodal physical field parameters of the target's location includes: Using the identity information and current task status of the detected target as the query entry point, relevant nodes and subgraphs are activated in the preset behavioral knowledge graph; The multimodal sociophysical field strength of the detected target's location is transformed into the contextual conditions of the preset behavioral knowledge graph. Conditional probability propagation and reasoning are then performed in the preset behavioral knowledge graph to obtain the expected behavioral spectrum.
6. The method according to claim 1, characterized in that, The step of calculating and outputting the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum includes: Calculate the distance between the real-time behavior field effect vector and the expected behavior spectrum in the feature space to obtain the initial deviation. Based on the current multimodal physical field state, several preset scenarios are enumerated, and the hypothetical behavioral field effect vector that the detected target should produce under each scenario is inferred. Calculate the deviation of each hypothetical behavior field effect vector from the expected behavior spectrum, and take the minimum value as the reasonable deviation baseline; The initial deviation is compared with the reasonable deviation baseline. If the former is significantly greater than the latter, a high safety risk is determined, and the safety risk entropy is calculated based on the degree of deviation. Otherwise, the initial deviation is reduced before calculating the safety risk entropy.
7. The method according to claim 6, characterized in that, After calculating and outputting the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum, the method further includes: Based on the multimodal physical field state of the previous moment, the field perturbation caused by the real-time behavior of the detected target is calculated. The field perturbation includes the redistribution of the local probability density of the movement probability field or the redirection of high-intensity regions in the attention field. Using the multimodal physical field after the field perturbation as input, the expected behavior spectrum of the affected entity is recalculated, and the degree of spectrum change is evaluated. The risk entropy of cooperation efficiency is obtained by weighted aggregation of the spectral changes of all affected entities. The comprehensive risk entropy value is corrected based on the collaborative efficiency risk entropy to obtain the final risk entropy value.
8. A visual inspection device based on artificial intelligence, characterized in that, The apparatus for performing the artificial intelligence-based visual inspection method as described in any one of claims 1-7, the apparatus comprising: The first acquisition module allows the user to acquire field energy data in a target scenario, where the field energy data is used to characterize task behavior data in the target scenario. A construction module is used to construct a multimodal physical field based on the field energy data. The multimodal physical field includes at least two of the following: attention field, task potential gradient field, path probability field, and robot action field. The first generation module is used to generate the expected behavior spectrum of the detected target at the current moment based on the identity information, current task status and multimodal physical field parameters of the detected target's location in the target scene. The expected behavior spectrum is used to describe the behavior probability distribution of the detected target. The second acquisition module is used to acquire the real-time behavior raw vector of the detected target based on visual perception. The real-time behavior raw vector includes a sequence of position, velocity, orientation, and posture. The second generation module is used to substitute the original real-time behavior vector into the current multimodal physical field to generate a real-time behavior field effect vector. The calculation module is used to calculate and output the comprehensive risk entropy value of the detected target based on the coupling relationship between the real-time behavior field effect vector and the expected behavior spectrum. The comprehensive risk entropy value is used to describe the quantitative assessment of security risk and collaboration efficiency risk.
9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the artificial intelligence-based visual detection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the artificial intelligence-based visual detection method as described in any one of claims 1-7.