A target detection method driven by target temporal causality in multiple industrial scenarios

By improving the yolov9 model, a multi-scene object detection data set was constructed and time-sequence causality representation was used, which solved the problems of large time overhead and weak understanding ability in industrial multi-scene object detection, and achieved more efficient and in-depth object detection effects.

CN119672320BActive Publication Date: 2025-05-16CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510164005.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In industrial production, the single-scene object detection model faces the problems of large time overhead and weak multi-scene understanding ability when dealing with multi-scene object detection tasks, resulting in the impact of production real-time and the relevance of scene logic.

Method used

By improving the yolov9 object detection model, an industrial multi-scenario object detection data set is constructed, and the target timing causality representation, timing causality quantitative representation and target timing causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causality causal

Benefits of technology

It improves the real-time and understanding ability of industrial multi-scenario object detection, enhances the credibility of object detection and in-depth understanding of the scenario, and reduces the frequency and time overhead of model switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672320B_ABST
    Figure CN119672320B_ABST
Patent Text Reader

Abstract

The present invention proposes a target detection method driven by target temporal causal relationship in industrial multiple scenarios, aiming at the problem of multi-scenario target detection in real flexible circuit board production. In order to overcome the problem of large time overhead and weak multi-scenario understanding ability of the target detection model for multi-scenario target detection, the method proposes target temporal causal cascade Mamba. The target temporal causal cascade Mamba processes the features after the Backbone layer and the Neck layer of the yolov9 model to solve the parameter values ​​in the multi-scenario target temporal causal quantitative representation diagram. The method also proposes a relational attention module to enhance the accuracy of relational feature conversion. The beneficial effects of the present invention are: the target detection method driven by target temporal causal relationship in multiple scenarios can model the production timing of targets between the same scene and the causality of targets between different scenes, understand the logical correlation between industrial production scenes, and thus improve the real-time and accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and deep learning, and specifically relates to a target detection method driven by target temporal causal relationship in multiple industrial scenarios. Background Art

[0002] With the continuous development of artificial intelligence technology, target detection technology has been widely used in many fields, such as industrial production, traffic management, medical diagnosis, national defense security, etc. In industrial production, target detection technology is mainly used in industrial safety monitoring, production line detection, and quality inspection scenarios. Safety monitoring scenario: Use target detection technology to detect targets such as personnel and personnel clothing (safety helmets, masks, work clothes), determine whether they are abnormal, and issue alarms and take measures in time to avoid accidents and damage. Production line detection scenario: Use target detection technology to detect targets such as equipment, products, and materials on the production line to achieve perception and regulation of the production line. Quality inspection scenario: Use target detection technology to detect targets such as defective products and product defects for quality control and quality management. It can be seen that the application of target detection technology can detect target objects in the production process, realize autonomous perception of the production process, and achieve the purpose of ensuring production safety, improving production efficiency and quality, and reducing production costs. At present, industrial target detection is to form a data set with target samples of a single scene, and then train a model for each data set to solve the single-scene target detection task. Single-scene target detection focuses on target feature analysis and detection tasks within a single scene. Therefore, it may achieve better detection results for targets in a single scene. However, single-scene object detection faces the following challenges when dealing with multi-scene object detection tasks:

[0003] (1) Multi-scenario target detection has a high time overhead. In order to ensure that the target detection task meets the continuity requirements of production, it is necessary to frequently call the target detection models of different task scenarios. This increases the switching frequency between single-scenario target detection models, brings additional time overhead, and affects the real-time performance of production.

[0004] (2) Weak multi-scenario comprehension capability. Since each single-scenario target detection model is relatively independent, the model ignores the relationship between scene targets. This not only breaks the logical connection between production scenes, but also reduces the in-depth cognition and understanding of industrial scenes.

[0005] In summary, to address the above problems, a target detection method driven by target temporal causality in multiple industrial scenarios is proposed. Summary of the invention

[0006] In view of the above problems, the purpose of the present invention is to provide a target detection method driven by target temporal causality in industrial multiple scenarios. This method extracts (captures, models, and applies) the production temporality of targets in the same scenario and the causality of targets in different scenarios by improving the yolov9 target detection model, thereby enhancing the credibility of target detection and the understanding of the scenario.

[0007] A target detection method driven by target temporal causal relationship in industrial multi-scenario, comprising the following steps:

[0008] S1. Constructing an industrial multi-scenario target detection dataset: Obtain images of three production scenarios of flexible printed circuit (FPC) production, including personnel wear, production line equipment, and product defects, and manually annotate each RGB image to obtain annotation information, wherein the annotation information includes target location and target category. The targets include personnel clothing (correct clothing, incorrect clothing), production line equipment (cutting machine, drilling machine, copper sinking wire, copper plating wire, laminating machine, laminating machine, industrial oven, punching machine, exposure machine, printing machine, optical automatic inspection machine), and product defects (common crushing, common pollution, CVL foreign matter, common scratches, common wrinkles, ink shedding, under-film oxidation, plating leakage, gold surface roughness, ink foreign matter, ink pollution). The annotated images constitute an industrial multi-scenario target detection dataset, which also implies target temporal sequence and target causality.

[0009] S2. Target temporal causal relationship representation: including causal relationship representation, temporal relationship representation and temporal causal relationship representation;

[0010] Causal relationship representation, that is, to find all safety monitoring scenario targets and production line inspection scenario targets related to quality inspection scenario targets in a flexible manufacturing environment, and screen them. Based on industrial mechanism knowledge, find out the possible causal relationship types between all safety monitoring scenario targets, production line inspection scenario targets and quality inspection scenario targets, and express the causal influence of the cause variable on the result variable after removing the confounding factors; Temporal relationship representation, first find out the safety monitoring scenarios and production line inspection scenarios with process timing, and construct a temporal relationship representation that satisfies symmetry and transitivity between targets in the same scenario; Temporal causal relationship representation, jointly represent the constructed causal relationship representation and the temporal relationship representation, and obtain the target temporal causal structure diagram based on the joint representation;

[0011] S3. Quantitative representation of temporal causal relationship: According to the target temporal causal structure diagram, after determining the confounding factor, the target feature probability distribution of the three scenarios at any time in the target temporal causal structure diagram can be obtained, so that the temporal causal relationship between the targets can be quantified, and thus the temporal causal relationship quantitative representation diagram can be established;

[0012] S4. Perform quantitative learning of target temporal causal relationships: Construct the target temporal causal cascade Mamba to perform relational feature conversion and solve the parameters in the temporal causal relationship quantitative representation graph, and use the relational attention module to optimize the relational feature conversion;

[0013] S5. Perform industrial multi-scenario target detection: Integrate the target temporal causal cascade Mamba into the yolov9 target detection model to obtain the final detection model, input the multi-scene image into the final model, and generate the final industrial multi-scenario target detection prediction results.

[0014] The industrial multi-scenario target detection dataset constructed in S1 above comes from three production scenarios of real flexible circuit board production, where the target timing comes from the process timing during product production (corresponding to the timing of personnel clothing requirements and the timing of equipment required for production), and the target causality comes from the fact that personnel clothing and current production equipment under specific processes will produce corresponding product defects.

[0015] In the above S2, it is known that the safety monitoring scenario target and the production line inspection scenario target as the cause variables will have a causal impact on the quality inspection scenario target. At this time, there are five types of causal relationships. The nonlinear Granger causality prediction model is used to qualitatively identify the causal relationship, and the following is obtained:

[0016] Deterministic causal relationship security monitoring target class ( ) to the quality inspection target class ( )、Production line detection target class( ) to the quality inspection target class ( ) under the condition of quality detection target class at time t is expressed as:

[0017] ,

[0018] ,

[0019] In the formula, , and They represent the target class of safety monitoring scene, the target class of production line inspection scene and the target class of quality inspection scene respectively. represents a nonlinear function, , Respectively In relationship and The lag order of , Respectively In relationship and The lag order of and They represent the error terms of two nonlinear functions respectively.

[0020] Uncertain causality or , and its predicted values ​​are expressed as:

[0021]

[0022] or ,

[0023] The parameters are defined similarly to the prediction method of deterministic causal relationships.

[0024] The above prediction values ​​were tested using non-parametric test methods to determine whether the causal variable made a significant contribution to the prediction of the outcome variable.

[0025] At the same time, the transfer entropy is used to qualitatively distinguish the causal relationship type and calculate the transfer entropy TE to quantify the information contribution of the cause variable to the result variable:

[0026] ,

[0027] In the formula, represents the cause variable, represents the outcome variable, represents the joint probability density function, and represents the conditional probability density function, and Represent the history length respectively, that is, in the calculation How far back in history should we look to determine the future state? represents the significance level parameter. If , it indicates that there is Nonlinear causal effects.

[0028] By comprehensively judging the causal relationship type based on the nonlinear Granger causality test and information entropy judgment results, the causal influence of the cause variable on the result variable after removing the confounding factors is expressed as:

[0029] ,

[0030] ,

[0031] ,

[0032] ,

[0033] In the formula, Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represent safety goals, production line goals and quality goal instances respectively, and Indicates possible confounding factors, do(A=a) means forcing the external variable A to be assigned a. and They represent transfer entropy and The weight of the causal influence after conversion to the do operator.

[0034] The temporal relationship between the safety monitoring scenario and the production line detection scenario is expressed as follows: .

[0035] The constructed temporal relationship satisfies the symmetry and transitivity constraints, which can be expressed as:

[0036] ,

[0037] In the formula, , Represents the distribution characteristics of the target temporal relationship set (such as mean, variance), Represents the target temporal relationship transfer characteristics.

[0038] Then, the target causal relationship and the target temporal relationship are jointly represented, and the constructed temporal causal relationship model is:

[0039]

[0040] In the formula, , Represent the temporal and causal search spaces of the target, respectively, and depend on the optimal temporal relationship and optimal causality , and Represents the target temporal relationship pair and the target causal relationship pair, and They represent the search space of causal-temporal relationship pairs and the search space of causal-causal relationship pairs respectively. represents the target pair space, and They are respectively the scoring function of the relationship between targets and the scoring function of the target search space.

[0041] , is defined as follows:

[0042] ,

[0043] In the formula, and They represent the temporal relationship and causal relationship between targets respectively. The above formula shows that the search space of the target depends on the corresponding relationship.

[0044] and is defined as follows:

[0045] ,

[0046] ,

[0047] In the formula, express The feature vector of the pair.

[0048] By solving the above formula, we can get the optimal timing relationship and optimal causality , and Target search space, and obtain the target temporal causal structure diagram.

[0049] The feature vectors of the targets in the three scenarios observed at time t in S3 above are , the target characteristic state value observed at time t=1 is recorded as the initial state distribution , then in the same batch of data streams In the example, at the t+1th moment after the initial moment, the target state distributions of the safety monitoring scene target and the production line inspection scene target are , ,in, Represents the feature vector of the security monitoring scene target at time t, Represents the temporal relationship between the targets at time t and time t+1 in the security monitoring scenario, Represents the feature vector of the production line inspection scene target at time t, Represents the temporal relationship between the target at time t and the target at time t+1 in the production line inspection scene.

[0050] When the confounding factor is determined, the target feature probability distribution representation of the three target data sets at time t in the target time series causal structure diagram can be obtained according to the causal relationship expression and the do operator:

[0051] ,

[0052] In the formula, represents the fusion operation of the target features, t and t-1 represent the time t and time t-1 of the scene respectively, Indicates the causal relationship between scenes, Represents the target class features of the corresponding scene, Represents the probability of the target class feature of the corresponding scene.

[0053] According to the target temporal causal structure diagram and the above formula, the temporal causal relationship between targets can be quantitatively represented, thereby establishing a temporal causal relationship quantitative representation diagram.

[0054] In the above S4, the target temporal causal cascade Mamba is constructed to solve the above equation , , , , , The values ​​of the 6 parameters.

[0055] Target temporal causal cascade Mamba receives the features of the image after passing through the Backbone layer and Neck layer of the yolov9 model. The specific operation steps are high-confidence target screening operation, temporal scene relationship feature conversion operation, and causal scene relationship feature conversion operation:

[0056] High confidence target screening operation: Perform Top-K screening on the features after passing through the Backbone layer and Neck layer of the yolov9 model, and output the top k high confidence target features:

[0057] ,

[0058] in, For input features For each feature vector at position (h,w): , It means taking the first k values ​​of the maximum value.

[0059] Time series scene relationship feature conversion operation: When predicting the safety monitoring scene target at time t+1, the features of the safety monitoring scene target at time t and the production line inspection scene target at time t+1 are concatenated after the high confidence target screening operation. Pass it into the projection layer and output it , , Three matrices, at the same time Time-invariant state matrix and the time-varying matrix Discretize to get and , and then Further with input Multiply, and original state Multiply and add the first two items to get the current state , and finally the current state With the matrix Multiply as input to get the final temporal relationship feature conversion result , when predicting the target of the production line detection scene at time t+1, the temporal scene relationship feature conversion operation corresponds to the temporal causal relationship quantitative representation diagram of the target feature probability distribution solution process of the security monitoring scene and the production line detection scene. The formula is as follows:

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] in, , They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents the concatenation operation of features, represents a linear projection operation, and Indicates the use Using zero-order hold technique and To discretize, represents matrix multiplication, represents matrix addition, and They represent the temporal relationship feature conversion results of the safety monitoring scene target and the production line inspection scene target respectively.

[0071] Causal scene relationship feature conversion operation: When predicting the quality inspection scene target at time t, the features of the safety monitoring scene target and the production line inspection scene target at time t are passed to the two projection layers after the high confidence target screening operation, and the output is obtained , , , , , Six matrices, simultaneously and For the time-invariant state matrix A, the time-varying matrix and Discretize to get , and , and then and Further separation and input and Multiply, and original state Multiply and add the first three items to get the current state , and finally the current state Respectively with the matrix and Multiply as input to get the final causal feature conversion result and ,The causal scene relationship feature conversion operation corresponds to the target feature probability distribution solution process of the quality detection scene in the temporal causal relationship quantitative representation diagram. The formula is as follows:

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] in, , They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents a linear projection operation, and Respectively indicate the use and Using zero-order hold technique To discretize, and Respectively indicate the use Using zero-order hold technique Discretize and use Using zero-order hold technique To discretize, represents matrix multiplication, represents matrix addition, and They respectively represent the causal feature conversion results of the safety monitoring scenario target and the production line inspection scenario target to the quality inspection scenario target.

[0079] Relational attention module: The obtained security monitoring scene target temporal relationship feature conversion results and production line inspection scene target temporal relationship feature conversion results at time t , , the causal relationship feature conversion results of the safety monitoring scene target and the production line inspection scene target to the quality inspection scene target , The input is sent to the relational attention module, and the attention weight is output, so that the model can pay more attention to the correct relational expression. The relational conversion result is then multiplied by the attention weight and added to the feature to obtain the relation-enhanced feature.

[0080] The above computational relation attention module first uses the Gaussian kernel to calculate the high confidence target features at time t With the previous The similarity of the high confidence target feature mean at each moment is then multiplied by the activation function and the result of the relational feature transformation to obtain the result after attention. The formula is as follows:

[0081]

[0082]

[0083]

[0084] in, represents average pooling, represents the bandwidth parameter of the Gaussian kernel, represents the silu activation function, Represents the weight index of relation attention.

[0085] Then, the obtained weight index is multiplied by the corresponding relationship conversion result and added to the feature to obtain the final prediction feature. The operation of the time series scenario is as follows:

[0086]

[0087]

[0088] in, and They represent the features of the security monitoring scene image and the production line detection scene image at time t+1 after the high confidence target screening operation, represents matrix addition, represents matrix dot product, and They represent the relational attention weights of the safety monitoring scene relational features and the production line inspection scene relational features at time t respectively.

[0089] The operation of the cause and effect scenario is as follows:

[0090]

[0091]

[0092] in, Represents the features of the quality inspection scene image at time t after the high confidence target screening operation, represents a multi-layer perceptron, represents matrix addition, represents matrix dot product, and They respectively represent the relationship attention weights from the safety monitoring scene target to the quality inspection scene target and from the production line inspection scene target to the quality inspection scene target at time t.

[0093] In the training process of the target temporal causal cascade Mamba, the mean square error loss is used, and the formula is as follows:

[0094]

[0095] Among them, m represents the total training time, , , They represent the prediction features at time t, , , They represent the high confidence target features obtained from the true labels respectively.

[0096] Finally, the obtained prediction features are predicted to obtain the final prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is the type of causal relationship that exists between the three scenario goals.

[0098] Figure 2 It is the target temporal causal structure diagram.

[0099] Figure 3 It is a quantitative representation diagram of temporal causal relationships.

[0100] Figure 4 It is the target timing causal cascade Mamba framework diagram.

[0101] Figure 5 This is a schematic diagram of the yolov9 model.

[0102] Figure 6 It is a schematic diagram of the temporal scene relationship feature conversion operation.

[0103] Figure 7 It is a schematic diagram of the causal scene relationship feature conversion operation.

[0104] Figure 8 This is the effect diagram of the target detection method driven by the target temporal causal relationship in multiple industrial scenarios. DETAILED DESCRIPTION

[0105] The technical solution of the present invention is described in detail below with reference to the accompanying drawings.

[0106] A target detection method driven by target temporal causal relationship in multiple industrial scenarios specifically includes the following steps:

[0107] S1. Constructing industrial multi-scenario target detection dataset: Obtain images of three production scenarios under the production of flexible printed circuit (FPC), including personnel wear, production line equipment, and product defects, and manually annotate each RGB image to obtain annotation information, the annotation information includes target location, target category, the target includes personnel clothing (correct clothing, incorrect clothing), production line equipment (cutting machine, drilling machine, copper sinking wire, copper plating wire, laminating machine, laminating machine, industrial oven, punching machine, exposure machine, printing machine, optical automatic inspection machine), product defects (public crushing, public pollution, CVL foreign matter, public scratches, public wrinkles, ink shedding, film oxidation, leakage plating, gold surface roughness, ink foreign matter, ink pollution), and the annotated images constitute an industrial multi-scenario target detection dataset with an image resolution of 640×640 pixels. The number of FPCIPMSD datasets is 5979, and the annotation software Labelimg is used to annotate the flexible circuit board defect images;

[0108] S2. Target temporal causal relationship representation: as shown in the attached Figure 1 As shown in the figure, under the premise that the safety monitoring scenario target and the production line inspection scenario target as cause variables will have a causal impact on the quality inspection scenario target, there are five types of causal relationships at this time. The nonlinear Granger causality prediction model is used to qualitatively identify the causal relationship, and the following is obtained:

[0109] Deterministic causal relationship security monitoring target class ( ) to the quality inspection target class ( )、Production line detection target class( ) to the quality inspection target class ( ) under the condition of quality detection target class at time t is expressed as:

[0110] ,

[0111] ,

[0112] In the formula, , and They represent the target class of safety monitoring scene, the target class of production line inspection scene and the target class of quality inspection scene respectively. represents a nonlinear function, , Respectively In relationship and The lag order of , Respectively In relationship and The lag order of and They represent the error terms of two nonlinear functions respectively.

[0113] Uncertain causality or , and its predicted values ​​are expressed as:

[0114]

[0115] or ,

[0116] The parameters are defined similarly to the prediction method of deterministic causal relationships.

[0117] The above prediction values ​​were tested using non-parametric test methods to determine whether the causal variable made a significant contribution to the prediction of the outcome variable.

[0118] At the same time, the transfer entropy is used to qualitatively distinguish the causal relationship type and calculate the transfer entropy TE to quantify the information contribution of the cause variable to the result variable:

[0119]

[0120] In the formula, represents the cause variable, represents the outcome variable, represents the joint probability density function, and represents the conditional probability density function, and Represent the history length respectively, that is, in the calculation How far back in history should we look to determine the future state? represents the significance level parameter. If , it indicates that there is Nonlinear causal effects.

[0121] By comprehensively judging the causal relationship type based on the nonlinear Granger causality test and information entropy judgment results, the causal influence of the cause variable on the result variable after removing the confounding factors is expressed as:

[0122] ,

[0123] ,

[0124] ,

[0125] ,

[0126] In the formula, Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represents the removal of the influence of confounding factors , Represent safety goals, production line goals and quality goal instances respectively, and Indicates possible confounding factors, do(A=a) means forcing the external variable A to be assigned a. and They represent transfer entropy and The weight of causal influence after conversion to the do operator, and They represent transfer entropy and The weight of the causal influence after conversion to the do operator.

[0127] The temporal relationship between the safety monitoring scenario and the production line detection scenario is expressed as follows: .

[0128] The constructed temporal relationship satisfies the symmetry and transitivity constraints, which can be expressed as:

[0129] ,

[0130] In the formula, , Represents the distribution characteristics of the target temporal relationship set (such as mean, variance), Represents the target temporal relationship transfer characteristics.

[0131] Then, the target causal relationship and the target temporal relationship are jointly represented, and the constructed temporal causal relationship model is:

[0132] ,

[0133] In the formula, , Represent the temporal and causal search spaces of the target, respectively, and depend on the optimal temporal relationship and optimal causality , and Represents the target temporal relationship pair and the target causal relationship pair, and They represent the search space of causal-temporal relationship pairs and the search space of causal-causal relationship pairs respectively. represents the target pair space, and They are respectively the scoring function of the relationship between targets and the scoring function of the target search space.

[0134] , is defined as follows:

[0135]

[0136] In the formula, and They represent the temporal relationship and causal relationship between targets respectively. The above formula shows that the search space of the target depends on the corresponding relationship.

[0137] and is defined as follows:

[0138] ,

[0139] ,

[0140] In the formula, express The feature vector of the pair.

[0141] By solving the above formula, we can get the optimal timing relationship and optimal causality , and The target search space is obtained, and the target temporal causal structure diagram is obtained, as shown in the attached Figure 2 shown.

[0142] S3. Quantitative representation of temporal causal relationship: According to the target temporal causal structure diagram, after determining the confounding factor, the target feature probability distribution of the three scenarios at any time in the target temporal causal structure diagram can be obtained, so that the temporal causal relationship between the targets can be quantified, and thus the temporal causal relationship quantitative representation diagram can be established;

[0143] The feature vectors of the targets in the three scenarios observed at time t in S3 above are , the target characteristic state value observed at time t=1 is recorded as the initial state distribution , then in the same batch of data streams In the example, at the t+1th moment after the initial moment, the target state distributions of the safety monitoring scene target and the production line inspection scene target are , ,in, Represents the feature vector of the security monitoring scene target at time t, Represents the temporal relationship between the targets at time t and time t+1 in the security monitoring scenario, Represents the feature vector of the production line inspection scene target at time t, Represents the temporal relationship between the target at time t and the target at time t+1 in the production line inspection scene.

[0144] When the confounding factor is determined, the target feature probability distribution representation of the three target data sets at time t in the target time series causal structure diagram can be obtained according to the causal relationship expression and the do operator:

[0145]

[0146] In the formula, represents the fusion operation of the target features, t and t-1 represent the time t and time t-1 of the scene respectively, Indicates the causal relationship between scenes, Represents the target class features of the corresponding scene, Represents the probability of the target class feature of the corresponding scene.

[0147] According to the target time-series causal structure diagram and the above formula, the time-series causal relationship between targets can be quantitatively represented, thereby establishing a quantitative representation diagram of the time-series causal relationship, as shown in the attached figure. Figure 3 shown.

[0148] S4. Quantitative learning of target temporal causal relationships: as shown in the attached Figure 4As shown, the target temporal causal cascade Mamba is constructed to perform relational feature conversion and solve the parameters in the temporal causal relationship quantitative representation graph, and the relational attention module is used to optimize the relational feature conversion;

[0149] Target temporal causal cascade Mamba receives the features of the image after passing through the Backbone layer and Neck layer of the yolov9 model. The network structure of the yolov9 target detection model is shown in the attached figure. Figure 5 As shown, the specific operation steps are high confidence target screening operation, temporal scene relationship feature conversion operation, and causal scene relationship feature conversion operation:

[0150] High confidence target screening operation: Perform Top-K screening on the features after passing through the Backbone layer and Neck layer of the yolov9 model, and output the top k high confidence target features:

[0151]

[0152] in, For input features For each feature vector at position (h,w): , It means taking the first k values ​​of the maximum value.

[0153] Temporal scene relationship feature conversion operation: as shown in the attached Figure 6 As shown in the figure, when predicting the safety monitoring scene target at time t+1, the features of the safety monitoring scene target at time t and the production line inspection scene target at time t+1 are concatenated after the high confidence target screening operation. Pass it into the projection layer and output it , , Three matrices, at the same time Time-invariant state matrix and the time-varying matrix Discretize to get and , and then Further with input Multiply, and original state Multiply and add the first two items to get the current state , and finally the current state With the matrix Multiply as input to get the final temporal relationship feature conversion result , when predicting the target of the production line detection scene at time t+1, the temporal scene relationship feature conversion operation corresponds to the temporal causal relationship quantitative representation diagram of the target feature probability distribution solution process of the security monitoring scene and the production line detection scene. The formula is as follows:

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164] in, , They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents the concatenation operation of features, represents a linear projection operation, and Indicates the use Using zero-order hold technique and To discretize, represents matrix multiplication, Represents matrix addition.

[0165] Causal scene relationship feature conversion operation: as shown in the attached Figure 7 As shown in the figure, when predicting the quality inspection scene target at time t, the features of the safety monitoring scene target and the production line inspection scene target at time t are respectively passed to the two projection layers after the high confidence target screening operation, and the output is obtained , , , , , Six matrices, simultaneously and For the time-invariant state matrix A, the time-varying matrix and Discretize to get , and , and then and Further separation and input and Multiply, and original state Multiply and add the first three items to get the current state , and finally the current state Respectively with the matrix and Multiply as input to get the final causal feature conversion result and ,The causal scene relationship feature conversion operation corresponds to the target feature probability distribution solution process of the quality detection scene in the temporal causal relationship quantitative representation diagram. The formula is as follows:

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] in, , They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents a linear projection operation, and Respectively indicate the use and Using zero-order hold technique To discretize, and Respectively indicate the use Using zero-order hold technique Discretize and use Using zero-order hold technique To discretize, represents matrix multiplication, represents matrix addition, and They respectively represent the causal feature conversion results of the safety monitoring scenario target and the production line inspection scenario target to the quality inspection scenario target.

[0173] Relational attention module: The obtained security monitoring scene target temporal relationship feature conversion results and production line inspection scene target temporal relationship feature conversion results at time t , , the causal relationship feature conversion results of the safety monitoring scene target and the production line inspection scene target to the quality inspection scene target , The input is sent to the relational attention module, and the attention weight is output, so that the model can pay more attention to the correct relational expression. The relational conversion result is then multiplied by the attention weight and added to the feature to obtain the relation-enhanced feature.

[0174] The above operation of calculating relational attention first uses the Gaussian kernel to calculate the high confidence target feature at time t With the previous The similarity of the high confidence target feature mean at each moment is then multiplied by the activation function and the result of the relational feature transformation to obtain the result after attention. The formula is as follows:

[0175]

[0176]

[0177]

[0178] in, represents average pooling, represents the bandwidth parameter of the Gaussian kernel, represents the silu activation function, Represents the weight index of relation attention.

[0179] Then, the obtained weight index is multiplied by the corresponding relationship conversion result and added to the feature to obtain the final prediction feature. The operation of the time series scenario is as follows:

[0180]

[0181]

[0182] in, and They represent the features of the security monitoring scene image and the production line detection scene image at time t+1 after the high confidence target screening operation, represents matrix addition, represents matrix dot product, and They represent the relational attention weights of the safety monitoring scene relational features and the production line inspection scene relational features at time t respectively.

[0183] The operation of the cause and effect scenario is as follows:

[0184]

[0185]

[0186] in, Represents the features of the quality inspection scene image at time t after the high confidence target screening operation, represents a multi-layer perceptron, represents matrix addition, represents matrix dot product, and They respectively represent the relationship attention weights from the safety monitoring scene target to the quality inspection scene target and from the production line inspection scene target to the quality inspection scene target at time t.

[0187] In the training process of the target temporal causal cascade Mamba, the mean square error loss is used, and the formula is as follows:

[0188]

[0189] Among them, m represents the total training time, , , They represent the prediction features at time t, , , They represent the high confidence target features obtained from the true labels respectively.

[0190] S5. Perform industrial multi-scenario target detection: Integrate the target temporal causal cascade Mamba into the yolov9 target detection model to generate the final industrial multi-scenario target detection prediction results. The mean average precision (MAP) is used as the evaluation index for defect detection. The MAP of various targets of the trained model on the industrial multi-scenario target detection dataset is shown in Table 1. The prediction effect of the industrial multi-scenario target detection dataset is visualized, as shown in the attached figure. Figure 8 shown.

[0191] Table 1: Test results of the trained target detection model on the industrial multi-scene target detection dataset

[0192]

Claims

1. A target detection method driven by target temporal causal relationship in industrial multi-scenario, characterized in that: The following steps are involved: S1. Constructing an industrial multi-scenario target detection dataset: The industrial multi-scenario target detection dataset includes images of personnel wear, production line equipment, and product defects in three production scenarios of flexible circuit board production; S2. Target temporal causal relationship representation: The target temporal causal relationship representation process includes causal relationship representation, temporal relationship representation and temporal causal relationship representation, and obtains a target temporal causal structure diagram; S3. Quantitative representation of temporal causal relationship: According to the target temporal causal structure diagram, after determining the confounding factor, the target feature probability distribution of the three scenarios at any time in the target temporal causal structure diagram can be obtained, so that the temporal causal relationship between the targets can be quantified, and thus the temporal causal relationship quantitative representation diagram can be established; S4. Perform quantitative learning of target temporal causal relationships: Construct the target temporal causal cascade Mamba to perform relational feature conversion and solve the parameters in the temporal causal relationship quantitative representation graph, and use the relational attention module to optimize the relational feature conversion; S5. Perform industrial multi-scenario target detection: Integrate the target temporal causal cascade Mamba into the yolov9 target detection model to obtain the final detection model, input the multi-scene image into the final model, and generate the final industrial multi-scenario target detection prediction results; In step S4, the target temporal causal cascade Mamba is constructed to perform relational feature conversion and solve the parameters in the temporal causal relationship quantitative representation diagram. The specific operation steps are high confidence target screening operation, temporal scene relationship feature conversion operation, and causal scene relationship feature conversion operation. High confidence target screening operation: Perform Top-K screening on the features after passing through the Backbone layer and Neck layer of the yolov9 model, and output the top k high confidence target features: in, For input features For each feature vector at position (h,w): , It means taking the first k values ​​of the maximum value; Time series scene relationship feature conversion operation: When predicting the safety monitoring scene target at time t+1, the features of the safety monitoring scene target at time t and the production line inspection scene target at time t+1 are concatenated after the high confidence target screening operation. Pass it into the projection layer and output it , , Three matrices, at the same time Time-invariant state matrix and the time-varying matrix Discretize to get and , and then Further with input Multiply, and original state Multiply and add the first two items to get the current state , and finally the current state With the matrix Multiply as input to get the final temporal relationship feature conversion result , the same is true when predicting the production line inspection scene target at time t+1. The formula is as follows: in, and They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents the concatenation operation of features, represents a linear projection operation, and Indicates the use Using zero-order hold technique and To discretize, represents matrix multiplication, represents matrix addition, and They represent the temporal relationship feature conversion results of the safety monitoring scene target and the production line inspection scene target respectively; Causal scene relationship feature conversion operation: When predicting the quality inspection scene target at time t, the features of the safety monitoring scene target and the production line inspection scene target at time t are passed to the two projection layers after the high confidence target screening operation, and the output is obtained , , , , , Six matrices, simultaneously and For the time-invariant state matrix A, the time-varying matrix and Discretize to get , and , and then and Further separation and input and Multiply, and original state Multiply and add the first three items to get the current state , and finally the current state Respectively with the matrix and Multiply as input to get the final causal feature conversion result and , the formula is as follows: in, and They represent the characteristics of the security monitoring scene and the production line inspection scene after the high-confidence target screening operation at the tth moment, represents a linear projection operation, and Respectively indicate the use and Using zero-order hold technique To discretize, and Respectively indicate the use Using zero-order hold technique Discretize and use Using zero-order hold technique To discretize, represents matrix multiplication, represents matrix addition, and They respectively represent the causal feature conversion results of the safety monitoring scenario target and the production line inspection scenario target to the quality inspection scenario target.

2. According to the target detection method driven by target temporal causal relationship in industrial multi-scenario according to claim 1, it is characterized in that: The causal relationship representation is to find all the safety monitoring scene targets and production line detection scene targets related to the quality inspection scene targets in the flexible manufacturing environment and screen them; find out the possible causal relationship types between all safety monitoring scene targets, production line detection scene targets and quality inspection scene targets based on industrial mechanism knowledge, and express the causal influence of the cause variable on the result variable after removing the confounding factors; the temporal relationship representation is to first find the safety monitoring scene and production line detection scene with process timing, and construct a temporal relationship representation that satisfies symmetry and transitivity between the targets in the same scene; The temporal causal relationship representation combines the constructed causal relationship representation and the temporal relationship representation, and obtains the target temporal causal structure diagram according to the combined representation.

3. According to the target detection method driven by target temporal causal relationship in industrial multi-scenario according to claim 1, it is characterized in that: The temporal causal relationship is quantitatively expressed. When the confounding factor is determined, the target feature probability distribution representation of the three target data sets at time t in the target temporal causal structure diagram can be obtained according to the causal relationship expression and the do operator: , In the formula, represents the fusion operation of the target features, t and t-1 represent the time t and time t-1 of the scene respectively, Indicates the causal relationship between scenes, , and They represent the security monitoring scene target class, the production line inspection scene target class and the quality inspection target class respectively. Represents the target class features of the corresponding scene, Represents the probability of the target class feature of the corresponding scene; According to the target temporal causal structure diagram and the above formula, the temporal causal relationship between targets can be quantitatively represented, thereby establishing a temporal causal relationship quantitative representation diagram.

4. According to the target detection method driven by target temporal causal relationship in industrial multi-scenario according to claim 1, it is characterized in that: The relational attention module converts the temporal relational feature conversion results of the safety monitoring scene target at time t and the temporal relational feature conversion results of the production line inspection scene target at time t. , , the causal relationship feature conversion results of the safety monitoring scene target and the production line inspection scene target to the quality inspection scene target , Input it into the relational attention module, output the attention weight, so that the model can pay more attention to the correct relational expression, and then multiply the relational conversion result by the attention weight and add it to the feature to obtain the relation-enhanced feature: The computational relation attention module first uses a Gaussian kernel to calculate the high-confidence target features at time t. With the previous The similarity of the high confidence target feature mean at each moment is then multiplied by the activation function and the result of the relational feature transformation to obtain the result after attention. The formula is as follows: in, represents average pooling, represents the bandwidth parameter of the Gaussian kernel, represents the silu activation function, The weight index representing the relational attention; Then, the obtained weight index is multiplied by the corresponding relationship conversion result and added to the feature to obtain the final prediction feature. The operation of the time series scenario is as follows: in, and They represent the features of the security monitoring scene image and the production line detection scene image at time t+1 after the high confidence target screening operation, represents matrix addition, represents matrix dot product, and They represent the relational attention weights of the safety monitoring scene relational features and the production line inspection scene relational features at time t respectively; The operation of the cause and effect scenario is as follows: in, Represents the features of the quality inspection scene image at time t after the high confidence target screening operation, represents a multi-layer perceptron, represents matrix addition, represents matrix dot product, and They respectively represent the relationship attention weights from the safety monitoring scene target to the quality inspection scene target and from the production line inspection scene target to the quality inspection scene target at time t.

5. According to the target detection method driven by target temporal causal relationship in industrial multi-scenario according to claim 1, it is characterized in that: The target temporal causal cascade Mamba is trained using mean square error loss: The mean square error loss is calculated: Among them, m represents the total training time, , , They represent the prediction features at time t, , , They represent the high confidence target features obtained from the true labels respectively; Finally, the target temporal causal cascade Mamba is integrated into the yolov9 target detection model to obtain the final detection model.

Citation Information

Patent Citations

  • Lightweight cervical cancer image cell detection system based on causal attention

    CN115761218A

  • Time sequence causal discovery method and device for digital agricultural information, medium and equipment

    CN116401291A