Lightweight human-computer interaction model system for industrial scene application

Through the full-process knowledge distillation of teacher and student model architecture, the problems of high model complexity and ambiguous recognition of interactive areas in industrial scenarios are solved, realizing lightweight and real-time response industrial human-computer interaction that is compatible with edge devices.

CN121213892BActive Publication Date: 2026-02-13HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511697785.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-13
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing human-computer interaction methods in industrial scenarios suffer from defects in model architecture and lightweight strategies, resulting in large model parameter scale, low computational efficiency, difficulty in deployment on small industrial equipment with limited resources, and ambiguous recognition of interaction areas, failing to meet real-time response requirements.

Method used

By adopting a teacher and student model architecture, and through a full-process knowledge distillation process including initial feature separation, regional feature enhancement distillation, action interaction association distillation, and interaction feature merging and parsing distillation, the student model is optimized to adapt to edge devices, achieving accurate localization of human-computer interaction areas and efficient discrimination of action interaction results.

Benefits of technology

It significantly reduces model complexity, making it embeddable in small industrial equipment, improving the accuracy of interactive area recognition and real-time response capabilities, meeting the requirements of lightweight and real-time performance, and adapting to edge device resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121213892B_ABST
    Figure CN121213892B_ABST
Patent Text Reader

Abstract

The application provides a lightweight human-computer interaction model system for industrial scene application, and relates to the technical field of industrial intelligence and computer vision, and comprises the following steps: inputting a to-be-processed image into a teacher model and a student model respectively, extracting initial image features through a backbone network, a feature flattening layer and a Transformer encoder, and then separating the initial image features through a Transformer decoder to obtain three features corresponding to a person path, an object path and an action interaction path of the teacher model and the student model respectively.The application realizes accurate positioning of a human-computer interaction area in an industrial scene, efficient discrimination of an action interaction result, and at the same time, guarantees the lightweight and real-time performance of the model, and adapts to the deployment requirements of edge industrial equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of industrial intelligence and computer vision technology, in particular to a lightweight human-computer interaction model system for industrial scene application. BACKGROUND

[0002] In the process of industrial intelligent transformation, human-computer interaction discrimination as the core link of industrial scene intelligent perception and automatic control, its performance directly affects the production efficiency improvement, operation safety guarantee and process optimization, for example, in the pipeline assembly, equipment operation and maintenance scene, accurate identification of the interaction area and action type of the operator and the tool, workpiece grabbing, tightening and placing, etc. is the key prerequisite to realize the process automation verification and risk warning. The following are the problems existing in the prior art: the core problems of the prior art are concentrated in the defects of model architecture and lightweight strategy. The current mainstream human-computer interaction method mainly relies on a single complex large model to realize human-computer interaction and action interaction discrimination. The existing technology generally realizes human-computer interaction detection technology based on computer vision and deep learning.

[0003] However, there is a significant adaptation gap between the existing technology and the actual needs of the industrial scene. The traditional human-centered interaction detection method has many shortcomings when dealing with industrial scene characteristics. The existing human-computer interaction method generally does not use knowledge distillation technology to realize model lightweight, or the constructed distillation system has significant limitations. The existing human-computer interaction method emphasizes precision and neglects efficiency, lacks systematization and adaptability, and cannot adapt to edge device resources. For the model, a large parameter size and a complex calculation process are designed to pursue high precision, which directly leads to a sharp increase in computing power consumption and makes it difficult to deploy on resource-limited small industrial devices, such as edge terminals and embedded controllers. This forms a sharp conflict with the demand for low cost and lightweight in industrial scenes. For the lack of sufficient strengthening of the relevance between human, object and action interaction in the interaction area, the accurate interaction area recognition of the operating tool and the workpiece in the industrial scene is blurred. The operation process is not designed in a standardized manner, and there is a lack of effective feature separation and strengthening mechanism, which makes it difficult for the model to adapt to standardized and repetitive operation characteristics in the industrial scene and affects the accurate extraction of interaction features. On the one hand, traditional methods are mostly limited to single feature or single stage knowledge transfer even if they try distillation, and cannot cover multi-link distillation of the whole process of human-computer interaction detection. On the other hand, they completely ignore the specific fine-grained feature interaction requirements of the industrial scene, and do not design special distillation strategies for industrial-specific tasks such as accurate positioning of human-computer interaction area and standardized interaction action discrimination, resulting in gaps in the distillation process and missing feature interaction distillation links, which makes it difficult for the model to meet the real-time response requirements of the industrial scene, and it cannot adapt to edge device resources. SUMMARY

[0004] To solve the problems of low interaction recognition accuracy caused by insufficient correlation of people, objects and action interaction, missing feature processing mechanism, limitations of distillation system, poor model adaptability and insufficient real-time response in the prior art in industrial scenarios, the present application provides a lightweight human-computer interaction model for industrial scenario applications, which realizes accurate positioning of human-computer interaction areas in industrial scenarios, efficient discrimination of action interaction results, and guarantees lightweight and real-time of the model, and adapts to the deployment requirements of edge industrial equipment.

[0005] To solve the above technical problems, the technical solutions of the present application are as follows:

[0006] In a first aspect, the lightweight human-computer interaction model system for industrial scenario applications comprises:

[0007] The acquisition module is configured to input the to-be-processed image into the teacher model and the student model respectively, extract initial image features through the backbone network, the feature flattening layer and the Transformer encoder, and then separate the initial image features through the Transformer decoder to obtain three features corresponding to the person path, the object path and the action interaction path of the teacher model and the student model respectively.

[0008] The optimization module is configured to process the three features obtained at the human-computer interaction first distillation of the region feature strengthening distillation module through the intermediate feature optimization unit and the tail distillation unit of the teacher model and the student model respectively, input the person path feature 2S and the object path feature 2S obtained by processing the student model through the intermediate feature optimization unit into the tail distillation unit of the teacher model, obtain the mapping person path feature and the mapping object path feature, and calculate the first feature classification distillation loss and the first feature regression distillation loss of the mapping feature and the person path feature 3T and the object path feature 3T output by the tail distillation unit of the teacher model, to obtain the loss at the human-computer interaction first distillation, so as to optimize the corresponding person path feature 3S and the object path feature 3S output by the student model.

[0009] The strengthening module is configured to input the three features 3T of the teacher model into the region feature strengthening processing center based on the three features 3S of the optimized student model and the three features 3T of the teacher model, and obtain the three region feature strengthening representation blocks and the three features 4T of the teacher model by strengthening processing, so that the student model outputs the optimized three features 4S by sharing the three region feature strengthening representation blocks to the corresponding feature paths of the student model for learning and optimization.

[0010] The sharing module is configured to process the person path feature 4T, the object path feature 4T and the action interaction path feature 4T of the teacher model in pairs and among the three in the action interaction correlation distillation module based on the three features 4T of the teacher model and the three features 4S of the student model, to obtain four groups of correlation features, integrate the four groups of correlation features to obtain the teacher interaction correlation feature 5T, and share the teacher interaction correlation feature 5T with the student model.

[0011] The computing module is used for shared teacher interaction related features 5T. The features corresponding to each path in the three-path features 4T of the teacher model and the three-path features 4S of the student model at the end of the action interaction related distillation module are respectively subjected to minimum feature distribution distillation loss calculation. The student model is optimized based on the loss to obtain the three-path features 5S of the student model;

[0012] The analysis module is used for the obtained three-path features 5S of the student model. The three-path features 4T of the teacher model and the three-path features 5S of the student model are respectively subjected to feature merging and analysis processing in the interaction feature merging and analysis distillation module with the teacher interaction related features 5T to obtain the three-path features 6T of the teacher model and the three-path features 6S of the student model. Subsequently, the object prediction distillation loss calculation is performed on the human path features 6T and the object path features 6T of the teacher model and the student model, and the action interaction related distillation loss calculation is performed on the action interaction path features 6T and 6S of the teacher model and the student model. The student model is jointly optimized based on the loss, and the three-path features 7S of the student model are output.

[0013] The processing module is used for the three-path features 7S of the student model. The three-path features 7S of the student model are input into the detection head for processing to output the interaction region positioning information of the human path and the object path, the human-machine interaction judgment result in the industrial scene, and the total distillation loss of the model.

[0014] Further, the to-be-processed image is input into the backbone networks of the teacher model and the student model to obtain respective initial image features, and the teacher model human path feature 1T, the object path feature 1T, and the action interaction path feature 1T are further obtained from the respective initial image features. The student model human path feature 1S, the object path feature 1S, and the action interaction path feature 1S are obtained, and the formulas are as follows:

[0015] ;

[0016] wherein, represents a feature vector, m is a feature path, m=R represents a human path feature path, m=W represents an object path feature path, m=J represents an action interaction path feature path, A is a model index, A=T is a teacher model, and A=S is a student model, is an initial image feature, is a human path main body intersection edge under the A model, is an object path main body intersection edge under the A model, is an action interaction path main body intersection edge under the A model, is an adaptive dynamic parameter of the m feature path in the industrial scene under the A model; is a feedforward calculation, is a scaled dot-product attention weight calculation; is a query parameter matrix; is a key parameter matrix; is a value parameter matrix; is the transpose of the teacher model human road key feature matrix, is an attention scaling factor; is a graph neural network operator, wherein, is a set of subject intersection edges, including human road and object, human road and action road, and object road and action road subject intersection edges, is an edge convolution operation.

[0017] Further, the human-computer interaction first distillation of the region feature strengthening distillation module is processed, and the mapped human road and object road features are obtained by processing, and the first feature classification and regression distillation loss calculation is performed, to obtain the human-computer interaction first distillation loss, as follows:

[0018] The human road feature 2S and the object road feature 2S of the student model processed by the intermediate feature optimization unit are transported to the tail distillation unit of the teacher model to generate the mapped human road and object road features, as follows:

[0019] ;

[0020] ;

[0021] wherein, , are the mapped human road feature and the mapped object road feature, is a convolution calculation of the tail distillation unit of the teacher model, represents a multi-layer perception machine, and here is a scene gating operator, specifically , and here is a Sigmoid activation function, and SceneEmbed is a scene encoding; is a human road feature multi-scale attention, is a different window scale size; is a dynamic biasing operator, specifically,

[0022] , and here is an initialized fixed bias vector, is an object road feature global average pooling, is a self-attention calculation;

[0023] The first feature classification and regression distillation loss calculation is performed, and the human-computer interaction first distillation loss is obtained, as follows:

[0024] ;

[0025] ;

[0026] ;

[0027] is the first feature classification distillation loss, B is the number of samples for single distillation calculation, k is the sample index, is the divergence calculation, is the classification prediction probability of the teacher model under the m feature channel, is the standard representation of KL divergence, the left side is the target distribution, the right side is the distribution to be matched, the whole represents the divergence of the latter relative to the former, which measures the difference between the two probability distributions; is the classification prediction probability of the student model under the m feature channel, is the classification quality score of the student model under the m feature channel, is the focusing parameter of the QFL loss, , is the indicator function, that is, sample k is 1 for positive example and 0 for negative example; is the first feature regression distillation loss, is the boundary positioning function, is the boundary positioning regression prediction of the teacher model under the m feature channel, is the boundary positioning regression prediction of the student model under the m feature channel, is the boundary positioning probability distribution of the teacher model under the m feature channel, is the boundary positioning probability distribution of the student model under the m feature channel; is the first distillation loss of human-computer interaction, is the square of the L2 norm, is the classification distillation loss weight, is the regression distillation loss weight, is the mapping human road feature L2 loss weight, is the mapping object road feature L2 loss weight, respectively, the human road feature 3T of the teacher model, the object road feature 3T.

[0028] Further, the regional feature enhancement processing center of the regional feature enhancement distillation module is as follows:

[0029] The regional feature enhancement representation block is as follows:

[0030] ;

[0031] Among them, is the human road regional feature enhancement representation block, is an action-guided attention operator, is a channel attention feature aggregation operator, C is the number of channels, is a human road weight parameter, denotes a parameterized transformation operator, denotes the L2 norm;

[0032] ;

[0033] wherein, is an object road region feature enhancement representation block, is a rectified linear unit activation function, is an object road feature noise gating operator for suppressing non-human object interactions, is an object road weight parameter, denotes a parameterized transformation operator, denotes the L1 norm;

[0034] ;

[0035] wherein, is an action interaction road region feature enhancement representation block, is a spatial position encoding, is an action interaction road weight parameter, is a teacher model action interaction road feature 3T, denotes a spatial dimension convolution operation, is a scaling factor;

[0036] ;

[0037] wherein, m=R, m=W, m=J respectively represent the teacher model human road feature 4T, object road feature 4T, action interaction road feature 4T, m=R, m=W, m=J respectively represent the human road region feature enhancement representation block, the object road region feature enhancement representation block, the action interaction road region feature enhancement representation block, is a residual weight;

[0038] The region feature enhancement representation block is shared into the student model, and the formula is as follows:

[0039] ;

[0040] wherein, m=R, m=W, m=J respectively represent the student model human road feature 4S, object road feature 4S, action interaction road feature 4S, and when m=R, m=W, m=J respectively represent the student model human road feature 3S, object road feature 3S, action interaction road feature 3S, The dynamic channel priority attention operator is specifically , is a global pooling operation, which is used to reduce the dimension of the feature map and extract global semantic features.

[0041] Further, the action interaction correlation distillation module specifically includes the following:

[0042] Correlation features 1-4 are as follows:

[0043] ;

[0044] wherein, is correlation feature 1, is a spatial first weight, is a spatial first bias, is a spatial second weight, is a spatial second bias, is a feature concatenation operation of the teacher model person path feature 4T and the object path feature 4T;

[0045] ;

[0046] wherein, is correlation feature 2, is layer normalization processing, is a feature concatenation operation of the teacher model object path feature 4T and the action interaction path feature 4T, is a feature cross attention calculation; similarly, the correlation feature 3 obtained by the teacher model person path feature 4T and the action interaction path feature 4T is represented by ;

[0047] ;

[0048] wherein, is correlation feature 4, is two-layer perception machine processing, is a feature concatenation operation of the teacher model object path feature 4T, object path feature 4T and action interaction path feature 4T, is a bit-encoding sine-cosine function, loc is a bit code, D is a model dimension, and i is a loop variable;

[0049] The teacher interaction correlation feature 5T is as follows:

[0050] ;

[0051] wherein, n = 1, 2, 3, 4, represents the index of the interaction features 1-4, is the teacher interaction correlation feature 5T, is an exponential function, is calculated for the multi-layer perceptron of attention weight, is calculated for the multi-layer perceptron of interaction feature n;

[0052] Further, the end processing of the action interaction correlation distillation module is as follows:

[0053] ;

[0054] wherein, When m = R, m = W, m = J represent the minimum feature distribution distillation loss of the person path feature 4T, the object path feature 4T, and the action interaction path feature 4T between the teacher model and the student model respectively, is the three-path feature 4T of the teacher model under the k sample index, is the three-path feature 4S of the student model under the k sample index, is the square of the Frobenius norm distance of the Gram matrix, and D is the dimension, is the transpose of the three-path feature of the teacher model, is the transpose of the three-path feature of the student model, is the temperature scaling value, are the three-path feature sample-by-sample L2 loss weight, the Frobenius norm distance loss weight, and the KL divergence loss weight between the teacher model and the student model respectively;

[0055] The three-path feature 5S of the student model is obtained by optimization through the minimum feature distribution distillation loss calculation between the three-path features of the teacher model and the student model, and the formula is as follows:

[0056] ;

[0057] wherein the optimization correction term is expanded as:

[0058] ;

[0059] wherein, is the learning rate parameter, is the optimization correction term, is the back propagation calculation, and when m = R, m = W, m = J represent the person path feature 5S, the object path feature 5S, and the action interaction path feature 5S of the student model respectively, represents the dimension scale of the feature.

[0060] Further, the interactive feature merging and analysis distillation module is as follows:

[0061] The object prediction distillation loss calculation is performed on the features 6T, 6S of the person path and the object of the teacher and student models respectively, and the formula is as follows:

[0062] ;

[0063] wherein, wherein, m = R, m = W, m = J represent the teacher model person path feature 6T, object path feature 6T, action interaction path feature 6T respectively, wherein, m = R, m = W, m = J represent the student model person path feature 6T, object path feature 6T, action interaction path feature 6T respectively, wherein, m = R, m = W represent the object prediction distillation loss of the person path feature 6T, 6S between the teacher model and the student model, and the object path feature 6T, 6S, are the teacher model person path feature 6T, object path feature 6T under the k sample index, and are the student model person path feature 6S, object path feature 6S under the k sample index, are the person path feature and object path feature sample-by-sample L2 loss weight between the teacher model and the student model, and the Frobenius norm distance loss weight of the person path feature and object path feature between the teacher model and the student model;

[0064] The action interaction correlation distillation loss calculation is performed between the action interaction path features 6T, 6S of the teacher and student models, and the formula is as follows:

[0065] ;

[0066] wherein, is the action interaction correlation distillation loss, is the two-layer MLP action interaction feature extraction, is the teacher model action interaction path feature 6T under the k sample index, is the student model action interaction path feature 6S under the k sample index, is the action interaction path feature covariance calculation.

[0067] Further, the total distillation loss calculation is as follows:

[0068] ;

[0069] wherein, is the total distillation loss, , , , are the person-machine interaction first distillation loss weight, minimum feature distribution distillation loss weight, object prediction distillation loss weight, and action interaction correlation distillation loss calculation respectively.

[0070] The person-machine interaction discrimination result of the industrial scene is as follows:

[0071] ;

[0072] wherein, is the human-computer interaction discrimination result of the industrial scene, is the action interaction discrimination classification multilayer perception, is dynamically weighted fusion, is a dynamic fusion weight, and C is the number of action interaction categories, is the optimized student model three-path feature, and when m=R, m=W, and m=J, 7S, 7S, and 7S represent the student model person path feature, the object path feature, and the action interaction path feature, respectively.

[0073] In view of the problems that the prior art relies on a single complex large model, the parameter scale is large, the calculation efficiency is low, and it is difficult to deploy in resource-limited small industrial equipment, which conflicts with the lightweight demand, the present application constructs a teacher and student model architecture, through a full-process knowledge distillation covering initial feature separation, regional feature strengthening distillation, action interaction correlation distillation, and interactive feature merging and analysis distillation, the complex interactive knowledge of the teacher model is migrated to the student model in stages, which significantly reduces the model complexity, so that it can be embedded in small industrial equipment, meeting the low-cost and lightweight demand, in view of the problems that the prior art lacks sufficient correlation of human, object and action interaction in industrial scenarios, lacks effective feature separation and strengthening mechanism, leading to blurred interactive area recognition and difficulty in adapting to standardized operation characteristics, the present application separates the initial image features into three features of human, object and action interaction, realizes targeted extraction and strengthening, designs a regional feature strengthening distillation module to optimize human and object features, and improves the interactive area recognition accuracy, with the help of an action interaction correlation distillation module, the three features are correlated and shared, the correlation modeling is strengthened, and the industrial operation strong process demand is adapted, in view of the problems that the prior art does not use knowledge distillation or the distillation system is limited to a single feature or a single stage, does not cover the whole process, and ignores the industrial fine-grained feature interaction demand, leading to breakpoints in the distillation process, the present application designs a multi-link distillation system covering the whole process, including regional feature strengthening, action interaction correlation, and interactive feature merging and analysis stages, and in view of the industrial fine-grained demand, special strategies are designed at each stage, such as calculating the human-computer interaction first distillation loss at the regional feature strengthening stage, calculating the minimum feature distribution distillation loss, object prediction distillation loss, and action interaction correlation distillation loss at the action interaction correlation stage, to ensure that the distillation has no breakpoints, fill the gap of industrial scene fine-grained feature interaction distillation, in view of the problems that the prior art emphasizes precision and ignores efficiency, lacks systematization and adaptability, leading to difficulty in meeting the real-time response requirements of industrial scenarios, and inability to adapt to edge device resources, resulting in poor scene adaptability and high deployment cost, the present application simplifies the student model parameters and calculation process through full-process knowledge distillation, improves the reasoning speed to meet real-time response, the student model can adapt to edge devices to reduce deployment cost, and at the same time, the scene adaptability is enhanced through the design of strengthening interactive area recognition and action correlation modeling, promoting the landing of human-computer interaction technology in industrial intelligence. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 is a system schematic diagram of a lightweight human-computer interaction model system for industrial scene application provided by an embodiment of the present application. DETAILED DESCRIPTION

[0075] Exemplary embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is to be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0076] As shown in Figure 1 Exemplary embodiments of the present disclosure propose a lightweight human-computer interaction model system for industrial scene applications, which includes:

[0077] The acquisition module is configured to input the to-be-processed image into the teacher model and the student model respectively, extract initial image features through a backbone network, a feature flattening layer, and a Transformer encoder, and separate the initial image features through a Transformer decoder to obtain three features corresponding to the human path, the object path, and the action interaction path of the teacher model and the student model respectively.

[0078] The optimization module is configured to process the three features obtained at a human-computer interaction first distillation of a regional feature strengthening distillation module through an intermediate feature optimization unit and a tail distillation unit for the features of the teacher model and the student model. The human path feature 2S and the object path feature 2S obtained by processing the student model through the intermediate feature optimization unit are input into the tail distillation unit of the teacher model to obtain a mapping human path feature and a mapping object path feature. The mapping features and the human path feature 3T and the object path feature 3T output by the tail distillation unit of the teacher model are subjected to first feature classification distillation loss and first feature regression distillation loss calculation to obtain a human-computer interaction first distillation loss, so as to optimize the corresponding human path feature 3S and the object path feature 3S output by the student model.

[0079] The strengthening module is configured to input the three features 3T of the teacher model into a regional feature strengthening processing center based on the three features 3S of the student model and the three features 3T of the teacher model after optimization, and obtain three regional feature strengthening representation blocks and the three features 4T of the teacher model through strengthening processing. The three regional feature strengthening representation blocks are shared to the corresponding feature paths of the student model for learning and optimization, so as to output the three features 4S of the student model after optimization.

[0080] The sharing module is configured to perform interaction correlation processing between the three features 4T of the teacher model and the three features 4S of the student model in an action interaction correlation distillation module, and obtain four groups of correlation features, and obtain the teacher interaction correlation features 5T after integration, and share the teacher interaction correlation features 5T to the student model.

[0081] The computing module is used for shared teacher interaction related features 5T. The features corresponding to each path in the three-path features 4T of the teacher model and the three-path features 4S of the student model at the end of the action interaction related distillation module are respectively subjected to minimum feature distribution distillation loss calculation. The student model is optimized based on the loss, and the three-path features 5S of the student model are obtained.

[0082] The analyzing module is used for the obtained three-path features 5S of the student model. The three-path features 4T of the teacher model and the three-path features 5S of the student model are respectively subjected to feature merging and analyzing processing in the interaction feature merging and analyzing distillation module with the teacher interaction related features 5T, so as to obtain the three-path features 6T of the teacher model and the three-path features 6S of the student model. Subsequently, the object prediction distillation loss calculation is performed on the human-path features 6T and the object-path features 6T of the teacher model and the student model, and the action interaction related distillation loss calculation is performed on the action interaction path features 6T and 6S of the two models. The student model is jointly optimized based on the loss, and the three-path features 7S of the student model are output.

[0083] The processing module is used for the three-path features 7S of the student model. The three-path features 7S of the student model are input into a detection head for processing. The interaction region positioning information of the human-path and the object-path, the human-machine interaction judgment result in the industrial scene, and the total distillation loss of the model are output.

[0084] In the embodiment of the present application, because the present application adopts a teacher-student dual model collaborative architecture, the initial image feature extraction and three-path feature splitting of human path, object path and action interaction path are realized by the acquisition module with backbone network, feature flattening layer and Transformer encoder-decoder, the student feature is optimized by the intermediate feature optimization unit and the tail distillation unit in the regional feature strengthening distillation module in the optimization module, the three-path regional feature strengthening representation blocks are generated by the regional feature strengthening processing center in the strengthening module and shared to the student model, four groups of associated features are obtained by the two-by-two and three-path interaction of the three-path features of the teacher model in the action interaction associated distillation module in the sharing module, and the four groups of associated features are integrated into the teacher interaction associated feature 5T shared, the student feature is optimized by calculating the minimum feature distribution distillation loss at the end of the action interaction associated distillation module in the calculation module, the student feature is optimized by feature merging analysis and object prediction or action interaction associated distillation loss calculation in the interactive feature merging analysis distillation module in the analysis module, and the three-path features 7S of the student model are input to the detection head for processing by the processing module, so the problems of large parameter size and difficulty in deploying edge devices caused by relying on a single complex large model in the existing industrial human-computer interaction technology are overcome, the interactive region recognition ambiguity problem caused by the insufficient interaction association between human and object and action is solved, the defects of the distillation system being limited to single feature or single stage and having breakpoints are made up, the technical problems of model precision being light and efficiency being light, and the inability to meet the industrial real-time response and poor scene adaptability are solved, and then the model lightweight, human-computer interaction region accurate positioning in industrial scenes, and action interaction result efficient discrimination are realized, while the model precision and real-time performance are taken into account, thereby providing reliable technical support for human-computer interaction applications in industrial intelligent scenes.

[0085] In a preferred embodiment of the present application, the to-be-processed image is input into the backbone network of the teacher and student models to obtain respective initial image features, and the respective initial image features are further processed to obtain the human path feature 1T, the object path feature 1T, and the action interaction path feature 1T of the teacher model, and the human path feature 1S, the object path feature 1S, and the action interaction path feature 1S of the student model, and the formula is as follows:

[0086] ;

[0087] wherein, represents a feature vector, m is a feature path, m=R represents a human path feature path, m=W represents an object path feature path, m=J represents an action interaction path feature path, A is a model index, A=T is a teacher model, and A=S is a student model, is an initial image feature, is a human path main body intersection edge of the A model, is an object path main body intersection edge of the A model, is an A model under action interaction road main body intersection edge, is an A model under m feature path in industrial scene adaptive dynamic parameter; is a feedforward calculation, is a scaled dot-product attention weight calculation; is a query parameter matrix; is a key parameter matrix; is a value parameter matrix; is a teacher model person road key feature matrix transpose, is an attention scaling factor; is a graph neural network operator, wherein, is a main body intersection edge set, there are person road and object, person road and action road, object road and action road main body intersection edges, is an edge convolution operation.

[0088] In the embodiment of the present application, because in a preferred embodiment of the present application, the to-be-processed image is respectively input into the teacher and student model backbone network to obtain the respective initial image features, the initial features are processed by a formula containing explicit feature paths and model indexes, the formula incorporates feedforward calculation, scaled dot-product attention weight calculation, graph neural network operator, and sets the industrial scene adaptive dynamic parameters of m feature paths under the A model, and the teacher and student models each have person road, object road, and action interaction road query matrix, key matrix transpose, value matrix calculation, and attention scaling factor, so the technical problems of the prior art, such as lack of feature hierarchical splitting mechanism for industrial scene, resulting in that person, object, and action interaction features are difficult to distinguish, not effectively capturing the correlation between person, object, and action, and lack of industrial scene exclusive adaptive parameters, are effectively overcome, and the feature processing is disconnected from the industrial demand, so that the teacher and student models can accurately extract independent person road features 1T / 1S, object road features 1T / 1S, and action interaction road features 1T / 1S, and the interaction correlation modeling between the three types of features is strengthened, and the adaptability of the features to the industrial scene is improved through the industrial adaptive dynamic parameters, providing accurate feature input for subsequent regional feature reinforcement distillation and other links.

[0089] In a preferred embodiment of the present application, the human-computer interaction first distillation of the regional feature reinforcement distillation module is processed, and the mapped person road and object road features are obtained, and the first feature classification and regression distillation loss calculation is performed, to obtain the human-computer interaction first distillation loss, which is as follows:

[0090] The person road features 2S and object road features 2S of the student model processed by the intermediate feature optimization unit are transported to the tail distillation unit of the teacher model, to generate mapped person road and object road features, and the formula is as follows:

[0091] ;

[0092] ;

[0093] wherein, , are the mapping human road features, the mapping object road features, is a teacher model tail distillation part convolution calculation, denotes a multi-layer perception, here is a scene gating operator, specifically , here is a Sigmoid activation function, and SceneEmbed is a scene encoding; is a human road feature multi-scale attention, is a different window scale size; is a dynamic bias operator, specifically,

[0094] , here is an initialized fixed bias vector, is an object road feature global average pooling, is a self-attention calculation, and the object road feature 2S of the student model is calculated inside the whole ;

[0095] The first feature classification and regression distillation loss is calculated, and the human-computer interaction first distillation place loss is obtained, and the formula is as follows:

[0096] ;

[0097] ;

[0098] ;

[0099] is the first feature classification distillation loss, B is the number of samples for single distillation calculation, k is the sample index, is a divergence calculation, is the classification prediction probability of the teacher model under the m feature channel, is the standard expression of KL divergence, the left side is the target distribution, the right side is the distribution to be matched, the whole represents the divergence of the latter relative to the former, which measures the difference between the two probability distributions; is the classification prediction probability of the student model under the m feature channel, the mapping road feature is mapped in the teacher model, is the classification quality score of the student model under the m feature channel, the mapping road feature is mapped in the teacher model, is the focusing parameter of the QFL loss, , is an indicator function, that is, sample k is 1 for positive example, and sample k is 0 for negative example. is a first feature regression distillation loss, is a boundary location function, is a boundary location regression prediction of the teacher model under the m feature channel, is a boundary location regression prediction of the student model under the m feature channel, is a boundary location probability distribution of the teacher model under the m feature channel, is a boundary location probability distribution of the student model under the m feature channel. is a human-machine interaction first distillation place loss, is a square of an L2 norm, is a classification distillation loss weight, is a regression distillation loss weight, is a mapping human road feature L2 loss weight, is a mapping object road feature L2 loss weight, are respectively a human road feature 3T and an object road feature 3T of the teacher model.

[0100] In the embodiment of the present application, because in a preferred embodiment of the present application, the human-machine interaction first distillation place of the regional feature reinforcement distillation module adopts the human road feature 2S and the object road feature 2S after the student model is processed by the intermediate feature optimization unit, and the human road feature 2S and the object road feature 2S are transported to the tail distillation unit of the teacher model, the mapping human road feature and the mapping object road feature are generated through the convolution calculation of the tail distillation unit of the teacher model, the scene gating operator, the human road feature multi-scale attention, and the dynamic bias operator, and the first feature classification distillation loss and the first feature regression distillation loss are used, and the two types of losses and the mapping human road or object road feature L2 loss are integrated to obtain the technical means of the human-machine interaction first distillation place loss, so the technical problems of the prior art are overcome, such as the fuzzy regional feature of human-object interaction in the industrial scene, the poor consistency of the student model and the teacher model in the human road / object road feature classification and spatial positioning, and the lack of noise filtering and multi-dimensional distillation loss constraint for the industrial scene, which leads to inaccurate interactive feature optimization, and thus the mapping human road / object road feature that adapts to the feature distribution of the teacher model is accurately generated, the human road feature 3S and the object road feature 3S output by the student model are optimized through multi-dimensional loss cooperation, the spatial positioning accuracy of the student model to the human-object interaction region in the industrial scene and the consistency of human-object classification are improved, and a key foundation is laid for the regional feature reinforcement processing and the lightweight and high-precision adaptation of the overall model to the industrial scene.

[0101] In a preferred embodiment of the present application, the regional feature reinforcement processing center of the regional feature reinforcement distillation module is as follows:

[0102] The regional feature reinforcement representation block is as follows:

[0103] ;

[0104] wherein, is a human path region feature enhancement representation block, is an action guidance attention operator, is a channel attention feature aggregation operator, C is the number of channels, is a human path weight parameter, denotes a parameterized transformation operator, denotes L2 norm;

[0105] ;

[0106] wherein, is an object path region feature enhancement representation block, is a rectified linear unit activation function, is an object path feature noise gating operator for suppressing non-human interaction, is an object path weight parameter, denotes a parameterized transformation operator, denotes L1 norm;

[0107] ;

[0108] wherein, is an action interaction path region feature enhancement representation block, is a spatial position encoding, is an action interaction path weight parameter, is a teacher model action interaction path feature 3T, denotes a spatial dimension convolution operation, is a scaling factor;

[0109] ;

[0110] wherein, when m = R, m = W, m = J respectively represent the teacher model human path feature 4T, object path feature 4T, action interaction path feature 4T, when m = R, m = W, m = J respectively represent the human path region feature enhancement representation block, the object path region feature enhancement representation block, the action interaction path region feature enhancement representation block, is a residual weight;

[0111] The region feature enhancement representation block is shared into the student model, and the formula is as follows:

[0112] ;

[0113] wherein, when m = R, m = W, m = J respectively represent the student model human path feature 4S, object path feature 4S, action interaction path feature 4S, when m=R, m=W, m=J represent the student model's human path feature 3S, object path feature 3S, and action interaction path feature 3S, respectively, is a dynamic path priority attention operator, specifically , is a global pooling operation, used for dimension reduction of the feature map to extract global semantic features.

[0114] In the embodiment of the present application, because in a preferred embodiment of the present application, the regional feature enhancement processing center of the regional feature enhancement distillation module adopts a regional feature enhancement representation block generated differently for the three path features 3T of the teacher model, for the human path feature 3T, a human path regional feature enhancement representation block is generated through the action guidance attention operator, the channel attention feature aggregation operator, the human path weight parameter, and the L2 norm; for the object path feature 3T, an object path regional feature enhancement representation block is generated through the rectified linear unit activation function, the CrossPathGate gating operator for suppressing non-human interaction noise, the object path weight parameter, and the L1 norm; for the action interaction path feature 3T, an action interaction path regional feature enhancement representation block is generated through spatial position encoding, the action interaction path weight parameter, and spatial dimension convolution operation. At the same time, each RFECB block and the corresponding three path features 3T of the teacher model are fused through residual weight to obtain the three path features 4T of the teacher model, and the three path RFECB blocks are shared with the student model. Combined with the dynamic path priority attention operator, the student model three path features 3S are obtained through global pooling and MLP to obtain the student model three path features 4S. Therefore, the present application overcomes the problems in the prior art that the different features of the human and the object and the action in the industrial scene are not differentiated for feature enhancement, the object path feature is easily disturbed by non-interaction noise, the action interaction path feature lacks spatial correlation modeling, and the key enhanced features of the teacher model cannot be efficiently transmitted to the student model, resulting in insufficient precision of the student features. The present application precisely and differentially enhances each path feature of the teacher model, effectively filters irrelevant noise of the object path, enhances the spatial correlation of the action interaction path, efficiently transmits the key enhanced feature knowledge of the teacher model to the student model, and makes the student model generate optimized three path features 4S, providing high-precision feature input for the action interaction correlation distillation link, and improving the ability of the overall model to capture and distinguish the human-machine interaction features in the industrial scene.

[0115] In a preferred embodiment of the present application, the action interaction correlation distillation module specifically includes the following:

[0116] The correlation features 1 to 4 are as follows:

[0117] ;

[0118] wherein, is the correlation feature 1, is the first spatial weight, a spatial first bias, a spatial second weight, a spatial second bias, a feature concatenation operation of the teacher model person track feature 4T and the object track feature 4T.

[0119] ;

[0120] wherein, is the associated feature 2, is a layer normalization process, is a feature concatenation operation of the teacher model object track feature 4T and the action interaction track feature 4T, is a feature cross-attention calculation. Similarly, the associated feature 3 obtained by the teacher model person track feature 4T and the action interaction track feature 4T can also be represented as ;

[0121] ;

[0122] wherein, is the associated feature 4, is a two-layer perception machine process, is a feature concatenation operation of the teacher model object track feature 4T, the object track feature 4T and the action interaction track feature 4T, is a bit-encoding sine-cosine function, loc is a bit code, D is a model dimension, and i is a loop variable.

[0123] The teacher interaction associated feature 5T is as follows:

[0124] ;

[0125] wherein, n = 1, 2, 3, 4, represents the index of the interaction feature 1 to 4, is the teacher interaction associated feature 5T, is an exponential function, is a multi-layer perception machine calculation of the attention weight, is a multi-layer perception machine calculation of the interaction feature n.

[0126] In the embodiment of the present application, because in a preferred embodiment of the present application, the interactive feature merging and parsing distillation module adopts split-path targeted loss calculation, for the human path feature 6T / 6S and the object path feature 6T / 6S of the teacher model and the student model, the object prediction distillation loss is calculated through the sample-by-sample L2 loss and the Frobenius norm distance loss, to constrain the consistency of the features of the two on the human and object core interactive objects; for the action interactive path feature 6T / 6S of the teacher model and the student model, the key action features are extracted through the two-layer MLP action interactive feature extraction operator, combined with the action interactive path feature covariance calculation, to build the action interactive correlation distillation loss, to constrain the consistency of the two in the action interactive logic and feature distribution, and all loss calculations are based on the sample index k corresponding to the matching features of the teacher and student models, so as to overcome the problems in the prior art that the distillation loss is not designed for different attributes of the human and object and action interaction in the split path, the object prediction result deviates greatly from the teacher model, and the action interaction logic transmission is not accurate, resulting in that the recognition of the key interactive objects is blurred in the industrial scene, and the action judgment is disconnected from the actual operation process, so as to achieve the consistency of the results of the student model and the teacher model in the human-object prediction, to strengthen the recognition accuracy of the student model to the core objects of the operating personnel, tools or workpieces in the industrial scene, and at the same time, to ensure that the student model learns the precise modeling ability of the teacher model to the human and object and the action interaction logic, to provide high-precision feature support for outputting the industrial scene human-computer interaction judgment result.

[0127] In a preferred embodiment of the present application, the end processing of the action interactive correlation distillation module is as follows:

[0128] ;

[0129] Among them, When m=R, m=W, m=J respectively represent the minimum feature distribution distillation loss of the human path feature 4T, 4S, the object path feature 4T, 4S, and the action interactive path feature 4T, 4S between the teacher model and the student model, is the three-path feature 4T of the teacher model under the k sample index, is the three-path feature 4S of the student model under the k sample index, is the Frobenius norm distance square of the Gram matrix, and D is the dimension, is the transpose of the three-path feature of the teacher model, is the transpose of the three-path feature of the student model, is the temperature scaling value, are respectively the three-path feature sample-by-sample L2 loss weight, the Frobenius norm distance loss weight, and the KL divergence loss weight between the teacher model and the student model;

[0130] The three-path feature 5S of the student model is obtained by optimizing the minimum feature distribution loss calculation between the teacher model and the student model three-path features, and the formula is as follows:

[0131] ;

[0132] The optimization correction term is expanded as follows:

[0133] ;

[0134] Wherein, is a learning rate parameter, is an optimization correction term, is a back propagation calculation, and m=R, m=W, m=J respectively represent the student model's human path feature 5S, the object path feature 5S, and the action interaction path feature 5S.

[0135] In the embodiment of the application, the action interaction correlation distillation module generates four groups of correlation features and integrates them into the teacher interaction correlation feature 5T to the teacher model's human path feature 4T and object path feature 4T, obtains the correlation feature through the spatial first weight, the spatial first bias, the spatial second weight, the spatial second bias, and the feature splicing operation; the correlation feature 2 is obtained by layer normalization processing, feature splicing operation, and feature cross attention calculation to the teacher model's object path feature 4T and action interaction path feature 4T, and the correlation feature 3 is generated by the same logic to the human path feature 4T and action interaction path feature 4T; the correlation feature 4 is obtained by two-layer perception machine processing, feature splicing operation, and sine-cosine function of bit encoding to the teacher model's human path, object path, and action interaction path features 4T, and the teacher interaction correlation feature 5T is obtained by weighting and integrating the four groups of correlation features through the exponential function, the multi-layer perception machine calculation of attention weight, and the multi-layer perception machine calculation of interaction feature n, so it overcomes the lack of correlation modeling of human and object, object and action, action and object, and full-dimensional interaction relationship in the prior art, and the lack of hierarchical integration mechanism of interaction features, which leads to the inability of the teacher model to capture the multi-agent collaborative interaction logic in the industrial operation process, and the difficulty of efficiently transferring the core interaction correlation knowledge, achieving comprehensive coverage of the multi-dimensional correlation relationship of human-machine interaction in the industrial scene, accurately capturing the collaborative logic of "operation subject, operation object, and operation behavior", and integrating the core interaction correlation knowledge of the teacher model into a unified teacher interaction correlation feature 5T to provide structured and high-value correlation knowledge support for student model learning, and improve the discriminative logic and accuracy of the model for complex interaction actions in the industrial scene.

[0136] In a preferred embodiment of the application, the interaction feature merging and parsing distillation module is as follows:

[0137] The teacher and student model human route, object feature 6T, 6S respectively carries out object prediction distillation loss calculation, the formula is as follows:

[0138] ;

[0139] Wherein, wherein, When m=R, m=W, m=J respectively represent the teacher model human route feature 6T, object route feature 6T, action interaction route feature 6T, When m=R, m=W, m=J respectively represent the student model human route feature 6T, object route feature 6T, action interaction route feature 6T, When m=R, m=W respectively represent the object prediction distillation loss of human route feature 6T, 6S between the teacher model and student model, object route feature 6T, 6S, The teacher model human route feature 6T, object route feature 6T under the k sample index, the student model human route feature 6S, object route feature 6S under the k sample index, Respectively, the L2 loss weight of human route, object route feature between the teacher model and student model, the Frobenius norm distance loss weight of human route, object route feature between the teacher model and student model;

[0140] The teacher and student model action interaction route feature 6T, 6S carries out action interaction correlation distillation loss calculation, the formula is as follows:

[0141] ;

[0142] Wherein, The action interaction correlation distillation loss, The two-layer MLP action interaction feature extraction, The teacher model action interaction route feature 6T under the k sample index, The student model action interaction route feature 6S under the k sample index, The action interaction route feature covariance calculation, Represent the dimension scale of feature.

[0143] In the embodiment of the present application, the end processing of the action interaction related distillation module adopts multi-dimensional collaborative constraint for the three-path feature 4T of the teacher model and the three-path feature 4S of the student model, corresponding matching according to the sample index k, calculating the minimum feature distribution distillation loss containing the sample-by-sample L2 loss, the Gram matrix Frobenius norm distance square, and the KL divergence loss, and then determining the optimization correction term based on the loss through back propagation and the learning rate parameter η, fusing it with the three-path feature 4S of the student model to obtain the optimized three-path feature 5S of the student model. Therefore, the technical problems of the prior art, such as only single-dimensional constraint of the feature deviation between the student model and the teacher model, not covering the feature value, global distribution, and probability distribution multi-level differences, and lacking of explicit optimization correction mechanism leading to the difficulty of the student model in accurately learning the interactive feature rules of the teacher model and poor consistency of the feature distribution, are overcome. The feature distribution of the student model and the teacher model in the person path, the object path, and the action interaction path is accurately aligned from multiple dimensions, the student model obtains high-precision three-path feature 5S through explicit back propagation optimization, provides high-quality feature input for the interactive feature merging and analysis distillation module, and ensures the precision and robustness of the lightweight student model in capturing the human-machine interaction features in the industrial scene.

[0144] In a preferred embodiment of the present application, the total distillation loss is calculated, and the specific formula is as follows:

[0145] ;

[0146] Among them, is the total distillation loss, , , , , respectively, are the human-machine interaction first distillation loss weight, the minimum feature distribution distillation loss weight, the object prediction distillation loss weight, and the action interaction related distillation loss calculation.

[0147] The human-machine interaction discrimination result in the industrial scene is calculated, and the specific formula is as follows:

[0148] ;

[0149] Among them, is the human-machine interaction discrimination result in the industrial scene, is the action interaction discrimination classification multi-layer perception, is the dynamic weighted fusion, is the dynamic fusion weight, C is the number of action interaction categories, is the optimized three-path feature of the student model, and m=R, m=W, m=J respectively represent the person path feature 7S, the object path feature 7S, and the action interaction path feature 7S of the student model.

[0150] In the embodiment of the present application, because the global loss integration and dynamic feature fusion are adopted to globally integrate the human-computer interaction first distillation loss, the minimum feature distribution distillation loss, the object prediction distillation loss, and the action interaction correlation distillation loss by the total distillation loss formula according to the corresponding weights, the collaborative optimization of each distillation stage of the model is realized; at the same time, through the industrial scene human-computer interaction judgment result formula, the dynamic weighted fusion of the optimized student model three-path features 7S is adopted, the action interaction judgment classification multi-layer perception and the Softmax function output interaction category result are combined, so the problems in the prior art that the global overall planning of the multi-stage distillation loss is lacked, the weight control is not good, the optimization goals of each link are disconnected, the overall precision of the model is limited, and the importance difference of the human and object and the action features in the industrial scene is not dynamically allocated the weight, the feature fusion is not flexible, the matching degree of the judgment result and the actual industrial operation is low, the model whole-process collaborative optimization is realized through the total distillation loss, the overall precision of the lightweight student model is ensured, at the same time, through the dynamic weighted fusion of the three-path features and the special classification MLP, the human-computer interaction judgment result meeting the needs of the industrial scene is accurately output, the recognition accuracy and the scene adaptability of the model to the industrial exclusive interaction action are improved.

[0151] In the specific implementation of the present application, the industrial scene image to be processed is respectively input into the backbone network of the teacher model and the student model. After processing by the backbone network, the teacher model and the student model each obtain initial image features. Based on these initial image features, further processing is performed to obtain the human path feature 1T, the object path feature 1T, and the action interaction path feature 1T of the teacher model, as well as the human path feature 1S, the object path feature 1S, and the action interaction path feature 1S of the student model. The feature paths are divided into three categories. The human path feature path is represented by a specific identifier and mainly focuses on the features related to the operators in the industrial scene. The object path feature path is represented by another specific identifier and mainly focuses on the features related to the tools or workpieces in the industrial scene. The action interaction path feature path is represented by a third specific identifier and mainly focuses on the features related to the interactive actions between the operators and the tools or workpieces in the industrial scene. The model index is used to distinguish between the teacher model and the student model. When the model index is a specific value, it represents the teacher model, and the corresponding initial image feature is the teacher model initial image feature. When the model index is another specific value, it represents the student model, and the corresponding initial image feature is the student model initial image feature. In different models, each feature path has a corresponding main intersection edge. The human path main intersection edge corresponds to the main connection relationship related to the human path feature in the model. The object path main intersection edge corresponds to the main connection relationship related to the object path feature in the model. The action interaction path main intersection edge corresponds to the main connection relationship related to the action interaction path feature in the model. At the same time, each feature path in different models is also provided with dynamic parameters adapted to the industrial scene, which are used to adjust the feature processing process to better adapt to the requirements of the industrial scene. In the feature processing process, two key operations, feedforward calculation and scaled dot-product attention weight calculation, are used. Feedforward calculation is used for further linear transformation and nonlinear processing of features, and scaled dot-product attention weight calculation is used to capture the attention association between features. The human path, object path, and action interaction path features of the teacher model and the student model are subjected to query matrix calculation, key matrix transpose calculation, and value matrix calculation, respectively. These calculations provide a basis for the implementation of the attention mechanism. An attention scaling factor is used to scale the numerical values in the attention calculation process to avoid the influence of excessively large numerical values on the calculation results. A graph neural network operator is also used, which is specifically implemented by performing edge convolution operations on the main intersection edge set. The main intersection edge set includes the main intersection edges between the human path and the object, the human path and the action path, and the object path and the action path. Edge convolution operations can effectively capture the feature association relationships represented by these intersection edges, improving the relevance and effectiveness of the features.

[0152] In the implementation of the present application, in the human-computer interaction first distillation of the regional feature reinforced distillation module, the features of the student model are first processed. The human path feature 2S and the object path feature 2S of the student model are obtained after processing by the intermediate feature optimization unit. The two features are delivered to the tail distillation unit of the teacher model, and the mapped human path feature and the mapped object path feature are generated through the calculation of the tail distillation unit of the teacher model. The tail distillation unit of the teacher model will perform convolution calculation, and the human path feature 2S of the student model can be obtained inside the convolution calculation. Scene gating operators will be used in the process. The operator first processes the scene encoding through a multi-layer perception, then obtains the activation result through a Sigmoid activation function, and finally performs element-wise multiplication operation on the activation result and the human path feature of the student model to achieve the filtering and strengthening of the features related to the industrial scene. The human path feature multi-scale attention will process the human path feature with different window scale sizes to capture the details of the human path feature at different scales. The dynamic bias operator first performs global average pooling processing on the object path feature of the student model, then combines the initialized fixed bias vector, and obtains the dynamic bias result through the multi-layer perception calculation. At the same time, the object path feature is optimized through the self-attention calculation. The object path feature 2S of the student model is obtained in the calculation process of the multi-layer perception inside. After obtaining the mapped human path feature and the mapped object path feature, the first feature classification distillation loss and the first feature regression distillation loss are calculated. The calculation of the first feature classification distillation loss needs to consider the number of samples and the sample index of single distillation calculation. The difference between the classification prediction probability of the teacher model under the corresponding feature path and the classification prediction probability of the student model under the corresponding feature path is calculated through the KL divergence. Combined with the focusing parameter of the QFL loss and the indicator function, the indicator function takes the value of 1 when the sample is a positive example and takes the value of 0 when the sample is a negative example. The first feature classification distillation loss is calculated. The calculation of the first feature regression distillation loss is to calculate the difference between the boundary positioning regression prediction of the teacher model under the corresponding feature path and the boundary positioning regression prediction of the student model under the corresponding feature path. At the same time, the KL divergence between the boundary positioning probability distribution of the teacher model under the corresponding feature path and the boundary positioning probability distribution of the student model under the corresponding feature path is calculated. The two are combined to obtain the first feature classification distillation loss. Combined with the square of the L2 norm, the first feature classification distillation loss is multiplied by the classification distillation loss weight, the first feature regression distillation loss is multiplied by the regression distillation loss weight, the difference between the mapped human path feature and the teacher model human path feature 3T is multiplied by the mapped human path feature L2 loss weight, and the difference between the mapped object path feature and the teacher model object path feature 3T is multiplied by the mapped object path feature L2 loss weight. The sum of these results is the loss at the human-computer interaction first distillation, and the human path feature 3S and the object path feature 3S output by the student model are optimized according to the loss.

[0153] In the embodiment of the present application, in the regional feature enhancement processing center of the regional feature enhancement distillation module, three types of regional feature enhancement representation blocks are first generated. In the generation process of the human path regional feature enhancement representation block, an action-guided attention operator and a channel attention feature aggregation operator are used. The channel attention feature aggregation operator processes the combination result of the human path weight parameter and the human path feature, normalizes it in combination with the L2 norm, and further optimizes it through the action-guided attention operator. Finally, the human path regional feature enhancement representation block is obtained, wherein the channel number is used to determine the channel dimension in the feature processing process. The generation of the object path regional feature enhancement representation block first processes the combination result of the object path weight parameter and the object path feature through the channel attention feature aggregation operator, normalizes it in combination with the L1 norm, filters irrelevant noise through the object path feature noise gating operator that suppresses the object path feature noise, and finally performs activation processing through the rectified linear unit activation function to obtain the object path regional feature enhancement representation block. The object path weight parameter is used to adjust the importance of each part of the object path feature. The generation of the action interaction path regional feature enhancement representation block needs to combine spatial position encoding to process the combination result of the action interaction path weight parameter and the action interaction path feature 3T of the teacher model, calculate it in combination with the attention scaling factor, and then strengthen the spatial dimension of the feature through the spatial dimension convolution operation to obtain the action interaction path regional feature enhancement representation block. The action interaction path weight parameter is used to adjust the importance of each part of the action interaction path feature. After obtaining the three types of regional feature enhancement representation blocks, the human path feature 3T, the object path feature 3T, and the action interaction path feature 3T of the teacher model are combined with the corresponding human path regional feature enhancement representation block, object path regional feature enhancement representation block, and action interaction path regional feature enhancement representation block, respectively, and multiplied by the residual weight to obtain the human path feature 4T, object path feature 4T, and action interaction path feature 4T of the teacher model. Subsequently, the three types of regional feature enhancement representation blocks are shared to the student model. The student model processes its own human path feature 3S, object path feature 3S, and action interaction path feature 3S in combination with the dynamic channel priority attention operator. The dynamic channel priority attention operator first performs global pooling processing on the features of each feature channel of the student model, then calculates the priority weight of each channel through the multilayer perceptron and the Sigmoid activation function, and finally fuses and optimizes the features of the student model and the shared regional feature enhancement representation blocks according to the weight to obtain the human path feature 4S, object path feature 4S, and action interaction path feature 4S of the student model.

[0154] In the embodiment of the present application, the obtaining of the association feature 1 is based on the person path feature 4T and the object path feature 4T of the teacher model, the two features are first subjected to feature splicing operation, and then sequentially subjected to the processing of the spatial first weight, the spatial first bias, the rectified linear unit activation function, the spatial second weight and the spatial second bias to obtain the association feature 1, wherein the spatial first weight and the spatial second weight are used to adjust the importance of the features in different spatial dimensions, the spatial first bias and the spatial second bias are used to adjust the offset of the feature values, the generation of the association feature 2 is first subjected to feature splicing operation on the object path feature 4T and the action interaction path feature 4T of the teacher model, and then subjected to self-attention calculation, while calculating the feature cross-attention between the object path feature 4T and the action interaction path feature 4T, adding the self-attention calculation result, the feature cross-attention calculation result and the spliced feature, and then performing layer normalization processing to obtain the association feature 2. The layer normalization processing is used to normalize the distribution of the features to improve the stability of the features; the feature cross-attention calculation is used to capture the association relationship between different feature paths, the generation logic of the association feature 3 is consistent with that of the association feature 2, except that the input feature is replaced by the person path feature 4T and the action interaction path feature 4T of the teacher model, and the same feature splicing, self-attention calculation, feature cross-attention calculation, feature addition and layer normalization processing are performed to obtain the association feature 3, the generation of the association feature 4 requires feature splicing operation on the person path feature 4T, the object path feature 4T and the action interaction path feature 4T of the teacher model, and then processing combined with the bit-coded sine-cosine function, the bit-coded sine-cosine function calculates the bit coding result according to the bit code, the model dimension and the loop variable, adds the result to the spliced feature, and then processes through two layers of perception machine to obtain the association feature 4, the two layers of perception machine processing are used for more complex nonlinear transformation and feature extraction of the features, after obtaining the four groups of association features, the teacher interaction association feature 5T is obtained by integrating the four groups of association features, in the integration process, for each group of association features, the association feature is first processed through the multi-layer perception machine for calculating the attention weight to obtain the attention weight of the association feature in the group, and then the association feature is processed through the multi-layer perception machine for processing the interaction feature to obtain the processing result of the association feature in the group, then the weight proportion of the attention weight of each group of association features is calculated through the exponential function, the weight proportion is multiplied by the processing result of the corresponding group of association features, and finally the weighted results of the four groups of association features are added to obtain the teacher interaction association feature 5T.

[0155] In the implementation of the present application, during the calculation process, the sample index needs to be considered. For each sample, first, the sample-by-sample L2 loss between the teacher model corresponding feature and the student model corresponding feature is calculated, and then multiplied by the sample-by-sample L2 loss weight. Then, the Frobenius norm distance square between the Gram matrix constructed by the teacher model corresponding feature and the Gram matrix constructed by the student model corresponding feature is calculated, and then multiplied by the Frobenius norm distance loss weight after being divided by the square of the model dimension. Then, the KL divergence is calculated between the teacher model corresponding feature processed by the Softmax function and divided by the temperature scaling value, and the student model corresponding feature processed by the Softmax function and divided by the temperature scaling value, and multiplied by the KL divergence loss weight. The three parts are added to obtain the minimum feature distribution distillation loss for the sample. The average of the losses of all samples is taken to finally obtain the minimum feature distribution distillation loss under the teacher model and student model corresponding feature channel.

[0156] After obtaining the minimum feature distribution distillation loss, the features of the student model are optimized based on the loss to obtain the three-path feature 5S of the student model. In the optimization process, the optimization correction term under the student model corresponding feature channel is calculated by back propagation. The back propagation calculation adjusts the gradient direction and size of the feature according to the minimum feature distribution distillation loss. The calculation of the optimization correction term needs to be combined with the learning rate parameter. The learning rate parameter is used to control the step size of each optimization to avoid the occurrence of shock or slow convergence in the optimization process. The optimization correction term is added to the human path feature 4S, the object path feature 4S and the action interaction path feature 4S of the student model respectively to obtain the optimized human path feature 5S, the object path feature 5S and the action interaction path feature 5S of the student model.

[0157] In the specific implementation of the present application, in the interactive feature merging and parsing distillation module, first, the features of the teacher model and the student model are subjected to merging and parsing processing, the person path feature 4T, the object path feature 4T and the action interaction path feature 4T of the teacher model are subjected to feature merging and parsing processing with the teacher interactive association feature 5T respectively, to obtain the person path feature 6T, the object path feature 6T and the action interaction path feature 6T of the teacher model; similarly, the person path feature 5S, the object path feature 5S and the action interaction path feature 5S of the student model are subjected to the same feature merging and parsing processing with the teacher interactive association feature 5T, to obtain the person path feature 6S, the object path feature 6S and the action interaction path feature 6S of the student model. After the feature merging and parsing processing is completed, the object prediction distillation loss is calculated. This loss calculation is carried out respectively for the person path features 6T and 6S and the object path features 6T and 6S of the teacher model and the student model, and the sample index needs to be considered. For each sample, the sample-by-sample L2 loss between the corresponding features of the teacher model and the corresponding features of the student model is calculated, multiplied by the person path or object path feature sample-by-sample L2 loss weight; the Frobenius norm distance between the matrix constructed by the corresponding features of the teacher model and the matrix constructed by the corresponding features of the student model is calculated, multiplied by the person path or object path feature Frobenius norm distance loss weight, and the two results are added together to obtain the object prediction distillation loss for the sample. The losses of all samples are averaged to finally obtain the object prediction distillation loss of the corresponding features of the teacher model and the student model. The action interaction association distillation loss is calculated. This loss calculation is carried out for the action interaction path features 6T and 6S of the teacher model and the student model. The action interaction path feature 6T of the teacher model and the action interaction path feature 6S of the student model are processed respectively through two layers of MLP action interaction feature extraction to extract more representative action interaction features. The difference between the covariance of the processed teacher model action interaction feature and the covariance of the processed student model action interaction feature is calculated. The two differences are combined to obtain the action interaction association distillation loss. According to the object prediction distillation loss and the action interaction association distillation loss, the person path feature 6S, the object path feature 6S and the action interaction path feature 6S of the student model are jointly optimized to adjust the parameters and distribution of the features, and finally the optimized student model person path feature 7S, object path feature 7S and action interaction path feature 7S are output.

[0158] In the implementation of the present application, first, the total distillation loss is calculated, which is composed of four parts of loss, namely, the human-computer interaction first distillation loss, the minimum feature distribution distillation loss, the object prediction distillation loss and the action interaction correlation distillation loss. When calculating, the human-computer interaction first distillation loss is multiplied by the human-computer interaction first distillation loss weight, the minimum feature distribution distillation loss is multiplied by the minimum feature distribution distillation loss weight, the object prediction distillation loss is multiplied by the object prediction distillation loss weight, and the action interaction correlation distillation loss is multiplied by the action interaction correlation distillation loss weight. The four weighted loss results are added to obtain the total distillation loss, which can comprehensively reflect the optimization effect of each distillation link of the model, provide a basis for the overall optimization of the model, obtain the human-computer interaction discrimination result of the industrial scene, and optimize the student model human road feature 7S, object road feature 7S and action interaction road feature 7S. Based on the three features, dynamic weighting fusion processing is performed on the three features. The dynamic weighting fusion will allocate a dynamic fusion weight to each feature according to the actual needs of the industrial scene and the importance of each feature path, and the size of the weight represents the contribution degree of the corresponding feature in the interaction discrimination. After multiplying each feature by the corresponding dynamic fusion weight and adding them, the fused features are obtained. The fused features are input into the action interaction discrimination classification multilayer perceptron. The multilayer perceptron classifies the fused features according to the number of possible action interaction categories in the industrial scene, and outputs the prediction probability of each action interaction category. According to these prediction probabilities, the human-computer interaction discrimination result in the industrial scene is determined, and the interaction action type between the operator and the tool or the workpiece in the current industrial scene is clarified.

Claims

1. A lightweight human-computer interaction model system for industrial applications, characterized in that, include: The acquisition module is used to input the images to be processed into the teacher model and the student model respectively. Initial image features are extracted through the backbone network, feature flattening layer, and Transformer encoder. Then, the Transformer decoder separates the initial image features into three paths: human path, object path, and action interaction path, corresponding to the teacher model and the student model respectively. Specifically, the images to be processed are input into the backbone networks of the teacher and student models to obtain their respective initial image features. Further processing of these initial image features yields the following: 1T human path features, 1T object path features, and 1T action interaction path features for the teacher model; and 1S human path features, 1S object path features, and 1S action interaction path features for the student model. The formulas are as follows: ; in, Let m represent a feature vector, where m=R represents the human path feature vector, m=W represents the object path feature vector, and m=J represents the action interaction path feature vector. A is the model index, where A=T represents the teacher model and A=S represents the student model. As initial image features, For the intersection edges of pedestrian paths under Model A, For the intersection edges of the main body of the object path under Model A, For the intersection edge of the main body of the action interaction path under Model A, The dynamic parameters for adapting the m-feature pathway under Model A to industrial scenarios; For feedforward calculation, Calculate the attention weights for scaling the dot product; For querying the parameter matrix; This is the key parameter matrix; For value parameter matrix; This is the transpose of the path key feature matrix of the teacher model. This is the attention scaling factor; For graph neural network operators, ,in, The set of main intersection edges includes the main intersection edges of pedestrian paths and object paths, pedestrian paths and action paths, and object paths and action paths. This is an edge convolution operation; The optimization module is used to process the features of the teacher model and the student model at the first distillation of the human-computer interaction in the region feature enhancement distillation module based on the obtained three-way features. The human path features 2S and object path features 2S obtained by the student model after processing by the intermediate feature optimization unit are input into the tail distillation unit of the teacher model to obtain the mapped human path features and mapped object path features. The mapped features and the human path features 3T and object path features 3T output by the tail distillation unit of the teacher model are used to calculate the first feature classification distillation loss and the first feature regression distillation loss to obtain the loss at the first distillation of the human-computer interaction, so as to optimize the corresponding human path features 3S and object path features 3S output by the student model. The enhancement module is used to enhance the three-way features 3S of the optimized student model and the three-way features 3T of the teacher model by inputting the three-way features 3T of the teacher model into the regional feature enhancement processing center for enhancement processing, thereby obtaining the three-way regional feature enhancement representation blocks and the three-way features 4T of the teacher model; the three-way regional feature enhancement representation blocks are shared with the corresponding feature paths of the student model for learning and optimization, so that the student model outputs the optimized three-way features 4S. The sharing module is used to perform pairwise and three-way interaction correlation processing on the teacher model's human path features 4T, object path features 4T, and action interaction path features 4S in the action interaction correlation distillation module to obtain four sets of correlation features. After integration, the teacher interaction correlation features 5T are obtained and shared with the student model. The calculation module is used to calculate the minimum feature distribution distillation loss for each path of the three-way features 4T of the teacher model and the three-way features 4S of the student model based on the shared teacher interaction association features 5T. Based on this loss, the student model is optimized to obtain the three-way features 5S of the student model. The parsing module is used to perform feature merging and parsing processing on the teacher model's three-way features 4T and the student model's three-way features 5S, respectively, with the teacher interaction association features 5T, in the interaction feature merging and parsing distillation module, to obtain the teacher model's three-way features 6T and the student model's three-way features 6S. Subsequently, the object prediction distillation loss is calculated on the human path features 6T and 6S and the object path features 6T and 6S of the teacher model and student model, and the action interaction path features 6T and 6S of the two are calculated on the action interaction association distillation loss. Based on the loss, the student model is jointly optimized, and the student model's three-way features 7S are output. The processing module is used to process the input detection head based on the three-way feature 7S of the student model, output the interaction area positioning information of the human path and the object path, the human-computer interaction discrimination result in the industrial scene, and calculate the total distillation loss of the model.

2. The lightweight human-computer interaction model system for industrial applications according to claim 1, characterized in that, The human-computer interaction first distillation step of the region feature enhancement distillation module is processed to obtain mapped human path and object path features. At the same time, the first feature classification and regression distillation loss are calculated to obtain the loss of the first human-computer interaction distillation step, as follows: The human path features 2S and object path features 2S, processed by the intermediate feature optimization unit of the student model, are fed to the tail distillation unit of the teacher model to generate mapped human path and object path features, as shown in the following formula: ; ; in, , These represent mapping human path features and mapping object path features, respectively. Convolution calculation for the tail distillation section of the teacher model. Representing a multilayer perceptron, in Its internal calculations yield the pedestrian-path features 2S of the student model. For scene gating operators, specifically , The Sigmoid activation function is used, and SceneEmbed is the scene encoding. Multi-scale attentional characteristics of human pathways For different window sizes; For dynamic bias operators, , To initialize a fixed bias vector, Global average pooling for path features. For self-attention calculation; The first feature classification and regression distillation loss calculation, along with the loss at the first distillation point in human-computer interaction, are obtained using the following formula: ; ; ; The first feature is the distillation loss for classification, B is the number of samples calculated in a single distillation, and k is the sample index. For divergence calculation, This represents the classification prediction probability of the teacher model under the m-feature pathway. It is the standard representation of KL divergence. The left side is the target distribution and the right side is the distribution to be matched. Overall, it represents the divergence of the latter relative to the former, measuring the difference between the two probability distributions. The student model maps the feature maps of the m-feature path to the classification prediction probabilities of the teacher model. Map the student model's features along the m-feature path to the classification quality score of the teacher model. The focusing parameters of the QFL loss, , This is the indicator function, where 1 is set for positive samples k and 0 is set for negative samples k. The first characteristic regression distillation loss, For boundary location functions, For the teacher model, boundary localization regression prediction under the m-feature path. To map the feature map of the student model onto the m-feature path and then onto the boundary localization regression prediction of the teacher model, The probability distribution for boundary localization of the teacher model under the m-feature path. The probability distribution of the boundary localization of the student model's feature mapping path under the m-feature path and the teacher model; Loss at the first distillation stage for human-computer interaction The square of the L2 norm, To classify the distillation loss weights, To regress the distillation loss weights, To map the L2 loss weights for human path features, The L2 loss weights are used to map the path features. The three features are the human path features (3T) and the object path features (3T) of the teacher model.

3. The lightweight human-computer interaction model system for industrial applications according to claim 2, characterized in that, The region feature enhancement processing center of the region feature enhancement distillation module is as follows: The formula for the region feature enhancement representation block is as follows: ; in, Human-road area feature enhancement representation block Attention operator for action guidance Here, C is the channel attention feature aggregation operator, where C is the number of channels. Human path weight parameters, Represents a parameterized transformation operator. Represents the L2 norm; ; in, Enhanced representation blocks for physical path region features. To modify the activation function of the linear unit, To suppress noise gating operators for object-path features without human interaction, For path weight parameters, Represents the L1 norm; ; in, Enhance the representation block for the action interaction path region features. Encoding spatial location, For action interaction path weight parameters, The action interaction path features of the teacher model are 3T. It is the scaling factor; ; in, When m=R, m=W, and m=J represent the human path features (4T), object path features (4T), and action interaction path features (4T) of the teacher model, respectively. When m=R, m=W, and m=J represent the enhanced representation block for human-path region features, the enhanced representation block for object-path region features, and the enhanced representation block for action-interaction-path region features, respectively. For residual weights; The region feature enhancement representation block is shared in the student model, as shown in the following formula: ; Where m=R, m=W, and m=J represent the human path feature 4S, object path feature 4S, and action interaction path feature 4S of the student model, respectively. m=R, m=W, and m=J represent the human path features (3S), object path features (3S), and action interaction path features (3S) of the student model, respectively. For dynamic path priority attention operators, specifically: , It is a global pooling operation used to reduce the dimensionality of feature maps and extract global semantic features.

4. The lightweight human-computer interaction model system for industrial applications according to claim 3, characterized in that, The action-interaction associated distillation module specifically includes the following: The formulas for associating features 1 to 4 are as follows: ; in, For associated feature 1, As the first weight in space, As the first offset in space, As the second weight in space, This is the second spatial offset. The feature concatenation operation is performed on the human path feature 4T and the object path feature 4T of the teacher model. ; in, For associated feature 2, For layer normalization processing, This involves the feature concatenation operation between the 4T object path features and the 4T action interaction path features of the teacher model. For feature cross-attention calculation; similarly, the associated feature 3 obtained from the teacher model's human path feature 4T and action interaction path feature 4T is used... Indicates; of which: ; in, For association feature 4, This is processed by a two-layer perceptron. This involves the feature concatenation operation of the teacher model's object path features 4T, object path features 4T, and action interaction path features 4T. loc is the sine and cosine function of the bit encoding, D is the model dimension, and i is the loop variable; The teacher interaction association feature 5T, the formula is as follows: ; Where n = 1, 2, 3, 4, represents the indices of interaction features 1 to 4. For teacher interaction-related features 5T, It is an exponential function. For the calculation of attention weights in a multilayer perceptron, The multilayer perceptron is computed for the interaction features n.

5. The lightweight human-computer interaction model system for industrial applications according to claim 4, characterized in that, The end-processing of the action interaction-related distillation module is defined by the following formula: + ; in, When m=R, m=W, and m=J represent the minimum feature distribution distillation loss of the human path features 4T and 4S, the object path features 4T and 4S, and the action interaction path features 4T and 4S between the teacher model and the student model, respectively. For the teacher model with three-way features 4T under the index of the k-th sample, For the student model with three-way features 4S under the index of the k-th sample, The squared distance is the Frobenius norm of the Gram matrix, where D is the dimension. Transpose the three-way features of the teacher model. Student model three-way feature transpose This is a temperature scaling value. These are the sample-by-sample L2 loss weights for the three-way features between the teacher model and the student model, the Frobenius norm distance loss weights, and the KL divergence loss weights, respectively. The optimal 5S feature of the student model is obtained by calculating the minimum feature distribution distillation loss between the three features of the teacher model and the student model, as shown in the following formula: ; The optimization correction terms are expanded as follows: + ; in, The learning rate parameter, To optimize the correction items, For backpropagation calculations, when m=R, m=W, and m=J represent the student model's human path features (5S), object path features (5S), and action interaction path features (5S), respectively. The dimensional scale representing the feature.

6. The lightweight human-computer interaction model system for industrial applications according to claim 5, characterized in that, The interactive feature merging and parsing distillation module is detailed below: The teacher and student models use the features 6T and 6S of the pedestrian path and object to calculate the object prediction distillation loss, as shown in the following formula: ; in, When m=R, m=W, and m=J represent the teacher model's human path feature 6T, object path feature 6T, and action interaction path feature 6T, respectively. When m=R, m=W, and m=J represent the student model's human path features 6T, object path features 6T, and action interaction path features 6T, respectively. When m=R and m=W represent the object prediction distillation losses of the human path features 6T and 6S, and the object path features 6T and 6S, respectively, between the teacher model and the student model. Let 6T be the human path feature and 6T be the object path feature of the teacher model under the k-th sample index, and let 6S be the human path feature and 6S be the object path feature of the student model under the k-th sample index. These are the sample-by-sample L2 loss weights for human-path and object-path features between the teacher model and the student model, and the Frobenius norm distance loss weights for human-path and object-path features between the teacher model and the student model, respectively. The loss calculation for action interaction correlation distillation between the teacher-student model action interaction path features 6T and 6S is performed using the following formula: ; in, For action interaction-related distillation loss, For two-layer MLP action interaction feature extraction, For the teacher model action interaction path feature 6T under the k-th sample index, For the student model action interaction path feature 6S under the k-th sample index, Calculate the covariance of action interaction path features.

7. The lightweight human-computer interaction model system for industrial applications according to claim 6, characterized in that, The total distillation loss is calculated using the following formula: ; in, For total distillation losses, , , , , respectively, are the loss weights at the first distillation stage of human-computer interaction, the loss weights at the minimum feature distribution distillation stage, the loss weights at the object prediction distillation stage, and the calculation of the loss at the action interaction association distillation stage; The specific formula for judging human-computer interaction in industrial scenarios is as follows: ; in, For human-computer interaction discrimination results in industrial scenarios, Multilayer perceptron for action interaction discrimination and classification, For dynamic weighted fusion, For dynamic fusion weights, C is the number of action interaction categories. For the optimized student model's three-way features, m=R, m=W, and m=J represent the student model's human path feature 7S, object path feature 7S, and action interaction path feature 7S, respectively.

Citation Information

Patent Citations

  • Robust online vector map construction method for space-time simplified query

    CN120426983A

  • Target detection position pixel coupling distillation method, system, equipment and medium

    CN120931889A