Full-performance detection, recognition and perception method and system based on machine vision

By employing feature encoding, fusion, and recursive processing of multimodal data, the challenge of multimodal feature fusion was solved, enabling high-precision detection in complex environments and improving the robustness and detection accuracy of machine vision systems.

CN120808077APending Publication Date: 2025-10-17GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510647209.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies lack a unified method for constructing feature spaces, making it difficult for information between modalities to work collaboratively and achieve efficient fusion of multimodal features. They also fail to make sufficient use of the high-dimensional correlation information of multimodal data, resulting in insufficient detection accuracy of the system in complex scenarios and making it difficult to meet high reliability requirements.

Method used

By collecting multimodal data for feature encoding and mapping to a unified feature space, the fusion weights are dynamically calculated using the L2 norm of the multimodal features for weighted normalization fusion to generate initial input features. The features are then layer by layer concatenated and the interaction between local details and global morphology is calculated to generate recursive features. High-dimensional space projection and gradient calculation are performed using a nonlinear perspective operator to generate enhanced fusion features. These features are then matched with the target template features in multiple dimensions, and the feature representation parameters are dynamically adjusted to optimize the detection results.

Benefits of technology

It improves the accuracy and robustness of detection and recognition, enhances the effectiveness of feature representation and detection efficiency, reduces the false detection rate, and ensures the stability and reliability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808077A_ABST
    Figure CN120808077A_ABST
Patent Text Reader

Abstract

The invention discloses a full-performance detection, recognition and perception method and system based on machine vision, and relates to the technical field of machine vision, and the method comprises the steps: collecting multi-modal data of a target object, carrying out the feature coding of the multi-modal data, mapping the multi-modal data to a feature space with a unified dimension, and carrying out the feature coding of the multi-modal data; the method comprises the following steps: dynamically calculating a fusion weight based on a norm of a multi-modal feature, carrying out weighted normalized fusion on the encoded feature, generating an initial input feature, splicing the initial input feature layer by layer, calculating and extracting an interaction relationship between local details and a global form of the feature, generating a recursive feature, and calculating and generating an enhanced fusion feature; and performing multi-dimensional matching on the fusion features and predefined target template features, calculating a comprehensive performance evaluation value, and iteratively optimizing a detection result. According to the method, the accuracy, robustness and efficiency of detection and recognition are improved, the expression ability of features is enhanced, and the stability and reliability of a detection result are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, and in particular to a full-performance detection and recognition perception method and system based on machine vision. BACKGROUND

[0002] With the diversification and complexity of application requirements, the demand for comprehensive perception and high-precision detection of target objects is also increasing. Multimodal data fusion is an important direction of machine vision technology development, which can more comprehensively describe the performance characteristics of target objects by combining the color and texture information of RGB images, the geometric shape information of depth images, and the thermal distribution information of infrared images. For example, in industrial detection, RGB images can reflect the appearance defects of products, depth images can help judge the shape error of products, and infrared images can detect the thermal properties of materials. The fusion of these multi-dimensional information can greatly improve the accuracy and coverage of detection. However, the processing complexity of multimodal data is high, and the data characteristics of each modality are significantly different. How to realize effective fusion and unified feature expression has become a major challenge in the current technical field.

[0003] In addition, with the expansion of the application range of machine vision technology, the complexity of the environment where the detection object is located is also increasing. In complex scenes, such as changes in lighting conditions, occlusions, and the presence of environmental noise, traditional feature extraction and target detection methods often show insufficient robustness, and detection accuracy is difficult to guarantee. Therefore, it is of great research value to develop a visual detection method that can maintain high robustness in complex scenes and dynamically adapt to environmental changes. The existing technology has great limitations in the use of multimodal data, and the system performs poorly when dealing with complex scenes; there is a lack of unified feature space construction method, and the information between modalities is difficult to work together, making it impossible to achieve efficient fusion of multimodal features; the use of high-dimensional associated information of multimodal data is insufficient, further limiting the recognition ability and comprehensiveness of performance detection of the system; the existing technology lacks a dynamic performance evaluation and optimization mechanism, and it is difficult to dynamically adjust according to different detection tasks and environmental changes, the detection result is easy to be disturbed and inaccurate, and it is difficult to meet the high reliability requirements of actual application. SUMMARY

[0004] In view of the above existing problems, the application provides a full-performance detection and recognition perception method and system based on machine vision, to solve the problems that there is no unified feature space construction method in the prior art, the information between modalities is difficult to work together, the efficient fusion of multi-modal features cannot be realized, the high-dimensional associated information of multi-modal data is not fully utilized, which further limits the recognition ability and the comprehensiveness of performance detection of the system, and there is no dynamic performance evaluation and optimization mechanism, it is difficult to dynamically adjust according to different detection tasks and environmental changes, the detection result is easy to be disturbed and inaccurate, and it is difficult to meet the high reliability requirements of practical applications.

[0005] To solve the above technical problems, a full-performance detection and recognition perception method based on machine vision is provided, comprising,

[0006] The multi-modal data of the target object is collected, the multi-modal data is respectively encoded, and the multi-modal data is mapped to a feature space of a unified dimension; the fusion weight is dynamically calculated based on the norm of the multi-modal feature, the encoded features are weighted and normalized, the initial input features are generated, the initial input features are spliced layer by layer, the interaction relationship between the local details and the global morphology of the extracted features is calculated, the recursive features are generated, the enhanced fusion features are calculated, and the fusion features are matched with the predefined target template features in multiple dimensions; the comprehensive performance evaluation value is calculated, the feature representation parameter is dynamically adjusted according to the gradient direction of the evaluation value, the detection result is iteratively optimized until the preset performance threshold is reached.

[0007] As a preferred scheme of the full-performance detection and recognition perception method based on machine vision, wherein: the feature encoding includes collecting multi-modal data of the target object through a sensor, respectively encoding the multi-modal data, and mapping the multi-modal data to a feature space of a unified dimension through linear transformation and a nonlinear activation function;

[0008] The multi-modal data includes RGB images, depth images and infrared images, respectively representing the color, three-dimensional geometric shape and thermal characteristics of the target;

[0009] The feature encoding further includes applying a nonlinear activation function to the linearly transformed features, mapping the multi-modal features to an encoding space of the same dimension, and the dimension of the encoding space is determined by the preset task complexity.

[0010] As a preferred scheme of the full-performance detection and recognition perception method based on machine vision, wherein: the weighted and normalized fusion includes dynamically calculating the fusion weight based on the L2 norm of the multi-modal feature, and weighting and normalizing the encoded features to generate the initial input features;

[0011] Generating the initial input features includes using the L2 norm of the multi-encoded features as the importance weight of the modal feature, performing weighted summation on the encoded RGB, depth, and infrared features according to the weight, and generating the fused initial input features.

[0012] As a preferred embodiment of the full-performance detection, recognition, and perception method based on machine vision described in the present invention, the layer-by-layer splicing includes inputting the initial input features into a recursive projection function, splicing the current input features with the historical recursive output features layer by layer, and calculating the interaction between the local details of the extracted features and the global morphology to generate recursive features;

[0013] The recursive projection function includes, in each layer of recursive calculation, dimensionally splicing the input features of the current layer with the historical features of the output of the previous layer to form an extended feature vector, applying linear transformation and composite activation function to the extended features to generate the projection features of the current layer, and superimposing the projection features with the historical features to form the output of the current layer. Through recursive iteration, the correlation between local details and global morphology in the features is extracted.

[0014] As a preferred embodiment of the full-performance detection, recognition, and perception method based on machine vision described in the present invention, initializing the weight matrix of the recursive projection function includes generating an initial transformation matrix based on the principal component analysis results of the training data, and dynamically adjusting the parameters of the weight matrix through a back-propagation algorithm during an iterative optimization process to minimize the comprehensive performance evaluation value;

[0015] The output of the recursive projection layer is defined as:

[0016] R k =P k (F)+R k-1

[0017] Among them, P k (F) is the projection feature of the kth layer, R k is the output of the k-th layer recursive feature, R k-1 is the output of the k-1th layer recursive feature, k is the variable index;

[0018] The calculation and generation of enhanced fusion features includes inputting the recursive features into a nonlinear perspective operator, and generating enhanced fusion features through high-dimensional space projection and feature gradient calculation; wherein the nonlinear perspective operator processing includes high-dimensional space projection of the recursive features, calculating the similarity between the features and the preset projection center through a Gaussian kernel function for each projection, and performing normalized weighted summation on all projection results to generate preliminary fusion features, and calculating the gradient of the fusion features with respect to the input features, and superimposing the gradient information into the fusion features.

[0019] As a preferred scheme of the machine vision-based full-performance detection and recognition perception method, the multi-dimensional matching comprises: performing dimension-by-dimension difference calculation on the enhanced fusion feature and a target template feature to generate a matching degree vector, applying a nonlinear mapping function to the matching degree vector, assigning weights to the multi-dimensions according to preset task requirements, performing weighted summation on the mapped matching degrees, and generating a comprehensive performance evaluation value.

[0020] The output feature of the nonlinear perspective operator is defined as:

[0021]

[0022] wherein, T(R k ) is the output feature of the nonlinear perspective operator, R k is the output of the kth layer recursive feature, N is the number of perspective projections, W i is the weight matrix of the ith perspective projection, b i is the bias vector of the ith perspective projection, a i is the weight coefficient of the ith perspective projection, i and k are variable indexes, and ||W i || F is the Frobenius norm of W i .

[0023] The formula for calculating the generated enhanced fusion feature is:

[0024]

[0025] wherein, F Fusion is the enhanced feature representation of the nonlinear perspective operator, T(R k ) is the output feature of the nonlinear perspective operator, R k is the output of the kth layer recursive feature, and k is a variable index.

[0026] As a preferred scheme of the machine vision-based full-performance detection and recognition perception method, the dynamic adjustment of the feature representation parameter comprises: calculating the update amount of the fusion feature according to the gradient direction of the comprehensive performance evaluation value, controlling the update amplitude through a preset dynamic learning rate, inputting the updated feature into the recursive projection function again, and repeatedly performing the feature extraction, perspective processing and performance evaluation steps until the evaluation value is lower than a preset threshold.

[0027] The performance index function is defined as:

[0028]

[0029] wherein, Q is the performance evaluation value, M is the number of performance indexes, w j is the weight of the jth performance index, and fj (F Fusion ) is a quantitative function of the jth performance index, j is a variable index, F Fusion is an enhanced feature representation of the non-linear perspective operator, F Template is a pre-defined target template feature, which is obtained in advance from a sample library;

[0030] Based on the performance evaluation result, the enhanced feature representation of the non-linear perspective operator is optimized, and the formula is represented as:

[0031]

[0032] Wherein, F Updated is the global feature after optimization, δ is the learning rate of feature update, F Fusion is the enhanced feature representation of the non-linear perspective operator, and Q is the performance evaluation value.

[0033] As a preferred scheme of the machine vision-based full performance detection and recognition perception system, it is characterized in that it comprises a multi-modal data acquisition and feature coding module, a multi-modal feature fusion module, a recursive feature projection and depth extraction module, and a non-linear perspective processing and optimization module.

[0034] The multi-modal data acquisition and feature coding module is used for acquiring multi-modal data of a target object, pre-processing the original data, and mapping the multi-modal data to a unified dimensional feature space through linear transformation and non-linear activation function, so as to eliminate the differences between modalities.

[0035] The multi-modal feature fusion module is used for dynamically calculating fusion weights, fusing the coded RGB, depth and infrared features into unified initial input features through a weighted normalization strategy, and generating fusion features with high expression capacity.

[0036] The recursive feature projection and depth extraction module is used for splicing the current input features and the historical recursive output features layer by layer through a recursive projection function, extracting the interaction relationship between local details and global morphology by using linear transformation and composite activation function, and generating deep recursive features.

[0037] The non-linear perspective processing and optimization module is used for high-dimensional space projection and gradient enhancement of the deep recursive features, generating enhanced fusion features, matching the fusion features with the target template features, calculating the comprehensive performance evaluation value, dynamically adjusting the feature representation parameters based on the gradient direction of the evaluation value, and iteratively optimizing the detection result through a closed-loop feedback mechanism.

[0038] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method for full-performance detection and recognition perception based on machine vision when executing the computer program.

[0039] A computer readable storage medium stores a computer program, and the computer program implements the steps of the method for full-performance detection and recognition perception based on machine vision when executed by a processor.

[0040] The present application has the following advantages: the present application collects multi-modal data and encodes features, comprehensively characterizes the color, three-dimensional geometry and thermal characteristics of a target object, provides a rich and multi-dimensional information basis for subsequent feature fusion and recognition, thereby improving the accuracy and robustness of detection and recognition; the L2 norm based on multi-modal features is used to dynamically calculate fusion weights and perform weighted normalization fusion, the multi-modal features are reasonably integrated, the initial input features are generated, the importance of different modal features is highlighted and redundant information is reduced, thereby improving the effectiveness and detection efficiency of feature representation; the initial input features are input into a recursive projection function, the features are spliced layer by layer and the interaction between local details and global morphology of the features is calculated, the recursive features are generated, the hierarchical structure and associated information in the features are deeply mined, and the expression ability and detection accuracy of the features are enhanced; the recursive features are input into a nonlinear perspective operator, high-dimensional space projection and feature gradient calculation are performed, the enhanced fusion features are generated, the features are further refined and strengthened, and the matching accuracy and detection performance are improved; the enhanced fusion features and target template features are matched in multiple dimensions, and a comprehensive performance evaluation value is calculated, the quantitative evaluation of the detection result is realized, an explicit feedback index is provided for iterative optimization, and the stability and reliability of the detection result are ensured; the feature representation parameters are dynamically adjusted according to the gradient direction of the comprehensive performance evaluation value, the iterative optimization of the detection result is realized, the detection accuracy and efficiency are improved, and the false detection rate is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0042] Figure 1 A general flowchart of a method for full-performance detection and recognition perception based on machine vision is provided for an embodiment of the present application.

[0043] Figure 2 A system scheme flowchart of a system for full-performance detection and recognition perception based on machine vision is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the above objectives, features and advantages of the present application more clear and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the protection scope of the present application.

[0045] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details presented herein. In other instances, well-known methods have not been described in detail in order to avoid obscuring aspects of the present application.

[0046] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. The "in one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is mutually exclusive with other embodiments.

[0047] The present application is described in detail with reference to the accompanying drawings. In the detailed description of the embodiments of the present application, the cross-sectional view of the device structure is partially enlarged without the general proportion for the convenience of description, and the schematic diagram is only an example, which should not limit the scope of protection of the present application. In addition, the three-dimensional spatial dimensions of length, width and depth should be included in actual manufacture.

[0048] Meanwhile, in the description of the present application, it should be noted that the terms "upper, lower, inner and outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first, second or third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0049] Unless otherwise specifically defined and limited, the terms "mounting, connecting, connection" in the present application should be understood broadly, for example: it can be fixed connection, detachable connection or integral connection; it can also be mechanical connection, electrical connection or direct connection, it can also be indirectly connected through intermediate medium, or it can be the communication inside two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0050] Example 1, with reference to Figure 1For the first embodiment of the present application, the embodiment provides a full performance detection and recognition perception method based on machine vision, comprising:

[0051] S1: Collecting multi-modal data of the target object, respectively encoding the features of the multi-modal data, and mapping the multi-modal data to a unified dimensional feature space.

[0052] Further, the multi-modal data is collected by sensors, including RGB images, depth images and infrared images, which correspond to the color distribution, three-dimensional geometric shape and thermal characteristic information of the target object, covering the main feature dimensions representing the performance of the target; the RGB image is collected by an industrial camera, which is used to capture the color and texture information of the target; the depth image is generated by a structured light camera, which is used to extract the spatial position and three-dimensional shape features of the target; the infrared image is obtained by thermal imaging, which is used to reflect the surface temperature distribution and thermal radiation characteristics of the target.

[0053] The multi-modal data respectively captures the color, three-dimensional geometric shape and thermal characteristic information of the target, but due to the differences in characteristics and dimensions between modalities, the multi-modal data needs to be encoded and projected to a consistent feature space. Linear transformation and nonlinear activation operations are applied to each modality data, which is expressed by the formula:

[0054] E(F m )=σ(W m ·F m +b m )

[0055] Wherein, E(F m ) is the encoded modality feature, F m is the multi-modal data, including RGB image F RGB , depth image F Depth and infrared image F IR , m∈{RGB, Depth, IR}, RGB is the image, Depth is the depth image, IR is the infrared image, W m is the modality mapping weight matrix, which is obtained by pre-training from the training samples, used to select the feature subspace with the most information content, b m is the bias term, which is obtained by linear regression fitting in the training data, used to correct the feature mapping center, and σ is the activation function, used to improve the nonlinear expression ability of the feature.

[0056] It should be noted that in order to integrate the multi-modal data features, the encoded modality features are directly assembled into a multi-modal feature set, which is expressed as:

[0057] F Encoded ={E(F RGB ),E(F Depth ),E(FIR )

[0058] wherein F Encoded is a multi-modal feature set, E(F RGB ) is the encoded RGB image modal feature, E(F Depth ) is the encoded depth image modal feature, and E(F IR ) is the encoded infrared image modal feature.

[0059] It should be noted that in order to integrate multi-modal data features, the encoded modal features are directly assembled into a multi-modal feature set, denoted as:

[0060] F Encoded ={E(F RGB ), E(F Depth ), E(F IR )}

[0061] wherein F Encoded is a multi-modal feature set, E(F RGB ) is the encoded RGB image modal feature, E(F Depth ) is the encoded depth image modal feature, and E(F IR ) is the encoded infrared image modal feature.

[0062] S2: dynamically calculating fusion weights based on the norm of multi-modal features, weighting and normalizing the encoded features to generate initial input features.

[0063] In the embodiments of the present application, the weighting and normalization fusion includes dynamically calculating fusion weights based on the L2 norm of multi-modal features, weighting and normalizing the encoded features to generate initial input features.

[0064] In an alternative embodiment, the weighting and normalization fusion includes assigning fixed weights to RGB, depth, and infrared modalities according to prior knowledge, RGB: 0.5, depth: 0.3, infrared: 0.2, linearly superimposing the encoded multi-modal features according to the fixed weights, and using the static weight fusion result as the input of the recursive projection.

[0065] In another alternative embodiment, the weighting and normalization fusion includes L2 normalization processing of the encoded modal features to eliminate dimensional differences, equal-weight averaging of the normalized RGB, depth, and infrared features, and inputting the average result into the recursive projection module.

[0066] The generating initial input features comprises: taking L2 norm of the multi-coding features as importance weights of the modal features, weighting and summing the coded RGB, depth and infrared features according to the weights to generate the fused initial input features.

[0067] The modal features in the multi-modal feature set need to be reconstructed into input features of the recursive feature projection function, and the modal features in the multi-modal feature set are linearly combined and normalized by the multi-modal feature fusion operation, and the formula is represented as:

[0068]

[0069] Wherein, F is the fused multi-modal feature, that is, the input feature of the recursive feature projection function, ||E(F m )|| is the L2 norm of the coded modal feature, used to measure the relative importance of the modal feature, E(F m ) is the coded modal feature, ∑ m' ||E(F m' )|| is the sum of the L2 norms of all modal features, used for normalization processing, and m' is an index variable used to traverse all modal types.

[0070] S3: The initial input features are spliced layer by layer, the interaction relationship between the local details and the global morphology of the extracted features is calculated, the recursive features are generated, the enhanced fusion features are calculated, and the multi-dimensional matching between the fusion features and the predefined target template features is performed.

[0071] Further, the layer-by-layer splicing comprises: inputting the initial input features into the recursive projection function, splicing the current input features with the historical recursive output features layer by layer, and calculating the interaction relationship between the local details and the global morphology of the extracted features to generate recursive features;

[0072] The recursive projection function comprises: in each recursive calculation, the input features of the current layer are dimensionally spliced with the historical features output by the last layer to form an extended feature vector, a linear transformation and a composite activation function are applied to the extended feature to generate the projection features of the current layer, and the projection features are superimposed with the historical features to form the output of the current layer, through recursive iteration, the correlation relationship between the local details and the global morphology of the extracted features is extracted, and the formula is represented as:

[0073] P k (F)=Ψ(W k ·Concat(F,R k-1 )+b k )

[0074] Ψ(x)=tanh(x)+ln(1+e x )

[0075] Among them, P k (F) is the projection feature of the kth layer, R k-1 is the output of the k-1th layer recursive feature, k is the variable index, W k is the weight matrix of the kth layer, Concat(F,R k-1 ) is the output R of the input feature F and the k-1 layer recursive feature k-1 Splicing in the feature dimension, F is the fused multimodal feature, that is, the input feature of the recursive feature projection function, b k is the bias term of the kth layer, adjusts the centralization of the recursive projection, Ψ is the activation function, x is the input of the activation function, tanh(x) is the output range of the smoothed recursive feature, ln(1+e x ) is used to increase the dynamic range of positive features.

[0076] In the embodiment of the present application, generating recursive features includes recursively concatenating input and historical features in multiple layers, and extracting local and global interaction relationships through a composite activation function.

[0077] In an optional embodiment, generating the recursive features includes performing a recursive projection, taking the single-layer recursive features as the final output, and inputting the final output into the nonlinear perspective module.

[0078] In another optional embodiment, generating recursive features includes using a ReLU function to perform multi-layer recursion, with each layer outputting through a ReLU activation function to generate shallow recursive features, which are input into subsequent modules.

[0079] Furthermore, the weight matrix initialization of the recursive projection function includes generating an initial transformation matrix based on the principal component analysis results of the training data, and dynamically adjusting the parameters of the weight matrix through the back-propagation algorithm during the iterative optimization process to minimize the comprehensive performance evaluation value;

[0080] The output of the recursive projection layer is defined as:

[0081] R k =P k (F)+R k-1

[0082] Among them, P k (F) is the projection feature of the kth layer, R k is the output of the k-th layer recursive feature, R k-1 is the output of the k-1th layer recursive feature, k is the variable index;

[0083] The initial conditions are expressed as:

[0084] R0=0

[0085] P0(F)=W0·F+b0

[0086] wherein R0is R k , P0(F) is the initialization of P k , W0is the weight matrix of the recursive initial projection, b0is the bias term of the recursive initial projection, and F is the fused multi-modal feature, i.e., the input feature of the recursive feature projection function.

[0087] It should be noted that the calculation of the enhanced fusion feature includes inputting the recursive feature into a nonlinear perspective operator to generate the enhanced fusion feature through high-dimensional space projection and feature gradient calculation; wherein the nonlinear perspective operator processing includes high-dimensional space projection of the recursive feature, similarity between the recursive feature and a preset projection center is calculated through a Gaussian kernel function each time, and the preliminary fusion feature is generated by normalizing and weighted summing all projection results, and the gradient information of the fusion feature to the input feature is calculated and superimposed into the fusion feature.

[0088] The deep recursive feature is input into the nonlinear perspective operator for further processing of the deep recursive feature, and the fusion feature is generated through multiple projections in high-dimensional space and feature interaction; the design core of the nonlinear perspective operator is to perform multi-dimensional nonlinear transformation on the deep recursive feature to capture the potential complex relationship between the features, and to perform dimensionality reduction operation in the feature dimension, and the nonlinear perspective operator generates feature interaction results through multiple projections.

[0089] wherein the output feature of the nonlinear perspective operator is defined as:

[0090]

[0091] wherein T(R k ) is the output feature of the nonlinear perspective operator, R k is the output of the kth recursive feature, N is the number of perspective projections, W i is the weight matrix of the ith perspective projection, b i is the bias vector of the ith perspective projection, a i is the weight coefficient of the ith perspective projection, i and k are variable indices, and ||W i || F is the Frobenius norm of W i .

[0092] The ln(1+||R k || 2 ) term in the nonlinear perspective operator formula performs nonlinear mapping on the L2 norm of the deep recursive feature, considers the global distribution of the recursive feature, and further enhances the expression ability of the feature; on the basis of T(R k ), the gradient information is further combined The constructed enhanced representation, improved sensitivity and discrimination ability, and the ability to capture subtle feature differences are particularly suitable for scenarios in multi-modal data fusion that require distinguishing subtle changes between modalities. The formula for generating enhanced fusion features is calculated as follows:

[0093]

[0094] where F Fusion is the enhanced feature representation of the non-linear perspective operator, T(R k ) is the output feature of the non-linear perspective operator, R k is the output of the k-th layer recursive feature, and k is a variable index. The enhanced feature representation of the non-linear perspective operator includes a high-dimensional representation of the multi-modal features of the target object, reflecting the comprehensive information of the global morphology, local texture, and physical properties of the target object, including color, geometric depth, and thermal imaging characteristics.

[0095] In the embodiments of the present application, generating enhanced fusion features includes Gaussian kernel projection combined with gradient enhancement and dynamic optimization of feature representation.

[0096] In an alternative embodiment, generating enhanced fusion features includes using a linear function for projection calculation and assigning a fixed weight to each projection, and taking the linear projection result as the enhanced feature.

[0097] S4: Calculate the comprehensive performance evaluation value, dynamically adjust the feature representation parameters according to the gradient direction of the evaluation value, and iteratively optimize the detection result until the preset performance threshold is reached.

[0098] Further, the enhanced feature of the non-linear perspective operator is used for performance evaluation to obtain the matching degree of the enhanced feature to the pre-defined target template feature, and the accuracy, robustness, and comprehensive performance of the system detection are realized through the quantification of various performance indicators related to the detection task; through comprehensive quantitative analysis of the enhanced feature representation in the entire detection and recognition process, the performance level of the enhanced feature representation in specific tasks is determined.

[0099] The performance indicator function is defined as:

[0100]

[0101] where Q is the performance evaluation value, representing the comprehensive performance of the system in the target detection and recognition task, M is the number of performance indicators, w j is the weight of the j-th performance indicator, f j (F Fusion ) is the quantization function of the j-th performance indicator, j is a variable index, F Fusion is the enhanced feature representation of the non-linear perspective operator, and F TemplatePre-obtain from the sample library for the predefined target template features;

[0102] Based on the performance evaluation result, the enhanced feature representation of the nonlinear perspective operator is optimized, and the formula is represented as:

[0103]

[0104] Wherein, F Updated is the optimized global feature, and δ is the learning rate of feature update, which is dynamically adjusted through the training data to ensure the convergence of the optimization process, F Fusion is the enhanced feature representation of the nonlinear perspective operator, Q is the performance evaluation value, and the whole process forms a complete closed loop from feature input to recursive projection, feature enhancement, to performance evaluation and feedback optimization.

[0105] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and they should be covered in the scope of the claims of the present application.

[0106] Embodiment 2, refer to Figure 2 , is a second embodiment of the present application, which provides a full performance detection and recognition perception system based on machine vision, including a multi-modal data acquisition and feature encoding module, a multi-modal feature fusion module, a recursive feature projection and depth extraction module, and a nonlinear perspective processing and optimization module.

[0107] The multi-modal data acquisition and feature encoding module is used to acquire multi-modal data of the target object, preprocess the original data, and map the multi-modal data to a unified dimensional feature space through linear transformation and nonlinear activation function, eliminating the differences between modalities.

[0108] The multi-modal feature fusion module is used to dynamically calculate the fusion weight, fuse the encoded RGB, depth and infrared features into unified initial input features through weighted normalization strategy, and generate high expression ability fusion features.

[0109] The recursive feature projection and depth extraction module is used to splice the current input feature and the historical recursive output feature layer by layer through the recursive projection function, extract the interaction relationship between local details and global morphology using linear transformation and composite activation function, and generate deep recursive features.

[0110] The nonlinear perspective processing and optimization module is configured to project deep recursive features into a high-dimensional space and perform gradient enhancement, generate enhanced fusion features, match the fusion features with target template features, calculate a comprehensive performance evaluation value, dynamically adjust feature representation parameters based on the gradient direction of the evaluation value, and iteratively optimize the detection result through a closed-loop feedback mechanism.

[0111] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

[0112] Embodiment 3, as a third embodiment of the present application, provides a full performance detection and recognition perception method based on machine vision. In order to verify the beneficial effects of the present application, scientific demonstration is carried out through experiments.

[0113] Applied to industrial quality inspection, the purpose is to detect and identify defects (such as scratches, stains, pits) on the surface of products, and to quantitatively evaluate the impact; the system combines RGB images, depth images and infrared images, and realizes accurate defect detection and performance evaluation through multi-modal data fusion and recursive feature learning.

[0114] Data acquisition: RGB image (color, texture information): resolution is 1920x1080 pixels; depth image (surface shape information): resolution is 640x480 pixels; infrared image (surface thermal characteristics): resolution is 320x240 pixels; and the RGB, depth and infrared images are denoised, aligned and normalized to ensure image dimension consistency.

[0115] Suppose the color feature value of the RGB image is F RGB =[150, 200, 255], the depth image feature is F Depth =[0.5], and the infrared image feature is F IR =[1.2].

[0116] RGB image feature encoding: the RGB image is encoded according to linear transformation and nonlinear activation function; set: weight matrix W RGB is:

[0117]

[0118] The bias term b RGB =[0.1, -0.05, 0.2]; the activation function is Sigmoid:

[0119] The eigenvalues of the RGB image are substituted into the formula to encode:

[0120]

[0121]

[0122] The encoding feature of the RGB image is E(F RGB ) = [1.0, 1.0, 1.0]; the feature of the depth image is encoded as W Depth = 0.5, b Depth = 0.1, F Depth = 0.5,

[0123] E(F Depth ) = σ(0.5·0.5 + 0.1) = σ(0.35) = 0.586

[0124] The encoding feature of the depth image is E(F Depth ) = 0.586; the same processing is performed on the infrared image, assuming W IR = 1.0, b IR = 0.1, F IR = 1.2,

[0125] E(F IR ) = σ(1.0·1.2 + 0.1) = σ(1.3) = 0.785

[0126] The encoding feature of the infrared image is E(F IR ) = 0.785.

[0127] The features of the RGB, depth and infrared images are fused by weighted summation;

[0128] ||E(F RGB )|| = 1.732, ||E(F Depth )|| = 0.586, ||E(F IR )|| = 0.785

[0129] The formula for calculating the fused feature F is:

[0130]

[0131] F = 0.625·[1.0, 1.0, 1.0] + 0.214·[0.586] + 0.289·[0.785]

[0132] F = [0.625, 0.625, 0.625] + [0.125, 0.125, 0.125] + [0.227, 0.227, 0.227]

[0133] = [0.977, 0.977, 0.977]

[0134] Recursive feature projection with nonlinear perspective operator, assuming the weight matrix of the first layer recursion is:

[0135]

[0136] Assuming the initial recursive output R0 = [0, 0], then:

[0137]

[0138] Add bias term:

[0139]

[0140] Processed by activation function:

[0141]

[0142] Therefore, the first layer recursive output is

[0143] Performance evaluation and identification: compare the recursive output with the preset threshold to judge the severity of defects; set a threshold vector T = [0.5, 0.5], if P1(F) > T, it indicates that there is a defect, and the severity is 0.687 and 0.599.

[0144] The detected defect severity is: the first dimension: 0.687 (indicating mild defect), the second dimension: 0.599 (indicating moderate defect); input the deep recursive features into the nonlinear perspective operator for multiple projections and feature interactions to further enhance the feature expression ability, thereby improving the effect of multi-modal data fusion; the specific steps include nonlinear transformation, dimensionality reduction of features, and optimization of feature interaction to capture potential complex relationships.

[0145] Assuming that the recursive feature projection has been completed in the previous stage, and the recursive feature R k = [0.45, 0.32] is obtained, and nonlinear perspective transformation is performed, with the following parameters: projection times N = 3, weight matrix W1 = [0.7, 0.2], W2 = [0.5, 0.5], W3 = [0.3, 0.8], bias vector b1 = 0.1, b2 = 0.05, b3 = 0.2, R k = [0.45, 0.32];

[0146] Calculate ||W1|| F Frobenius norm, the Frobenius norms of (W1, W2, W3) are respectively:

[0147]

[0148] Based on the obtained parameters, calculate the weight coefficients α1, α2, α3 for each projection:

[0149]

[0150] Compute the output of the perspective operator:

[0151] T(R k )=0.296·exp(-||[0.7,0.2]·[0.45,0.32]-0.1|| 2 )+ln(1+||[0.45,0.32]|| 2 )

[0152] Using gradient information to enhance the perspective operator output, by calculating Get enhanced feature representation F Fusion , the gradient is calculated as The enhanced features are:

[0153]

[0154] The performance evaluation function is as follows: the number of performance indicators is M = 2, and the weight coefficients are w1 = 0.6 and w2 = 0.4 respectively. The quantization function is defined as:

[0155]

[0156] Assume that the target template feature F Template =[0.65,0.68], then:

[0157] ||F Fusion -F Template || 2 =0.0001

[0158]

[0159] The comprehensive performance evaluation value is:

[0160]

[0161] In the defect detection task of the automated production line, the machine vision system needs to accurately identify possible subtle defects such as surface cracks, bubbles, color differences, etc.; the present invention can improve the perception of small defects in complex environments and reduce the false positive rate. The low value of the comprehensive performance evaluation value shows that the method of the present invention can accurately identify the slight differences between normal products and defective products.

[0162] Example 4, the fourth embodiment of the present invention, is different from the first three embodiments in that:

[0163] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0164] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, which can be embodied in any computer readable medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instruction execution systems, apparatus or devices. For the purpose of this specification, the "computer readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by or in connection with an instruction execution system, apparatus or device, or in conjunction with these instruction execution systems, apparatus or devices.

[0165] More specific examples (non-exhaustive list) of the computer readable medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer readable medium can even be paper or other suitable medium on which the program can be printed, because the program can be obtained electronically, for example, by optical scanning of the paper or other medium, followed by editing, interpreting or processing as necessary, and then stored in the computer memory. The software program can be transmitted from a website, server or other remote source using a modem or other means for electronic processing.

[0166] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the embodiments described above, various steps or methods can be implemented, for example, by software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following technology, known in the art, or combinations thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions upon an application of data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

Claims

1. A full-performance detection, recognition and perception method based on machine vision, characterized by: include, Collect multimodal data of the target object, perform feature encoding on the multimodal data, and map the multimodal data into a feature space of unified dimension; Dynamically calculate the fusion weight based on the norm of multimodal features, perform weighted normalization fusion on the encoded features, and generate the initial input features; The initial input features are spliced ​​layer by layer, and the interaction between the local details of the extracted features and the global morphology is calculated to generate recursive features. The enhanced fusion features are then calculated and matched with the predefined target template features in multiple dimensions. Calculate the comprehensive performance evaluation value, dynamically adjust the feature representation parameters according to the gradient direction of the evaluation value, and iteratively optimize the detection results until the preset performance threshold is reached.

2. The method for full performance detection, recognition, and perception based on machine vision according to claim 1, characterized in that: The feature encoding includes collecting multimodal data of the target object through a sensor, performing feature encoding on the multimodal data respectively, and mapping the multimodal data to a feature space of uniform dimension through linear transformation and nonlinear activation function; The multimodal data includes RGB images, depth images, and infrared images, which respectively represent the color, three-dimensional geometric shape, and thermal properties of the target; The feature encoding also includes applying a nonlinear activation function to the linearly transformed features to map the multimodal features to a coding space of the same dimension, where the dimension of the coding space is determined by a preset task complexity.

3. The method for full performance detection, recognition and perception based on machine vision according to claim 2, characterized in that: The weighted normalization fusion includes dynamically calculating the fusion weight based on the L2 norm of the multimodal features, performing weighted normalization fusion on the encoded features, and generating initial input features; Generating the initial input features includes using the L2 norm of the multi-encoded features as the importance weight of the modal feature, performing weighted summation on the encoded RGB, depth, and infrared features according to the weight, and generating the fused initial input features.

4. The method for full performance detection, recognition, and perception based on machine vision according to claim 3, characterized in that: The layer-by-layer splicing includes inputting the initial input features into a recursive projection function, splicing the current input features with the historical recursive output features layer by layer, and calculating the interaction between the local details of the extracted features and the global morphology to generate recursive features; The recursive projection function includes, in each layer of recursive calculation, dimensionally splicing the input features of the current layer with the historical features of the output of the previous layer to form an extended feature vector, applying linear transformation and composite activation function to the extended features to generate the projection features of the current layer, and superimposing the projection features with the historical features to form the output of the current layer. Through recursive iteration, the correlation between local details and global morphology in the features is extracted.

5. The method for full performance detection, recognition and perception based on machine vision according to claim 4, characterized in that: Initializing the weight matrix of the recursive projection function involves generating an initial transformation matrix based on the principal component analysis results of the training data, and dynamically adjusting the parameters of the weight matrix through the back-propagation algorithm during the iterative optimization process to minimize the comprehensive performance evaluation value. The output of the recursive projection layer is defined as: R k =P k (F)+R k-1 Among them, P k (F) is the projection feature of the kth layer, R k is the output of the k-th layer recursive feature, R k-1 is the output of the k-1th layer recursive feature, k is the variable index; The calculation and generation of enhanced fusion features includes inputting the recursive features into a nonlinear perspective operator, and generating enhanced fusion features through high-dimensional space projection and feature gradient calculation; wherein the nonlinear perspective operator processing includes high-dimensional space projection of the recursive features, calculating the similarity between the features and the preset projection center through a Gaussian kernel function for each projection, and performing normalized weighted summation on all projection results to generate preliminary fusion features, and calculating the gradient of the fusion features with respect to the input features, and superimposing the gradient information into the fusion features.

6. The method for full performance detection, recognition and perception based on machine vision according to claim 5, characterized in that: The multi-dimensional matching includes calculating the difference between the enhanced fusion features and the target template features dimension by dimension to generate a matching degree vector, applying a nonlinear mapping function to the matching degree vector, assigning weights to multiple dimensions according to preset task requirements, and performing weighted summation of the mapped matching degrees to generate a comprehensive performance evaluation value; The output features of the nonlinear perspective operator are defined as: Among them, T(R k ) is the output feature of the nonlinear perspective operator, R k is the output of the k-th layer recursive feature, N is the number of perspective projections, W i is the weight matrix of the i-th perspective projection, b i is the bias vector of the i-th perspective projection, α i is the weight coefficient of the i-th perspective projection, i and k are variable indices, ||W i || F W i The Frobenius norm of The formula for calculating and generating enhanced fusion features is expressed as: Among them, F Fusion is the enhanced feature representation of the nonlinear perspective operator, T(R k ) is the output feature of the nonlinear perspective operator, R k is the output of the k-th layer recursive feature, and k is the variable index.

7. The method for full performance detection, recognition and perception based on machine vision according to claim 6, characterized in that: The dynamic adjustment of the feature representation parameters includes calculating the update amount of the fusion feature according to the gradient direction of the comprehensive performance evaluation value, controlling the update amplitude through a preset dynamic learning rate, re-inputting the updated feature into the recursive projection function, and repeating the feature extraction, perspective processing and performance evaluation steps until the evaluation value is lower than a preset threshold; The performance indicator function is defined as: Among them, Q is the performance evaluation value, M is the number of performance indicators, and w j is the weight of the jth performance indicator, f j (F Fusion ) is the quantitative function of the jth performance indicator, j is the variable index, F Fusion is the enhanced feature representation of the nonlinear perspective operator, F Template It is a predefined target template feature, which is pre-obtained from the sample library; Based on the performance evaluation results, the enhanced feature representation of the nonlinear perspective operator is optimized, and the formula is expressed as follows: Among them, F Updated is the optimized global feature, δ is the learning rate of feature update, F Fusion is the enhanced feature representation of the nonlinear perspective operator, and Q is the performance evaluation value.

8. A system using the full-performance detection, recognition, and perception method based on machine vision according to any one of claims 1 to 7, characterized in that: It includes multimodal data acquisition and feature encoding module, multimodal feature fusion module, recursive feature projection and depth extraction module, and nonlinear perspective processing and optimization module; The multimodal data acquisition and feature encoding module is used to collect multimodal data of the target object, preprocess the original data, and map the multimodal data to a feature space of unified dimension through linear transformation and nonlinear activation function to eliminate differences between modalities; The multimodal feature fusion module is used to dynamically calculate the fusion weight, fuse the encoded RGB, depth and infrared features into a unified initial input feature through a weighted normalization strategy, and generate a highly expressive fusion feature; The recursive feature projection and depth extraction module is used to splice the current input features and the historical recursive output features layer by layer through the recursive projection function, and use linear transformation and composite activation function to extract the interactive relationship between local details and global morphology to generate deep recursive features; The nonlinear perspective processing and optimization module is used to perform high-dimensional spatial projection and gradient enhancement on deep recursive features, generate enhanced fusion features, match the fusion features with the target template features, calculate the comprehensive performance evaluation value, dynamically adjust the feature representation parameters based on the gradient direction of the evaluation value, and iteratively optimize the detection results through a closed-loop feedback mechanism.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a full-performance detection, recognition and perception method based on machine vision described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a full-performance detection, recognition and perception method based on machine vision described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Space-time quality closed-loop control method for vehicle infrastructure collaborative perception fusion and planning

    CN121564979A