Fault detection method and apparatus, device, and storage medium

Through the integration of the feature of the attention network and the encoder network, combined with the multi-layer perceptron, the automatic fault detection of false twisted parts of the ammunition machine is achieved, solving the problem of difficulty in monitoring wear in the prior art, and improving the accuracy and efficiency of detection.

WO2025156351A1PCT designated stage Publication Date: 2025-07-31ZHEJIANG HENGYI PETROCHEMICAL CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/078743
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2024-02-27
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The prior art is difficult to realize automated monitoring of false twisted parts of the elastic-adding machine, especially fault detection in wear situations.

Method used

The fault detection method based on attention network, encoder network and multi-layer perceptron is adopted to realize fault detection of false twisted parts of the ammunition machine through image acquisition, feature fusion and decoding operations.

Benefits of technology

Automatic fault detection of false twisted parts of the ammunition machine is realized, and the accuracy and efficiency of detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024078743_31072025_PF_FP_ABST
    Figure CN2024078743_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, and provides a fault detection method and apparatus, a device, and a storage medium. The method comprises: performing image acquisition on a false twisting part of a draw texturing machine to obtain a target image; constructing a first fused feature on the basis of at least one attention module in an attention network, wherein for each attention module, the attention module comprises a first sub-module and a second sub-module, the first sub-module constructs a first sub-feature for a fine-grained feature within a first field of view range and a coarse-grained feature within a second field of view range, and the second sub-module uses a residual constructed on the basis of second input information to obtain a mask image; inputting the first fused feature into an encoder network to obtain an encoded feature; performing a decoding operation on the encoded feature on the basis of a decoder network to obtain a decoded feature; obtaining a second fused feature on the basis of the decoded feature; and inputting the second fused feature into a multi-layer perceptron to obtain a fault detection result for the false twisting part. The present disclosure can implement automatic false twisting part detection.
Need to check novelty before this filing date? Find Prior Art

Description

Fault detection method, device, equipment and storage medium Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as artificial intelligence, computer vision, and image processing. Background Art

[0002] In the industrial context of texturing, texturing machines play a crucial role in the spinning process. In related technologies, the yarns used in texturing workshops include POY (Pre-Oriented Yarn), also known as POY precursor, and DTY (Draw Textured Yarn).

[0003] Texturing machines process POY raw yarn to produce DTY yarn. These machines contain a false twist unit that applies the false twist to the POY raw yarn. Frequent use of these machines can lead to wear and tear, which can affect the texturing process. Therefore, automated monitoring of these units remains a challenge.

[0004] Summary of the Invention

[0005] The present disclosure provides a fault detection method, apparatus, device and storage medium for monitoring false twist components in a texturing process.

[0006] In a first aspect, the present disclosure provides a fault detection method, comprising:

[0007] Capture images of the false twist components of the texturing machine to obtain target images;

[0008] Constructing a first fusion feature based on at least one attention module in the attention network; wherein, for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructing a first sub-feature based on fine-grained features within a first field of view in the first input information and coarse-grained features within a second field of view, and the second submodule using a residual constructed based on the second input information to obtain a mask map; the mask map is multiplied by the first sub-feature to obtain an output feature of the attention module; the first field of view is smaller than the second field of view;

[0009] Inputting the first fused feature into an encoder network to obtain an encoding feature; wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoding sub-feature, and the encoding feature includes the encoding sub-feature of at least one encoder;

[0010] Performing a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoding sub-feature, and the decoding feature includes the decoding sub-features of the plurality of decoders;

[0011] Fusing the decoding sub-features in the decoding feature to obtain a second fused feature;

[0012] The second fusion feature is input into a multi-layer perceptron to obtain a fault detection result of the false twist component of the texturing machine.

[0013] In a second aspect, the present disclosure provides a fault detection device, comprising:

[0014] An acquisition unit is used to acquire images of the false twist components of the texturing machine to obtain a target image;

[0015] A construction unit is configured to construct a first fusion feature based on at least one attention module in the attention network; wherein, for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructing a first sub-feature based on fine-grained features within a first field of view in the first input information and coarse-grained features within a second field of view, and the second submodule using a residual constructed based on the second input information to obtain a mask map; the mask map is multiplied by the first sub-feature to obtain an output feature of the attention module; the first field of view is smaller than the second field of view;

[0016] an encoding unit, configured to input the first fused feature into an encoder network to obtain an encoding feature; wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoding sub-feature, and the encoding feature includes the encoding sub-feature of at least one encoder;

[0017] a decoding unit, configured to perform a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein the decoder network includes a plurality of decoders, each decoder outputting a corresponding decoding sub-feature, and the decoding feature includes the decoding sub-features of the plurality of decoders;

[0018] a fusion unit, configured to fuse the decoding sub-features in the decoding feature to obtain a second fused feature;

[0019] The prediction unit is used to input the second fusion feature into a multi-layer perceptron to obtain a fault detection result of the false twist component of the texturing machine.

[0020] According to a third aspect, an electronic device is provided, including:

[0021] at least one processor; and

[0022] a memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0024] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0025] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0026] Based on the method proposed in the embodiment of the present disclosure, an automatic detection process for the false twist components of the texturing machine is realized.

[0027] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0029] FIG1 is a schematic diagram of a texturizing machine according to an embodiment of the present disclosure.

[0030] FIG2 is a schematic diagram of a false twist component according to an embodiment of the present disclosure.

[0031] FIG3 is a flow chart of a fault detection method according to an embodiment of the present disclosure.

[0032] FIG4 is a schematic diagram of the architecture of the entire neural network model according to an embodiment of the present disclosure.

[0033] FIG5 is a schematic diagram of a first submodule of an attention module according to an embodiment of the present disclosure.

[0034] FIG6 is a schematic diagram of a second submodule of an attention network according to an embodiment of the present disclosure.

[0035] FIG7 is a schematic diagram of an encoder network and a decoder network according to an embodiment of the present disclosure.

[0036] FIG8 is a schematic diagram of an encoder according to an embodiment of the present disclosure.

[0037] FIG9 is a schematic structural diagram of a fault detection device according to an embodiment of the present disclosure.

[0038] FIG10 is a block diagram of an electronic device for implementing the fault detection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The present disclosure will be described in further detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0040] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.

[0041] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0042] In the industrial scenario of the texturing process, the false twist components of the texturing machine may wear out and need to be monitored.

[0043] Among them, a schematic diagram of the key structure of a possible texturing machine is shown in Figure 1. The key structure of the texturing machine includes: a raw yarn rack 101, a wire cutter 102, a first roller 103, a first heat box 104, a cooling plate 105, a false twist component 106, a nozzle 107, a second roller 108, a second heat box 109, a third roller 110, a broken end detection device 111, and a winding component 112.

[0044] The first roller 103, the second roller 108, and the third roller 110 are used to ensure that the silk thread is processed along a predetermined path. During the texturing process, the speeds of the first roller 103, the second roller 108, and the third roller 110 are matched to each other to ensure that the silk fabric is not broken or stacked.

[0045] According to the requirements of product form, the nozzle 107 can be used to process the silk thread processed by the false twist component into a network-shaped silk thread product, so as to ensure that different silk thread products have the required touch and form.

[0046] When producing highly elastic yarn, the second heat box 109 may not be used. When producing medium elastic yarn, the temperature of the second heat box 109 may be adjusted to a first preset temperature, for example, about 140°C. When producing low elastic yarn, the temperature of the second heat box 109 may be adjusted to a second preset temperature, for example, between 165°C and 195°C.

[0047] When the broken end detection device 111 detects the broken end, it will trigger the wire cutter 102 to cut the wire, so as to prevent the POY raw yarn from piling up in the subsequent process.

[0048] When processing single-strand yarn, the corresponding false twist unit, shown as 106 in Figure 1, primarily uses two rubber rollers to twist the silk fabric. When processing multi-strand yarn, the false twist unit 106 processes the multiple strands into a combined yarn. One possible example, as shown in Figure 2, involves a first strand passing through from the upper left, and a second strand passing through from the upper right. These two strands are then combined to form a single strand.

[0049] In order to automatically and accurately detect the false twist components of the texturing machine, a fault detection method is proposed in the embodiment of the present disclosure, as shown in FIG3 , including the following contents:

[0050] S301, capturing images of the false twist components of the texturing machine to obtain a target image.

[0051] A drone can be used to conduct regular inspections of the false twist components of the texturing machine, or images of the false twist components can be captured based on monitoring. Any method that can obtain images of the false twist components is applicable to the embodiments of the present disclosure.

[0052] S302, constructing a first fusion feature based on at least one attention module in the attention network; wherein, for each attention module, the attention module includes a first sub-module and a second sub-module, the first sub-module constructs a first sub-feature for the fine-grained features within the first field of view in the first input information and the coarse-grained features within the second field of view, and the second sub-module uses the residual constructed based on the second input information to obtain a mask map; the mask map is multiplied by the first sub-feature to obtain the output feature of the attention module; the first field of view is smaller than the second field of view.

[0053] S303: Input the first fused feature into an encoder network to obtain an encoding feature; wherein the encoder network includes multiple encoders, each encoder outputs a corresponding encoding sub-feature, and the encoding feature includes the encoding sub-feature of at least one encoder.

[0054] S304, performing a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein the decoder network includes multiple decoders, each decoder outputs a corresponding decoding sub-feature, and the decoded feature includes the decoding sub-features of the multiple decoders.

[0055] S305: Fusing the decoding sub-features in the decoding feature to obtain a second fused feature.

[0056] S306: Input the second fused feature into a multi-layer perceptron to obtain a fault detection result of the false twist component of the texturing machine.

[0057] The attention network, encoder network, decoder network, and multilayer perceptron can use real images of historical false twist components as sample images to train the parameters of these models (including the attention network, encoder network, decoder network, and multilayer perceptron). Specifically, the fault location and fault type in the sample image are annotated, and the fault type can be the degree of wear. The sample image is sequentially input into the attention network to be trained, the encoder network to be trained, the decoder network to be trained, and the multilayer perceptron to be trained to obtain a predicted fault location and a predicted degree of wear. A first loss value is obtained based on the loss between the predicted fault location and the actual fault location, and a second loss value is obtained based on the loss between the predicted degree of wear and the actual degree of wear. The first and second loss values ​​are weighted and summed to obtain a target loss. Based on the target loss, the parameters of the attention network to be trained, the encoder network to be trained, the decoder network to be trained, and the multilayer perceptron to be trained are adjusted. When convergence conditions are met, the attention network, encoder network, decoder network, and multilayer perceptron are obtained, which are used to detect faults in the false twist components of the texturing machine in the target image.

[0058] In the disclosed embodiment, the attention network includes at least one attention module. Fine-grained features represent the local details of interest, while coarse-grained features represent the surrounding features of these local details. The fine-grained and coarse-grained features can capture information consistent with visual characteristics and fuse features within different granularity ranges to improve the expressiveness of the extracted features, thereby facilitating accurate detection of false twist components. Furthermore, the residual constructed from the second input information of the false twist component is also fused within the same attention module, further incorporating soft attention features into the features output by the attention module. This allows the same attention model to extract features of the false twist component from multiple perspectives, improving feature expressiveness and, consequently, the accuracy of the detection results. Furthermore, as the number of attention modules in the attention network increases, the expressiveness of the output features extracted from the target image becomes stronger. This output feature is then input into a subsequent encoder and decoder. Each encoder in the encoder network has a corresponding decoder in the decoder network. The decoded sub-features output by the multiple decoders are then fused to generate a second fused feature. This second fused feature is then input into a multi-layer perceptron to obtain a fault detection result for the false twist component of the texturing machine, thereby implementing an automated detection process for the false twist component of the texturing machine.

[0059] The architecture of the entire neural network model proposed in the embodiment of the present disclosure is shown in FIG4 . The model includes an attention network, an encoder network, a decoder network, and a multi-layer perceptron. The specific structure of each network is described in detail below:

[0060] 1) Attention Network

[0061] In some embodiments, the attention network, as shown in FIG4 , may include N attention modules (including the first attention module, the second attention module, ..., the Nth attention module in FIG4 , where N is a positive integer). The specific input information and output information of each attention module can be described as follows:

[0062] 1) For the first attention module in the attention network:

[0063] The first input information of the first submodule of the first attention module is the target image.

[0064] The second input information of the second submodule of the first attention module includes the target image and at least one historically collected image of the false twist component, and the collection time interval between the target image and any of the historically collected images is less than a preset time length.

[0065] The historically captured images may be m images sampled at equal intervals within a preset time duration before the target image is captured. For example, if the target image is sampled at time t, the preset time duration is a, and one image is captured at intervals of time b, then the first historically captured image is captured at time ta, the second historically captured image is captured at time t-a+b, and so on. Of course, historically captured images may also be obtained by sampling at unequal intervals within the preset time duration. Any method for obtaining historically captured images is applicable to the embodiments of the present disclosure.

[0066] 2) For any attention module other than the first attention module in the attention network:

[0067] The first input information of the first submodule of any attention module is the output feature output by the previous attention module of any attention module;

[0068] The second input information of the second submodule of any of the attention modules is the residual output of the second submodule of the previous attention module of any of the attention modules.

[0069] As shown in Figure 4, the input of the first submodule of the first attention module is the target image, and the output is the first first sub-feature. The input of the second submodule of the first attention module is the historically acquired image, and the output is the first residual. Multiply the first first sub-feature and the first residual to obtain the output of the first attention module, which is the first output feature. The first output feature is used as the input of the first submodule of the second attention module, and the output is the second first sub-feature. The first residual is used as the input of the second attention module, and the output is the second residual. This is executed sequentially until N attention modules are executed, and the output of the attention network is the first fusion feature.

[0070] In the disclosed embodiments, a target image can be used to detect the current state of a false twist component, thereby enabling fault detection. Historically captured images of the target image can be used to determine changes in the false twist component over a preset period of time. By comprehensively considering the current state of the target image and its historical state, a feature expression of the false twist component is extracted, facilitating the detection of wear on the false twist component.

[0071] In some embodiments, constructing the first sub-feature for the fine-grained features within the first field of view and the coarse-grained features within the second field of view in the first input information may be implemented as follows:

[0072] Step A1: Determine multiple target points in the first input information, and perform the following operations for each target point:

[0073] Step A11: extracting fine-grained features of the target point within the first field of view centered at the target point.

[0074] In some embodiments, extracting fine-grained features of the target point within the first field of view centered at the target point may be implemented as follows:

[0075] Step B1: Determine a feature value of the target point in the first input information as a first query feature of the target point.

[0076] The false twist component of the texturing machine is cropped in the target image to obtain the pixel point containing the false twist component, which is the target point. The black rectangular box in FIG5 is the target point.

[0077] Step B2: determining the first field of view of the target point with the position coordinates of the target point in the first input information as the center.

[0078] The first field of view can be 3×3 or 5×5, which can be determined based on actual conditions and is not limited in this embodiment. The first field of view in this embodiment is the gray rectangular box around the target point in FIG5 .

[0079] Step B3: constructing a first key feature and a first value feature of the target point based on the feature values ​​of feature points other than the target point within the first field of view.

[0080] Step B4: determining the fine-grained features of the target point based on the first query feature, the first key feature, and the first value feature.

[0081] In one possible implementation, the set of pixels in the sliding window centered at (i, j) on the feature map constructed by the first input information is defined as ρ(i, j). For a fixed window size k×k, ‖ρ(i, j)‖=k 2 Then the relationship between the first query feature and the first key feature can satisfy the equation (1), so as to obtain the fine-grained feature S (i,j)~ρ(i,j) :

[0082] In another possible implementation, position offset and mask filling can be introduced to obtain another type of fine-grained feature. As shown in Figure 5, in the fine-grained feature acquisition path (hereinafter referred to as the first path), the position offset is the relative position relationship between fine-grained tokens within the first field of view, which can also be understood as the relative position relationship between each pixel within the first field of view.

[0083] In the first pass, pixels at the edge of the feature map inevitably have similarities calculated with zero padding outside the boundary. To prevent these zero similarities from affecting the softmax operation, a padding mask is used to set these results to -∞.

[0084] Therefore, as shown in FIG5 , on the first path, the first query feature and the first key feature are cross-multiplied as in formula (1) to obtain the first feature to be fused, and the first feature to be fused and the position offset after mask filling processing on the first path are added pixel by pixel to obtain the fine-grained feature.

[0085] In the disclosed embodiment, the fine-grained features extracted by combining the features of the target point with the features of the nearby pixel points in the first field of view have strong expressive power and can lay a foundation for the subsequent determination of the fault detection result.

[0086] Step A12: extracting coarse-grained features of the target point within the second field of view centered at the target point.

[0087] In some embodiments, extracting the coarse-grained features of the target point within the second field of view centered at the target point may be implemented as follows:

[0088] Step C1 : determining the second field of view of the target point with the position coordinates of the target point in the first input information as the center.

[0089] As shown in FIG5 , the second field of view is the pixel points in a larger range centered on the target point and covering the first field of view, which is indicated by the white rectangular frame shown in FIG5 .

[0090] Step C2: constructing a second key feature and a second value feature of the target point based on the feature values ​​of feature points other than the target point within the second field of view.

[0091] Step C3: determining the coarse-grained feature of the target point based on the first query feature, the second key feature, and the second value feature of the target point; wherein the feature value of the target point in the first input information is the first query feature of the target point.

[0092] In one possible implementation, the set of pixels in the second field of view centered at (i, j) on the feature map constructed by the first input information is defined as ρ′(i, j). The set of pixels pooled from the second field of view is defined as σ(X). For the pooling size H p ×W p , ‖σ(X)‖=H p W p , in order to obtain coarse-grained features:

[0093] In another possible implementation, position offset and mask filling can be introduced to obtain another coarse-grained feature. In the coarse-grained feature acquisition path (hereinafter referred to as the second path), the position offset is the relative position relationship between coarse-grained tokens.

[0094] In order to further enhance the scalability of multi-scale image inputs for pixel-focused attention, different methods are used to calculate the position offset, which can be B within the second field of view. (i,j)~σ(X) .

[0095] As shown in Figure 5, the second path is to perceive global features near the target point. Pooling can be performed within the second field of view. The pooling window and step size can be determined based on the actual situation. The pooling operation moves the pooling window across the target image, selects the pixels with the maximum or average value, and generates a new feature map.

[0096] In the second path, logarithmically spaced continuous position bias (log-CPB) is used, where ReLU (activation function) is used to extract the position coordinates (Q (ij) ) and the pixel set of the second key feature The relative spatial coordinates Δ (i,j) ~σ(X), and then calculate B (i,j) ~σ(X).

[0097] Therefore, as shown in FIG5 , the first query feature and the second key feature are cross-multiplied as in formula (2) to obtain the second feature to be fused, and the second feature to be fused and the position offset are added pixel by pixel to obtain the coarse-grained feature.

[0098] In the disclosed embodiment, the coarse-grained features extracted by combining the features of the target point with the pixels in the surrounding second field of view have stronger expressive power, which can lay a foundation for the subsequent determination of the fault detection result.

[0099] Step A13: concatenate the fine-grained features and the coarse-grained features to obtain the initial features of the target point.

[0100] The fine-grained features and the coarse-grained features are concatenated in an additive manner to obtain the initial features of the target point.

[0101] Step A14: Mapping the initial features of the target point to feature values ​​of the target point through a nonlinear mapping method.

[0102] Step A2: splicing the feature values ​​of the multiple target points according to the position information of each target point in the target image to obtain the first sub-feature.

[0103] As shown in Figure 5, the coarse-grained features and fine-grained features are input into the connection layer and the activation layer to obtain the connection features. The separation layer is used to separate the connection features to separate the coarse-grained features from the fine-grained features, ultimately obtaining intermediate coarse-grained features and intermediate fine-grained features. The first-valued features and the intermediate fine-grained features are cross-multiplied to obtain the target fine-grained features, and the second-valued features and the intermediate coarse-grained features are cross-multiplied to obtain the target coarse-grained features. The target fine-grained features and the target coarse-grained features are summed to obtain the initial features, which are then passed through the nonlinear mapping layer to obtain the feature value of the target point. This operation is performed sequentially on each target point to obtain the first sub-feature.

[0104] In the disclosed embodiment, coarse-grained features are used to capture the overall structure of the image, while fine-grained features are more specific and accurate. Capturing both coarse-grained and fine-grained features of the target point can better identify faults in the false twist components of the texturing machine.

[0105] In some embodiments, the residual constructed based on the second input information to obtain a mask map can be implemented as follows:

[0106] Step D1: Perform Fourier transform on the second input information using a Fourier filter to obtain a time-varying component.

[0107] The second input information for the second submodule of the first attention module in the attention network is the previously acquired historical image. The second input information for the second submodules of the second attention module through the Nth attention module in the attention network is the residual output of the previous attention module. The framework diagram of the second submodule of each attention network is shown in Figure 6.

[0108] Step D2: input the time-varying component into an estimation module constructed based on a neural network to obtain an estimated value of the time-varying component.

[0109] Step D3: determining the residual between the time-varying component and the estimated value of the time-varying component to obtain the mask image.

[0110] Among them, the estimation module constructed based on the neural network can be a time-varying Koopman predictor (Koopa), and the Koopa model is composed of multiple layers of stackable Koopa basic modules.

[0111] Each Koopa basic module focuses on learning the dynamic characteristics of a specific level. By stacking Koopa basic modules, the model can capture the multi-level and complex dynamic changes of the time series. Each Koopa basic module uses the residual of the previous block's dynamic fitting as input to hierarchically learn operators. The approach proposed in the disclosed embodiments improves the prediction accuracy of the time-varying Koopman predictor and enhances the model's adaptability to complex, nonlinear, and non-stationary time series, thereby facilitating the extraction of dynamic characteristics of false twist components.

[0112] 2) Decoder Network

[0113] In some embodiments, the encoders in the encoder network correspond to the decoders in the decoder network in a one-to-one manner; as shown in FIG7 , the encoder network shows a case where there are four encoders, and the decoder network shows a case where there are four decoders. Regardless of the number of encoders and decoders, the decoding operation based on the decoder network on the encoded features to obtain the decoded features can be implemented as follows:

[0114] Step E1: for each target decoder in the decoder network, perform the following operations:

[0115] Step E11: Obtain the encoding sub-feature output by the encoder corresponding to the target decoder as the third query feature of the target decoder.

[0116] Taking the four encoders shown in the encoder network in Figure 7 as an example, the target image of size H×W×3 is processed, where H is the height of the target image and W is the width of the target image. The four encoders i∈{1,2,…,4} generate hierarchical and multi-resolution sub-features E respectively. i ,in C i Represents the weight corresponding to the i-th encoder. That is, it can be understood that the first encoder processes The second encoder processes the image The third encoder processes the image The fourth encoder processes the image image.

[0117] In Figure 7, the encoders in the encoder network execute in the order of first encoder, second encoder, third encoder, and fourth encoder. After the encoder network completes execution, the decoder network executes in the order of fourth decoder, third decoder, second decoder, and first decoder.

[0118] As shown in Figure 7, the third query feature of the first decoder is the first coding sub-feature output by the first encoder; the third query feature of the second decoder is the second coding sub-feature output by the second encoder; the third query feature of the third decoder is the third coding sub-feature output by the third encoder; and the third query feature of the fourth decoder is the fourth coding sub-feature output by the fourth encoder.

[0119] Step E12, when the target decoder is the first decoder, obtain the encoding sub-features of all encoders in the encoder network, and construct the third value feature and the third key feature of the target decoder.

[0120] In step E13, when the target decoder is any decoder other than the first decoder, the decoding sub-features output by each previous decoder before the target decoder are obtained as preferred sub-features; and based on the feature set constructed by the encoding sub-features of each encoder, the encoding sub-features of the encoder corresponding to the previous decoder are replaced with the preferred sub-features to obtain the third value feature and the third key feature of the target decoder.

[0121] As shown in Figure 7, after the encoder network completes processing, the first encoding sub-feature, the second encoding sub-feature, the third encoding sub-feature, and the fourth encoding sub-feature are obtained. Based on the first encoding sub-feature, the second encoding sub-feature, the third encoding sub-feature, and the fourth encoding sub-feature, the third value feature and the third key feature of the fourth decoder input are determined. Based on the fourth decoding sub-feature, the fourth decoding sub-feature is used to replace the fourth encoding sub-feature in the original third value feature and the third key feature as the input of the third decoder to obtain the third decoding sub-feature of the third decoder. Then, the first decoding sub-feature, the second decoding sub-feature, the third decoding sub-feature, and the fourth decoding sub-feature are obtained as the output of the decoder network.

[0122] Step E14: input the third query feature, the third value feature, and the third key feature into the target decoder to obtain a decoding sub-feature output by the target decoder.

[0123] In some embodiments, the architecture of each decoder is shown in FIG8 , including a mixed attention mechanism module (Mix-Attention), a layer normalization module (LN), and a feed-forward network (FFN). Specifically, the method for obtaining the decoding sub-features output by the target decoder can be implemented as follows:

[0124] Step F1: input the third query feature, the third key feature, and the third value feature into a hybrid attention mechanism module to obtain a first intermediate feature.

[0125] Step F2: After fusing the first intermediate feature and the third query feature, the first intermediate feature and the third query feature are input into a layer normalization module to obtain a second intermediate feature.

[0126] Step F3: input the second intermediate feature into the feedforward network to obtain the third intermediate feature.

[0127] Step F4: Fusing the second intermediate feature and the third intermediate feature to obtain a decoding sub-feature output by the target decoder.

[0128] Step E2: constructing the decoding feature based on the decoding sub-features of each decoder.

[0129] Based on the four decoding sub-features obtained, they are fused to obtain the decoding feature. The fusion method can be weighted summation, splicing, etc.

[0130] In the disclosed embodiment, compared to self-attention, the query features, key features, and value features generated are identical and originate from the same source, namely, the same encoder / decoder stage. In contrast, the disclosed embodiment employs a hybrid attention mechanism module that utilizes mixed features from multiple stages, each originating from a separate encoder. This allows query features to find matches across all different stages, enabling matching at a higher level of contextual granularity, thereby enhancing the accuracy of fault detection in the target image.

[0131] Based on the same technical concept, an embodiment of the present disclosure proposes a fault detection device 900, as shown in FIG9 , comprising:

[0132] The acquisition unit 901 is used to acquire images of the false twist components of the texturing machine to obtain a target image;

[0133] A construction unit 902 is configured to construct a first fusion feature based on at least one attention module in the attention network; wherein, for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructing a first sub-feature based on fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, and the second submodule using a residual constructed based on the second input information to obtain a mask map; the mask map is multiplied by the first sub-feature to obtain an output feature of the attention module; the first field of view is smaller than the second field of view;

[0134] The encoding unit 903 is configured to input the first fused feature into an encoder network to obtain an encoding feature; wherein the encoder network includes multiple encoders, each encoder outputs a corresponding encoding sub-feature, and the encoding feature includes the encoding sub-feature of at least one encoder;

[0135] A decoding unit 904 is configured to perform a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein the decoder network includes multiple decoders, each decoder outputs a corresponding decoding sub-feature, and the decoding feature includes the decoding sub-features of the multiple decoders;

[0136] A fusion unit 905 is configured to fuse the decoding sub-features in the decoding feature to obtain a second fused feature;

[0137] The prediction unit 906 is configured to input the second fusion feature into a multi-layer perceptron to obtain a fault detection result of the false twist component of the texturing machine.

[0138] In some embodiments, for the first attention module in the attention network:

[0139] The first input information of the first submodule of the first attention module is the target image;

[0140] The second input information of the second submodule of the first attention module includes the target image and at least one historically collected image of the false twist component, and the collection time interval between the target image and any of the historically collected images is less than a preset time length;

[0141] For any attention module other than the first attention module in the attention network:

[0142] The first input information of the first submodule of any attention module is the output feature output by the previous attention module of any attention module;

[0143] The second input information of the second submodule of any of the attention modules is the residual output of the second submodule of the previous attention module of any of the attention modules.

[0144] In some embodiments, the building block comprises:

[0145] The determining subunit is configured to determine a plurality of target points in the first input information, and perform the following operations for each target point:

[0146] Extracting fine-grained features of the target point within the first field of view centered at the target point; and

[0147] Extracting coarse-grained features of the target point within the second field of view centered at the target point;

[0148] Splicing the fine-grained features and the coarse-grained features to obtain the initial features of the target point;

[0149] Mapping the initial features of the target point to feature values ​​of the target point through a nonlinear mapping method;

[0150] The splicing subunit is configured to splice the feature values ​​of the plurality of target points according to the position information of each target point in the target image to obtain the first sub-feature.

[0151] In some embodiments, the determining subunit is specifically configured to:

[0152] determining a feature value of the target point in the first input information as a first query feature of the target point;

[0153] Determining the first visual field of the target point with the position coordinates of the target point in the first input information as the center;

[0154] constructing a first key feature and a first value feature of the target point based on feature values ​​of feature points outside the target point within the first field of view;

[0155] Based on the first query feature, the first key feature, and the first value feature, a fine-grained feature of the target point is determined.

[0156] In some embodiments, the determining subunit is specifically configured to:

[0157] Determining the second field of view of the target point with the position coordinates of the target point in the first input information as the center;

[0158] constructing a second key feature and a second value feature of the target point based on the feature values ​​of the feature points other than the target point within the second field of view;

[0159] Based on the first query feature, the second key feature, and the second value feature of the target point, a coarse-grained feature of the target point is determined; wherein the feature value of the target point in the first input information is the first query feature of the target point.

[0160] In some embodiments, the building block is specifically used to:

[0161] a transform subunit, configured to perform a Fourier transform on the second input information using a Fourier filter to obtain a time-varying component;

[0162] an estimation subunit, configured to input the time-varying component into an estimation module constructed based on a neural network to obtain an estimated value of the time-varying component;

[0163] The residual determination subunit is configured to determine the residual between the time-varying component and the estimated value of the time-varying component to obtain the mask map.

[0164] In some embodiments, the encoders in the encoder network correspond to the decoders in the decoder network in a one-to-one manner; and the decoding unit comprises:

[0165] A decoding sub-feature determination sub-unit is used for each target decoder in the decoder network, each target decoder also including a processing sub-unit, including:

[0166] Obtaining the encoding sub-feature output by the encoder corresponding to the target decoder as the third query feature of the target decoder; and

[0167] When the target decoder is the first decoder, obtaining encoding sub-features of all encoders in the encoder network, and constructing a third value feature and a third key feature of the target decoder;

[0168] When the target decoder is any decoder other than the first decoder, obtaining decoding sub-features output by each previous decoder before the target decoder as preferred sub-features; and replacing the encoding sub-features of the encoder corresponding to the previous decoder with the preferred sub-features for the feature set constructed from the encoding sub-features of each encoder, to obtain a third value feature and a third key feature of the target decoder;

[0169] Inputting the third query feature, the third value feature, and the third key feature into the target decoder to obtain a decoding sub-feature output by the target decoder;

[0170] The decoding feature construction subunit is configured to construct the decoding feature based on the decoding sub-features of each decoder.

[0171] In some embodiments, the decoding sub-feature determination subunit is specifically configured to:

[0172] Inputting the third query feature, the third key feature, and the third value feature into a hybrid attention mechanism module to obtain a first intermediate feature;

[0173] After fusing the first intermediate feature and the third query feature, the resultant feature is input into a layer normalization module to obtain a second intermediate feature;

[0174] Inputting the second intermediate feature into the feedforward network to obtain a third intermediate feature;

[0175] The second intermediate feature and the third intermediate feature are fused to obtain a decoding sub-feature output by the target decoder.

[0176] For the description of specific functions and examples of each module, sub-module\unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0177] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0178] Figure 10 is a block diagram of the structure of an electronic device according to an embodiment of the present disclosure. As shown in Figure 10, the electronic device includes: a memory 1010 and a processor 1020, and the memory 1010 stores a computer program that can be run on the processor 1020. The number of memories 1010 and processors 1020 can be one or more. The memory 1010 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes the method provided by the above method embodiment. The electronic device may also include: a communication interface 1030 for communicating with external devices and performing data exchange and transmission.

[0179] If the memory 1010, processor 1020, and communication interface 1030 are implemented independently, the memory 1010, processor 1020, and communication interface 1030 can be interconnected via a bus and communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG10 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0180] Optionally, in a specific implementation, if the memory 1010, the processor 1020 and the communication interface 1030 are integrated on a chip, the memory 1010, the processor 1020 and the communication interface 1030 can communicate with each other through an internal interface.

[0181] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0182] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).

[0183] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, data subscriber line (DSL)) or wireless (e.g., infrared, Bluetooth, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)). It is worth noting that the computer-readable storage medium mentioned in the present disclosure may be a non-volatile storage medium, in other words, a non-transient storage medium.

[0184] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0185] In the description of the embodiments of the present disclosure, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0186] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" in this document is only a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0187] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.

[0188] The above description is merely an exemplary embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.

Claims

1. A fault detection method, comprising: Performing image acquisition on the false-twist component of a texturing machine to obtain a target image; Constructing a first fusion feature based on at least one attention module in an attention network; wherein, for each attention module, the attention module includes a first sub-module and a second sub-module, the first sub-module constructs a first sub-feature for the fine-grained features within a first visual field range and the coarse-grained features within a second visual field range in the first input information, and the second sub-module obtains a mask map by using the residual based on the second input information; the mask map is multiplied by the first sub-feature to obtain the output feature of the attention module; the first visual field range is smaller than the second visual field range; Inputting the first fusion feature into an encoder network to obtain an encoded feature; wherein, the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes the encoded sub-features of at least one encoder; Performing a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein, the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of a plurality of decoders; Fusing the decoded sub-features in the decoded feature to obtain a second fusion feature; Inputting the second fusion feature into a multi-layer perceptron to obtain the fault detection result of the false-twist component of the texturing machine.

2. The method according to claim 1, wherein For the first attention module in the attention network: The first input information of the first sub-module of the first attention module is the target image; The second input information of the second sub-module of the first attention module includes the target image and at least one historical acquisition image of the false-twist component, and the time interval between the target image and any one of the historical acquisition images is less than a preset duration; For any attention module other than the first attention module in the attention network: The first input information of the first sub-module of the any attention module is the output feature output by the previous attention module of the any attention module; The second input information of the second sub-module of the any attention module is the residual output by the second sub-module of the previous attention module of the any attention module.

3. The method according to claim 2, wherein The constructing a first sub-feature for the fine-grained features within a first visual field range and the coarse-grained features within a second visual field range in the first input information includes: Determining a plurality of target points in the first input information, and respectively performing the following operations for each target point: Extracting the fine-grained feature of the target point within the first visual field range centered on the target point; and, Extracting the coarse-grained feature of the target point within the second visual field range centered on the target point; Concatenating the fine-grained feature and the coarse-grained feature to obtain the initial feature of the target point; Mapping the initial feature of the target point to the feature value of the target point by a non-linear mapping method; Concatenating the feature values of the plurality of target points according to the position information of each target point in the target image to obtain the first sub-feature.

4. The method according to claim 3, wherein, Extracting the fine-grained features of the target point within the first field of view centered at the target point includes: determining a feature value of the target point in the first input information as a first query feature of the target point; Determining the first visual field of the target point with the position coordinates of the target point in the first input information as the center; constructing a first key feature and a first value feature of the target point based on feature values of feature points outside the target point within the first field of view; Based on the first query feature, the first key feature, and the first value feature, a fine-grained feature of the target point is determined.

5. The method according to claim 3, wherein The extracting the coarse-grained features of the target point within the second field of view centered at the target point includes: Determining the second field of view of the target point with the position coordinates of the target point in the first input information as the center; constructing a second key feature and a second value feature of the target point based on the feature values of the feature points other than the target point within the second field of view; Based on the first query feature, the second key feature, and the second value feature of the target point, a coarse-grained feature of the target point is determined; wherein the feature value of the target point in the first input information is the first query feature of the target point.

6. The method according to claim 2, wherein The residual constructed based on the second input information obtains a mask map, including: Performing a Fourier transform on the second input information using a Fourier filter to obtain a time-varying component; Inputting the time-varying component into an estimation module constructed based on a neural network to obtain an estimated value of the time-varying component; A residual between the time-varying component and an estimate of the time-varying component is determined to obtain the mask map.

7. The method according to any one of claims 1-6, wherein, The encoder in the encoder network and the decoder in the decoder network correspond one to one; the decoding operation on the encoded features based on the decoder network to obtain the decoded features includes: For each target decoder in the decoder network, perform the following operations: Obtaining the encoding sub-feature output by the encoder corresponding to the target decoder as the third query feature of the target decoder; and When the target decoder is the first decoder, obtaining encoding sub-features of all encoders in the encoder network, and constructing a third value feature and a third key feature of the target decoder; When the target decoder is any decoder other than the first decoder, obtaining decoding sub-features output by each previous decoder before the target decoder as preferred sub-features; and replacing the encoding sub-features of the encoder corresponding to the previous decoder with the preferred sub-features for the feature set constructed from the encoding sub-features of each encoder, to obtain a third value feature and a third key feature of the target decoder; Inputting the third query feature, the third value feature, and the third key feature into the target decoder to obtain a decoding sub-feature output by the target decoder; The decoding feature is constructed based on the decoding sub-features of each decoder.

8. The method according to claim 7, wherein Inputting the third query feature, the third value feature, and the third key feature into the target decoder to obtain the decoded sub-feature output by the target decoder includes: Inputting the third query feature, the third key feature, and the third value feature into a hybrid attention mechanism module to obtain a first intermediate feature; After fusing the first intermediate feature and the third query feature, inputting them into a layer normalization module to obtain a second intermediate feature; Inputting the second intermediate feature into a feed-forward network to obtain a third intermediate feature; Fusing the second intermediate feature and the third intermediate feature to obtain the decoded sub-feature output by the target decoder.

9. A fault detection device, comprising: An acquisition unit configured to acquire an image of a false-twist component of a texturing machine to obtain a target image; A construction unit configured to construct a first fusion feature based on at least one attention module in an attention network; for each attention module, the attention module includes a first sub-module and a second sub-module, the first sub-module constructs a first sub-feature for fine-grained features within a first visual field range and coarse-grained features within a second visual field range in a first input information, the second sub-module uses a residual based on a second input information to obtain a mask map; the mask map is multiplied by the first sub-feature to obtain the output feature of the attention module; the first visual field range is smaller than the second visual field range; An encoding unit configured to input the first fusion feature into an encoder network to obtain an encoded feature; wherein, the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes the encoded sub-features of at least one encoder; A decoding unit configured to perform a decoding operation on the encoded feature based on a decoder network to obtain a decoded feature; wherein, the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of a plurality of decoders; A fusion unit configured to fuse the decoded sub-features in the decoded feature to obtain a second fusion feature; A prediction unit configured to input the second fusion feature into a multi-layer perceptron to obtain a fault detection result of the false-twist component of the texturing machine.

10. The device according to claim 9, wherein, Regarding the first attention module in the attention network: The first input information of the first sub-module of the first attention module is the target image; The second input information of the second sub-module of the first attention module includes the target image and at least one historical acquisition image of the false-twist component, and the time interval between the target image and any historical acquisition image is less than a preset duration; Regarding any attention module other than the first attention module in the attention network: The first input information of the first sub-module of the any attention module is the output feature output by the previous attention module of the any attention module; The second input information of the second sub-module of the any attention module is the residual output by the second sub-module of the previous attention module of the any attention module.

11. The apparatus according to claim 10, wherein, The construction unit includes: Determine sub-units, which are used to determine a plurality of target points in the first input information, and respectively perform the following operations for each target point: Extract the fine-grained features of the target point within the first field of view centered on the target point; and, Extract the coarse-grained features of the target point within the second field of view centered on the target point; Concatenate the fine-grained features and the coarse-grained features to obtain the initial features of the target point; Map the initial features of the target point to the feature values of the target point through a non-linear mapping method; A concatenation sub-unit, which is used to concatenate the feature values of the plurality of target points according to the position information of each target point in the target image to obtain the first sub-feature.

12. The apparatus according to claim 11, wherein, The determination sub-unit is specifically used for: Determine the feature value of the target point in the first input information as the first query feature of the target point; Determine the first field of view of the target point centered on the position coordinates of the target point in the first input information; Based on the feature values of the feature points other than the target point within the first field of view, construct the first key feature and the first value feature of the target point; Based on the first query feature, the first key feature, and the first value feature, determine the fine-grained features of the target point.

13. The device according to claim 11, wherein The determination sub-unit is specifically used for: Determine the second field of view of the target point centered on the position coordinates of the target point in the first input information; Based on the feature values of the feature points other than the target point within the second field of view, construct the second key feature and the second value feature of the target point; Based on the first query feature of the target point, the second key feature, and the second value feature, determine the coarse-grained features of the target point; wherein, the feature value of the target point in the first input information is the first query feature of the target point.

14. The apparatus according to claim 10, wherein, The construction unit includes: A transformation sub-unit, which is used to perform Fourier transform on the second input information by using a Fourier filter to obtain a time-varying component; An estimation sub-unit, which is used to input the time-varying component into an estimation module constructed based on a neural network to obtain an estimated value of the time-varying component; A residual determination sub-unit, which is used to determine the residual between the time-varying component and the estimated value of the time-varying component to obtain the mask graph.

15. The device according to any one of claims 9 - 14, wherein, The encoders in the encoder network and the decoders in the decoder network are in one-to-one correspondence; the decoding unit includes: A decoding sub-feature determination sub-unit, which is used to perform the following operations for each target decoder in the decoder network: Obtain the encoded sub-feature output by the encoder corresponding to the target decoder as the third query feature of the target decoder; and, In the case where the target decoder is the first decoder, obtain the encoded sub-features of all the encoders in the encoder network and construct the third value feature and the third key feature of the target decoder; In the case where the target decoder is any decoder other than the first decoder, obtain the decoded sub- Features, as preferred sub-features; and for the feature sets constructed from the encoding sub-features of each encoder, replace the encoding sub-features of the encoder corresponding to the prior decoder with the preferred sub-features to obtain the third value feature and the third key feature of the target decoder; Input the third query feature, the third value feature, and the third key feature into the target decoder to obtain the decoded sub-feature output by the target decoder; A decoding feature construction subunit, configured to construct the decoding feature based on the decoded sub-features of each decoder.

16. The apparatus according to claim 15, wherein, The decoded sub-feature determination subunit is specifically configured to: Input the third query feature, the third key feature, and the third value feature into the hybrid attention mechanism module to obtain a first intermediate feature; After fusing the first intermediate feature and the third query feature, input them into a layer normalization module to obtain a second intermediate feature; Input the second intermediate feature into a feed-forward network to obtain a third intermediate feature; Fuse the second intermediate feature and the third intermediate feature to obtain the decoded sub-feature output by the target decoder.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Conveyor belt edge detection method and device, electronic equipment and storage medium

    CN115359055A

  • Noisy plant point cloud semantic segmentation method and system based on self-attention feature fusion

    CN116311218A

  • Insulator image quality evaluation method and system based on multi-task learning

    CN116433647A

  • Multi-source data collaboration and fusion perception method for network automatic driving

    CN117237772A

  • Methods and apparatus to perform dense prediction using transformer blocks

    US20220012848A1