Fault detection method, apparatus, device, and storage medium
The method uses an attention network with fine-grained and coarse-grained feature extraction, combined with encoder and decoder networks, to automate the detection of wear in false twisting members, improving the false twisting process reliability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ZHEJIANG HENGYI PETROCHEMICAL CO LTD
- Filing Date
- 2025-01-21
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies lack effective methods for automatically monitoring the wear of false twisting members in false twisters, which can affect the false twisting process flow.
A fault detection method utilizing an attention network with fine-grained and coarse-grained feature extraction, followed by an encoder and decoder network, and a multilayer perceptron to analyze images of the false twisting member, enabling automated detection of wear.
Enables accurate and automated detection of wear in false twisting members, enhancing the reliability of the false twisting process.
Smart Images

Figure 0007854530000018 
Figure 0007854530000019 
Figure 0007854530000020
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of computers, and particularly to technical fields such as artificial intelligence, computer vision, and image processing.
Background Art
[0002] In the industrial scene of the false twisting process, the false twister plays an important role in the spinning process flow. In related technologies, the yarn related to the false twisting workplace includes POY (Pre Oriented Yarn), abbreviated as POY raw yarn, and DTY (Draw Textured Yarn).
[0003] The DTY yarn is obtained by the false twister processing the POY raw yarn. The false twister includes a false twisting member. The false twisting member performs a false twisting operation on the POY raw yarn. When the false twister is frequently used, the false twisting member of the false twister may wear, which may in turn affect the false twisting process flow. Therefore, realizing the automatic monitoring of the false twisting member of the false twister is one of the conventional problems.
Summary of the Invention
[0004] The present disclosure provides a fault detection method, apparatus, device, and storage medium for realizing the monitoring of a false twisting member in a false twisting process flow.
[0005] In a first aspect, the present disclosure performs image acquisition on the false twisting member of the false twister to obtain a target image, and A first fused feature is constructed based on at least one attention module in the attention network, wherein each attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map and the first sub-feature, and the first field of view is smaller than the second field of view. The first fused feature is input to an encoder network to obtain an encoded feature, wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes an encoded sub-feature of at least one encoder. The decoding operation is performed on the encoded feature based on the decoder network to obtain the decoded feature, wherein the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the plurality of decoders. The process involves fusing each decode sub-feature in the aforementioned decode feature to obtain a second fused feature, We propose a fault detection method characterized by including inputting the second fusion feature into a multilayer perceptron to obtain a fault detection result for the false-twisted member of the false-twist machine.
[0006] In the second phase, this disclosure is: An acquisition unit for acquiring target images by acquiring images of the false-twisted members of a false-twist machine, A construction unit for constructing a first fused feature based on at least one attention module in an attention network, wherein for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map and the first sub-feature, the first field of view is smaller than the second field of view, An encoding unit for inputting the first fused feature into an encoder network to obtain an encoded feature, the encoder network comprising a plurality of encoders, each encoder outputting a corresponding encoded sub-feature, the encoded feature comprising an encoded sub-feature of at least one encoder, A decoding unit for obtaining decoded features by performing a decoding operation on the encoded features based on a decoder network, wherein the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the plurality of decoders. A fusion unit for obtaining a second fused feature by fusing each decoded sub-feature in the aforementioned decoded feature, The present invention provides a fault detection device that includes a prediction unit for inputting the second fusion feature into a multilayer perceptron to obtain a fault detection result for the false-twisted member of the false-twist machine.
[0007] In the third phase, At least one processor, An electronic device including a memory that is communicably connected to at least one processor, The present invention provides an electronic device in which a memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to perform any of the methods of the embodiments of this disclosure.
[0008] In the fourth aspect, a non-temporary computer-readable storage medium is provided on which computer instructions are stored, causing the computer to perform any of the methods of the embodiments of the present disclosure.
[0009] In the fifth phase, a computer program product is provided, which includes a computer program, and which, when executed by a processor, implements any of the methods of the embodiments of the present disclosure.
[0010] Based on the method according to the embodiment of this disclosure, an automatic detection flow for false-twisted members in a false-twist machine was realized.
[0011] The information provided in the Summary of the Invention should be understood as not limiting the key points or important features of the embodiments of this disclosure, nor limiting the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description.
[0012] In the drawings, unless otherwise specified, the same reference numeral in multiple drawings indicates the same or similar member or element. These drawings are not necessarily drawn to scale. These drawings are merely illustrations of some embodiments provided in this disclosure and should not be considered to limit the scope of this disclosure. [Brief explanation of the drawing]
[0013] [Figure 1] Figure 1 is a schematic diagram of a false twisting machine in one embodiment of the present disclosure. [Figure 2] Figure 2 is a schematic diagram of a false-twist member in one embodiment of the present disclosure. [Figure 3] Figure 3 is a flow schematic diagram of a fault detection method according to an embodiment of the present disclosure. [Figure 4] Figure 4 is an architecture schematic diagram of an entire neural network model according to an embodiment of the present disclosure. [Figure 5] Figure 5 is a schematic diagram of a first sub-module of an attention module according to an embodiment of the present disclosure. [Figure 6] Figure 6 is a schematic diagram of a second sub-module of an attention network according to an embodiment of the present disclosure. [Figure 7] Figure 7 is a schematic diagram of an encoder network and a decoder network according to an embodiment of the present disclosure. [Figure 8] Figure 8 is a schematic diagram of an encoder according to an embodiment of the present disclosure. [Figure 9] Figure 9 is a configuration schematic diagram of a fault detection device according to an embodiment of the present disclosure. [Figure 10] Figure 10 is a block diagram of an electronic device for implementing the fault detection method of the embodiment of the present disclosure.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, the present disclosure will be described in more detail with reference to the drawings. In the drawings, the same reference numerals indicate functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0015] Also, in order to better explain the present disclosure, a number of specific details are described in the following specific embodiments. Those skilled in the art should understand that the present disclosure can be implemented similarly without specific details. In some examples, well-known methods, means, elements, circuits, etc. by those skilled in the art are not described in detail so as to emphasize the gist of the present disclosure.
[0016] Note that the terms "first" and "second" are for the purpose of distinction and should not be understood as indicating or implying relative importance or implying the number of indicated components. Thus, the features described with "first" and "second" mean explicitly or implicitly including one or more of such features. In the description of the present disclosure, unless otherwise specified, "plural" means two or more.
[0017] In the industrial scene of the false twisting process, the false twisting member of the false twister may be worn, and it is necessary to monitor the false twisting member of the false twister.
[0018] As a kind of false twister, the schematic diagram of its core structure is as shown in FIG. 1. The core structure of the false twister includes a raw yarn rack 101, a yarn cutter 102, a first roller 103, a first heating box 104, a cooling plate 105, a false twisting member 106, a nozzle 107, a second roller 108, a second heating box 109, a third roller 110, a yarn breakage detection device 111, and a winding member 112.
[0019] Among them, the first roller 103, the second roller 108, and the third roller 110 are for ensuring that the yarn is processed along a predetermined path. In the false twisting process flow, the first roller 103, the second roller 108, and the third roller 110 are coordinated with each other in speed so that the yarn product is not torn or stacked.
[0020] According to the needs of the product form, the nozzle 107 may be used to process the yarn processed through the false twisting member into a network-form yarn product so that different yarn products have the required touch and form.
[0021] When producing high-elasticity yarn, the second heating box 109 does not need to be used. When producing medium-elasticity yarn, the temperature of the second heating box 109 may be adjusted to the first preset temperature, for example, the first preset temperature may be around 140°C. When producing low-elasticity yarn, the temperature of the second heating box 109 may be adjusted to the second preset temperature, for example, the second preset temperature may be between 165°C and 195°C.
[0022] When the thread break detection device 111 detects a thread break, it triggers the thread cutter 102 to cut the thread so that the POY yarn does not accumulate in subsequent flows.
[0023] When processing a single thread, the corresponding false twisting member, as shown in Figure 1, 106, mainly employs two rubber rolls to perform a twisting motion on the thread product. When processing multiple threads, the false twisting member 106 processes the multiple threads into a combined thread. As one example, as shown in Figure 2, the first thread passes through from the upper left, and the second thread passes through from the upper right, and these two threads are then combined to obtain a single thread.
[0024] In order to automatically and accurately detect false-twist members in a false-twist machine, an embodiment of this disclosure provides a fault detection method, which, as shown in Figure 3, includes the following: S301: Image acquisition is performed on the false-twisted member of the false-twist machine to obtain the target image.
[0025] Unmanned aerial vehicles may be used to perform periodic inspections of the false-twist members of the false-twist machine, or images of the false-twist members may be acquired based on a monitor; however, any method for acquiring images of the false-twist members can be applied to the embodiments of this disclosure.
[0026] S302: A first fused feature is constructed based on at least one attention module in the attention network, where each attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map by the first sub-feature, and the first field of view is smaller than the second field of view.
[0027] S303: The first fused feature is input to an encoder network to obtain an encoded feature, where the encoder network includes multiple encoders, each encoder outputting a corresponding encoded sub-feature, and the encoded feature includes an encoded sub-feature of at least one encoder.
[0028] S304: A decoding operation is performed on the encoded feature based on the decoder network to obtain a decoded feature, where the decoder network includes multiple decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the multiple decoders.
[0029] S305: A second fused feature is obtained by fusing each decoded sub-feature in the decoded feature.
[0030] S306: The second fusion feature is input to the multilayer perceptron to obtain the failure detection result for the false twisted member of the false twisting machine.
[0031] For the attention network, encoder network, decoder network, and multilayer perceptron, actual images of past false-twist members can be used as sample images to train the parameters of these models (including the attention network, encoder network, decoder network, and multilayer perceptron). Specifically, the sample images are labeled with the fault location and fault type, and the fault type may be the degree of wear. The sample images are sequentially input into the attention network, encoder network, decoder network, and multilayer perceptron awaiting training to obtain the predicted fault location and predicted degree of wear. A first loss value is obtained based on the loss between the predicted fault location and the actual fault location, and a second loss value is obtained based on the loss between the predicted degree of wear and the actual degree of wear. The target loss is obtained by weighting the first and second loss values. Based on the target loss, the parameters of the attention network, encoder network, decoder network, and multilayer perceptron awaiting training are adjusted. When the convergence conditions were met, we were able to obtain an attention network, encoder network, decoder network, and multilayer perceptron for fault detection in the false-twisted members of the false-twisting machine in the target image.
[0032] In the embodiments of this disclosure, the attention network includes at least one attention module, where fine-grained features indicate local details of interest, and coarse-grained features indicate features surrounding those local details. The combination of fine-grained and coarse-grained features allows for the acquisition of information that matches the visual features, and features within different granularity ranges can be merged to enhance the expressiveness of the extracted features, thereby contributing to the accurate detection of false-twisted members. Furthermore, within the same attention module, residuals constructed based on second input information of the false-twisted member are merged so that soft attention features are further merged with the features output by the attention module. This allows the same attention model to extract features of the false-twisted member from multiple viewpoints, thereby enhancing the expressiveness of the features and, consequently, the accuracy of the detection results. Moreover, the more layers of attention modules there are in the attention network, the stronger the expressiveness of the output features of the extracted target image. Furthermore, the output features are input to subsequent encoders and decoders, and for each encoder in the encoder network, there exists one decoder corresponding to that encoder in the decoder network. Subsequently, the decoded sub-features output by multiple decoders are fused to obtain a second fused feature, and this second fused feature is input to a multilayer perceptron to obtain the fault detection result for the false-twisted members of the false-twist machine, thereby realizing an automated detection flow for the false-twisted members of the false-twist machine.
[0033] The overall architecture of the neural network model submitted in the embodiments of this disclosure is shown in Figure 4, and the model includes an attention network, an encoder network, a decoder network, and a multilayer perceptron, the specific structure of each network described in detail below. (1) Attention Network In some embodiments, the attention network may include N attention modules (including the first attention module, the second attention module, ..., the Nth attention module in Figure 4, where N is a positive integer), as shown in Figure 4. The specific input and output information for each attention module may be described as follows. [1] With respect to the first attention module in the attention network, The first input information of the first submodule of the first attention module is the target image. The second input information of the second submodule of the first attention module includes the target image and at least one previously acquired image of the false twist member, and the acquisition time interval between the target image and any one of the previously acquired images is smaller than a preset time length.
[0034] Among these, the previously acquired images may be m images obtained by sampling at equal intervals within a predetermined time length before acquiring the target image. To illustrate with an example, if the time at which the target image is sampled is time t, the predetermined time length is a, and the time interval for acquiring each image is b, then the first previously acquired image is acquired at time ta, the second previously acquired image is acquired at time t-a+b, and so on. Of course, previously acquired images may also be acquired by sampling at non-equal intervals within the predetermined time length. All methods for acquiring previously acquired images can be applied to the embodiments of this disclosure.
[0035] [2] For any attention module other than the first attention module in the attention network, The first input information of the first submodule of any of the aforementioned attention modules is the output feature output by the attention module immediately preceding any of the aforementioned attention modules. The second input information of the second submodule of any of the aforementioned attention modules is the residual output by the second submodule of the attention module immediately preceding any of the aforementioned attention modules.
[0036] As shown in Figure 4, the first submodule of the first attention module takes the target image as input and outputs the first first sub-feature. The second submodule of the first attention module takes a previously acquired image as input and outputs the first residual. The output of the first attention module, i.e., the first output feature, is obtained by multiplying the first first sub-feature by the first residual. The first output feature is used as input to the first submodule of the second attention module, and its output is the second first sub-feature. The first residual is used as input to the second submodule of the second attention module, and its output is the second residual. Similarly, the process is carried out sequentially as described above until the execution of N attention modules is complete, and the output of the attention network is obtained as the first fused feature.
[0037] In the embodiments of this disclosure, the target image can be used to detect the immediate state of the false-twisted member, thereby enabling fault detection. Past images of the target image can be used to determine the changes in the false-twisted member within a predetermined time period, and the immediate state and past state of the target image are considered together to comprehensively extract the characteristics of the false-twisted member in order to detect the degree of wear of the false-twisted member.
[0038] In some embodiments, constructing a first sub-feature with respect to the fine-grained features within a first field of view and the coarse-grained features within a second field of view in the first input information can be carried out as follows, that is, it can be carried out by including the following steps. Step A1: Determine multiple target points in the first input information, and perform the following operations for each target point. Step A11: Extract the fine-grained features of the target point within the first field of view centered on the target point.
[0039] In some embodiments, extracting the fine-grained features of the target point within the first field of view centered on the target point can be carried out as follows, that is, it can be carried out by including the following steps. Step B1: Determine the feature value of the target point in the first input information as the first query feature of the target point.
[0040] In order to obtain the pixel points that encompass the false-twisted member, i.e., the target points, the false-twisted member of the false-twisting machine is cut out from the target image, and the black rectangular frame in Figure 5 represents the target points.
[0041] Step B2: Determine the first field of view of the target point, centering on the position coordinates of the target point in the first input information.
[0042] The first field of view may be 3x3 or 5x5, and can be determined according to the actual situation; the embodiments of this disclosure are not limited thereto. The first field of view in the embodiments of this disclosure is as shown by the gray rectangular frame around the target point in Figure 5.
[0043] Step B3: Based on the feature values of feature points other than the target point within the first field of view, a first key feature and a first value feature of the target point are constructed.
[0044] Step B4: Based on the first query feature, the first key feature, and the first value feature, the fine-grained features of the target point are determined.
[0045] In one exemplary embodiment, the set of pixels in a sliding window centered at (i,j) in the feature diagram constructed from the first input information is defined as ρ(i,j). For a fixed window size of k×k, The filename is JPEG0007854530000001.jpg1345. For the first query feature and the first key feature, the fine-grained feature is calculated using the following formula (1). You can obtain JPEG0007854530000002.jpg1130. JPEG0007854530000003.jpg12102 In another exemplary embodiment, a positional offset and mask filling method may be introduced to obtain a different fine-grained feature. As shown in Figure 5, in the path for obtaining the fine-grained feature (hereinafter also referred to as the first path), the positional bias is the relative positional relationship between fine-grained tokens within the first field of view, and can also be understood as the relative positional relationship between each pixel point within the first field of view.
[0046] In the first path, the similarity of the zero-filling outside the boundaries of the feature map edges to the pixels is necessarily calculated. To prevent the calculated similarity values from affecting the softmax operation, a padding mask is used to set these results to -∞.
[0047] As a result, as shown in Figure 5, in the first path, the first fusion-awaited feature is obtained after performing an cross product operation on the first query feature and the first key feature as shown in equation (1), and then a fine-grained feature is obtained by performing an addition operation on each pixel with respect to the first fusion-awaited feature and the position offset after mask filling in the first path.
[0048] In the embodiments of this disclosure, the fine-grained features extracted by combining the features of the target point and the pixel points in the first field of view surrounding it have strong expressive power, which can then serve as a basis for determining the fault detection result.
[0049] Step A12: Extract the coarseness characteristics of the target point within the second field of view centered on the target point.
[0050] In some embodiments, extracting the coarseness characteristics of the target point within the second field of view centered on the target point can be carried out as follows, that is, it can be carried out by including the following steps. Step C1: Determine the second field of view of the target point, centering on the position coordinates of the target point in the first input information.
[0051] As shown in Figure 5, the second field of view is the pixel points within a larger area encompassing the first field of view, centered on the target point, i.e., the area shown in the white rectangular frame in Figure 5.
[0052] Step C2: Based on the feature values of feature points other than the target point within the second field of view, a second key feature and a second value feature of the target point are constructed.
[0053] Step C3: Based on the first query feature, the second key feature, and the second value feature of the target point, the coarse-grained feature of the target point is determined, and among these, the feature value in the first input information of the target point is the first query feature of the target point.
[0054] In one exemplary embodiment, the set of pixels in a second field of view centered at (i,j) in the feature diagram constructed from the first input information is defined as ρ'(i,j). The set of pixels obtained by pooling within the second field of view is defined as σ(X). The pooling size is Regarding JPEG0007854530000004.jpg1226, The filename is JPEG0007854530000005.jpg1351, and the following coarseness characteristics are obtained. JPEG0007854530000006.jpg13102 In another exemplary embodiment, a positional offset and mask filling method may be introduced to obtain a different coarseness characteristic. In the path for obtaining the coarseness characteristic (hereinafter referred to as the second path), the positional bias is the relative positional relationship between the coarseness tokens.
[0055] To further enhance the multi-scale image input capability of pixel focus attention, a different method is employed to calculate the position bias, and this position bias is within the second field of view. JPEG0007854530000007.jpg621 is also acceptable.
[0056] As shown in Figure 5, in the second path, a pooling operation may be performed first within the second field of view to sense the general features around the target point, and the pooling window and step size can be determined according to the actual situation. The pooling operation generates a new feature map by selecting the maximum or average pixel by moving the pooling window in the target image.
[0057] In the second path, a log-clustered continuous position bias (log-CPB) is used, and within that, the ReLU (activation function) is used to obtain the position coordinates (Q) of the first query feature. (i,j) ) and the pixel collection of the second key feature ( Spatial relative coordinates between JPEG0007854530000008.jpg1217) From JPEG0007854530000009.jpg1334, further Calculate JPEG0007854530000010.jpg934.
[0058] As a result, as shown in Figure 5, a second fusion-await feature is obtained by performing an cross product operation on the first query feature and the second key feature as shown in equation (2), and a coarse-grained feature is obtained by adding the second fusion-await feature and the position offset pixel by pixel.
[0059] In the embodiments of this disclosure, the ability to represent coarse-grained features extracted by combining the features of the target point and the pixel points of a second field of view surrounding it is enhanced, which then provides a basis for determining the fault detection result.
[0060] Step A13: The initial characteristics of the target point are obtained by combining the fine particle size characteristics and the coarse particle size characteristics.
[0061] The initial characteristics of the target point are obtained by combining the fine-grained and coarse-grained characteristics in an additive manner.
[0062] Step A14: The initial features of the target point are mapped to the feature values of the target point using a nonlinear mapping method.
[0063] Step A2: The feature values of the multiple target points are combined according to the positional information of each target point in the target image to obtain the first sub-feature.
[0064] As shown in Figure 5, coarse-grained and fine-grained features are input to a bonding layer and an activation layer to obtain a bonding feature, and a separation operation is performed on the bonding feature using a separation layer to separate the coarse-grained and fine-grained features, ultimately obtaining intermediate coarse-grained and intermediate fine-grained features. The target fine-grained feature is obtained by performing a cross product operation on the first value feature and the intermediate fine-grained feature, and the target coarse-grained feature is obtained by performing a cross product operation on the second value feature and the intermediate coarse-grained feature. An addition operation is performed on the target fine-grained feature and the target coarse-grained feature to obtain an initial feature, and then the feature value of the target point is obtained by passing the initial feature through a nonlinear mapping layer. The first sub-features are obtained by performing this operation sequentially for each target point.
[0065] In the embodiments of this disclosure, coarse grain features are used to capture the overall structure of the image, while fine grain features are relatively specific and accurate. By simultaneously capturing both coarse and fine grain features at a target point, the failure status of the false twisting member of the false twisting machine can be monitored more effectively.
[0066] In some embodiments, obtaining a mask map using residuals constructed based on the second input information can be carried out as follows, i.e., by including the following steps: Step D1: A Fourier transform is performed on the second input information using a Fourier filter to obtain the time-varying component.
[0067] In the attention network, the second input information for the second submodule of the first attention module is the previously acquired image mentioned above. For the second submodule of the second attention module to the Nth attention module in the attention network, the second input information is the residual output by the previous attention module. Figure 6 shows the framework diagram of the second submodule of each attention network.
[0068] Step D2: Input the time-varying component into an estimation module built on a neural network to obtain an estimate of the time-varying component.
[0069] Step D3: Determine the residual between the time-varying component and the estimated value of the time-varying component, and obtain the mask map.
[0070] Among these, the estimation module built on a neural network may be a time-varying Koopman predictor (Koopa), and the Koopa model consists of multiple stackable Koopa foundation modules.
[0071] Each Koopa foundation module focuses on learning dynamic characteristics at a specific level, and by stacking Koopa foundation modules, the model can capture complex, multi-layered dynamic changes in the time series. Each Koopa foundation module learns by using the residuals obtained by fitting the previous Koopa foundation module as input, ultimately obtaining an estimate of the time-varying component. The method presented in the embodiments of this disclosure can improve the prediction accuracy of the time-varying Koopman predictor and enhance the adaptability of the model to complex nonlinear and transient time series, such as extracting the dynamic change features of the false-twist member.
[0072] 2) Decoder Network In some embodiments, the encoders in the encoder network correspond one-to-one with the decoders in the decoder network, and Figure 7 shows that there are four encoders in the encoder network and four decoders in the decoder network. Regardless of the number of encoders and decoders, the decoding operation performed on the encoded features based on the decoder network to obtain the decoded features can be carried out as follows, that is, it can be carried out including the following steps. Step E1: Perform the following operation for each target decoder in the decoder network. Step E11: The encoded sub-features output by the encoder corresponding to the target decoder are obtained as the third query features of the target decoder.
[0073] As an example, the four encoders in the encoder network shown in Figure 7 have different sizes. The target image, JPEG0007854530000011.jpg936, is processed, where H is the height of the target image, W is the width of the target image, and four encoders are used. JPEG0007854530000012.jpg1455 is a layered, multi-resolution sub-feature E i Each of these results in, Regarding JPEG0007854530000013.jpg1954, C i is the weight corresponding to the i-th encoder. That is, the first encoder is This processes the image JPEG0007854530000014.jpg1737, and the second encoder is: This processes the image JPEG0007854530000015.jpg1637, and the third encoder is: This processes the image JPEG0007854530000016.jpg1546, and the fourth encoder is: It can be understood that this process is for the image JPEG0007854530000017.jpg1546.
[0074] In Figure 7, the execution order of the encoders in the encoder network is the first encoder, the second encoder, the third encoder, and the fourth encoder. After the encoder network has finished executing, the execution order of the decoder network is the fourth decoder, the third decoder, the second decoder, and the first decoder.
[0075] As shown in Figure 7, the third query feature of the first decoder is the first encoded sub-feature output by the first encoder, the third query feature of the second decoder is the second encoded sub-feature output by the second encoder, the third query feature of the third decoder is the third encoded sub-feature output by the third encoder, and the third query feature of the fourth decoder is the fourth encoded sub-feature output by the fourth encoder.
[0076] Step E12: If the target decoder is the first decoder, obtain the encoded sub-features of all encoders in the encoder network and construct the third value feature and third key feature of the target decoder.
[0077] Step E13: If the target decoder is any decoder other than the first decoder, the decode sub-features output by each preceding decoder before the target decoder are obtained as preferred sub-features, and the set of features constructed from the encoded sub-features of each encoder is replaced with the encoded sub-features of the encoder corresponding to the preceding decoder using the preferred sub-features, thereby obtaining the third value feature and the third key feature of the target decoder.
[0078] As shown in Figure 7, after the encoder network processing is complete, the first, second, third, and fourth encoded sub-features are obtained, and based on the first, second, third, and fourth encoded sub-features, the third value feature and the third key feature to be input to the fourth decoder are determined. After obtaining the fourth decode sub-features, the fourth decode sub-features are used to replace the fourth encoded sub-features in the original third value feature and third key feature, and this is used as input to the third decoder. In this way, the third decode sub-features of the third decoder are obtained, and by processing them sequentially, the first, second, third, and fourth decode sub-features output by the decoder network are obtained.
[0079] Step E14: The third query feature, the third value feature, and the third key feature are input to the target decoder, and the decoded sub-features output by the target decoder are obtained.
[0080] In some embodiments, the architecture of each decoder includes a Mix-Attention mechanism module, a Layer Normalization module (LN), and a Feedforward Network (FFN), as shown in Figure 8, and a specific method for obtaining the decoded sub-features output by the target decoder can be implemented as follows, that is, it can be implemented including the following steps. Step F1: The third query feature, the third key feature, and the third value feature are input into the mixed attention mechanism module to obtain the first intermediate feature.
[0081] Step F2: The first intermediate feature and the third query feature are merged and input into the layer normalization module to obtain the second intermediate feature.
[0082] Step F3: Input the second intermediate feature into the feedforward network to obtain a third intermediate feature.
[0083] Step F4: The second intermediate feature and the third intermediate feature are merged to obtain the decoded sub-features output by the target decoder.
[0084] Step E2: Construct the decoding features based on the decoding sub-features of each decoder.
[0085] Four decoded sub-features are obtained, and then these are merged to obtain a decoded feature. The merging method can be a weighted sum, a concatenation, or any other method.
[0086] In the embodiments of this disclosure, the self-attention mechanism uses the same source for generating query features, key features, and value features, i.e., from the same encoder / decoder. However, embodiments of this disclosure employ a mixed-attention mechanism module and employ a multi-scale staircase mixed feature, where each feature originates from an independent encoder. That is, different query features are permitted to originate from different decoding staircases, enabling different degrees of matching to contextual granularity, thereby enhancing the accuracy of the fault detection function from the target image.
[0087] Based on the same technical concept, an embodiment of the present disclosure provides a fault detection device 900, as shown in Figure 9, An acquisition unit 901 for acquiring images of a false twist member of a false twist machine and obtaining a target image, A construction unit 902 constructs a first fused feature based on at least one attention module in an attention network, wherein each attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map and the first sub-feature, the first field of view is smaller than the second field of view, An encoding unit 903 for inputting the first fused feature into an encoder network to obtain an encoded feature, wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes an encoded sub-feature of at least one encoder. A decoding unit 904 for obtaining decoded features by performing a decoding operation on the encoded features based on a decoder network, wherein the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the plurality of decoders. A fusion unit 905 for obtaining a second fused feature by fusing each decoded sub-feature in the aforementioned decoded feature, The system includes a prediction unit 906 that inputs the second fusion feature into a multilayer perceptron to obtain a fault detection result for the false-twisted member of the false-twist machine.
[0088] In some embodiments, the first attention module in the attention network is as follows: The first input information of the first submodule of the first attention module is the target image. The second input information of the second submodule of the first attention module includes the target image and at least one previously acquired image of the false twist member, and the acquisition time interval between the target image and any one of the previously acquired images is smaller than a preset time length. For any of the attention modules other than the first attention module in the aforementioned attention network, The first input information of the first submodule of any of the aforementioned attention modules is the output feature output by the attention module immediately preceding any of the aforementioned attention modules. The second input information of the second submodule of any of the aforementioned attention modules is the residual output by the second submodule of the attention module immediately preceding any of the aforementioned attention modules.
[0089] In some embodiments, the construction unit is A determination subunit for determining multiple target points in the first input information, and for each target point, performing the following operations: extracting fine-grained features of the target point within a first field of view centered on the target point; extracting coarse-grained features of the target point within a second field of view centered on the target point; obtaining initial features of the target point by combining the fine-grained features and the coarse-grained features; and obtaining feature values of the target point by mapping the initial features of the target point using a nonlinear mapping method. The system includes a composite subunit for obtaining the first sub-feature by combining the feature values of the plurality of target points according to the positional information of each target point in the target image.
[0090] In some embodiments, the defining subunit is specifically, The feature value of the target point in the first input information is determined as the first query feature of the target point. The first field of view range of the target point is determined, centered on the position coordinates of the target point in the first input information. Based on the feature values of feature points other than the target point within the first field of view, a first key feature and a first value feature of the target point are constructed. This is for determining the fine-grained characteristics of the target point based on the first query characteristics, the first key characteristics, and the first value characteristics.
[0091] In some embodiments, the defining subunit is specifically, The second field of view of the target point is determined with respect to the position coordinates of the target point in the first input information, Based on the feature values of feature points other than the target point within the second field of view, a second key feature and a second value feature of the target point are constructed. This is for determining the coarseness characteristics of the target point based on the first query characteristics, the second key characteristics, and the second value characteristics of the target point. The feature value in the first input information for the target point is the first query feature of the target point.
[0092] In some embodiments, the construction unit is specifically, A transformation subunit employing a Fourier filter to perform a Fourier transform on the second input information to obtain the time-varying component, An estimation subunit for obtaining an estimate of the time-varying component by inputting the time-varying component into an estimation module built on a neural network, The system includes a residual determination subunit for determining the residual between the time-varying component and the estimated value of the time-varying component, and for obtaining the mask map.
[0093] In some embodiments, the encoders in the encoder network correspond one-to-one with the decoders in the decoder network, and the decoding unit is A decode sub-feature determination subunit for performing the following operations for each target decoder in the decoder network: acquiring the encoded sub-features output by the encoder corresponding to the target decoder as the third query features of the target decoder; if the target decoder is the first decoder, acquiring the encoded sub-features of all encoders in the encoder network and constructing the third value features and third key features of the target decoder; if the target decoder is any decoder other than the first decoder, acquiring the decode sub-features output by each preceding decoder before the target decoder as preferred sub-features, replacing the encoded sub-features of the encoder corresponding to the preceding decoder with the preferred sub-features in the feature set constructed from the encoded sub-features of each encoder, and acquiring the third value features and third key features of the target decoder; and inputting the third query features, the third value features, and the third key features to the target decoder and acquiring the decode sub-features output by the target decoder. It includes a decode feature construction subunit for constructing the decode features based on the decode sub-features of each decoder.
[0094] In some embodiments, the decode sub-feature determination subunit is specifically, The third query feature, the third key feature, and the third value feature are input to the mixed attention mechanism module to obtain the first intermediate feature. The first intermediate feature and the third query feature are merged and input into the layer normalization module to obtain the second intermediate feature. The second intermediate feature is input into the feedforward network to obtain a third intermediate feature. This method is used to merge the second intermediate feature and the third intermediate feature to obtain the decoded sub-features output by the target decoder.
[0095] A description of the specific functions and examples of each module, submodule / unit of the apparatus according to the embodiments of this disclosure can be found in the relevant descriptions of the corresponding steps in the embodiments of the method described above, and is therefore omitted here.
[0096] In the proposed technology disclosed herein, the acquisition, storage, and use of users' personal information will comply with the provisions of applicable laws and regulations and will not violate public order and morals.
[0097] Figure 10 is a block diagram of the configuration of an electronic device according to one embodiment of the present disclosure. As shown in Figure 10, the electronic device includes a memory 1010 and a processor 1020, the memory 1010 storing a computer program that can be executed by the processor 1020. The number of memories 1010 and processors 1020 may be one or more. The memory 1010 may store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device can perform the method provided in the embodiment of the above method. The electronic device may further include a communication interface 1030 for communicating with external devices and exchanging and transmitting data.
[0098] If the memory 1010, processor 1020, and communication interface 1030 are separate components, then the memory 1010, processor 1020, and communication interface 1030 are connected to each other via a bus and can communicate with each other. This bus may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For convenience, Figure 10 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0099] Selectively, as a concrete implementation, if the memory 1010, processor 1020, and communication interface 1030 are integrated onto a single chip, the memory 1010, processor 1020, and communication interface 1030 can communicate with each other via an internal interface.
[0100] The processor may be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any ordinary processor. Furthermore, the processor may be a processor capable of supporting an Advanced RISC Machine (ARM) architecture.
[0101] Furthermore, the memory may selectively include read-only memory and random access memory, or non-volatile random access memory. The memory may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may include ROM (Read-Only Memory), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically EPROM (EEPROM), or flash memory. Volatile memory may include Random Access Memory (RAM) used as an external cache. The above description is illustrative and not restrictive. Many forms of RAM are available. For example, static random access memory (Static RAM, SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Date SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchlink dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus random access memory (Direct RAM BUS RAM, DR RAM) may be used.
[0102] In the embodiments described above, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. If implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the flows or functions described in the embodiments of this disclosure are generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer instructions may be stored on a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., coaxial cable, optical fiber, digital subscriber line, DSL) or wireless connection (e.g., infrared, Bluetooth®, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer, or it may be a data storage device such as a server or data center that includes one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, or magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), semiconductor media (e.g., Solid State Disks (SSDs)), etc. The computer-readable storage medium according to this disclosure may also be a non-volatile storage medium, in other words, a non-temporary storage medium.
[0103] Those skilled in the art will understand that all or part of the steps for realizing the above embodiment may be completed by hardware, or by a program that instructs the relevant hardware, and that the program may be stored in a computer-readable storage medium, the storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.
[0104] In the descriptions of the embodiments of this disclosure, the terms “one embodiment,” “several embodiments,” “example,” “specific example,” or “several examples” mean that the specific features, structures, materials, or characteristics described in relation to such embodiment or example are included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in an appropriate manner in any one or more embodiments or examples. Furthermore, a person skilled in the art may combine different embodiments or examples and features in different embodiments or examples described herein, provided that they are not inconsistent.
[0105] In the description of the embodiments of this disclosure, unless otherwise specified, " / " means "or," for example, "A / B" can represent "A" or "B." The "and / or" statements in this specification are merely related relationships that describe related subjects and mean that there are three possible relationships, for example, "A and / or B" can indicate three situations: "A" exists alone, "A" and "B" exist together, and "B" exists alone.
[0106] In the description of the embodiments of this disclosure, the terms “first” and “second” are for distinction purposes only and should not be understood to indicate or imply relative importance or the number of designated constituent elements. Thus, features limited by “first” and “second” may explicitly or implicitly include one or more such features. In the description of the embodiments of this disclosure, unless otherwise specified, “multiple” means two or more.
[0107] The foregoing are merely illustrative examples of the Disclosure and do not limit the Disclosure. Any modifications, equivalent substitutions, or improvements made to the spirit and principles of the Disclosure should be included within the scope of the claims of the Disclosure.
Claims
1. A computer having a processor, This involves acquiring images of the false-twisted members of the false-twist machine to obtain the target image, Constructing a first fused feature based on at least one attention module in the attention network (where, for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map and the first sub-feature, and the first field of view is smaller than the second field of view), The first fused feature is input to an encoder network to obtain an encoded feature (wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes an encoded sub-feature of at least one encoder), The process involves performing a decoding operation on the encoded features based on a decoder network to obtain decoded features (wherein the decoder network includes multiple decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the multiple decoders), The process involves fusing each decode sub-feature in the aforementioned decode feature to obtain a second fused feature, A fault detection method characterized by inputting the second fusion feature into a multilayer perceptron to obtain a fault detection result for the false twist member of the false twist machine.
2. Regarding the first attention module in the aforementioned attention network, The first input information of the first submodule of the first attention module is the target image. The second input information of the second submodule of the first attention module includes the target image and at least one previously acquired image of the false twist member, and the acquisition time interval between the target image and any one of the previously acquired images is smaller than a preset time length. For any of the attention modules other than the first attention module in the aforementioned attention network, The first input information of the first submodule of any of the aforementioned attention modules is the output feature output by the attention module immediately preceding any of the aforementioned attention modules. The second input information of the second submodule of any of the aforementioned attention modules is the residual output by the second submodule of the attention module immediately preceding any of the aforementioned attention modules. The method according to feature 1.
3. Constructing a first sub-feature with respect to the fine-grained features within the first field of view and the coarse-grained features within the second field of view in the first input information is: The first input information determines a plurality of target points, and for each target point, the fine-grained features of the target point are extracted within the first field of view centered on the target point, the coarse-grained features of the target point are extracted within the second field of view centered on the target point, the fine-grained features and the coarse-grained features are combined to obtain the initial features of the target point, and the initial features of the target point are mapped to the feature values of the target point using a nonlinear mapping method. This includes combining the feature values of the plurality of target points according to the positional information of each target point in the target image to obtain the first sub-feature, The method according to feature 2.
4. Extracting the fine-grained characteristics of the target point within the first field of view centered on the target point is: The characteristic value in the first input information of the target point is determined as the first query feature of the target point, The first field of view of the target point is determined with respect to the position coordinates of the target point in the first input information, Based on the feature values of feature points other than the target point within the first field of view, a first key feature and a first value feature of the target point are constructed. This includes determining the fine-grained characteristics of the target point based on the first query characteristics, the first key characteristics, and the first value characteristics. The method according to feature 3.
5. Extracting the coarseness characteristics of the target point within the second field of view centered on the target point is: The second field of view of the target point is determined with respect to the position coordinates of the target point in the first input information, Based on the feature values of feature points other than the target point within the second field of view, a second key feature and a second value feature of the target point are constructed. This includes determining the coarseness characteristics of the target point based on the first query characteristics, the second key characteristics, and the second value characteristics of the target point, The feature value in the first input information of the target point is the first query feature of the target point. The method according to feature 3.
6. Obtaining a mask map using the residuals constructed based on the second input information is: By employing a Fourier filter and performing a Fourier transform on the second input information to obtain the time-varying component, The time-varying component is input to an estimation module constructed based on a neural network to obtain an estimated value of the said time-varying component, This includes determining the residual between the time-varying component and the estimated value of the time-varying component to obtain the mask map, The method according to feature 2.
7. The encoders in the encoder network correspond one-to-one with the decoders in the decoder network. Performing a decoding operation on the encoded features based on the decoder network to obtain the decoded features is, For each target decoder in the decoder network, the following operations are performed: acquiring the encoded sub-features output by the encoder corresponding to the target decoder as the third query features of the target decoder; if the target decoder is the first decoder, acquiring the encoded sub-features of all encoders in the encoder network and constructing the third value features and third key features of the target decoder; if the target decoder is any decoder other than the first decoder, acquiring the decoded sub-features output by each preceding decoder before the target decoder as preferred sub-features, replacing the encoded sub-features of the encoder corresponding to the preceding decoder with the preferred sub-features in the feature set constructed from the encoded sub-features of each encoder, and acquiring the third value features and third key features of the target decoder; and inputting the third query features, the third value features, and the third key features to the target decoder and acquiring the decoded sub-features output by the target decoder. This includes constructing the decoding features based on the decoding sub-features of each decoder, The method according to any one of claims 1 to 6, characterized by...
8. Inputting the third query feature, the third value feature, and the third key feature into the target decoder and obtaining the decoded sub-features output by the target decoder is: The third query feature, the third key feature, and the third value feature are input to the mixed attention mechanism module to obtain the first intermediate feature, The first intermediate feature and the third query feature are merged and input into the layer normalization module to obtain the second intermediate feature, The second intermediate feature is input into the feedforward network to obtain a third intermediate feature, This includes fusing the second intermediate feature and the third intermediate feature to obtain a decoded sub-feature output by the target decoder, The method according to feature 7.
9. A fault detection device, An acquisition unit for acquiring target images by acquiring images of the false-twisted members of a false-twist machine, A construction unit for constructing a first fused feature based on at least one attention module in an attention network (where, for each attention module, the attention module includes a first submodule and a second submodule, the first submodule constructs a first sub-feature for fine-grained features within a first field of view and coarse-grained features within a second field of view in the first input information, the second submodule obtains a mask map using residuals constructed based on the second input information, the output feature of the attention module is obtained by multiplying the mask map and the first sub-feature, and the first field of view is smaller than the second field of view), An encoding unit for inputting the first fused feature into an encoder network to obtain an encoded feature (wherein the encoder network includes a plurality of encoders, each encoder outputs a corresponding encoded sub-feature, and the encoded feature includes an encoded sub-feature of at least one encoder), A decoding unit for obtaining decoded features by performing a decoding operation on the encoded features based on a decoder network (wherein the decoder network includes a plurality of decoders, each decoder outputs a corresponding decoded sub-feature, and the decoded feature includes the decoded sub-features of the plurality of decoders), A fusion unit for obtaining a second fused feature by fusing each decoded sub-feature in the aforementioned decoded feature, The system includes a prediction unit for inputting the second fusion feature into a multilayer perceptron to obtain a fault detection result for the false twist member of the false twist machine, A fault detection device characterized by the following features.
10. Regarding the first attention module in the aforementioned attention network, The first input information of the first submodule of the first attention module is the target image. The second input information of the second submodule of the first attention module includes the target image and at least one previously acquired image of the false twist member, and the acquisition time interval between the target image and any one of the previously acquired images is smaller than a preset time length. For any of the attention modules other than the first attention module in the aforementioned attention network, The first input information of the first submodule of any of the aforementioned attention modules is the output feature output by the attention module immediately preceding any of the aforementioned attention modules. The second input information of the second submodule of any of the aforementioned attention modules is the residual output by the second submodule of the attention module immediately preceding any of the aforementioned attention modules. The apparatus according to feature 9.
11. The aforementioned construction unit is A determination subunit for determining multiple target points in the first input information, and for each target point, performing the following operations: extracting fine-grained features of the target point within a first field of view centered on the target point; extracting coarse-grained features of the target point within a second field of view centered on the target point; combining the fine-grained features and the coarse-grained features to obtain initial features of the target point; and mapping the initial features of the target point to feature values of the target point using a nonlinear mapping method. Includes a combining subunit for obtaining the first sub-feature by combining the feature values of the plurality of target points according to the positional information of each target point in the target image, The apparatus according to feature 10.
12. The aforementioned definite subunit is, The feature value of the target point in the first input information is determined as the first query feature of the target point. The first field of view range of the target point is determined, centered on the position coordinates of the target point in the first input information. Based on the feature values of feature points other than the target point within the first field of view, a first key feature and a first value feature of the target point are constructed. This is for determining the fine-grained characteristics of the target point based on the first query characteristics, the first key characteristics, and the first value characteristics. The apparatus according to feature 11.
13. The aforementioned definite subunit is, The second field of view of the target point is determined with respect to the position coordinates of the target point in the first input information, Based on the feature values of feature points other than the target point within the second field of view, a second key feature and a second value feature of the target point are constructed. This is for determining the coarseness characteristics of the target point based on the first query characteristics, the second key characteristics, and the second value characteristics of the target point. The feature value in the first input information of the target point is the first query feature of the target point. The apparatus according to feature 11.
14. The aforementioned construction unit is A transformation subunit employing a Fourier filter to perform a Fourier transform on the second input information to obtain time-varying components, An estimation subunit for obtaining an estimate of the time-varying component by inputting an estimation module constructed based on a neural network, A residual determination subunit for determining the residual between the time-varying component and the estimated value of the time-varying component and obtaining the mask map, is included. The apparatus according to feature 10.
15. The encoders in the encoder network correspond one-to-one with the decoders in the decoder network. The decoding unit is, A decode sub-feature determination subunit for performing the following operations for each target decoder in the decoder network: acquiring the encoded sub-features output by the encoder corresponding to the target decoder as the third query features of the target decoder; if the target decoder is the first decoder, acquiring the encoded sub-features of all encoders in the encoder network and constructing the third value features and third key features of the target decoder; if the target decoder is any decoder other than the first decoder, acquiring the decode sub-features output by each preceding decoder before the target decoder as preferred sub-features, replacing the encoded sub-features of the encoder corresponding to the preceding decoder with the preferred sub-features in the feature set constructed from the encoded sub-features of each encoder, and acquiring the third value features and third key features of the target decoder; and inputting the third query features, the third value features, and the third key features to the target decoder and acquiring the decode sub-features output by the target decoder. A decode feature construction subunit for constructing the decode feature based on the decode sub-features of each decoder, The apparatus according to any one of claims 9 to 14.
16. The aforementioned decode sub-feature determination subunit is, The third query feature, the third key feature, and the third value feature are input to the mixed attention mechanism module to obtain the first intermediate feature. The first intermediate feature and the third query feature are merged and input into the layer normalization module to obtain the second intermediate feature. The second intermediate feature is input into the feedforward network to obtain a third intermediate feature. This method is for fusing the second intermediate feature and the third intermediate feature to obtain the decoded sub-features output by the target decoder. The apparatus according to feature 15.
17. It is an electronic device, At least one processor, Includes memory that is communicably connected to at least one processor, The memory stores a command that can be executed by the at least one processor, and the command is executed by the at least one processor to cause the at least one processor to perform the method according to any one of claims 1 to 6. electronic equipment.
18. A non-temporary, computer-readable storage medium storing computer instructions for causing a computer to perform the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
False twisted yarn of polyester composite fiber and method for production thereof
JP2007186844A
Coarse-to-fine attention networks for optical signal detection and recognition
JP2023540989A
Method, apparatus, electronic device and computer readable storage medium for image searching
US20200242153A1
Platform-aware transformer-based performance prediction
US20230306083A1