Spinning fault prediction method and device, equipment and storage medium

By using multimodal data fusion and large language models, the problem of untimely fault detection in spinning equipment has been solved, enabling intelligent and refined management of the spinning process and improving the accuracy of fault prediction and production stability.

CN121935706APending Publication Date: 2026-04-28ZHEJIANG HENGYI PETROCHEMICAL CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG HENGYI PETROCHEMICAL CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing spinning processes, fault detection of spinning equipment relies on manual experience and simple sensors, which leads to untimely fault detection and affects product quality and production schedule.

Method used

A multimodal data fusion method is adopted, which collects video streams and time series of monitoring parameters of target components in the spinning process, and uses deep learning and large language models to predict spinning faults, including video feature extraction, text feature extraction and optimization of visual feature representation, to achieve accurate fault prediction.

Benefits of technology

It improves the accuracy and stability of spinning fault prediction, reduces resource waste, ensures smooth production and product quality, and realizes intelligent control of the spinning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935706A_ABST
    Figure CN121935706A_ABST
Patent Text Reader

Abstract

The invention provides a spinning fault prediction method and device, equipment and a storage medium, and relates to the field of data processing. Collecting a video stream for the target component, and collecting a time sequence; acquiring a video clip of the current time window from the video stream, and acquiring a parameter clip of the current time window from the time sequence; inputting the video clip and the parameter clip into a parameter initialization model to obtain a parameter initial value of a controller of the fault prediction model; inputting the video clip into a visual encoder of the fault prediction model to obtain an initial visual feature; inputting the parameter fragments into a text encoder of the fault prediction model to obtain text features; inputting the initial visual features and the text features into a large language model of the fault prediction model to construct key value pair information of a last decoding layer of the large language model; optimizing a visual feature representation in the key value pair information based on a parameter initial value of the controller; and determining a spinning fault prediction result based on the optimized visual feature representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and in particular to the technical fields of neural networks and large language models. Background Technology

[0002] Spinning is an industrial process that extrudes melt through a spinneret to form fibers, and it is widely used in textiles, chemicals, and other fields. The stable operation of spinning equipment during the spinning process plays a decisive role in product quality. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for predicting spinning failures to solve or mitigate one or more technical problems in the prior art.

[0004] According to one aspect of this disclosure, a method for predicting spinning defects is provided, comprising: Video streams are acquired from target components in the spinning process, and a time sequence of preset monitoring parameters is also acquired; the target components include spinning assemblies and / or winding mechanisms. Obtain the video segment of the current time window from the video stream, and obtain the parameter segment of the current time window from the time sequence; The video clip and the parameter clip are input into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model; The video clip is input into the visual encoder of the fault prediction model to obtain initial visual features; The parameter fragments are input into the text encoder of the fault prediction model to obtain text features; The initial visual features and the text features are input into the large language model of the fault prediction model to construct the key-value pair information of the last decoding layer of the large language model; The visual feature representation in the key-value pair information is optimized based on the initial parameter values ​​of the controller; Based on the optimized visual feature representation, the spinning fault prediction result for the current time window is determined: According to another aspect of this disclosure, a device for predicting spinning defects is provided, comprising: The acquisition module is used to acquire video streams of target components in the spinning process and to acquire time sequences of preset monitoring parameters; the target components include spinning assemblies and / or winding mechanisms. The acquisition module is used to acquire a video segment of the current time window from the video stream, and to acquire a parameter segment of the current time window from the time sequence; The first processing module is used to input the video clip and the parameter clip into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model; The visual feature extraction module is used to input the video clip into the visual encoder of the fault prediction model to obtain initial visual features; The text feature extraction module is used to input the parameter fragments into the text encoder of the fault prediction model to obtain text features; The second processing module is used to input the initial visual features and the text features into the large language model of the fault prediction model to construct the key-value pair information of the last decoding layer of the large language model; An optimization module is used to optimize the visual feature representation in the key-value pair information based on the initial parameter values ​​of the controller; The determination module is used to determine the spinning fault prediction result for the current time window based on the optimized visual feature representation.

[0005] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods of any embodiment of the present disclosure.

[0006] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method according to any embodiment of this disclosure.

[0007] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method according to any embodiment of this disclosure.

[0008] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.

[0010] Figure 1 This is a flowchart illustrating a method for predicting spinning failures according to an embodiment of the present disclosure. Figure 2 This is a flowchart illustrating a method for predicting spinning failures according to another embodiment of the present disclosure; Figure 3 This is a flowchart illustrating a method for predicting spinning failures according to another embodiment of the present disclosure; Figure 4 This is a schematic diagram of obtaining the initial channel features of a first channel according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of a spinning failure prediction device according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of an electronic device structure for a method for predicting spinning failures according to the ninth embodiment of this disclosure. Detailed Implementation

[0011] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0012] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0013] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "multiple" means two or more, unless otherwise explicitly specified.

[0014] It should be noted that, unless it is explicitly stated that there is a sequential order of execution between different operations, or that there is a sequential order of execution between different operations in terms of technical implementation, the execution order between multiple operations may not be significant, and multiple operations may be executed simultaneously.

[0015] In industrial applications of spinning processes, key components include spinning assemblies and winding mechanisms.

[0016] The spinning assembly includes a heating device, melt distribution pipe, auxiliary material conveying pipe, metering pump, spinneret assembly, cooling blowing system, and oil nozzle oiling.

[0017] The melt distribution pipe has an insulation layer, and the heating device needs to ensure the temperature of the melt inside the melt distribution pipe so that the melt can be distributed to the metering pump. The metering pump evenly distributes the melt to each component of the spinning box, and the melt is formed into a fine stream through the fine holes of the spinneret component in the spinning assembly. The fine stream of melt is then formed into a filament by the cooling air blowing system.

[0018] The melt needs to be kept warm in the piping of the spinning box to ensure the temperature is as uniform as possible.

[0019] The spinneret assembly is the core forming component of the spinning process, including filter sand, distribution plate, spinneret, etc.

[0020] The winding mechanism includes a guide roller system, a network nozzle, a traverse winding mechanism, a roll forming control component, and a tapered tension control component.

[0021] The guide roller system typically includes 2-3 pairs of normal temperature / micro-heated rollers for precise control of spinning tension.

[0022] GR1 (First Guide Roller): Pulls the filament bundle coming out of the nozzle at a speed slightly lower than the spinneret extrusion speed (creating negative or micro-stretching) and controls the pretension.

[0023] GR2 (Second Guide Roller): Main traction roller, its speed determines the winding speed (2500-4000 m / min).

[0024] GR3 (Third guide roller, optional): For additional tension adjustment or network nozzle mounting position.

[0025] Network nozzle: A device that uses compressed air to impact the filament bundle, causing the monofilaments to partially open and entangle, forming periodic network nodes.

[0026] Traverse winding mechanism: friction roller drive + precision traverse guide Roll forming control: An intelligent control system that achieves uniform roll hardness and no overlapping defects by dynamically adjusting tension and lateral movement parameters.

[0027] Taper Tension Control: Actively reduces winding tension as the roll diameter increases.

[0028] In traditional spinning processes, monitoring and control of each component relies primarily on human experience and simple physical sensors. Manual inspection is not only inefficient but also struggles to capture subtle changes in each component in real time and accurately, easily leading to delayed fault detection and impacting product quality and production schedule. Furthermore, simple physical sensors can only acquire limited physical parameters, such as temperature and vibration, failing to comprehensively and intuitively reflect the actual working status of each component.

[0029] During the spinning process, any malfunction of the aforementioned components may lead to a decline in product quality or production interruption. Therefore, to improve the prediction of malfunctions in the spinning process, this disclosure proposes a method for predicting spinning malfunctions, such as... Figure 1 As shown, it includes: S101, acquire video streams of target components in the spinning process, and acquire time sequence of preset monitoring parameters; target components include spinning assemblies and / or winding mechanisms.

[0030] The spinning assembly and the winding mechanism can be used as target components, and any one of the sub-assemblies can also be used as a target component individually.

[0031] For example, in specific implementation, the guide roller system in the winding mechanism can be used as the target component, and special attention should be paid to GR2 in the guide roller system.

[0032] It should be noted that the main types of spinning products involved in the spinning process of the embodiments of this disclosure may include one or more of the following: partially oriented yarns (POY), fully drawn yarns (FDY), and polyester staple fiber. For example, the specific type of yarn may include polyester partially oriented yarns, polyester fully drawn yarns, polyester drawn yarns, and polyester staple fiber.

[0033] In addition, spinning products can also include polyester chips (PET), polycaprolactam chips (PA6), and other chip products.

[0034] S102, obtain the video segment of the current time window from the video stream, and obtain the parameter segment of the current time window from the time sequence.

[0035] During implementation, a dynamic image sequence generated by continuously capturing the real-time operation of the target component using visual acquisition devices such as industrial cameras and high-speed video equipment can be obtained as a video stream, which is used to capture the appearance of the target component.

[0036] The timing sequence of monitoring parameters for the target component can be collected using sensors. The specific sensors used and the parameters collected will be explained later.

[0037] The current time window can be the current time point plus a specified duration prior to it. The duration of the time window can be set according to the detection requirements.

[0038] S103, input the video clip and parameter clip into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model.

[0039] The parameter initialization model is used to map multimodal inputs (video segments + parameter segments) to the initial parameter matrix of the value buffer controller in the fault prediction model, so as to reduce the gradient vanishing problem caused by traditional zero initialization, while maintaining the physical interpretability of the low-rank structure.

[0040] S104, input the video clip into the visual encoder of the fault prediction model to obtain the initial visual features.

[0041] A visual encoder is a feature extraction module built on deep learning or traditional machine vision algorithms. Its core function is to extract and transform features from image information in videos through neural network models, and to convert video data into information that can be understood and processed by large language models.

[0042] S105, input the parameter fragment into the text encoder of the fault prediction model to obtain text features.

[0043] In this context, each parameter in the parameter fragment is often in the form of text.

[0044] A text encoder is used to extract text features from parameter fragments, converting the parameter sequence in the parameter fragments into parameter feature vectors as text features. The specific parameters included in the parameter sequence and their order depend on the parameter organization method during model training, and this disclosure does not limit this aspect.

[0045] S106, input the initial visual features and text features into the large language model of the fault prediction model to construct the key-value pair information of the final decoding layer of the large language model.

[0046] Large Language Models (LLMs) are deep learning models trained on large amounts of text data that can generate natural language text or understand the meaning of language text. LLMs can handle a variety of natural language tasks, such as text classification, question answering, and dialogue.

[0047] In practice, initial visual features and text features can be input into the large language model. The large language model consists of multiple decoding layers connected in series. The output of the i-th decoding layer is the input data of the (i+1)-th decoding layer. The attention module of each decoding layer will calculate and generate corresponding key-value pairs based on these input features. As the number of layers increases, the large language model's understanding of the input data gradually deepens, and the semantic information becomes richer.

[0048] The key-value pair information in the final decoding layer contains the deepest semantic representation after processing through multiple decoding layers, which can better capture the core features and contextual information of the input data.

[0049] The large language model uses the key-value pair information from the last decoding layer to predict and analyze spinning fault prediction results.

[0050] Key-value pair information includes the following: Key Cache: An "index" that encodes input features for attention matching.

[0051] Value Cache: Encodes the "content" of input features for attention aggregation.

[0052] Video value caching ( ): The cached portion of the initial visual features in the last decoding layer. Deep semantic representations that store initial visual features can influence the attention weights of large language models for video clips.

[0053] S107, optimize the visual feature representation in key-value pair information based on the initial parameter values ​​of the controller.

[0054] In this embodiment of the disclosure, the controller is a trainable low-rank parameter matrix, which is the core control unit used to optimize the value cache corresponding to visual features in the decoding layer of a large model. Its dimension matches the value cache corresponding to the visual features (the highest-order semantic representation of the visual features after processing by the full-layer network), so that the visual feature representation can be iteratively optimized through the parameters of the controller to improve the accuracy of the final prediction result.

[0055] S108, Based on the optimized visual feature representation, determine the spinning fault prediction result for the current time window.

[0056] After obtaining the optimized visual feature representation, the fault prediction model performs a prediction process based on this optimized visual feature representation to obtain the corresponding spinning fault prediction result. The spinning fault prediction result may include the fault type and the confidence level corresponding to the fault type.

[0057] In this embodiment, the video stream provides intuitive information about the appearance of the target component, while the time sequence reflects the dynamic changes in the key process parameters of the target component. This fusion of multimodal data provides richer feature inputs for the fault prediction model. Simultaneously, data acquisition and processing within each time window enables real-time monitoring of equipment status and dynamic adjustment of controller parameters based on current data, enhancing the adaptability of the fault prediction model to different fault modes. Furthermore, by dynamically calculating the initial values ​​of the controller parameters through the parameter initialization model and optimizing the visual feature representation, the sensitivity of the fault prediction model to fault features can be further enhanced. The optimized visual feature representation can accurately reflect the abnormal state of the target component, reducing misjudgments caused by inaccurate features. At the same time, fault prediction based on the optimized visual feature representation can more accurately identify fault types, thereby improving the accuracy of fault prediction.

[0058] This invention leverages the powerful learning capabilities and prior knowledge of large language models, combined with parameters collected from vision and multiple sensors, to comprehensively analyze and detect the state of target components. This overcomes the limitations of single sensors in fully reflecting the operational status of target components, effectively reducing quality issues such as yarn breakage and drift caused by abnormalities in target components, thereby ensuring smooth and stable production. Furthermore, the detection results of target components can replace traditional periodic replacement strategies, reducing unnecessary component replacement frequency, lowering resource waste and production costs, and ultimately achieving intelligent and refined control of the spinning process.

[0059] In some embodiments, the visual feature representation in key-value pair information is optimized based on the initial values ​​of the controller parameters, such as... Figure 2 As shown, taking the last decoding layer as the target decoding layer, the following operations are performed iteratively: S201, For the current step, obtain the input information of the target decoding layer. The input information includes the current output of the large language model and the current key-value pair information of the target decoding layer.

[0060] Each round of optimization is considered as one step, and the current step refers to the current round of optimization.

[0061] The current output of the large language model refers to the current detection result of the large language model for the target component.

[0062] S202 uses a large language model to predict the input information and obtains the confidence level of the intermediate output result of the current step.

[0063] The intermediate output results are the non-final detection results of the output layer of the large language model for the current time window. This disclosure optimizes the detection results of the large language model for the target component by continuously optimizing the visual feature representation through an optimizer.

[0064] In implementation, after receiving input information, the large language model can first use its own attention mechanism and the operations of the decoding and output layers to perform feature extraction and association analysis on the input information, and calculate the logits corresponding to the next label in the current step. Among them, logits are the unnormalized raw prediction scores, which directly reflect the degree of preference of the large language model for candidate results of each category.

[0065] Next, check logits Performing Softmax normalization transforms the raw scores into a probability distribution that conforms to probability axioms. Each element in the probability distribution , which is the confidence level of the intermediate output result of the current step, and its value represents the reliability of the large language model in determining that the current monitoring data belongs to the i-th type (fault or normal state).

[0066] S203, determine the distribution entropy of the current step based on the confidence level.

[0067] Distribution entropy is an index used to quantify the uncertainty of an intermediate output result, calculated based on the probability distribution of the confidence level of the intermediate output result at the current step. The larger the distribution entropy, the more dispersed the probability distribution of the confidence level, and the higher the uncertainty of the intermediate result; the smaller the distribution entropy, the more concentrated the probability distribution of the confidence level, and the stronger the certainty of the intermediate output result.

[0068] During implementation, the distribution entropy of the current step is determined based on the confidence level, which can be described by expression (1): (1) In expression (1), V represents the distribution entropy at time t, which is also the distribution entropy at the current step; V represents the set of all possible values ​​of the confidence score of the intermediate output result at the current step. This indicates the confidence level of the intermediate output result in the current step; The logarithm represents the confidence level of the intermediate output of the current step, used to achieve a reasonable quantification of uncertainty.

[0069] S204, Determine the exponential moving average entropy of the current step based on the distribution entropy and the exponential moving average entropy of the previous step.

[0070] Exponential Moving Average Entropy (EMA) is a statistical indicator that combines the exponential moving average (EMA) with information entropy. By giving more weight to recent data and reducing the impact of historical data fluctuations, this indicator more accurately reflects the changing trend of entropy values.

[0071] During implementation, the exponential moving average entropy of the current step is determined and can be described by expression (2): (2) In expression (2), This represents the exponential moving average entropy of the current step; This represents the exponential moving average entropy of the previous step in the current step. This represents the distribution entropy of the current step; This represents the smoothing coefficient, with a value range of [0,1]. The closer the value is to 1, the higher the weight of the exponential moving average entropy of the previous step in the current step, and the stronger the smoothing effect.

[0072] S205, determine the target loss based on the exponential moving average entropy of the current step.

[0073] The target loss is the core indicator for measuring the degree of deviation between the model's prediction results and the ideal state; the magnitude of the target loss directly guides the optimization direction of the controller parameters.

[0074] During implementation, the target loss can be described by expression (3): (3) In expression (3), Indicates target loss; This represents the weighting coefficient, which takes a value of 1 or -1 and is used to guide the optimization direction. This represents the exponential moving average entropy of the current step. The controller parameters are optimized by minimizing the target loss, thereby improving the visual feature representation.

[0075] S206, Determine the gradient of the target loss with respect to the controller parameters.

[0076] The gradient is the partial derivative of the target loss function with respect to the controller parameters, reflecting the rate and direction of change of the target loss with respect to the controller parameters; the sign and magnitude of the gradient determine the direction of the controller parameter update.

[0077] In practice, the gradient of the target loss with respect to the controller parameters is determined, and can be described by expression (4): (4) In expression (4), Denote the gradient of the loss function with respect to the controller parameters ; Denote the target loss function of the controller; Denote the distribution entropy at the current step; Denote the weight coefficient for the k-th round of iteration; When , the gradient direction is the reverse of the distribution entropy decreasing direction, driving the controller parameters to be adjusted in the direction of increasing distribution entropy; When , the gradient direction is the same as the distribution entropy decreasing direction, driving the controller parameters to be adjusted in the direction of decreasing distribution entropy.

[0078] S207. Update the current parameters of the controller based on the gradient to obtain the updated result of the controller parameters.

[0079] That is, according to the gradient calculated in step S206, the controller parameters are updated by using an optimization algorithm (such as stochastic gradient descent, etc.) to obtain the updated result of the controller parameters.

[0080] In implementation, the controller parameters can also be adjusted by an optimizer (such as AdamW), and at the same time, the number of parameters is compressed by means of a low-rank matrix to reduce the computational complexity and accelerate feature extraction.

[0081] In implementation, when updating the controller parameters to obtain the updated result, it can be described by expression (5): (5) In expression (5), Denote the updated controller parameters; Denote the current controller parameters; Denote the learning rate, which controls the optimization step size; Denote the gradient of the loss function with respect to the controller parameters ;

[0082] Among them, compressing the number of parameters of the controller by means of a low-rank matrix can be described by expression (6): (6) In expression (6), Denote the original adjustable parameters of the controller, and are the trainable weight matrices inside the controller, ; among them, r << d, such that the number of parameters is significantly compressed, while reducing the computational amount and avoiding model overfitting; Denote the visual features; Denote the scaling factor.

[0083] S208. Update the visual feature representation in the current key-value pair information based on the updated result of the controller parameters.

[0084] That is, the visual feature representation in the current key-value pair information is updated using the updated controller parameters.

[0085] In practice, the visual feature representation in the current key-value pair information is updated based on the controller parameter update result, which can be expressed by expression (7): (7) In expression (7), This represents the visual feature representation of the updated key-value pair information; This represents the original visual feature representation, i.e., the visual features stored in key-value pairs before the update; This represents the visual features adjusted based on control parameters; The norm of the adjusted visual feature representation; The norm of the original visual feature representation is used to keep the feature energy of the updated visual features consistent with that of the original visual features, reducing the impact of feature scale fluctuations on the stability of subsequent key-value pair matching or fusion calculations; where feature energy is a quantitative description of the information carrying strength and numerical scale of the visual feature vector.

[0086] S209, if it is determined that the termination condition is not met, construct the input information for the next step of the current step based on the updated visual feature representation, take the next step as the current step, and return to execute the operation of obtaining the input information of the target decoding layer for the current step.

[0087] That is, after completing an update, check whether the termination condition is met. The termination condition may be that a preset number of iterations has been reached, or the target loss is less than a certain threshold. If the termination condition is not met, construct the input information for the next step based on the updated visual feature representation, take the next step as the current step, and then return to execute the operation of obtaining input information to continue iterative optimization.

[0088] In this embodiment of the disclosure, the distribution entropy is calculated by confidence level and the target loss is constructed by combining it with the exponential moving average entropy. Then, the controller parameters are optimized based on the loss gradient, and the visual feature representation in the key-value pairs is dynamically adjusted. This can continuously reduce the uncertainty of the model's prediction of the target component's state, and ultimately improve the accuracy and stability of the spinning fault prediction results.

[0089] In the spinning process, there are many possible forms of spinning faults. In this embodiment of the disclosure, the spinning fault prediction results include at least one of the following: spinneret blockage fault prediction value and network node anomaly prediction value. The spinneret blockage fault prediction value indicates whether the spinneret is blocked, the type of blockage, and the degree of blockage; the network node anomaly prediction value indicates whether a network node is abnormal and the degree of abnormality.

[0090] The spinneret blockage fault prediction value is based on a first sub-video stream in the video stream and at least one of the following preset monitoring parameters: melt pressure at the spinneret assembly inlet, rotational speed of the oil metering pump, and spinning temperature. The first sub-video stream is a video stream obtained by acquiring images of the spinneret. Key parameters are explained below: Melt pressure, specifically the melt pressure (MPa) at the inlet of the spinning assembly (i.e., the spinneret assembly), is used to monitor the filter sand clogging status. In practice, melt pressure can be monitored in real-time using pressure sensors. Spinneret clogging obstructs melt flow, leading to increased melt pressure. By monitoring changes in melt pressure, signs of spinneret clogging can be detected promptly. For example, a melt pressure exceeding the normal melt pressure threshold may indicate a clogging problem in the spinneret.

[0091] The rotational speed of the oil metering pump controls the oil content of the yarn. In implementation, this can be achieved by reading the speed of the metering pump's servo motor using a PLC (Programmable Logic Controller) system. The pump's speed affects the yarn's oil content. Abnormal speeds (e.g., exceeding a threshold speed) can lead to uneven oil distribution on the yarn surface, affecting the spinneret's lubrication and increasing the risk of clogging. Monitoring the speed ensures a normal oil supply and reduces the likelihood of clogging.

[0092] Spinning temperature refers to the setpoint heating temperature of the melt distribution tube in the spinning box. In practice, the temperature of the melt distribution tube can be collected using thermocouples as the spinning temperature. Spinning temperature directly affects the viscosity and flowability of the melt. Excessively high temperatures (e.g., above the first temperature threshold) or excessively low temperatures (e.g., below the second temperature threshold) can lead to abnormal melt flow, increasing the risk of spinneret blockage. By monitoring the spinning temperature, the melt can flow under optimal conditions, reducing the likelihood of blockage.

[0093] The spinneret clogging failure prediction value is a prediction of whether the spinneret will experience clogging failure, including but not limited to the type of clogging (such as hole edge defects, hole internal deposits, etc.) and the severity (such as mild, moderate, severe).

[0094] The first sub-video stream refers to the video stream obtained by acquiring images of the spinneret, which is used to visually observe the appearance of the spinneret, such as whether the orifices are blocked or whether there is melt deposition.

[0095] In this embodiment, by combining the spinneret video stream (first sub-video stream) and first sub-parameters such as melt pressure, oil metering pump speed, and spinning temperature, the fault prediction model can comprehensively evaluate the spinneret status from both visual and process parameter dimensions. The video stream can visually display the spinneret blockage, while an abnormal increase in melt pressure can further confirm the presence of blockage. The fusion of multimodal data can significantly improve the accuracy of fault detection.

[0096] The network node anomaly prediction value is based on a second sub-video stream in the video stream and at least one of the following preset monitoring parameters: Network air pressure, used to indicate the set value of compressed air pressure for the network nozzle; Network degree is used to represent the number of network nodes per meter of wire. The target guide roller linear speed is the linear speed of the guide roller used to determine the winding speed. The second sub-video stream is a video stream obtained by capturing images of the filament region directly below the nozzle outlet of the network.

[0097] Network air pressure refers to the setpoint of compressed air pressure for the network nozzles. During implementation, the compressed air pressure can be monitored in real time using a pressure sensor. Network air pressure can affect node formation. Insufficient air pressure (e.g., below the air pressure threshold) may result in an insufficient number of nodes, while air pressure fluctuations (e.g., fluctuations exceeding the normal air pressure range) may lead to uneven node distribution. Monitoring network air pressure ensures proper node formation and reduces the risk of network node anomalies.

[0098] Network density, or the number of network nodes per meter of yarn, ensures the yarn maintains good stability during winding, preventing it from becoming loose or tangled. In practice, a photoelectric counter can be used to count the number of network nodes. Insufficient network density (e.g., below the first network density threshold) may cause the yarn to become loose during subsequent processing, while excessive network density (e.g., above the second network density threshold) may cause the yarn to become too stiff. By monitoring the network density, the number of nodes can be kept within a reasonable range to ensure normal yarn production.

[0099] The target guide roller linear speed is the linear speed of the guide roller used to determine the winding speed; an exemplary target guide roller can be GR2 (the second guide roller), whose linear speed directly determines the winding speed. In implementation, the linear speed of GR2 can be obtained based on a linear speed sensor.

[0100] Linear velocity sensors are used to measure and record parameters of mechanical motion. In the spinning process, linear velocity sensors can be mounted on the drive shaft of the guide roller to measure the roller's rotational speed. Through the linear velocity sensor, mechanical motion is converted into an electrical signal, which is then used to calculate the linear velocity.

[0101] The linear speed of the target guide roller can affect the tension and cooling effect of the filament. Excessive linear speed (e.g., exceeding the first linear speed threshold) may lead to insufficient filament cooling and insufficient knot strength; excessive linear speed (e.g., below the second linear speed threshold) may lead to excessive filament tension and uneven knot distribution. By monitoring the linear speed of the guide roller, the filament can be wound in an optimal state, reducing the risk of network knot abnormalities.

[0102] The second sub-video stream refers to the video stream obtained by capturing images of the filament bundle area directly below the network nozzle outlet. It is used to visually observe the node status of the filament bundle, such as the number of nodes, distribution uniformity, and intensity.

[0103] In this embodiment of the disclosure, by combining the video stream (second sub-video stream) of the filament bundle area directly below the network nozzle outlet with second sub-parameters such as network air pressure, network degree, and target guide roller linear speed, the fault prediction model can more comprehensively assess the state of the network nodes. For example, the video stream can visually display the distribution and intensity of the nodes, while abnormal fluctuations in network air pressure can further confirm the cause of the node anomalies. The fusion of multimodal data improves the accuracy of fault detection.

[0104] In some embodiments, to avoid parameter collapse (initialization) caused by zero initialization of controller parameters... To achieve gradient vanishing at all zeros and maintain the interpretability of the low-rank structure (alignment of channel physical meaning with weight distribution), the controller parameters can be initialized dynamically. Video clips and parameter clips are input into the parameter initialization model to obtain the initial parameter values ​​for the controller of the fault prediction model, such as... Figure 3 As shown, it can be implemented as follows: S301, from the first images of multiple frames belonging to the first sub-video stream of the video segment, obtain the first target frame image that meets the preset quality requirements.

[0105] To ensure sufficiently high image quality and improve detection accuracy, during implementation, image frames that meet preset quality requirements can be selected from the first sub-video stream, starting from the first frame of the current time window.

[0106] The preset quality requirements may include at least one of the following indicators: Sharpness refers to the level of detail in an image or video, reflecting its sharpness, edge definition, and texture richness. High sharpness (e.g., exceeding a sharpness threshold to meet preset quality requirements) means that details such as object outlines and lines in an image or video can be displayed more clearly, making it easier to distinguish various elements in the image or video.

[0107] Peak signal-to-noise ratio (PSNR) reflects the relative magnitude of the maximum possible power in an image or signal to the power of interference noise affecting observation, and is usually expressed in decibels (dB). A high PSNR value (such as one above the PSNR threshold, which meets the preset quality requirements) generally indicates good image quality and minimal damage to the image during reconstruction or processing.

[0108] The Structural Similarity Index (SSIM) is a metric used to measure the similarity between two images, taking into account brightness, contrast, and structural information. A high SSIM value (e.g., above the SSIM threshold, indicating that a preset quality requirement is met) indicates that the image's structural information has been well preserved.

[0109] Mean Squared Error (MSE) is a metric used to measure the difference between predicted and true values. In image processing, a low MSE value (below the MSE threshold indicates that the preset quality requirements are met) indicates that the error in the image is small, which helps improve the accuracy of subsequent fault prediction models.

[0110] Mean Absolute Error (MAE) is the average of the absolute values ​​of the differences between predicted and true values. In image processing, a low MAE value (below the MAE threshold, which meets the preset quality requirements) indicates that the error in the image is small, which helps improve the accuracy of subsequent fault prediction models.

[0111] The Feature Similarity Index (FSIM) is a similarity metric based on image features, taking into account features such as phase congruency and gradient magnitude. A high FSIM value (above the FSIM threshold, indicating that the preset quality requirements are met) indicates that the visual features of the image have been well preserved, which helps subsequent fault prediction models to better identify the texture and structural features of target components.

[0112] Visual Information Fidelity (VIF) is a quality assessment metric based on visual information theory, which considers the information fidelity of an image across different visual subbands. A higher VIF value (e.g., above a VIF threshold to meet preset quality requirements) indicates better preservation of the image's visual information.

[0113] Each of the aforementioned indicators can be set with a corresponding threshold based on the actual situation, and this disclosure does not impose any restrictions on this.

[0114] S302, convert the first target frame image into a grayscale image to obtain the first grayscale image.

[0115] Converting the first target frame image to grayscale reduces data dimensionality while preserving key image information. Grayscale images help improve data processing efficiency.

[0116] S303, based on the parameter initialization model, extracts the initial channel features of multiple first channels of the controller from the first grayscale image.

[0117] The initial characteristics of the first channel are used to visually describe the state of the spinneret.

[0118] In some embodiments, extracting the initial channel features of multiple first channels of the controller from the first grayscale image based on the parameter initialization model can be implemented as follows: Step A1: Perform an opening operation on the first grayscale image to obtain the first reference image.

[0119] Among them, for the first grayscale image Perform an opening operation to obtain the first reference image ( As shown in equation (8): (8) Where B represents an n×n circular structuring element, and n is a positive integer. This indicates the opening operation. This indicates the closing operation.

[0120] Step A2: Perform a closing operation on the first grayscale image to obtain the second reference image.

[0121] Among them, for the first grayscale image Perform a closing operation to obtain the second reference image. As shown in equation (9): (9) The parameters involved in equation (9) have been explained above and will not be repeated here.

[0122] Opening and closing operations are used to extract morphological features from the edges and interiors of spinneret orifices. Opening operations highlight defects at the orifice edges, while closing operations highlight deposits inside the orifices. Combining these two operations allows for comprehensive capture of changes in the spinneret's state.

[0123] Step A3: Determine the pixel value difference at the same position between the first reference image and the first grayscale image to obtain the first difference feature.

[0124] Among them, the first difference feature As shown in equation (10): (10) in, Indicates the first reference figure. This represents a grayscale image.

[0125] The first distinguishing feature is used to highlight defects at the edges of the spinneret's holes.

[0126] Step A4: Determine the pixel value differences at the same locations between the first grayscale image and the second reference image to obtain the second difference feature.

[0127] Among them, the second difference feature As shown in equation (11): (11) in, Indicates the second reference figure. This represents a grayscale image.

[0128] The second distinguishing feature is used to highlight the deposition inside the orifices of the spinneret.

[0129] Step A5: Perform adaptive grayscale threshold segmentation on the difference map between the first reference map and the second reference map to obtain the maximum value of the inter-class variance.

[0130] Among them, the maximum value of the inter-class variance As shown in equation (12): (12) Here, Otsu represents adaptive thresholding to calculate the maximum inter-class variance.

[0131] The maximum inter-class variance can reflect the global clogging trend of the spinneret.

[0132] Step A6: Input the first difference feature, the second difference feature, and the maximum inter-class variance into the corresponding first multilayer perceptron network to obtain the initial channel features of multiple first channels; wherein, the first difference feature, the second difference feature, and the maximum inter-class variance each correspond to one first channel.

[0133] In implementation, each of the three features—the first difference feature, the second difference feature, and the maximum inter-class variance—corresponds to a first channel. For example, the first difference feature corresponds to first channel 1, which reflects the defects at the edge of the spinneret's orifice; the second difference feature corresponds to first channel 2, which reflects the deposition inside the spinneret's orifice; and the maximum inter-class variance corresponds to first channel 3, which reflects the global clogging trend of the spinneret.

[0134] Each feature corresponds to its own first multilayer perceptron network.

[0135] When implementing, such as Figure 4 As shown, a first reference image and a second reference image are obtained based on the first grayscale image; a first difference feature is obtained based on the first reference image and the first grayscale image; a second difference feature is obtained based on the first grayscale image and the second reference image; the maximum inter-class variance is obtained based on the first reference image and the second reference image; the first difference feature is processed by the corresponding first multilayer perceptron network 1 to obtain the initial channel feature of the first channel 1 corresponding to the first difference feature. The second difference feature is processed by the corresponding first multilayer perceptron network 2 to obtain the initial channel feature of the first channel 2 corresponding to the second difference feature. The maximum inter-class variance is processed by the corresponding first multilayer perceptron network 3 to obtain the initial channel features of the first channel 3 corresponding to the maximum inter-class variance. .

[0136] In this embodiment, by calculating the differences between the opening and closing operation results and the original grayscale image, abnormal changes at the edges and inside the spinneret orifices can be highlighted. The first and second difference features can intuitively reflect the clogging status of the spinneret and are key information for fault prediction. Using the maximum inter-class variance, the optimal threshold can be automatically determined to highlight the global clogging trend. By performing nonlinear mapping on the aforementioned features through a first multilayer perceptron network, deeper information about the features can be extracted, generating richer initial channel features. These features provide more comprehensive information support for subsequent fault prediction.

[0137] S304, based on the parameter initialization model, at least one first sub-parameter is processed to obtain the respective channel weights of multiple first channels.

[0138] The first sub-parameter is at least one of the aforementioned melt pressure at the spinneret assembly inlet, the rotational speed of the oil metering pump, and the spinning temperature. The weight of each first channel is affected by at least one first sub-parameter.

[0139] In some embodiments, processing at least one first sub-parameter based on the parameter initialization model to obtain the respective channel weights of multiple first channels can be implemented as follows: Step B1, execute for each first channel separately: Step B11: Obtain the aperture area of ​​the first channel for T time windows to obtain the aperture area sequence.

[0140] The orifice area sequence is a sequence of spinneret orifice area data over multiple time windows.

[0141] For each first channel, calculate the aperture area of ​​that channel within T time windows to obtain the aperture area sequence. .

[0142] Step B12: Determine the standard deviation and mean of the pore area sequence.

[0143] Wherein, the standard deviation of the T time windows As shown in equation (13): (13) Among them, the mean of T time windows As shown in equation (14): (14) Step B13: Determine the fluctuation intensity of the first channel based on the standard deviation and mean.

[0144] Among them, fluctuation intensity refers to an index calculated from the standard deviation and mean of the orifice area sequence, used to quantify the degree of fluctuation in orifice area. Fluctuation intensity can reflect the drastic change in spinneret orifice diameter and is an important basis for judging the clogging status.

[0145] Among them, the fluctuation intensity of the first channel As shown in equation (15): (15) in, It can be 'm', used to prevent division by zero; an example could be... .

[0146] Step B2: Take the fluctuation intensity of each first channel and at least one first sub-parameter as a first set, and for each first target parameter in the first set, obtain the first mapping feature corresponding to each first target parameter based on the neural network module.

[0147] The number of channels corresponding to each first target parameter is as follows: If the first target parameter includes the melt pressure at the inlet of the spinneret assembly, it may affect the first channel 1 and the first channel 3.

[0148] Excessive melt pressure (e.g., exceeding the normal melt pressure threshold) may cause excessive melt flow rate, resulting in damage to the orifice edges. Specifically, when the melt pressure exceeds a specified range above the normal melt pressure threshold, the high-speed flowing melt causes excessive shearing and abrasion to the spinneret orifice edges, leading to orifice edge breakage or deformation. This is a typical characteristic of sudden blockage, which will affect the first channel 1.

[0149] Furthermore, periodic fluctuations or continuous increases in melt pressure indicate increased flow resistance in the entire spinneret assembly, which may be caused by simultaneous blockage of multiple pores or filter sand, indicating a global blockage risk, which will also affect the first channel 3.

[0150] If the first target parameter includes the rotational speed of the oil metering pump, it may affect the first channel 2 and the first channel 3.

[0151] Specifically, if the speed of the metering pump is too high (such as exceeding the first speed threshold), it may cause uneven melt flow and easily form deposits in the orifice. Specifically, when the metering pump speed is too high, the melt distribution is uneven, some spinnerets have an excessive melt supply, the melt stays in the orifice for a longer time, and deposits are formed after cooling, which also affects the first channel 2.

[0152] In addition, when the speed of the oil metering pump and the melt pressure deviate from the set value at the same time (such as the speed being lower than the second speed threshold, but the melt pressure being higher than the normal melt pressure threshold), it indicates that the blockage has affected the overall melt delivery system and a global risk warning needs to be triggered, which means that it will affect the first channel 3.

[0153] When the first target parameters include the spinning speed and spinning temperature, it may affect the first channel 1, the first channel 2 and the first channel 3.

[0154] Specifically, if the spinning temperature is too low (e.g., below the second temperature threshold), the melt viscosity increases significantly, the flow resistance increases, and the shear force on the edge of the hole is enhanced, leading to defects, which affects the first channel 1; at the same time, the high-viscosity melt is more difficult to completely extrude and is more likely to remain and form deposits, which affects the first channel 2.

[0155] If the spinning temperature deviates from the set normal temperature range (such as being greater than the first temperature threshold or less than the second temperature threshold), it may affect the overall risk of blockage, that is, affect the first channel 3.

[0156] In implementation, each first target parameter is processed by a neural network module to obtain a first mapping feature. For example, for the j-th first target parameter... The first mapping feature is obtained through processing. As shown in equation (16): (16) in, Indicates the first target parameter The number of channels allocated is as follows: if the spinning temperature affects three first channels, then when the first target parameter is the spinning temperature, the corresponding number of first channels is 3. This represents the weight matrix of the linear mapping layer for the m-th channel of the neural network module. This represents the bias term of the linear mapping layer. This represents the activation function.

[0157] The first mapping feature represents the mapping relationship between the first target parameter and its corresponding first channel.

[0158] Step B3: Input each first mapping feature into the global attention mechanism network to obtain the weight contribution value of each first mapping feature to each first channel.

[0159] During implementation, each first mapping feature Input the corresponding global attention mechanism network and calculate the weight contribution value of the i-th first mapping feature to the j-th channel. As shown in equation (17): (17) in, Let d represent the learnable gating vector of the first channel j, and d represent the latent dimension. This represents the first mapping feature.

[0160] In practice, each first mapping feature is calculated based on equation (17) to obtain the weight contribution value of each first mapping feature to each first channel.

[0161] Step B4, execute for each first channel separately: Step B41: Select at least one first valid value from the weight contribution values ​​of each first mapping feature to the first channel based on the first indicator function; the first indicator function defines the parameters in the first set that have an impact on each first channel.

[0162] The set u of at least one first valid value is shown in equation (18): (18) in, ( ) represents the first indicator function, which is the channel set when the first channel j is affected by the i-th first target parameter. Returns 1 otherwise, returns 0. For example, if the melt pressure affects channel 1 and channel 3, then... .

[0163] Step B42: Determine the sum of at least one first valid value to obtain the channel weight of the first channel.

[0164] In practice, the sum of at least one first effective value will be calculated, and the sum will be normalized based on the softmax function to obtain the channel weight of the j-th first channel. As shown in equation (19): (19) The parameters involved in equation (19) have been described above, and will not be repeated here in the embodiments of this disclosure.

[0165] In this embodiment, the fluctuation of the orifice area can be quantified by calculating the mean and standard deviation of the orifice area sequence. The fluctuation intensity reflects the degree of change in the orifice area. Combining the fluctuation intensity with the first sub-parameter into a first set can reflect the state of the spinneret from multiple aspects. These parameters complement each other and can provide richer information to capture changes in the state of the spinneret. The global attention mechanism can dynamically adjust the parameter weights, enabling the fault prediction model to adapt to different process conditions and fault modes, thereby enhancing the adaptability of the fault prediction model. In summary, multi-parameter fusion and dynamic weight allocation can more accurately identify faults and significantly improve the accuracy of the fault prediction model.

[0166] S305, Based on the respective channel weights of multiple first channels, the initial channel features of multiple first channels are weighted and summed to obtain the parameter values ​​of multiple first channels; In one embodiment, obtaining parameter values ​​for multiple first channels can be achieved by: sequentially performing orthogonalization and fully connected layer processing on the initial features of each first channel to obtain the first weighted features of each first channel; and performing weighted summation on the first weighted features of each first channel based on their respective channel weights to obtain parameter values ​​for multiple first channels.

[0167] (20) in, This represents the first feature to be weighted in the first channel j. Let j be the initial channel feature of the first channel. It is a fully connected layer. Indicates orthogonalization. This represents the channel weight of the first channel j, and the parameter value of the first channel j is obtained. .

[0168] It should be noted that each first channel has corresponding parameter values.

[0169] In this embodiment of the disclosure, orthogonalization is used to make the feature vectors of different channels independent of each other, thereby reducing redundant information between features. Based on the fully connected layer processing, high-level information in the features can be further extracted. Based on the channel weights, the features are weighted and summed, so that the fault prediction model can dynamically adjust its contribution to the final output according to the importance of each channel, thereby improving the accuracy of the fault prediction model.

[0170] S306: Obtain a second target frame image that meets the preset quality requirements from the multiple frames of second images belonging to the second sub-video stream of the video segment.

[0171] To ensure sufficient image quality for subsequent processing, in practice, starting from the first frame of the current time window, an image meeting preset quality requirements can be selected from the second sub-video stream as the second target frame image. The preset quality requirements have been described above and will not be repeated here. The parameters included in the preset quality requirements of the second sub-video stream can be the same or different, and the thresholds for the same parameters can also be the same or different.

[0172] S307, based on the parameter initialization model, extracts the initial channel features of multiple second channels of the controller from the second target frame image.

[0173] The initial characteristics of the second channel indicate that there are quality defects at the nodes formed by the filament bundle at the network nozzle.

[0174] In some embodiments, the initial channel features of multiple second channels of the controller are extracted from the second target frame image based on the parameter initialization model, which can be specifically implemented as follows: Step C1: Extract the node density of the silk thread from the second target frame image.

[0175] During implementation, a binary mask is used to process the second target frame image, simplifying the complex thread image into a 0 / 1 representation, where 0 represents non-node pixels and 1 represents node pixels. Counting the number of 1s directly reflects the total number of nodes, thus obtaining the total number of nodes in the second target frame image. Combined with the image calibration coefficient (pixel to meter conversion), the actual detected node density is calculated. Node density features are calculated based on the actual detected node density. As shown in equation (21): (twenty one) in, This indicates the total number of nodes in the second target frame image. This indicates the length (in pixels) of the second target frame image. Represents the image calibration coefficients. Indicates standard density.

[0176] When the actual detection node density is lower than the standard density If the value is greater than 0, it may trigger a defect warning indicating insufficient node quantity. The greater the deviation in node density, the more likely it is to cause a problem. The higher the value, the better. If the actual detected node density is not lower than the standard density, it indicates that the number of nodes is normal.

[0177] Step C2: Determine the number of nodes in each image block in the second target frame image to obtain the number of local nodes in each image block.

[0178] During implementation, the second target frame image is divided into an m×n grid, with each grid block being an image block. The number of nodes in each grid block is counted to obtain the local node count. .

[0179] Step C3: Determine the node variation coefficient of the second target frame image based on the mean and standard deviation of the number of local nodes in each image block.

[0180] In practice, the nodal variation coefficient F2 is as shown in equation (22): (twenty two) in, Let represent the mean of the c-th image patch. Let represent the standard deviation of the c-th image patch, m and n represent the number of grids divided by each edge, k is the row index, and l is the column index. This represents the number of nodes within the image block in the k-th row and l-th column.

[0181] The coefficient of variation among nodes reflects whether the distribution of nodes is uniform. If the value exceeds the node variation threshold, it indicates uneven node distribution; If the value is not greater than the node variation threshold, it indicates that the node distribution is uniform.

[0182] Step C4: Extract the node mask image of the second target frame image.

[0183] During implementation, a binary mask can be used for processing to obtain a node mask image. .

[0184] Step C5: Determine the cumulative value of the pixel-by-pixel product of the grayscale image and the node mask image of the second target frame image to obtain the first value.

[0185] During implementation, the first value As shown in equation (23): (twenty three) in, The coordinates in the second target frame image are... grayscale value, The coordinates in the second target frame image are... The node mask value.

[0186] Step C6: Determine the sum of the accumulated values ​​of each pixel in the node mask image and the sum of the preset coefficients to obtain the second value.

[0187] In practice, the second value As shown in equation (24): (twenty four) in, The coordinates in the second target frame image are... Node mask diagram, This represents a preset coefficient used to prevent division by zero in subsequent operations.

[0188] Step C7: Determine the ratio of the first value to the second value to obtain the intensity candidate value.

[0189] During implementation, intensity candidate values As shown in equation (25): (25) Step C8: Select the maximum value between the intensity candidate frame and the default value to obtain the node intensity of the second target frame image.

[0190] During implementation, the node strength of the second target frame image As shown in equation (26): (26) in, The value of decay after the Otsu algorithm automatically calculates the global threshold can be determined experimentally.

[0191] Step C9: Input the node density, node variation coefficient, and node strength into the corresponding second multilayer perceptron network to obtain the initial channel features of multiple second channels; wherein, the node density, node variation coefficient, and node strength each correspond to a second channel.

[0192] In practice, each feature corresponds to a second channel. For example, node density corresponds to second channel 1, which reflects the situation of insufficient number of nodes; node variation coefficient corresponds to second channel 2, which reflects the situation of uneven node distribution; and node strength corresponds to second channel 3, which reflects the situation of insufficient node strength.

[0193] Each feature corresponds to a specific second-layer perceptron network.

[0194] S308, based on the parameter initialization model, processes at least one second sub-parameter to determine the respective channel weights of multiple second channels.

[0195] The second sub-parameter is at least one of the aforementioned network air pressure, network density, and target guide roller linear speed. The weight of each second channel is affected by at least one second sub-parameter.

[0196] In some embodiments, processing at least one second sub-parameter based on the parameter initialization model to determine the respective channel weights of multiple second channels can be implemented as follows: Step D1: For each second target parameter of at least one second sub-parameter, obtain the second mapping feature corresponding to each second target parameter based on the neural network module.

[0197] In practice, the second mapping feature is obtained in a similar manner to the first mapping feature obtained above, as in equation (16).

[0198] Step D2: Input multiple second mapping features into the global attention mechanism network to obtain the weight contribution value of each second mapping feature to each second channel.

[0199] In practice, based on a similar approach to the aforementioned equation (17), the weight contribution value of each second mapping feature to each second channel can be obtained.

[0200] Step D3, execute for each second channel separately: Step D31: Select at least one second valid value from the weight contribution values ​​of each second mapping feature to the second channel based on the second indicator function; the second indicator function defines the parameters in the second set that affect each second channel; The second set includes at least one second sub-parameter.

[0201] In practice, at least one second valid value can be obtained based on a similar approach to the aforementioned equation (18).

[0202] Step D32: Determine the sum of at least one second effective value to obtain the channel weight of the second channel.

[0203] In practice, the channel weights of each second channel can be obtained in a manner similar to the aforementioned equation (19).

[0204] In this embodiment of the disclosure, by comprehensively considering multimodal data and dynamic weight allocation, the fault prediction model can accurately predict network node anomalies, thereby reducing false alarms and missed alarms. In addition, through the global attention mechanism and indicator function, the fault prediction model only needs to focus on the features most important to the current task, thereby reducing unnecessary computation and further improving computational efficiency.

[0205] S309, The initial features of multiple second channels are weighted based on their respective channel weights to obtain the parameter values ​​of multiple second channels; The initial parameter values ​​of the controller include multiple parameter values ​​for the first channel and multiple parameter values ​​for the second channel.

[0206] In one embodiment, obtaining parameter values ​​for multiple second channels can be achieved by: sequentially performing orthogonalization and fully connected layer processing on the initial features of each second channel to obtain the second weighted features of each second channel; and performing weighted summation on the second weighted features of each second channel based on the respective channel weights of the multiple second channels to obtain parameter values ​​for the multiple second channels.

[0207] Specifically, the parameter values ​​for each second channel can be obtained based on equation (20). Similar to the first channel, each second channel has corresponding parameter values.

[0208] In this embodiment of the disclosure, the orthogonalization and fully connected layer processing enable the extraction of high-level information from features while achieving the independence of feature vectors between different channels. The features are weighted and summed based on channel weights to improve the accuracy of the fault prediction model.

[0209] In this embodiment, by extracting high-quality target frame images from the video stream and combining them with preset monitoring parameters, the subsequent large language model can more comprehensively capture fault features, thereby improving the accuracy of fault prediction. Furthermore, by using grayscale images to simplify image features while retaining key information, and by calculating channel weights and performing weighted summation, the importance of each channel feature can be dynamically adjusted to initialize the controller's parameter initial values, laying a strong foundation for subsequent fault prediction by the fault prediction model.

[0210] Based on the same technical concept, this disclosure also proposes a spinning failure prediction device 500, such as... Figure 5 As shown, it includes: The acquisition module 501 is used to acquire video streams of target components in the spinning process and acquire time sequences of preset monitoring parameters; the target components include spinning assemblies and / or winding mechanisms. The acquisition module 502 is used to acquire a video segment of the current time window from the video stream, and to acquire a parameter segment of the current time window from the time sequence; The first processing module 503 is used to input the video segment and the parameter segment into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model; The visual feature extraction module 504 is used to input the video segment into the visual encoder of the fault prediction model to obtain initial visual features; The text feature extraction module 505 is used to input the parameter fragment into the text encoder of the fault prediction model to obtain text features; The second processing module 506 is used to input the initial visual features and the text features into the large language model of the fault prediction model to construct the key-value pair information of the last decoding layer of the large language model; The optimization module 507 is used to optimize the visual feature representation in the key-value pair information based on the initial parameter values ​​of the controller; The determination module 508 is used to determine the spinning fault prediction result of the current time window based on the optimized visual feature representation.

[0211] In some embodiments, the optimization module is specifically used to take the last decoding layer as the target decoding layer and iteratively perform the following operations: For the current step, obtain the input information of the target decoding layer, which includes the current output result of the large language model and the current key-value pair information of the target decoding layer; The large language model is used to predict the input information to obtain the confidence level of the intermediate output result of the current step; The distribution entropy of the current step is determined based on the confidence level; The exponential moving average entropy of the current step is determined based on the distribution entropy and the exponential moving average entropy of the previous step. The target loss is determined based on the exponential moving average entropy of the current step; Determine the gradient of the target loss with respect to the controller parameters of the controller; The current parameters of the controller are updated based on the gradient to obtain the controller parameter update result; The visual feature representation in the current key-value pair information is updated based on the controller parameter update result; If the termination condition is not met, the input information for the next step of the current step is constructed based on the updated visual feature representation. The next step is then taken as the current step, and the operation of obtaining the input information of the target decoding layer for the current step is returned.

[0212] In some embodiments, the spinning fault prediction results include: spinneret blockage fault prediction value and network node anomaly prediction value; The spinneret blockage fault prediction value is predicted based on the first sub-video stream in the video stream and at least one first sub-parameter among the following preset monitoring parameters: melt pressure at the inlet of the spinneret assembly, rotation speed of the oil metering pump, and spinning temperature. The network node anomaly prediction value is predicted based on a second sub-video stream in the video stream and at least one of the following preset monitoring parameters: Network air pressure, used to indicate the set value of compressed air pressure for the network nozzle; Network degree is used to represent the number of network nodes per meter of wire. The target guide roller linear speed is the linear speed of the guide roller used to determine the winding speed. The first sub-video stream is a video stream obtained by acquiring images of the spinneret; the second sub-video stream is a video stream obtained by acquiring images of the filament bundle area directly below the outlet of the web nozzle.

[0213] In some embodiments, the first processing module includes: The first acquisition unit is used to acquire a first target frame image that meets preset quality requirements from multiple first images belonging to the first sub-video stream of the video segment. A conversion unit is used to convert the first target frame image into a grayscale image to obtain a first grayscale image; The first extraction unit is used to extract the initial channel features of multiple first channels of the controller from the first grayscale image based on the parameter initialization model. The first determining unit is used to initialize the model based on the parameters, process the at least one first sub-parameter, and obtain the respective channel weights of the plurality of first channels; The first calculation unit is used to perform a weighted summation of the initial channel features of the plurality of first channels based on their respective channel weights to obtain the parameter values ​​of the plurality of first channels. The second acquisition unit is used to acquire a second target frame image that meets preset quality requirements from multiple frames of second images belonging to the second sub-video stream of the video segment; The second extraction unit is used to extract the initial channel features of multiple second channels of the controller from the second target frame image based on the parameter initialization model; The second determining unit is used to process the at least one second sub-parameter based on the parameter initialization model to determine the respective channel weights of the plurality of second channels; The second calculation unit is used to weight the initial channel features of the plurality of second channels based on their respective channel weights to obtain the parameter values ​​of the plurality of second channels. The initial parameter values ​​of the controller include the parameter values ​​of the plurality of first channels and the parameter values ​​of the plurality of second channels.

[0214] In some embodiments, the first extraction unit is configured to: An opening operation is performed on the first grayscale image to obtain a first reference image; and, Perform a closing operation on the first grayscale image to obtain a second reference image; The pixel value difference at the same position between the first reference image and the first grayscale image is determined to obtain the first difference feature; The pixel value difference at the same position between the first grayscale image and the second reference image is determined to obtain the second difference feature; An adaptive grayscale threshold segmentation operation is performed on the difference map between the first reference map and the second reference map to obtain the maximum value of the inter-class variance. The first difference feature, the second difference feature, and the maximum inter-class variance are respectively input into the corresponding first multilayer perceptron network to obtain the initial channel features of the plurality of first channels; wherein, the first difference feature, the second difference feature, and the maximum inter-class variance each correspond to a first channel.

[0215] In some embodiments, the first determining unit is configured to: Execute separately for each first channel: Obtain the aperture area of ​​the first channel for T time windows to obtain the aperture area sequence; Determine the standard deviation and mean of the pore area sequence; Based on the standard deviation and mean, the fluctuation intensity of the first channel is determined; The fluctuation intensity of each first channel and the at least one first sub-parameter are taken as a first set. For each first target parameter in the first set, the first mapping feature corresponding to each first target parameter is obtained based on the neural network module. Each first mapping feature is input into the global attention mechanism network to obtain the weight contribution value of each first mapping feature to each first channel; Execute separately for each first channel: At least one first valid value is selected from the weight contribution values ​​of each first mapping feature to the first channel based on the first indicator function; the first indicator function defines the parameters in the first set that affect each first channel; The channel weight of the first channel is obtained by determining the sum of the at least one first effective value.

[0216] In some embodiments, the first computing unit is configured to: The initial features of each first channel are sequentially processed by orthogonalization and fully connected layers to obtain the first weighted features of each first channel; Based on the respective channel weights of the plurality of first channels, the first features to be weighted of each first channel are weighted and summed to obtain the parameter values ​​of the plurality of first channels.

[0217] In some embodiments, the second extraction unit is configured to: Extract the node density of the silk threads from the second target frame image; The number of nodes in each image block in the second target frame image is determined to obtain the number of local nodes in each image block; The node variation coefficient of the second target frame image is determined based on the mean and standard deviation of the number of local nodes in each image block. Extract the node mask image of the second target frame image; The first value is obtained by determining the cumulative value of the pixel-by-pixel product of the grayscale image of the second target frame image and the node mask image; The sum of the accumulated values ​​of each pixel in the node mask image and the sum of the preset coefficients are determined to obtain the second value; Determine the ratio of the first value to the second value to obtain candidate intensity values; The maximum value between the intensity candidate frames and the default value is selected to obtain the node intensity of the second target frame image; The node density, the node variation coefficient, and the node strength are respectively input into the corresponding second multilayer perceptron network to obtain the initial channel features of the plurality of second channels; wherein the node density, the node variation coefficient, and the node strength each correspond to a second channel.

[0218] In some embodiments, the second determining unit is configured to: For each second target parameter of the at least one second sub-parameter, a second mapping feature corresponding to each second target parameter is obtained based on the neural network module; Multiple second mapping features are input into the global attention mechanism network to obtain the weight contribution value of each second mapping feature to each second channel; Perform the following for each second channel: At least one second valid value is selected from the weight contribution values ​​of each second mapping feature to the second channel based on the second indicator function; the second indicator function defines the parameters in the second set that affect each second channel; The channel weight of the second channel is obtained by determining the sum of the at least one second effective value.

[0219] In some embodiments, the second computing unit is configured to: The initial features of each second channel are sequentially processed by orthogonalization and fully connected layers to obtain the second weighted features of each second channel; Based on the respective channel weights of the multiple second channels, the second features to be weighted of each second channel are summed in a weighted manner to obtain the parameter values ​​of the multiple second channels.

[0220] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 6As shown, the electronic device includes a memory 610 and a processor 620. The memory 610 stores a computer program that can run on the processor 620. There can be one or more memories 610 and processors 620. The memory 610 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in the above-described method embodiments. The electronic device may also include a communication interface 630 for communicating with external devices and performing data exchange and transmission.

[0221] If the memory 610, processor 620, and communication interface 630 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0222] Optionally, in a specific implementation, if the memory 610, processor 620, and communication interface 630 are integrated on a single chip, then the memory 610, processor 620, and communication interface 630 can communicate with each other through an internal interface.

[0223] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0224] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).

[0225] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0226] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0227] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0228] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.

Claims

1. A method for predicting spinning defects, comprising: Video streams are acquired from target components in the spinning process, and time sequences of preset monitoring parameters are also acquired. The target component includes a spinning assembly and / or a winding mechanism; Obtain the video segment of the current time window from the video stream, and obtain the parameter segment of the current time window from the time sequence; The video clip and the parameter clip are input into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model; The video clip is input into the visual encoder of the fault prediction model to obtain initial visual features; The parameter fragments are input into the text encoder of the fault prediction model to obtain text features; The initial visual features and the text features are input into the large language model of the fault prediction model to construct the key-value pair information of the last decoding layer of the large language model; The visual feature representation in the key-value pair information is optimized based on the initial parameter values ​​of the controller; Based on the optimized visual feature representation, the spinning fault prediction result for the current time window is determined.

2. The method according to claim 1, wherein, The optimization of the visual feature representation in the key-value pair information based on the initial parameter values ​​of the controller includes: Using the last decoding layer as the target decoding layer, iteratively perform the following operations: For the current step, obtain the input information of the target decoding layer, which includes the current output result of the large language model and the current key-value pair information of the target decoding layer; The large language model is used to predict the input information to obtain the confidence level of the intermediate output result of the current step; The distribution entropy of the current step is determined based on the confidence level; The exponential moving average entropy of the current step is determined based on the distribution entropy and the exponential moving average entropy of the previous step. The target loss is determined based on the exponential moving average entropy of the current step; Determine the gradient of the target loss with respect to the controller parameters of the controller; The current parameters of the controller are updated based on the gradient to obtain the controller parameter update result; The visual feature representation in the current key-value pair information is updated based on the controller parameter update result; If the termination condition is not met, the input information for the next step of the current step is constructed based on the updated visual feature representation. The next step is then taken as the current step, and the operation of obtaining the input information of the target decoding layer for the current step is returned.

3. The method according to claim 2, wherein the spinning failure prediction result includes: spinneret Congestion fault prediction value, network node anomaly prediction value; The spinneret blockage fault prediction value is predicted based on the first sub-video stream in the video stream and at least one first sub-parameter among the following preset monitoring parameters: melt pressure at the inlet of the spinneret assembly, rotation speed of the oil metering pump, and spinning temperature. The network node anomaly prediction value is predicted based on a second sub-video stream in the video stream and at least one of the following preset monitoring parameters: Network air pressure, used to indicate the set value of compressed air pressure for the network nozzle; Network degree is used to represent the number of network nodes per meter of wire. The target guide roller linear speed is the linear speed of the guide roller used to determine the winding speed. The first sub-video stream is a video stream obtained by acquiring images of the spinneret; the second sub-video stream is a video stream obtained by acquiring images of the filament bundle area directly below the outlet of the web nozzle.

4. The method according to claim 3, wherein, The step of inputting the video clip and the parameter clip into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model includes: From the first images of multiple frames belonging to the first sub-video stream of the video segment, obtain the first target frame image that meets the preset quality requirements; The first target frame image is converted into a grayscale image to obtain the first grayscale image; Based on the parameter initialization model, the controller's initial channel features of multiple first channels are extracted from the first grayscale image. The model is initialized based on the parameters to process at least one first sub-parameter, thereby obtaining the respective channel weights of the plurality of first channels; The initial channel features of the multiple first channels are weighted and summed based on their respective channel weights to obtain the parameter values ​​of the multiple first channels. From the multiple frames of second images belonging to the second sub-video stream of the video segment, obtain a second target frame image that meets the preset quality requirements; Based on the parameter initialization model, the controller's initial channel features of multiple second channels are extracted from the second target frame image; The model is initialized based on the parameters to process at least one second sub-parameter and to determine the respective channel weights of the plurality of second channels; The initial channel features of the multiple second channels are weighted based on their respective channel weights to obtain the parameter values ​​of the multiple second channels; The initial parameter values ​​of the controller include the parameter values ​​of the plurality of first channels and the parameter values ​​of the plurality of second channels.

5. The method according to claim 4, wherein, The step of extracting the initial channel features of multiple first channels of the controller from the first grayscale image based on the parameter initialization model includes: An opening operation is performed on the first grayscale image to obtain a first reference image; and, Perform a closing operation on the first grayscale image to obtain a second reference image; The pixel value difference at the same position between the first reference image and the first grayscale image is determined to obtain the first difference feature; The pixel value difference at the same position between the first grayscale image and the second reference image is determined to obtain the second difference feature; An adaptive grayscale threshold segmentation operation is performed on the difference map between the first reference map and the second reference map to obtain the maximum value of the inter-class variance. The first difference feature, the second difference feature, and the maximum inter-class variance are respectively input into the corresponding first multilayer perceptron network to obtain the initial channel features of the plurality of first channels; wherein, the first difference feature, the second difference feature, and the maximum inter-class variance each correspond to a first channel.

6. The method according to claim 4 or 5, wherein, The process of initializing the model based on the parameters to process at least one first sub-parameter and obtain the respective channel weights of the plurality of first channels includes: Execute separately for each first channel: Obtain the aperture area of ​​the first channel for T time windows to obtain the aperture area sequence; Determine the standard deviation and mean of the pore area sequence; Based on the standard deviation and mean, the fluctuation intensity of the first channel is determined; The fluctuation intensity of each first channel and the at least one first sub-parameter are taken as a first set. For each first target parameter in the first set, the first mapping feature corresponding to each first target parameter is obtained based on the neural network module. Each first mapping feature is input into the global attention mechanism network to obtain the weight contribution value of each first mapping feature to each first channel; Execute separately for each first channel: Based on the first indicator function, at least one first valid value is selected from the weight contribution values ​​of each first mapping feature to the first channel; the first indicator function defines the parameters in the first set that affect each first channel; The channel weight of the first channel is obtained by determining the sum of the at least one first effective value.

7. The method according to claim 4, wherein, The step of weighted summing of the initial channel features of the plurality of first channels based on their respective channel weights to obtain the parameter values ​​of the plurality of first channels includes: The initial features of each first channel are sequentially processed by orthogonalization and fully connected layers to obtain the first weighted features of each first channel; Based on the respective channel weights of the plurality of first channels, the first features to be weighted of each first channel are weighted and summed to obtain the parameter values ​​of the plurality of first channels.

8. The method according to claim 4, wherein, The extraction of channel initialization features of multiple second channels of the controller from the second target frame image based on the parameter initialization model includes: Extract the node density of the silk threads from the second target frame image; The number of nodes in each image block in the second target frame image is determined to obtain the number of local nodes in each image block; The node variation coefficient of the second target frame image is determined based on the mean and standard deviation of the number of local nodes in each image block. Extract the node mask image of the second target frame image; The first value is obtained by determining the cumulative value of the pixel-by-pixel product of the grayscale image of the second target frame image and the node mask image; The sum of the accumulated values ​​of each pixel in the node mask image and the sum of the preset coefficients are determined to obtain the second value; Determine the ratio of the first value to the second value to obtain candidate intensity values; The maximum value between the intensity candidate frames and the default value is selected to obtain the node intensity of the second target frame image; The node density, the node variation coefficient, and the node strength are respectively input into the corresponding second multilayer perceptron network to obtain the initial channel features of the plurality of second channels; wherein the node density, the node variation coefficient, and the node strength each correspond to a second channel.

9. The method according to claim 4 or 8, wherein, The step of initializing the model based on the parameters, processing the at least one second sub-parameter, and determining the respective channel weights of the plurality of second channels includes: For each second target parameter of the at least one second sub-parameter, a second mapping feature corresponding to each second target parameter is obtained based on a neural network module; Multiple second mapping features are input into the global attention mechanism network to obtain the weight contribution value of each second mapping feature to each second channel; Perform the following for each second channel: At least one second valid value is selected from the weight contribution values ​​of each second mapping feature to the second channel based on the second indicator function; the second indicator function defines the parameters in the second set that affect each second channel; The channel weight of the second channel is obtained by determining the sum of the at least one second effective value.

10. The method according to claim 4, wherein, The step of weighting the initial channel features of the plurality of second channels based on their respective channel weights to obtain the parameter values ​​of the plurality of second channels includes: The initial features of each second channel are sequentially processed by orthogonalization and fully connected layers to obtain the second weighted features of each second channel; Based on the respective channel weights of the multiple second channels, the second features to be weighted of each second channel are summed in a weighted manner to obtain the parameter values ​​of the multiple second channels.

11. A device for predicting spinning defects, comprising: The acquisition module is used to acquire video streams of target components in the spinning process and to acquire time-series sequences of preset monitoring parameters. The target component includes a spinning assembly and / or a winding mechanism; The acquisition module is used to acquire a video segment of the current time window from the video stream, and to acquire a parameter segment of the current time window from the time sequence; The first processing module is used to input the video clip and the parameter clip into the parameter initialization model to obtain the initial parameter values ​​of the controller of the fault prediction model; The visual feature extraction module is used to input the video clip into the visual encoder of the fault prediction model to obtain initial visual features; The text feature extraction module is used to input the parameter fragments into the text encoder of the fault prediction model to obtain text features; The second processing module is used to input the initial visual features and the text features into the large language model of the fault prediction model to construct the key-value pair information of the last decoding layer of the large language model; An optimization module is used to optimize the visual feature representation in the key-value pair information based on the initial parameter values ​​of the controller; The determination module is used to determine the spinning fault prediction result for the current time window based on the optimized visual feature representation.

12. The apparatus according to claim 11, wherein, The optimization module is specifically used to take the last decoding layer as the target decoding layer and iteratively perform the following operations: For the current step, obtain the input information of the target decoding layer, which includes the current output result of the large language model and the current key-value pair information of the target decoding layer; The large language model is used to predict the input information to obtain the confidence level of the intermediate output result of the current step; The distribution entropy of the current step is determined based on the confidence level; The exponential moving average entropy of the current step is determined based on the distribution entropy and the exponential moving average entropy of the previous step. The target loss is determined based on the exponential moving average entropy of the current step; Determine the gradient of the target loss with respect to the controller parameters of the controller; The current parameters of the controller are updated based on the gradient to obtain the controller parameter update result; The visual feature representation in the current key-value pair information is updated based on the controller parameter update result; If the termination condition is not met, the input information for the next step of the current step is constructed based on the updated visual feature representation. The next step is then taken as the current step, and the operation of obtaining the input information of the target decoding layer for the current step is returned.

13. The apparatus according to claim 12, wherein the spinning failure prediction result includes: spinneret Congestion fault prediction value, network node anomaly prediction value; The spinneret blockage fault prediction value is predicted based on the first sub-video stream in the video stream and at least one first sub-parameter among the following preset monitoring parameters: melt pressure at the inlet of the spinneret assembly, rotation speed of the oil metering pump, and spinning temperature. The network node anomaly prediction value is predicted based on a second sub-video stream in the video stream and at least one of the following preset monitoring parameters: Network air pressure, used to indicate the set value of compressed air pressure for the network nozzle; Network degree is used to represent the number of network nodes per meter of wire. The target guide roller linear speed is the linear speed of the guide roller used to determine the winding speed. The first sub-video stream is a video stream obtained by acquiring images of the spinneret; the second sub-video stream is a video stream obtained by acquiring images of the filament bundle area directly below the outlet of the web nozzle.

14. The apparatus according to claim 13, wherein, The first processing module includes: The first acquisition unit is used to acquire a first target frame image that meets preset quality requirements from multiple first images belonging to the first sub-video stream of the video segment. A conversion unit is used to convert the first target frame image into a grayscale image to obtain a first grayscale image; The first extraction unit is used to extract the initial channel features of multiple first channels of the controller from the first grayscale image based on the parameter initialization model. The first determining unit is used to initialize the model based on the parameters, process the at least one first sub-parameter, and obtain the respective channel weights of the plurality of first channels; The first calculation unit is used to perform a weighted summation of the initial channel features of the plurality of first channels based on their respective channel weights to obtain the parameter values ​​of the plurality of first channels. The second acquisition unit is used to acquire a second target frame image that meets preset quality requirements from multiple frames of second images belonging to the second sub-video stream of the video segment; The second extraction unit is used to extract the initial channel features of multiple second channels of the controller from the second target frame image based on the parameter initialization model; The second determining unit is used to process the at least one second sub-parameter based on the parameter initialization model to determine the respective channel weights of the plurality of second channels; The second calculation unit is used to weight the initial channel features of the plurality of second channels based on their respective channel weights to obtain the parameter values ​​of the plurality of second channels. The initial parameter values ​​of the controller include the parameter values ​​of the plurality of first channels and the parameter values ​​of the plurality of second channels.

15. The apparatus according to claim 14, wherein, The first extraction unit is used for: An opening operation is performed on the first grayscale image to obtain a first reference image; and, Perform a closing operation on the first grayscale image to obtain a second reference image; The pixel value difference at the same position between the first reference image and the first grayscale image is determined to obtain the first difference feature; The pixel value difference at the same position between the first grayscale image and the second reference image is determined to obtain the second difference feature; An adaptive grayscale threshold segmentation operation is performed on the difference map between the first reference map and the second reference map to obtain the maximum value of the inter-class variance. The first difference feature, the second difference feature, and the maximum inter-class variance are respectively input into the corresponding first multilayer perceptron network to obtain the initial channel features of the plurality of first channels; wherein, the first difference feature, the second difference feature, and the maximum inter-class variance each correspond to a first channel.

16. The apparatus according to claim 14 or 15, wherein, The first determining unit is configured to: Execute separately for each first channel: Obtain the aperture area of ​​the first channel for T time windows to obtain the aperture area sequence; Determine the standard deviation and mean of the pore area sequence; Based on the standard deviation and mean, the fluctuation intensity of the first channel is determined; The fluctuation intensity of each first channel and the at least one first sub-parameter are taken as a first set. For each first target parameter in the first set, the first mapping feature corresponding to each first target parameter is obtained based on the neural network module. Each first mapping feature is input into the global attention mechanism network to obtain the weight contribution value of each first mapping feature to each first channel; Execute separately for each first channel: Based on the first indicator function, at least one first valid value is selected from the weight contribution values ​​of each first mapping feature to the first channel; the first indicator function defines the parameters in the first set that affect each first channel; The channel weight of the first channel is obtained by determining the sum of the at least one first effective value.

17. The apparatus according to claim 14, wherein, The first computing unit is used for: The initial features of each first channel are sequentially processed by orthogonalization and fully connected layers to obtain the first weighted features of each first channel; Based on the respective channel weights of the plurality of first channels, the first features to be weighted of each first channel are weighted and summed to obtain the parameter values ​​of the plurality of first channels.

18. The apparatus according to claim 14, wherein, The second extraction unit is used for: Extract the node density of the silk threads from the second target frame image; The number of nodes in each image block in the second target frame image is determined to obtain the number of local nodes in each image block; The node variation coefficient of the second target frame image is determined based on the mean and standard deviation of the number of local nodes in each image block. Extract the node mask image of the second target frame image; The first value is obtained by determining the cumulative value of the pixel-by-pixel product of the grayscale image of the second target frame image and the node mask image; The sum of the accumulated values ​​of each pixel in the node mask image and the sum of the preset coefficients are determined to obtain the second value; Determine the ratio of the first value to the second value to obtain candidate intensity values; The maximum value between the intensity candidate frames and the default value is selected to obtain the node intensity of the second target frame image; The node density, the node variation coefficient, and the node strength are respectively input into the corresponding second multilayer perceptron network to obtain the initial channel features of the plurality of second channels; wherein the node density, the node variation coefficient, and the node strength each correspond to a second channel.

19. The apparatus according to claim 14 or 18, wherein, The second determining unit is used for: For each second target parameter of the at least one second sub-parameter, a second mapping feature corresponding to each second target parameter is obtained based on a neural network module; Multiple second mapping features are input into the global attention mechanism network to obtain the weight contribution value of each second mapping feature to each second channel; Perform the following for each second channel: At least one second valid value is selected from the weight contribution values ​​of each second mapping feature to the second channel based on the second indicator function; the second indicator function defines the parameters in the second set that affect each second channel; The channel weight of the second channel is obtained by determining the sum of the at least one second effective value.

20. The apparatus according to claim 14, wherein, The second computing unit is used for: The initial features of each second channel are sequentially processed by orthogonalization and fully connected layers to obtain the second weighted features of each second channel; Based on the respective channel weights of the multiple second channels, the second features to be weighted of each second channel are summed in a weighted manner to obtain the parameter values ​​of the multiple second channels.

21. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

23. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.