Hardware control method, storage medium and manufacturing system

Through a time series-based machine learning framework and Transformer model, combined with a photon programming processor, the problem of hardware device operation control and optimization is solved, and the intelligent control and remote management of the device are realized, which is suitable for a variety of complex environments.

CN120086526APending Publication Date: 2025-06-03潘婧
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148744.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-22
Filing Date
2025-02-11
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively control and optimize the operation of hardware devices, especially when processing large amounts of time series data and multi-objective optimization.

Method used

Using a time series-based machine learning framework, combining Transformer model and photon programming processor, we acquire and predict device sensor data, optimize target data and abnormal event data, and realize intelligent control and optimization of hardware devices.

Benefits of technology

This method can achieve efficient control and optimization of hardware equipment, ensure normal operation of equipment, improve the efficiency and energy efficiency of manufacturing processes, and provide remote configuration and programming capabilities, which are suitable for environments without Internet connections such as deep space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086526A_ABST
    Figure CN120086526A_ABST
Patent Text Reader

Abstract

The invention provides a hardware control method, a medium and a manufacturing system. The method comprises the following steps: acquiring equipment sensor data from a sensor according to a time sequence; acquiring equipment optimization target data from the plurality of optimization targets according to a time sequence; acquiring historical abnormal event and intervention event data of the equipment; acquiring static equipment input parameters; applying a time series model to the acquired equipment sensor and optimization data, historical event data and static equipment input parameters to obtain predicted equipment sensor data; the hardware operation is optimized and controlled based on the predicted equipment sensor data; and meanwhile, a predicted action is provided for abnormal event intervention according to the predicted equipment sensor data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 440,875, filed on February 13, 2024, with the invention title "Time series based machine learning framework for hardware equipment and its implementations with transformers and on optical programming processors", and U.S. Patent Application No. 18 / 991,642, filed on December 22, 2024, with the invention title "Time series based machine learning framework for hardware equipment and its implementations with transformers and on optical programming processors". The entire contents of the above - mentioned priority applications are incorporated herein by reference in their entirety. Background Art

[0003] Hardware devices such as assembly lines in manufacturing, power systems, and heating, ventilation, and air conditioning (HVAC) in buildings are usually controlled by computer processors running programs and are fed back by data obtained from sensors distributed in the hardware device system. Summary of the Invention

[0004] This application relates to a method for controlling hardware, including: obtaining device sensor data arranged in a time series from a plurality of sensors; obtaining device optimization target data arranged in a time series from a plurality of optimization targets; and obtaining historical data of device abnormal events and intervention events.

[0005] In some embodiments, the first and second time series can be combined into at least one multi - objective time series. In some other embodiments, the first and second time series are different time series.

[0006] In some embodiments, the method further includes: obtaining static device input parameters; applying a time series model to device sensor data, historical data, and static device input parameters to obtain predicted device sensor data over time; in some embodiments, performing manufacturing optimization based on the predicted device optimization target data; in other embodiments, ensuring the normal operation of the hardware device based on the predicted abnormal event data; and providing a predicted action for the predicted abnormal event intervention based on the action recommendation data.

[0007] In some embodiments, the method further includes at least one iteration among obtaining historical data of device abnormal events and intervention events, obtaining static device input parameters, and applying a time series model to device sensor data, historical data, and static device input parameters to obtain predicted device sensor data, optimization target values, and predicted abnormal events; and providing predicted actions and evaluation actions based on the results of the iteration.

[0008] In some embodiments, the "providing" step includes displaying the result on a display screen or sending a mobile phone alert to the user.

[0009] In some embodiments, the "providing" step includes sending a control signal to the control circuit for controlling the hardware according to the result to achieve manufacturing optimization.

[0010] In some embodiments, the time series model includes a Transformer model, i.e., the base model.

[0011] In some embodiments, the method further includes outputting multiple y variables from the time series model, including device sensor data y_(si - t) of the i-th sensor data changing with time t. Optionally, the sensor data y_(si - t) includes temperature data measured at a specified location. Optionally, the device sensor data y_(si - t) includes the amplitude, voltage, current, frequency, force, etc. of the motor.

[0012] In some embodiments, the method further includes outputting multiple y variables from the time series model, including device optimization target data y_(oj - t) of the j-th optimization target data changing with time t.

[0013] In some embodiments, the device optimization target data y_(oj - t) includes the energy output, power, torque (which can sometimes also be measured as y_(si - t)), energy efficiency, etc. of the motor.

[0014] In some embodiments, the machine learning architecture stacks at least two layers of models on top of the input time series data.

[0015] Optionally, the machine learning architecture stacks three layers of models on the input time series data.

[0016] Optionally, the machine learning service is based on at least a two - layer model architecture.

[0017] Optionally, the machine learning service is based on a three - layer model architecture.

[0018] In some embodiments, a method is provided for generating a sequence of target values beyond a single timestamp, and using the generated sequence of target values as input for the model of the downstream layer when real data is not available. Among them, the anomaly event model predicts anomaly events in the long term because the previous layer can predict sensor data in the long term, and the selection of the survival classifier can process historical anomaly event data row by row, thus overcoming the problem of fewer anomaly event labels in the supervised learning concept.

[0019] In some embodiments, the anomaly event model based on the Transformer model can predict the predicted anomaly events.

[0020] Optionally, the anomaly event model based on the Transformer model can also predict the predicted anomaly events when there are no previous historical anomaly events.

[0021] Optionally, the anomaly event model based on the Transformer model can also predict the predicted anomaly events when there is only one previous historical anomaly event.

[0022] Optionally, the anomaly event model based on the Transformer model can also predict the predicted anomaly events when there are multiple historical anomaly events.

[0023] In some embodiments, since the previous layer can predict sensor data in the long term, the action recommendation model can output predicted actions in the long term.

[0024] Optionally, the action recommendation model can output predicted actions in the long term based on the supervised learning method of the recommendation system.

[0025] Optionally, the action recommendation model can output predicted actions in the long term through a multi - category binary classification model.

[0026] Optionally, the action recommendation model can output predicted actions in the long term by forming predictions through graph links.

[0027] In some embodiments, the action recommendation model can output predicted actions in the long term from the Transformer model.

[0028] Optionally, the action recommendation model can output predicted actions in the long term from the Transformer model by combining multi - modal methods.

[0029] Optionally, the action recommendation model can combine with an action evaluation method to output long-term predicted actions from the Transformer model.

[0030] In some embodiments, the control method of the hardware provides hardware device parameter optimization through a time series model, where the optimization model predicts the optimization target value. The hardware parameters in the input of the model that can optimize the optimization target value are the optimal set of hardware parameters.

[0031] Optionally, a more efficient hardware device parameter search method is a machine learning-based search method.

[0032] Optionally, a more efficient hardware device parameter search method is a machine learning-based search method, such as random search and sequential model-based optimization (SMBO).

[0033] In some embodiments, a hardware control system is provided that executes an optimization strategy and predicted abnormal event intervention by adjusting device operation parameters based on a hierarchical machine learning model.

[0034] Optionally, the system includes an application service that compares the input signal of the hierarchical machine learning model with the current hardware state and generates an executable software control signal.

[0035] Optionally, these software control signals are converted into hardware control signals to change device operation, so that corresponding adjustments can be made according to the predicted insights.

[0036] Optionally, the system can switch between an autonomous mode and a human intervention mode; in the human intervention mode, human operations take precedence over automatic control for direct human intervention when necessary.

[0037] In some embodiments, the hardware device parameter optimization is based on a Transformer model.

[0038] In some embodiments, a hardware device control method based on a time series machine learning model or a Transformer model is provided, as well as a new numerical method that does not rely on explicit formulas.

[0039] In some embodiments, a large Transformer model (base model) is provided that can learn and organize a large amount of hardware device data from various application scenarios, different types of devices, and various input and output data types and sources, similar to the base model of a large language model (LLM), for generating various sequences such as sensing, optimization, abnormal events, action sequences, sequence prediction, specific task prediction, fine-tuning, transfer learning, embedding, retrieval, etc.

[0040] In some embodiments, the hardware control method provides an EquiFormer system design based on the concept of interconnected devices, wherein the EquiFormer system is applied to a remotely configurable and programmable photonic programmable processor (RCP_OPP); the design of the RCP_OPP is based on an Internet of Things (IoT) component assuming an Internet connection; assuming no Internet connection, when the RCP_OPP needs to be integrated into a larger system to control devices in deep space, it is based on photon entanglement and is called a photonic programmable processor (PC_OPP) controlled by photons; wherein the PC_OPP controls photon entanglement from the earth to deep space.

[0041] In another aspect, a non-transitory computer-readable medium storing instructions is provided, the instructions being executed by one or more processing circuits to implement the steps of a method of controlling hardware.

[0042] In yet another aspect, a manufacturing system is provided that includes one or more processing circuits, a manufacturing / assembly line, a sensor, a controller, and the above-described non-transitory computer-readable medium for optimizing a manufacturing process.

[0043] In some embodiments, three methods are provided to convert hardware input parameters into photonic output parameters, thereby enabling long-distance (e.g., deep space) communication and control.

[0044] In one embodiment, the hardware parameter values ​​are mapped to laser light wave parameters. A high-power laser source with precise aiming capability sends the mapped laser signal directly from the earth to a remote space location. The receiver in space converts the laser light wave parameters back into signals for controlling the hardware in the PC_OPP configured for this method.

[0045] In another embodiment, photon-based continuous variable quantum teleportation is used. The continuous amplitude and phase orthogonal quantities of the light field are used as information carriers. Through entangled states and classical communication, the unknown quantum state is transmitted from the sender to the receiver, thereby realizing the transmission of the hardware parameter values ​​encoded in the continuous variable.

[0046] In yet another embodiment, photon-based qubit encoding is used. The hardware parameter values ​​are converted into binary form and mapped to qubits through the quantum state of photons. A novel blockchain-based qubit communication network is proposed, in which the quantum state of photons is transmitted through a network of nodes, and the nodes are similar to blocks in the blockchain. This method can enhance security and error correction capabilities without relying on classical communication methods. The blockchain-based qubit communication network is original to the inventor.

[0047] Optionally, the blockchain-based qubit communication network includes multiple nodes. The nodes receive qubit-encoded information from the sender, locally compare and verify the received information, and correct errors and detect tampering through majority consensus, thus eliminating the need for classical communication for error correction.

[0048] This blockchain-based qubit communication network can serve as a general qubit communication mechanism and is not limited to the field of hardware control.

[0049] Optionally, after transmitting the hardware parameters to deep space via the above three methods, if necessary, other mechanisms can be used to replace the PC_OPP near the hardware device.

[0050] In some embodiments, a method for training a deep learning model (including the Transformer model) is provided, where ring-all reduce gradient updates are used based on model parallelism and data parallelism to improve training scalability without relying on a centralized parameter server.

[0051] In one embodiment, when both data parallelism and model parallelism are adopted simultaneously, the ring-all reduce technique is applied to perform gradient averaging in a point-to-point manner between computing units, thereby achieving efficient synchronization of gradients calculated from different data batches corresponding to the same model partition.

[0052] Optionally, this method supports partitioning matrices in Low-Rank Adaptation (LoRA) across different computing units and uses ring-all reduce gradient updates to synchronize the gradients of LoRA fine-tuning.

[0053] In some embodiments, a general solution for using a large language model to generate 2D or 3D visualization charts through general language representation and optional graph representation is proposed.

[0054] In one embodiment, the visualization chart is converted into a general language representation, where the visualization elements are defined by attributes such as type, color, position, and relationship. These representations are used to fine-tune existing large language models to generate visualization charts.

[0055] In another embodiment, the visualization elements are encoded as tokens with corresponding embeddings, and the sequence of visualization elements is organized based on the element relationships described in the graph.

[0056] In yet another embodiment, a hybrid model architecture is proposed, which includes a language modality and a visualization graph modality. The model uses modules based on the Transformer model in each modality and combines a fusion layer with cross-modal attention to generate visualization charts.

[0057] Optionally, a mechanism for verifying the generated representation is proposed to ensure that the generated visualization chart is legal; meanwhile, human or model-based feedback can be utilized to improve the generation quality.

[0058] In some embodiments, a method for enabling graph modality generation is proposed, enabling a model based on the Transformer model to generate a vertex ID sequence based on a graph path, thereby achieving richer connectivity between tokens than in a sequential manner.

[0059] Optionally, this method is applied to generate action sequences for AI agents, recommendation systems, and visualization charts, where the relationships between items are represented by a graph rather than just a sequence.

[0060] In some embodiments, a search and drop-ship system for enabling image generation is proposed to enhance product discovery and shopping experience.

[0061] In one embodiment, a model such as a convolutional neural network is used to analyze an image to identify attributes; a multimodal large model generates a product image based on a user's natural language query in combination with an image having specific attributes.

[0062] Optionally, even if the existing natural language description of the image lacks the query attribute, image-to-image search can be performed based on image similarity to find a matching product; if a matching product still cannot be found, drop-ship service can be provided.

[0063] In some embodiments, a multimodal automotive base model integrating EquiFormer is proposed, which includes modules for hardware devices, human behavior, energy management, and other aspects, adopting a structure with a fusion layer similar to a multimodal model. This multimodal automotive base model can enhance the interactivity with human drivers, pedestrians, and other entities, thereby improving the capabilities of the autonomous driving system.

[0064] Optionally, the hardware module uses the EquiFormer model structure as the backbone to achieve advanced prediction and control of the automotive system.

[0065] Optionally, the model includes sub-modules applicable to different types of vehicles such as internal combustion engine vehicles, hybrid vehicles, and electric vehicles, and can integrate various input services such as weather, traffic, and road conditions. Brief Description of the Drawings

[0066] Figure 1 Shows an overall system design of EquiFormer including machine learning services.

[0067] Figure 2 Shows a general machine learning architecture of EquiFormer.

[0068] Figure 3 Shows the EquiFormer sensor model that generates a sequence of predicted target values.

[0069] Figure 4 Shows an EquiFormer machine learning service based on the Internet of Things and the cloud.

[0070] Figure 5 Shows an anomaly event model based on a survival classifier.

[0071] Figure 6 Shows an action recommendation method based on a recommendation system model.

[0072] Figure 7 Shows an action recommendation method with graph link formation prediction.

[0073] Figure 8 Shows a hardware optimization model architecture based on the Transformer model.

[0074] Figure 9 Shows a hardware optimization method based on Transformer embedding input retrieval.

[0075] Figure 10 Shows a method for predicting anomaly events based on embedding vectors and Transformer, for the case when historical anomaly events occur only once.

[0076] Figure 11 Shows a conceptual diagram of the hardware components of a remotely programmable and configurable optical programming processor (RPC_OPP), but does not include specific detailed electrical diagrams.

[0077] Figure 12 Shows a control system for an optical programming processor (PC_OPP) based on optical entanglement.

[0078] Figure 13 Shows a blockchain-style qubit communication network.

[0079] Figure 14 Shows a ring-style global reduction gradient update operation for a deep learning model based on data parallelism and model parallelism.

[0080] Figure 15 Shows the model partitioning of low-rank adaptation.

[0081] Figure 16 Shows the correspondence between model partitioning and the partitioning of low-rank matrices A and B.

[0082] Figure 17Shows a double - loop global reduction operation for low - rank adaptation fine - tuning based on data parallelism and model parallelism.

[0083] Figure 18 Shows a general visualization chart generation process.

[0084] Figure 19 Shows a multi - modal modular automotive base model. Detailed implementation

[0085] Section 1: Introduction

[0086] The operating characteristics of the hardware device are as follows: a large amount of rapidly changing time - series data from different sensors is generated almost in real - time; the device preset / tuning parameters change relatively slowly; and there are multiple non - mutually exclusive events, such as maintenance, faults, adjustments, and production; and there are multiple optimization goals, such as best energy efficiency, maximum output, and minimum faults. This complexity of the hardware system makes it difficult to optimize traditional machine - learning methods.

[0087] Supervised - learning modeling techniques usually require labels. In time - series problems, a dependent variable that changes over time is needed. However, except for sensor data (which is also the modeling input), the actual manufacturing optimization goals (such as production output rate, etc.) usually do not change much. For events such as maintenance and faults, they are extremely rare in the actual data. This usually makes it difficult to apply machine - learning techniques to model hardware systems such as manufacturing processes, energy systems, and HVAC systems.

[0088] Artificial intelligence (AI) has made progress in the fields of natural language processing (NLP) and image / video processing. However, the hardware and manufacturing industries have not undergone major transformations due to AI. The reasons may include: (1) in the manufacturing industry, events such as abnormal events or intervention operations (such as maintenance) rarely occur, making it difficult for supervised - learning techniques to use these rare events as labels; (2) it is difficult to model the large amount of time - series data from device sensors with traditional time - series models, due to reasons such as the limited variability of the dependent and input variables, the long sequence of sensor data, multiple dependent variables, multiple correlated modeling goals, the collinearity of input variables, and the lack of integration of data sources within the industry (compared with Internet data on the cloud) due to the lack of a secure Internet of Things (IoT) hardware infrastructure and integration with cloud data platforms, etc.

[0089] In some implementations, traditional time series techniques can be used to model manufacturing equipment data, such as regression-based time series methods (Mendenhall, Sincich, et al., 2003). This regression method typically models one time series at a time. However, manufacturing equipment usually has a large number of sensors even for a single device / processor, and one sensor may sense data in multiple dimensions. One hopes to understand all future sensor values, which requires modeling multiple time series while dealing with issues such as long sequences of sensor data, multiple dependent variables, multiple correlated modeling objectives, and collinearity of input variables.

[0090] Another approach could be tree-based models (Liukis, 2020). Typically, this method also models one time series at a time. Among traditional machine learning methods, it is most suitable for modeling multiple mutually exclusive classifications as dependent variables with a time component in the modeling input. In the manufacturing process, many future target values of interest are either related or classified but not mutually exclusive. For example, the temperature of a hardware component of a device during use may be related to its volume. In another example, as a device component ages, it may exhibit defects in one area, which also increases the probability of defects in another area. Although this method may model continuous dependent variables, it lacks the ability to model complex correlations and variabilities between dependent and independent variables.

[0091] Deep learning, with its unique ability to predict non - mutually exclusive multi - label outputs and its universal approximation ability (i.e., it can model any relationship between dependent and independent variables), represents the next step in manufacturing process modeling. The Transformer is a deep - learning model that can model long sequences (Vaswani, Shazeer, et al., 2017). It has been applied in natural language processing and image processing and has achieved many breakthroughs. For example, the first step in large - language models is a pre - trained Transformer model based on a language corpus. However, the reasons why machine - learning modeling techniques are difficult to apply to hardware systems such as manufacturing processes include: (1) the lack of a large number of integrated data sources and IoT / cloud system designs to achieve this goal; (2) the lack of hardware designs to implement the above - mentioned system designs, especially when the distance from sensors to the data center is far or there is no Internet available, etc.; (3) the lack of dedicated time - series encoding for hardware systems for sequence modeling; (4) the lack of multi - objective model stacking / architecture design; (5) there is no dedicated machine - learning solution based on Transformer for specific application scenarios in hardware systems, while there are many similar solutions in the NLP and image fields. In the target hardware optimization section, traditional control theory relies on explicit mathematical formulas to achieve parameter optimization. Whether the mathematical formula deviates from reality or the actual data is difficult to approximate through explicit formulas, even if the machine - learning optimization model obtains the best parameters, it is difficult to use the existing control theory to correct the deviation of the parameter values.

[0092] Section 2: Implementation of the System Framework

[0093] Section 2.1: System Design

[0094] The machine - learning service part can include: (1) a pre - trained machine - learning model for historical time - series data; (2) a simulation method based on the pre - trained model to update the parameters of the optimization objective; (3) continuous online model updates; (4) continuous parameter optimization based on real - time device data; (5) event prediction and action triggering based on multi - modal models. In addition to the machine - learning service, a hardware system and software services necessary for the operation of the machine - learning service are also required.

[0095] For example, an air conditioner used for refrigeration may have the following inputs: Non - timestamp inputs: designed AC voltage input, designed AC current input, product size, designed power consumption, weight, etc.; Timestamp sensor data: indoor temperature, indoor humidity, indoor oxygen content, indoor carbon dioxide content, indoor air pressure, outdoor temperature, real - time AC power, real - time motor frequency, real - time motor torque, real - time motor temperature, indoor air flow rate, etc.; Optimization objectives: seasonal energy efficiency ratio (SEER), cooling capacity measured in British Thermal Units (BTU); Examples of abnormal events: motor failure, circuit failure, etc.

[0096] Section 2.2: Machine Learning Design Based on Time Series

[0097] In the above system design, the machine learning service is supported by multiple layers of machine learning models. First, the time series machine learning model architectures in parts (1) and (3) are very similar. This is part of the "first layer" machine learning architecture in the embodiments of this application.

[0098] First, time t is defined as a series of time points from 0 to n.

[0099] Second, multiple y variables can be specified to represent the output / target of the model. (1) In some embodiments, device sensor data y_(si-t) can be obtained, where i is the i-th sensor data from the device, which varies with time t. For example, the device temperature T at a certain time point t and a specific spatial location. (2) In some embodiments, continuous device optimization targets y_(oj-t) can be obtained, where j is the value of the j-th device optimization target, which varies with time t. For example, the manufacturing output rate at a certain time point. (3) In some embodiments, binary device anomaly events y_(ek-t) can be obtained, where k is the k-th device-related anomaly event. The value can only take 0 (indicating that the event has not occurred) or 1 (indicating that the event has occurred).

[0100] In actual data, the case where the value is 1 is very rare, such as events like hardware failures or hardware alarms.

[0101] Third, the model can be trained with input data (200). Figure 2 The overall machine learning architecture from the input data to the stacked model layers is shown. (1) In the first example, the time step itself t_(t < n) (such as Figure 2 206) can be input. (2) Other input variables, such as sensor data y_(si-t, t < n) (such as Figure 2 202) and optimization target y_(oj-t, t < n) (such as Figure 2 204), can also be included. The characteristic of the time series is that for the y variable at t = n, all its previous values t < n can be used as inputs. (3) In the third example, other time-based input variables x_(l-t, t < n) (such as Figure 2 208) can be provided. Examples include the temperature provided by the weather forecast service. (4) In the fourth example, other non-time-related input variables x_m (such as Figure 2of 209). Examples include the design specifications of the device, such as the designed coil diameter and the voltage input designed for the device. The input variables in the x_m group do not change over time, and not all time series models require each input variable to change over time, but they are generally kept in the general form of the model. When a time series machine learning model models many different but similar hardware devices simultaneously, the input variables in the x_m group may vary. When applying Transformer modeling techniques to handle a large number of similar but different hardware devices (such as large language corpora), the input variables in the x_m group are extremely useful and can provide valuable baseline information.

[0102] Figure 1 Shows an overall system design of EquiFormer including machine learning services. In the device data service 104, the system first determines in step 102 whether there is historical time series data applicable to a specific use case. If not, the device data acquisition and storage component 103 can be called. The data obtained from the device data service 104 can be exchanged with the machine learning service 106, especially at determination step 108 to determine whether there is a previous model. If not, the offline historical data model 110 can be called to provide offline hardware parameter optimization 120. If there is a previous model 108, the online model update 130 can be called, which facilitates online hardware parameter optimization 140 and task-specific abnormal event prediction and action recommendation 150.

[0103] The manual hardware administrator 154 and the application service 150 for device control can run independently or can be optionally run in parallel to manage the hardware layer 160. The manual hardware administrator 154 and the application service 150 exchange information and control instructions bidirectionally through the user interface 158. The user interface 158 includes visual / audio / tactile signals, touchscreens, keyboards, and mobile phone / tablet alerts or applications, etc. as components of the application service 150. The application service 150 receives inputs from the machine learning service 106 and the hardware layer 160 and generates software-based control signals for the hardware layer 160. The application service 150 includes two functional components: the hardware optimization service 156 and the abnormal event intervention service 152.

[0104] The hardware optimization service 156 receives static or dynamic optimization parameter inputs provided by the offline 120 or online 140 hardware parameter optimization components. It can autonomously exchange information and control signals with the hardware layer 160 and selectively interact with the human hardware administrator 154. The anomaly event prediction and recommended actions provided by the online hardware parameter optimization 140 can be displayed as visual information on the user interface 158 through the anomaly event intervention service 152, or transmitted as anomaly event alerts through other audio / visual / tactile signals. These alerts notify the human hardware administrator 154. Since anomaly event intervention is crucial, the human hardware administrator can directly interact with the hardware layer 160 to perform critical task operations. They can also bypass the application service 150 and directly obtain inputs from the anomaly event prediction and action recommendation component 150 to make independent judgments without the inputs of services 106 and 150.

[0105] The human hardware administrator 154 can directly intervene in the device 160 through the user interface 158, or intervene through the controller 166 via the user interface 158. The hardware optimization service 156 and the anomaly event intervention service can automatically control the device 165 through the controller 166. Optionally, the human hardware administrator 154 can modify the input of the controller 166 through the user interface 158 of the application service 150. The controller 166 processes the control signals generated by the software and converts them into physical signals, such as the amplitude of light waves, which the device 165 interprets to adjust its operating state.

[0106] The interface between the application service 150 and the controller 166 can be a direct physical connection, such as through a USB-C port or a connection block, which allows the software to communicate directly with the controller without an intermediate communication component; or optionally through the communication component 168. The communication component can transmit the software control signals over longer distances that cannot be achieved through a direct physical connection.

[0107] The Internet of Things (IoT) sensors 162 in the hardware system 160 collect data from the device 164 and transmit this data to the data acquisition and storage component 104. The IoT sensors 162 also provide real-time feedback to the human hardware administrator 154 and the application service 150 for device control, enabling iterative adjustment of the control signals provided to the hardware layer 160. For example, the input control signal to the hardware layer may specify "increase the speed to x revolutions per minute", but after a period of time, the motor speed may stabilize at x minus delta revolutions per minute. In this case, the application service 150 can selectively send updated control signals to the hardware layer 160 without relying on the computationally intensive device data service 104 and the machine learning service 105.

[0108] The machine learning architecture of EquiFormer as a system ( Figure 1)'s underlying architecture consists of three stacked models. The overall machine learning architecture is as shown in Figure 2 . In the first layer 210, there are two types of models: the sensor model 212 and the optimization model 215, which are used to model time series data, and their outputs are the predicted sensor values 217 and the optimization values 218 respectively. The optimization strategy 214 is built based on the optimization model 215, and the hardware control strategy 216 is a sub-component of the optimization strategy 214. In the second layer 220, there is an anomaly event model 222. In the third layer 230, there is an action recommendation model 232. Each layer is built using the output of the previous layer as input. Figure 2 The sensor model 212 and the optimization model 214 in Figure 1 support 110 and 130 in Figure 2 The hardware optimization 214 and control 216 in Figure 1 support 120 and 140 in Figure 2 The anomaly event model 222 and the action recommendation model 232 in Figure 1 support 150 in

[0109] Although Figure 2 organizes the model stack into three layers, the actual model stack organization is still flexible. Complex machine learning designs (such as EquiFormer) may require multiple layers of model stacking to have greater flexibility in predicting the future and to allow multiple target variables to be separated into multiple models or integrated into one model. Some machine learning techniques (such as the Transformer-based base model) are capable of generating any prediction sequence, including but not limited to sensor values, optimization targets, anomaly events, and actions. Therefore, Figure 2 the three-layer model in Figure 2 can be compressed into the base model, demonstrating the flexibility of the model stack organization. However, this application emphasizes the importance of the sensor model 212 in Figure 2 . When actual sensor data is not available and predicted values are required as input for subsequent models, the sensor model 212 predicts the sensor values. In some embodiments, the sensor model may be located in the independent layer 220 in Figure 2 , or integrated into the Transformer-based base model. In other embodiments, when actual sensor values can be directly used as input for the second layer, or the third layer, or both the second and third layer models, these models can ignore the output of the sensor model. However, the sensor model itself is still retained because it is still necessary in scenarios where actual sensor values are inaccessible.

[0110] Taking an industrial robot as an example, the non-timestamp specifications it measures may include: designed voltage and current inputs, the reach and load capacity of the robotic arm, degrees of freedom (joint flexibility), physical dimensions and weight, and the type of controller and programming interface. The timestamp sensor data may include: joint positions and velocities, the load on each joint or end effector, motor temperature, real-time energy consumption, ambient temperature and humidity (if applicable), and vibration and acoustic signals. The optimization objectives may include: operational efficiency (speed and accuracy of movement), energy efficiency, minimizing wear, and extending uptime and reliability. The abnormal events may include: mechanical joint failures, overheating of motors or electronics, calibration drift, accidental collisions or obstacles, and software or control system errors.

[0111] Section 2.3: Time Series Models for Sensor Data and Hardware Optimization

[0112] Some embodiments of this application model the following time series relationships: (1) The machine learning model uses the sequence data before time t < n, with inputs including time t_(t < n), all x_(l - t, t < n), all y_(si - t, t < n), and all x_m, and predicts the i-th sensor data y_(si - t, t = n) at time t = n. This is called the sensor model. (2) The machine learning model uses the sequence data before time t < n, with inputs including time t_(t < n), all x_(l - t, t < n), all y_(si - t, t < n), all additional y_(oj - t, t < n), and all x_m, and predicts the j-th optimization objective y_(oj - t, t = n) at time t = n. This is called the optimization model.

[0113] In the sensor model, there are i output variables at time t = n. In some embodiments of this application, the selection of the time series model is flexible. Different embodiments can select different types of time series models to represent these relationships.

[0114] Furthermore, some embodiments of this application can select a time series model with only one output variable for each model; thus, i models are established. Alternatively, some embodiments can skip establishing i models because a deep learning model can have i heads (i.e., i output variables). In some embodiments of this application, the system design does not fix the specific model selection. The same is true for the optimization model. For example, it is not necessary to establish j models as long as the modeling technique can predict j output variables.

[0115] Once the sensor model and the optimization model are constructed, regardless of whether the selected model technology is traditional time series technology or new deep learning technology (such as Transformer), these models can not only generate the target value at time t = n, but also generate the target values at times t = n+1, n+2, ..., n+Delta_t, applicable to y_(si-t) and y_(oj-t). This is achieved through traditional or deep learning-based generative artificial intelligence.

[0116] Although other time series technologies are also capable of generating such sequences, they have not been widely applied to manufacturing equipment data, which may be because: (1) Sequences generated by non-deep learning, non-Transformer methods tend to deviate from the real data of manufacturing equipment, which is to some extent due to (2) their lack of the ability to model long sequences of input data, and (3) their inability to model the collinearity of multiple inputs and the correlation of multiple outputs.

[0117] For example, in the case of anomaly detection, traditional statistical methods (such as z-score) only require 1 data point to test whether the data point deviates from a population metric such as the mean. Those skilled in the art generally believe that obtaining one target data point at time t is sufficient for decision-making, such as sending an alert.

[0118] For optimization objectives, those skilled in the art tend not to use sequence modeling techniques, but rather tend to compress the time dimension to simplify the modeling process.

[0119] In various embodiments of the present application, traditional time series technology can be used to generate sequences of manufacturing equipment data, and the generated sequences can then be used as inputs for the prediction of anomaly event models and action recommendation models. However, through innovative Transformer-based modeling technology, time series sequence modeling can be improved.

[0120] The method of generating the target value after time t = n (for example, at time t = n-1) is to recursively input the next predicted target value into the model, as Figure 3 shown. Figure 3Taking the sensor model as an example, the process of generating this sequence is described intuitively. For example, at time t = n - 1, a person skilled in the art can obtain the predicted value y_(si - t, t = n) at time t = n by inputting the real data 310 into the sensor model 330. In the next iteration, the predicted y_(si - t, t = n) is used as the value of the actual sensor data at time t = n and input into the sensor model again, so as to obtain the predicted value y_(si - t, t = n + 1) at time t = n + 1. Even when still at time t = n - 1, a person skilled in the art has not obtained the real value y_(si - t, t = n) at time t = n. If the value of a certain input variable x_(l - t, t = n) at time n is unknown when the actual time is n - 1, the technician can based on the best estimate value. For example, if the variable is the predicted temperature from a weather forecast service, the predicted values for the next few days can be used; or in traditional machine learning, it is regarded as missing data and filled with historical values (such as mean, maximum or minimum). By repeating this process for the next predicted time step, the technician will obtain a sequence of predicted values: y_(si - t, t = n), y_(si - t, t = n + 1), y_(si - t, t = n + 2), etc. This sequence 320 is the sequence of sensor target values generated by the machine learning model. The process of the optimization model generating the optimization target sequence is similar: just use the y_(si - t, t >= n) value predicted by the sensor model as the input and provide it to the optimization model.

[0121] In modern machine learning service practices, a person skilled in the art or an automated process needs to collect initial data for the initial training of the model, which corresponds to 110 in Figure 1 Over time, more data will be collected and the model will be updated, which corresponds to 130 in Figure 1 This is also the reason why the same sensor model and optimization model appear twice in Figure 1 The online model update part ( Figure 1 130 in

[0122] Section 2.4: Service Layer

[0123] To utilize Figure 1 the EquiFormer system in Figure 4The online service layer shown in [Figure] can serve as the application layer of the EquiFormer system. For the convenience of readers' understanding, the description of the service layer is arranged after the content of the first layer of the model, that is, it is described in Section 2.4.

[0124] The service layer 420 uses Internet of Things (IoT) facilities 410 instead of isolated device data collection, as described in Section 2.4. A large number of different hardware devices provide more information for the input variable group x_m ( Figure 2 in [Figure]). The input and output relationship models from a large number of devices provide baseline or large model information for specific devices through the service layer. It utilizes numerous devices connected to a larger model (EquiFormer) via the IoT and uses the large model to generate content (sensor data and minimized sequences) for specific inputs (specific devices). The sources of IoT devices and facilities can be diverse, ranging from manufacturing facilities, power systems, transportation systems, automobiles, building monitoring systems, security systems to any other devices. Figure 4 Only some examples of IoT devices are listed. However, this does not limit the scope of this application. The types of sensor information and sensors can also be very diverse and are not shown in detail in Figure 4 for simplicity.

[0125] The data service 430 first serves as a key component to provide data support for the modeling service. It is responsible for data processing 432, data encryption 434 (since security is always crucial), data storage 436, and data analysis 438. The sub-components of the data service 430 are not limited to the above examples (432, 434, 436, 438). Other appropriate data-related sub-components can also be included. In addition, it also provides its own utilities through, including but not limited to, a data panel 462, an API interface 464, a query function 466, etc.

[0126] Figure 4 The machine learning service 440 in [Figure] corresponds to Figure 1 106 in [Figure], and in particular, Figure 4 the online machine learning service 450 in [Figure] corresponds to Figure 1 107 in [Figure]. It provides utility functions through its sub-services. For example, the action service 452 and the exception event service 454 can send alerts and provide APIs. Some examples of API usage include but are not limited to: (a) controlling hardware devices when an abnormal event is predicted or a suggestion is provided; (b) passing feedback and data back to the machine learning and data services. The hardware control optimization service 456 not only determines and continuously updates the optimal hardware parameters, but more importantly, when the hardware parameters deviate from the optimal values, it calculates corrective measures and continuously controls the hardware through the API to minimize the deviation. Figure 4 The sensor and optimization service 458 in [Figure] is based on Figure 2Sequence modeling of 212 and 215 in the middle. Although their main task is to provide prediction data for downstream services (i.e., Figure 4 456, 454, 452 in it), they also need to provide their own utility services. They can provide functions through an interface called "EquiFormer large model API" 476, such as viewing the model, embeddings, model parameters, model weights, embedding comparison, model update, and fine-tuning, etc. A description of the embeddings will be further provided later.

[0127] In modern software architectures, modular design is usually adopted and based on a microservices architecture. Figure 4 Only limited API and service examples are shown in it. In actual applications, technicians can expand the API at any time or add more services when appropriate. Similarly, Figure 4 Only limited service sub-components are shown in it. In actual applications, those skilled in the art can expand, separate, or add more sub-components at any time to meet the requirements.

[0128] Section 2.5: Hardware Parameter Optimization

[0129] In the terms of machine learning, parameter optimization refers to searching for a combination of machine learning model parameters such that the predicted target value best matches the true target value. What is discussed in this section is not the parameter optimization of machine learning. In hardware, parameter optimization refers to searching for values of hardware-specific settings to maximize the optimization objective. The theme of this section is an innovative method to find the best hardware parameters in a virtual environment by using a machine learning model.

[0130] The following is the process of how to find the best hardware parameters by optimizing the model. In the optimization model in some embodiments of this application, the settings of the hardware are not the parameters of the model but the inputs of the model. They can be included in x_m, such as the designed number of turns of the coil; or sometimes included in x_(l-t), such as the temperature of a certain part of the hardware that can be set and changed at different times.

[0131] Once the optimization model is provided, for each target, the values of x_m and x_(l-t) related to the parameter settings can be virtually changed in a simulation environment, while combining the values of real, historical, simulated, or predicted t, x_(l-t,t<n), y_(si-t,t<n), and y_(oj-t,t<n). Then, the model will generate a sequence of predicted target values y_(o_j-t,t≥n). Then it can be observed which combinations of the values of x_m and x_(l-t) related to the parameter settings can produce the most ideal y_(oj-t,t≥n). A common application is to maximize the mean value of y_(oj-t,t≥n).

[0132] During the parameter optimization process, the strategy of changing its value within the space of x_m and x_(l-t) is called the parameter search strategy. In some embodiments of the present application, the parameter search strategy is flexible. It should be emphasized that some parameter search strategies that are not commonly used in the hardware / manufacturing industry but are commonly used in machine learning practices can also be applied to hardware / device optimization within the framework of some embodiments of the present application. There are three common machine learning parameter search techniques that can be applied to some embodiments of the present application.

[0133] First, it can be grid search, which is carried out by traditional equipment manufacturers in a physical laboratory environment and is usually used before the equipment is put into actual production.

[0134] Various embodiments of the present application provide the following innovative aspects: (1) Transferring this physical laboratory test process to machine learning-based virtual simulation theoretically saves time and resources. Traditional hardware engineers usually perform optimization in a physical laboratory. They actually change the design details or settings of the device and measure the actual optimization target value. After they exhaust the search space (i.e., all combinations of hardware parameter values that can be set and tested), they will select the combination of x_m and x_(l-t) that can generate the best y_(oj-t,t≥n) value. (2) Some embodiments of the present application use a time series model to predict the optimization target value. While traditional hardware engineers tend to use explicitly parameterized mathematical formulas to summarize their data rather than machine learning models. Then, they expand the search space based on the mathematical formulas they summarize. They usually conduct physical tests in a controlled laboratory environment with fixed conditions (such as temperature control intervals, etc.) and try to find the best settings for the optimization target, ignoring the fact that devices operating in the industry constantly face changing conditions in the actual environment. This is why the optimization target value of the device may perform well in the laboratory, but usually has different performances in terms of the optimization target in industrial use in the actual environment. Various embodiments of the present application consider the change of the y_(si-t) value when predicting the optimization target value y_(oj-t). Some machine learning techniques (such as deep learning) can achieve a general approximation of the relationship between the input and the target value, and this relationship is difficult to represent by explicitly parameterized mathematical formulas. Some embodiments of the present application are not restricted by the choice of specific modeling techniques and apply machine learning techniques to the hardware device optimization process.

[0135] The second strategy is random search. When there are more parameters to be searched than one, the number of combinations in the search space will increase rapidly according to the combination formula. If it is necessary to reduce the number of combinations of parameter values (x_m and x_(l-t)) used to obtain the predicted or actual optimized objective value, some embodiments of the present application propose to generate combinations of parameter values in a (pseudo) random manner. This random search strategy is common in machine learning optimization. Various embodiments of the present application innovatively apply random search to hardware parameter optimization.

[0136] The third search strategy is the Sequential Model-Based Optimization (SMBO) method, which is an improvement over random search. Sequential Model-Based Optimization (SMBO) is used to optimize functions with expensive evaluations. Especially in cases where each function evaluation requires a large amount of time or resources, such as tuning the hyperparameters of a machine learning model, SMBO is a very useful method. When the cost of function evaluation is extremely high, such as training a large machine learning model, SMBO is particularly powerful. It is widely applied to the hyperparameter optimization of machine learning models.

[0137] Different from the next set of combinations after randomly selecting seed combinations, the Sequential Model-Based Optimization (SMBO) method infers the next set of combinations based on a machine learning model. Some embodiments of the present application adopt this strategy because if the optimization model is based on deep learning, even virtual evaluation requires high computational costs. The search strategy based on SMBO is common in machine learning parameter optimization but has not been widely applied in hardware parameter optimization.

[0138] The Tree-structured Parzen Estimator (TPE) belongs to the family of Sequential Model-Based Optimization (SMBO) methods. TPE is applied to the Optuna TM framework (Akiba, Sano et al., 2019), while its variant, Adaptive TPE, is used in the HyperOpt library (Bergstra, Yamins et al., 2013).

[0139] The Tree-structured Parzen Estimator (TPE) is an algorithm for machine learning hyperparameter optimization. In a high-dimensional hyperparameter space, TPE is more efficient than other hyperparameter optimization methods (such as grid search or random search). TPE can also effectively handle hyperparameters with non-uniform and conditional distributions.

[0140] In some embodiments of the present application, Optuna TM (an open-source hyperparameter optimization framework) can be adopted for machine learning. Optuna TMProvides a user-friendly yet powerful way to automatically search for the optimal hyperparameters of a machine learning model and helps to efficiently find high-quality solutions in a relatively short time.

[0141] The term "adaptive" in Adaptive TPE means that the algorithm dynamically adjusts its approach as it learns more about the hyperparameter space. As the optimization process progresses, the algorithm becomes better at predicting which hyperparameters are likely to lead to better performance and focuses the search on the most promising regions of the hyperparameter space. In the HyperOpt library, Adaptive TPE is used to efficiently find the best hyperparameters for a given machine learning task. This is particularly useful when dealing with high-dimensional spaces and complex objective functions, as traditional methods (such as grid search) become computationally infeasible in these cases.

[0142] Users engaged in hardware-related work usually cannot directly use libraries such as HyperOpt or Optuna TM in hardware parameter optimization problems because these libraries are designed specifically for optimizing supported machine learning models.

[0143] However, various embodiments of the present application can apply the process principles of HyperOpt or Optuna TM to the hardware parameter search strategy. TPE is a sequential model-based optimization (SMBO) method. In machine learning hyperparameter tuning, TPE models P(x|y) and P(y), where x represents the hyperparameters and y represents the associated loss (objective function), and then selects the x that minimizes the expected value of y. It divides the parameter space into two regions based on the observed loss values and then preferentially samples from the region with lower loss. In some embodiments of the present application, y can be replaced by the objective function y_(oj-t), and x should be replaced by the parameter values x_m and x_(l-t). Depending on the specific j-th hardware optimization objective, some embodiments of the present application can aim to maximize or minimize y_(oj-t). Depending on the specific optimization objective, sampling can be preferentially from the region where y_(oj-t) is larger or smaller. Some embodiments of the present application can also use the principle of Adaptive TPE to dynamically adjust the number of samples and balance exploration and exploitation during the optimization process. Various embodiments of the present application innovatively apply SMBO to hardware optimization.

[0144] Typically, in traditional manufacturing processes, once the best set of parameters is found in a laboratory environment, these parameters are set in the hardware and remain unchanged. However, according to the system design of some embodiments of the present application, its online service part, such as Figure 1As shown. In real-time industrial production, an online machine learning model has the ability to suggest new parameter values that may go beyond the range verified in a physical laboratory because the optimization model is updated online. These newly suggested sets of hardware parameters may not be immediately trustworthy enough to be directly set into the hardware. The allowable update range of parameter values can be set in the machine learning service ( Figure 1 140 in) to ensure that parameter values within these ranges can be updated by the online machine learning model during production.

[0145] Rules can also be set to specify that parameter values within certain ranges are not allowed to be directly fed back to the hardware, but instead the newly suggested parameter values are tested in parallel and submitted to the physical laboratory to verify whether the suggestions are accurate. Through this method, the overall EquiFormer system can accelerate the iteration cycle of product development in the physical laboratory while reducing production risks. The development cycle of traditional manufacturing equipment lacks machine learning-recommended parameters and real-time industrial data feedback. The development process almost always goes from the laboratory to the industrial site, and it is difficult to achieve a parallel process between laboratory development and industrial applications.

[0146] Taking a chip manufacturing device (such as a lithography machine) as an example, its non-timestamped specifications that can be measured may include: power requirements (voltage, current), device size, wafer size compatibility (e.g., 300mm wafers), light source type and intensity (for lithography machines), resolution and overlay accuracy, throughput (number of wafers processed per hour). Timestamped sensor data may include: wafer temperature and humidity, vibration and stability measurements, light intensity and wavelength (for lithography machines), real-time power consumption, chamber pressure (for vacuum processes), and the positioning accuracy of the wafer stage. Optimization goals may include: yield (percentage of qualified chips on each wafer), accuracy and repeatability of patterns, throughput (maximizing the number of wafers processed), minimizing defects and contamination, and energy and resource efficiency. Exceptional events may include: misalignment or pattern errors, device vibrations affecting resolution, wafer contamination, light source failures (for lithography machines), and vacuum system failures (for deposition or etching equipment).

[0147] Section 2.6: Exception Event Model

[0148] In machine learning practice, industrial anomaly events (such as anomalies caused by faults, cyberattacks, natural disasters, etc.) have unique modeling challenges and characteristics: such anomaly events are relatively rare in historical data. In traditional supervised learning methods, this is manifested as the scarcity of labels for such anomaly events. In generative artificial intelligence methods, this is manifested as the scarcity of tokens representing such anomaly events in the corpus. An obstacle in the past was that devices were isolated from each other, resulting in insufficient anomaly events being collected to train machine learning models. According to some embodiments of the present application, the innovation in system design is mainly reflected in combining hardware and machine learning infrastructure and applying this infrastructure to anomaly event modeling. For manufacturing devices, device data collection services and machine learning services can be deployed in the cloud, and requests can be sent to the Internet of Things (IoT) infrastructure to send sensed event data (including anomaly events such as faults) to the cloud.

[0149] Through cloud sharing, more anomaly event data can be obtained to build models. After the anomaly event models are built, these models can be shared with other devices through the cloud. Therefore, even if a certain device in a factory has never had a fault (such as "fault" in anomaly events), but if another similar device has had a fault and the anomaly event model has been trained in the cloud, then this anomaly event model can still be applied to the device that has never had a fault to predict its future faults. This is similar to the application of machine learning models in cybersecurity. Even if a new type of attack has not been discovered in the United States, but as long as this attack has been discovered elsewhere, the attack model has been trained and shared in the cloud, this attack model can predict possible attacks in the United States. Figure 4 Shows the integration of IoT into the EquiFormer machine learning system.

[0150] The following is a machine learning method for a survival classifier for anomaly event models (not the Transformer method in Section 3). In machine learning, describing anomaly events in machine learning with the word "rare" is relative. The goal is to determine whether the anomaly event y_(ek - t, t = n) occurs at time n, where it takes the value 1 when it occurs and 0 when it does not occur. At each time step within a given time period, if an anomaly event does not occur most of the time, then this anomaly event is rare. If the number of anomaly events in the historical data is much less than 100, the new method described in Section 3 can be used to model and predict anomaly events. For anomaly events such as hardware failures, their occurrence frequencies may be low, but in the shared historical data, according to empirical estimates, they still occur at least more than 100 times, and traditional machine learning techniques can still be used, such as Figure 5The survival classifier model shown. During training, the input to the historical anomaly event model includes the sensor input y_(si-t, t=n-1) at the previous time step n-1, the time-sensitive non-sensor input x_(l-t, t=n-1), and x_m that varies across devices. The survival classifier needs to organize the input and output data in its own way. Each row in the data represents a particular device at a specific time step (t from 0 to n), its input variables, and the target historical anomaly event status. Once a device fails (does not survive) at time step n, it is removed from the data at time step n+1. Due to this removal requirement, each anomaly event will have its own model, rather than a single model that includes multiple anomaly event target labels (except for the new method described in Section 3). To address the problem of imbalanced target data, it is only necessary to carefully sample and weight the data points where historical anomaly events occur and those where no anomaly events occur. In addition, when using the survival classifier machine learning method, the training data actually spans the previous few time steps across rows, as shown in the Time Step column in Figure 5 . In a specific row, Figure 5 the label in the Target column in

[0151] is the direct next time step of the input time step. Each row in each time step may be important in this solution and will be discussed in more detail. Figure 2 The overall architecture that includes a multi-layer stacked machine learning model in some embodiments of the present application provides unique advantages for the use of the anomaly event model.

[0152] The architecture in some embodiments of the present application can predict anomaly events for multiple future time steps. This ability to predict anomaly events for multiple future time steps: (1) provides Figure 4The service layer and human administrators in [description] provide sufficient response time; (2) It provides the ability for the action recommendation model to output predicted actions for multiple future time steps. On the other hand, if the anomaly event model can only receive actual data (such as sensor data for the previous few time steps up to the most recent time step) as input, and a catastrophic anomaly event is predicted to occur only one time step later, then there will not be enough time to take corrective measures or actions to change the device state to avoid such anomaly events. In a machine learning architecture lacking the first layer model (especially the sensor model), or if the sensor model is not a sequence model containing a time series model, sensor input data for multiple future time steps cannot be obtained. In this case, relying solely on actual sensor input and attempting to predict anomaly events using survival classifier machine learning methods will inevitably encounter the problem of being unable to predict anomaly events for multiple future time steps. During training, if the anomaly event label is in the next immediate time step, it is obvious that the anomaly event prediction based on actual sensor data can only predict up to the next time stamp. During training, if the historical anomaly event label is several time steps later, the anomaly event model may not be able to capture significant changes in the input sensor data before or a short time step before the occurrence of the anomaly event, and thus may not be able to predict anomaly events several time steps later.

[0153] The novelty of applying the survival classifier model method to device anomaly event modeling includes: (1) It is applied to the sequence of device sensor data and uses a survival classifier to model anomaly events in the next immediate time step. Device sensor data is usually modeled using time series modeling methods. However, in the designs of others, the anomaly event is part of the time series model, that is, \(y_{(ek - t)}\) is regarded as \(y_{(oj - t)}\). From a machine learning perspective, \(y_{(oj - t)}\) usually has a value and varies within a reasonable range, that is, its observed values form a random distribution. In this case, there is no problem with time series modeling. However, applying time series methods to anomaly events may inevitably encounter the problem of target data imbalance and lack of a good solution. By using a survival classifier on the time series model, elimination and sampling can be performed at the row level, thus providing additional tools to handle the target data imbalance problem. Other aspects of novelty may include: (2) Generally, the survival classifier modeling method is applied to patient / disease survival data, rather than hardware device data. (3) A novel Transformer model specific modeling technique for the anomaly event model will be discussed in Section 3.3, which can even model fewer event occurrences.

[0154] Section 2.7: Action Recommendation Model

[0155] After predicting an abnormal event, subsequent actions can be provided through hard-coded logic or different machine learning models. Examples of such scenarios include, but are not limited to: if the predicted sensor value (such as temperature, amplification value, etc.) exceeds a certain threshold, or it is predicted that an abnormal event will occur (such as a certain component of the device will malfunction), then trigger an action (such as replacing the component, increasing the maintenance program, adding lubricating oil, etc.). Traditional hard-coded logic (not within the scope of the claims of this application) can only encode relatively simple "if-then" relationships. This application will skip the details of possible traditional hard-coded logic and focus on machine learning solutions.

[0156] Figure 6 Shows the organization of the input and target data of the action recommendation model. According to the content of this application, a person skilled in the art can apply the machine learning modeling techniques of the recommendation system in an innovative way to output predicted actions in manufacturing equipment. There are p available actions y_ap that actually occurred in the factory before, and they are target variables, only taking values of 0 or 1. The inputs of all sensor models and all historical abnormal events that occurred are input into the action recommendation model. The choice of actual machine learning modeling techniques is flexible. In this application scenario, these p available actions can be modeled by any multi-class binary classification model, such as tree-based models and logistic regression models, multi-layer perceptron (MLP) deep learning models widely used in the online advertising industry (such as DIN, Wide&Deep, FNN, DeepFM, AFM, NFM, FM, etc.) (Wang 2020), or sequence-based deep learning models such as LSTM (Pan, Sheng, etc., 2019), etc. After the multi-class binary classification model assigns a numerical value to a specific action, the strategy for actually completing the predicted action is flexible. For example, for actions to prevent critical failures, the prevention strategy can adopt an over-designed approach; for routine maintenance actions, the strategy can be optimized under cost-effectiveness and resource constraints.

[0157] Alternatively, a person skilled in the art can also regard this problem as a graph modeling problem and use deep learning techniques on the graph (Zhang, Cui, etc., 2020). The devices are used as nodes, and the same type of actions (such as oil change maintenance) actually performed by the same entity (such as a repair professional company) on different devices form edges. In addition to Figure 6Beyond the content described in [reference], the input of the multi-class binary classification model described in the previous paragraph can also be node embeddings generated from a graph modeling perspective. In this application, an embedding refers to a vector output during the process of generating a vector through deep learning. A person skilled in the art can compress the time dimension in the graph, that is, only construct a smaller number of graphs compared to the large number of time steps in the data, and each graph does not change over time relative to the next graph. Alternatively, a person skilled in the art can also construct a graph that changes over time, that is, a time-related graph. Although each snapshot of the graph may contain aggregated information from multiple time steps, a certain relationship is assumed to exist between the snapshots of the graph on the time axis. In this application scenario, instead of predicting node classification, the formation of links is predicted (DGL_Team2018), as Figure 7 shown. In the graph, there are devices 710, 712,..., 718. In a snapshot of a time-compressed graph or a time-related graph, it is known that devices 710 and 718 have formed link 720 because they have previously performed the same action. Similarly, there is also a link 722 between devices 710 and 712, but device 714 has not formed a link with other devices. If it is predicted that device 718 will form a new link with any of the devices with existing links (such as 710, 712, or 714), then device 718 is predicted to perform the same action as that for forming the link.

[0158] The novelty of some embodiments of this application lies in that this graph link formation prediction method has not been applied to the maintenance and fault prevention problems of hardware devices. Generally, graphs are mainly used in scenarios such as drug discovery, social networks, and web links.

[0159] After applying supervised modeling techniques with multiple output prediction actions, an ensemble layer can also be added to see which action received the majority vote among multiple models. More variations and implementations will be provided after introducing the Transformer in the section on EquiFormer, and specific Transformer modeling techniques will be applied in Section 3.4 to provide action recommendations.

[0160] Section 2.8: Hardware Control System

[0161] The hardware control system is as Figure 1 shown. It is a key component responsible for: (1) executing an optimization strategy by maintaining the hardware at optimal operating parameters, whether through real-time or scheduled adjustments; (2) performing predicted abnormal event intervention actions by pre-adjusting device operations to prevent or mitigate the impact of such events. These two execution strategies (sequences of actions or operations) are both formulated by machine learning architectures and services. By adjusting the operating parameters of the device, the control system can effectively change the operating state of the device to make it consistent with the predicted optimization goal and alleviate potential abnormal events.

[0162] Broadly speaking, Figure 1 the entire EquiFormer system in Figure 1 is a hardware control system that uses machine learning to control the hardware device 164. Narrowly speaking, the hardware control system includes a hardware administrator 154, an application service 150, an optional communication component 168, a controller 166, and an Internet of Things (IoT) sensor 162 that provides data feedback. In this article, the definition of the hardware control system adopts a narrow interpretation.

[0163] The application service 150 includes two main components: (1) an optimization component that receives input signals from the hardware optimization model; (2) an exception event intervention component that receives input signals from the exception event model and the action recommendation model. For example, the predicted exception events may be mechanical failures, overheating, or accidental collisions. The recommended actions may include, for example, replacing mechanical parts to avoid mechanical failures, reducing the motor speed to prevent overheating, adjusting the trajectory to avoid collisions, or recalibrating sensors to cope with environmental changes.

[0164] (1) Signal comparison: Compare the input signals from the machine learning model with the current state of the hardware. This involves evaluating the difference between the optimal parameters or intervention parameters predicted by the machine learning and the actual operating parameters of the device 164.

[0165] (2) Generate software control signals for the device: Based on the comparison results, the application service 150 generates executable software signals that can modify the operating parameters of the device. These software signals indicate to the controller how to change the state of the hardware operation so that it can make downstream adjustments based on the predicted insights. The software-generated control signals can, for example, take the form of key-value pairs in JSON or YAML format. For example: {"action": "increase", "object_id": "motor_xyz", "parameter_id": "motor_speed_abc", "parameter_unit": "RPM", "amount": 10}.

[0166] (3) Generate computer hardware control signals: To transfer the control signals generated by software to hardware, the application service 150 converts these signals into computer hardware control signals and sends them out through computer ports. Although the entire data service 104, machine learning service 106, and application service 150 can reside in the cloud, the terminal where the application service 150 interacts with the hardware layer 160 needs to generate signals that can be understood by the hardware layer. For example, the software control signal {"action":"increase", "object_id":"motor_xyz", "parameter_id":"motor_speed_abc", "parameter_unit":"RPM", "amount":10} must be converted by the application service 150 into a computer hardware control signal. For example: At pin a of port x, at time step 0, a high voltage (e.g., 5 volts) represents the "increase" action. At pin b of port x, at time step 0, a high voltage represents the operation object ID of "motor_xyz". At pin c of port x, at time step 0, a high voltage represents the adjusted parameter ID of "motor_speed_abc". At pin d of port x, at time step 0, a high voltage represents the parameter unit of RPM. At pin e of port x, at time steps 1 - 100, a high voltage (5 volts, equivalent to the binary "1" in the 0 and 1 system) at each time step represents an increase unit. Therefore, a signal sequence "111111111100000...0" represents an increase of 10 units. This is using a signal sequence with limited pins over a long time to represent a quantity. If time is limited but there are many pins, the decimal quantity can be represented by binary pins at time step 0, such as using "1010" on pins efgh to represent 10 units. It can also be mixed.

[0167] In the hardware layer 160, the optional communication component 168 performs the following functions:

[0168] (4) Relay software control signals to the controller: If the controller is not directly connected to the computer port, an optional communication unit 168 is added in the hardware layer 160. This section discusses communication using WiFi or local area network (LAN). Although Bluetooth can also be used to transmit control signals, it uses different protocols and mechanisms. For the sake of brevity, we will focus on WiFi and LAN and point out that the implementation of Bluetooth requires adaptation to the Bluetooth protocol stack. Communication components based on quantum communication will be discussed in Section 4.

[0169] After the application service 150 serializes the software control signal into a JSON, XML, or binary protocol suitable for transmission via WiFi or LAN, the communication component 168 encapsulates the signal into a network TCP or UDP packet with the necessary protocol headers (e.g., TCP / IP headers for LAN or WiFi communication). The communication component 168 routes the packet to the appropriate controller on the network using the IP address. These packets are transmitted via WiFi or LAN using standard wireless or wired signal protocols.

[0170] In the hardware layer 160, the controller component 166 performs the following functions:

[0171] (5) Generate physical device control signals: After receiving control signals from the communication component 168 or the application service 150, the controller 166 processes these signals to extract control instructions. For example, the controller 166 may first extract the signal: interpret the "action" field to determine that an increase operation needs to be performed; use the "object_id" to identify "motor_xyz" as the target device; confirm the "parameter_id" corresponding to "motor_speed_abc"; apply the "parameter_unit" of RPM to ensure the correct measurement unit is used; and the "amount" value to increase the motor speed by 10 units.

[0172] Subsequently, the controller converts these instructions into physical device control signals that directly control the device. A remotely configurable and programmable photon programming processor is a type of controller, which is described in detail in Section 4. In this example, if the device is a 4-pole three-phase AC motor, increasing the frequency of the AC power supply will increase the motor speed. The controller may consist of Variable Frequency Drives (VFDs), which increase the AC supply frequency by 0.333 Hz according to the synchronous speed formula by sending physical signals. Which specific motor is controlled may depend on the wiring and configuration of the controller.

[0173] The Internet of Things sensor (IoT sensor) 162 provides the following functions:

[0174] (7) Collect device data to form a feedback loop: The IoT sensor 162 within the hardware system 160 collects real-time data from the device 164 and feeds it back to the data acquisition and storage component 104, while directly feeding it back to the human hardware administrator 154 and the application service 150. This real-time feedback enables iterative adjustment of the control signal without relying on the computationally intensive processes of the device data service 104 and the machine learning service 105. For example, if the motor fails to reach the expected speed due to unforeseen limitations, the application service 150 can quickly send a new control signal to adjust the operation of the device accordingly.

[0175] The control system supports the following dual-mode operations.

[0176] (8) Operations in autonomous mode and manual intervention mode: The control system is versatile in operation and can perform operations / interventions without the explicit input of the manual hardware administrator 154, or allow manual intervention when needed. In scenarios with a priority on autonomy, the system can autonomously implement intervention operations to prevent equipment failures or optimize performance. Especially when the underlying machine learning service can foresee abnormal events and intervention operations in long-term predictions, many interventions can be automatically planned and executed without manual intervention. When the predicted failures or intervention operations require operations faster than human response, the control system can also execute independently and automatically. On the other hand, the system can display alerts and suggestions through the display screen 152 or other interfaces, allowing the manual operator to review and manually intervene if necessary. This dual ability ensures a balance between the automation efficiency and human judgment of the system, adapting to different scenario requirements. This dual-mode operation allows multiple ways to achieve the same operation. For example, if the recommended intervention operation for 152 is to reduce the speed of the motor to zero, there are the following three ways: (1) Manual direct operation: The hardware administrator can directly turn off the switch of the motor. In this case, the manual operation takes precedence over the automatic control system. This is similar to pressing the power-off button of a robotic cleaner. (2) Controller signal operation: The controller can send a signal to the motor to reduce the power voltage to zero. This is similar to the application program of a robotic cleaner. When it is determined that the cleaning is complete and the robotic cleaner is located at the base station, the application program automatically stops the motor. (3) Execution after manual confirmation: The application service can display the intervention operation to the manual administrator 154 and wait for a manual response, and then execute the operation with manual intervention. The manual administrator can confirm turning off the motor or modify the operation. This is similar to the application program of a robotic cleaner displaying a shutdown request and waiting for the user to confirm. The user can confirm shutdown or let the robotic cleaner continue to run without shutting down. However, in the content of this application, although the narrow-sense hardware control system is similar to a robotic cleaner in some aspects, the machine learning service and the entire EquiFormer system achieve broader optimization goals and abnormal event interventions.

[0177] The hardware control system utilizes Figure 1 the machine learning service 106 in Figure 2A hierarchical machine learning architecture consisting of a sensor model, an optimization model, hardware and control components, an anomaly event model, and an action recommendation model. These models provide predictive insights into the future behavior of the device, enabling the system to make informed adjustments. Specifically, the sensor model provides predicted sensor values, the optimization model provides an optimal set of device operating parameters, the hardware and control components provide the best action strategies, the anomaly event model predicts potential anomaly events, and the action recommendation model suggests specific intervention measures. Figure 1 The application service 150 in Figure 1 integrates these insights provided by the machine learning service 106 in

[0178] Section 2.9: Comparison with Equipment Service Monitoring

[0179] In some implementations, the "lookout-for-equipment" function is a service for warning of equipment failures. The modeling part of this function can be supported by online model refreshing and service provision. In terms of machine learning models, after predicting sensor values at future times, usually only statistical tests are used to determine whether there is a statistically significant difference between the predicted future sensor values and the sensor values at past times. If there is a difference, an alarm is issued.

[0180] According to some embodiments of the present application, the machine learning service is different from other "equipment monitoring" functions in many aspects.

[0181] For example, in some embodiments of the present application, the sensor model can select very flexible machine learning techniques. The application of the Transformer architecture in the sensor model of manufacturing data has not been foreseen before, and Section 3 of the present application further describes how to apply Transformer-based techniques to new applications and methods for manufacturing time series data.

[0182] For another example, in some embodiments of the present application, anomaly event alerts are determined by the machine learning model in Layer 2, rather than by statistical tests.

[0183] In addition, in some embodiments of the present application, an optimization model and an action recommendation model that do not exist in other implementations are also provided.

[0184] In terms of hardware, cloud services may lack specific designs for Internet of Things hardware in equipment monitoring services.

[0185] What chips / components need to be added to specific hardware (photon programming processor) to enable the normal operation of the Internet of Things machine learning service will be further described below.

[0186] In particular, to send anomaly alerts, some other implementations use severity scores and other scores, which are statistical test-based methods (e.g., when a new value is input, a one-sample test is performed to determine whether it belongs to a normal distribution). These implementations may claim that the accuracy of alerts can be improved by annotating anomalies, but this does not mean that they use the same machine learning modeling methods as some embodiments of the present application. Other statistical tests may also be used (e.g., when a new value is input, a two-sample test is performed to determine whether it belongs to a normal distribution or an abnormal distribution).

[0187] Some embodiments of the present application adopt machine learning modeling methods and provide a brand-new solution based on a large Transformer model for this specific problem in Section 3. At the same time, the solution provided here also takes into account the normal behavior of a group of hardware devices with similar but different specifications, rather than just the behavior of a single isolated device.

[0188] After detecting an anomaly, other types of services may not provide machine learning solutions at all. The user needs to provide downstream operations by themselves, either manually adding actions (e.g., sending a text message to a phone number) when an anomaly is detected in an anonymous function, or the user constructs a machine learning model by themselves. The present application provides a machine learning modeling solution for automatically performing actions that should be taken after detecting an anomaly.

[0189] Section 3: EquiFormer: Specific Application of Transformer in Hardware Devices

[0190] Section 3.1: Transformer as an Embedding and a Base Model

[0191] Transformer-based models have been used for modeling text, image sequences, and more recently for modeling time series (especially in the financial field, such as stock prices and retail data), but they may never have been applied to hardware device data and control. Since the sensor model and optimization model in the present application involve time series data, this powerful Transformer model can be applied to the hardware device modeling problem. The innovation of the present invention is that Transformer may not have been used in systems including hardware in some embodiments of the present application, especially in manufacturing data. Many embedding methods, fine-tuning techniques, and applications for modeling abnormal events will be modified for the first time in the present application to solve hardware device data or manufacturing problems that are difficult to solve by traditional methods or even the methods described in Section 2 of the present application. This section will further describe in detail the novel Transformer implementation in some embodiments of the present application.

[0192] Figure 8 shows a modeling architecture for manufacturing a Transformer-based base model for hardware modeling. (1) For illustration, in the Figure 8 input 820 and output 830, the window size of the time step is 8. In this application, the window size is flexible. Note that in Transformer, the position step is slightly different from the time step, but for simplicity, the existing subscript t can be used to represent the position step. In Transformer, the position step is the step size in the window size, and as the window moves over the time series data, the position step remains from 0 to 7 (taking the window size of 8 as an example), but the actual time step located in X_t will change as the window moves. For example, if the time steps of the time series are from 0 to 11 and the window size of the Transformer is 8, then: the first row of data in the Transformer window consists of time steps 0 to 7, and the second row of data in the Transformer window consists of time steps 1 to 8. (2) The architecture inside the Transformer is flexible, such as the number of layers, the position of LayerNorm, etc. (3) The input data X_t 850 is a vector or embedding of relevant and important input variables at its corresponding position (here, the time step in the time series). The exact order, selection, and embedding of the input variables assembled into X_t are flexible. (4) The target data Y_t 860 is a vector or embedding of relevant and important target variables at its corresponding position. The exact order, selection, and embedding of the target variables assembled into Y_t are also flexible. In Figure 8 for the purpose of illustration, the sensor information y_(si-t) is regarded as part of the target. However, given the flexibility of the Transformer model target vector, if the sensor model is not part of the use case, y_(si-t) can be part of the input vector, and the target vector can remove its y_(si-t) component. (5) The output 840 is the predicted target shifted one time step to the right.

[0193] This architecture refers to the Transformer model as an embedding model because, as a deep learning model, the output of its last layer (before the target output) is a vector that can be used as the embedding of the window starting from a specific position. In addition, this architecture also refers to the Transformer model as a base model for the following reasons: (1) As a deep learning model, Transformer is capable of modeling multiple labels in the target. This base model integrates Figure 1A model with a time series component, including a sensor model, an optimization model, an abnormal event model, and possibly an action recommendation model, forms a unified modeling structure. In the case where the Transformer model is used to predict sensor data and optimization objectives, regardless of whether the objectives in training also include objectives of abnormal events or action recommendations, Section 2.6 can be referred to for using non-Transformer-based methods for downstream abnormal event modeling and action recommendations. (2) It supports downstream application scenarios of abnormal event prediction and action triggering based on the Transformer method. It is the basis for all downstream application scenarios.

[0194] The embeddings of the input and target in some embodiments of this application have uniqueness: (1) In the Transformer model used in large language models (LLMs), both the input and the target are in the form of word embeddings. While in some embodiments of this application, Y can be a variable of a different category from X. For example, when X consists of a set of sensor data vectors, Y can consist of a set of optimization objective vectors. (2) Y can take various forms and embedding methods. This flexibility of Y in some embodiments of this application provides unique advantages compared with the application scenarios of large language models, enabling multi-modal methods to be applied to the application scenarios of EquiFormer. One of the current research directions of Transformer is multi-modal fusion, and the current research mainly focuses on targets composed of a mixture of text and image / video (Gal, Alaluf, etc., 2022). While the targets of EquiFormer are unique and different from existing multi-modal research: they include sensor data, optimization objectives, abnormal events, and actions. (3) In addition, if some sensor data can be converted into images, EquiFormer has the flexibility to use different forms of data as the input or target of the Transformer model. On the one hand, EquiFormer can use parameterized forms of sensor data; on the other hand, EquiFormer can adopt image snapshots, embed the images into a convolutional neural network (CNN), and then use the embeddings generated by the CNN as the input or target of the sensor data. Examples of such data can be sound waves, light waves, particle imaging, quantum state tomography, etc. Taking light waves as an example, the parameterized form can be the amplitude, frequency, etc. of each component wave, and the CNN embedding form can be the CNN embedding vector of the light wave image.

[0195] Compared with some multi-modal research, a CNN component is innovatively added in some embodiments of this application, which can convert non-image inputs / targets in manufacturing equipment data into image inputs / targets.

[0196] Section 3.2: Optimization and Its Control

[0197] In certain embodiments of the present application, one application scenario of EquiFormer is hardware optimization and control. Optimization can refer to finding the optimal input for a manufacturing optimization goal, while control can refer to adjusting the randomness of the input to the optimal input calculated from the optimization model.

[0198] There are various solutions to utilize the Transformer model to solve optimization problems:

[0199] (1) The method described in Section 2.4 can be used: Replace other models with the Transformer model described in Section 2.4, generate predicted optimization target values for future time steps by varying the hardware input, and use the predicted input values as the input to the optimization target when the input values for future time steps are not available. This method may be cumbersome when the search space is large.

[0200] (2) Another innovation for the Transformer model is to utilize Transformer embeddings. Figure 9 A simplified flowchart is shown to explain the process. This process can be referred to as retrieval optimization based on EquiFormer in certain embodiments of the present application. It should be noted that the embedding model can be the same as or different from the optimization Transformer model. However, both models need to be Transformer models. When the embedding Transformer model is different from the optimization Transformer model, the objective in the embedding Transformer model can be different from the objective in the optimization Transformer model. In the present application, the case where the current optimization objective is the same as the objective during the training of the optimization Transformer model is assumed.

[0201] (2.1) In the input 910 of the given search space, a person skilled in the art can generate input combinations using any method (grid search, random search, or SMBO) mentioned in Section 2.5 or higher. Then, instead of actually generating the target values from the Transformer model, a person skilled in the art can generate embedding vectors from the embedding Transformer model 920. Thereafter, these embedding vectors can be referred to as "candidate embeddings" 930. The reason for choosing to generate embeddings instead of target values may be that a person skilled in the art needs the embeddings and the target values to exist in different systems (e.g., an online device and an offline vector database 901), or may need to fine-tune the Transformer embeddings trained from one set of target values to a different set of target values, such as a changed optimization objective. Additionally, reducing one step in the Transformer calculation (from embedding to target value) may save some computational resources.

[0202] (2.2) In the existing historical input 990, the input variable vector can be input into the embedding Transformer model 980 to obtain the embedding vector. Thereafter, these embedding vectors can be referred to as "baseline embeddings". It should be noted that the embedding Transformer models 980 and 920 are the same embedding Transformer model. Again, note that the optimization Transformer model 950 can be the same as or different from the embedding Transformer model. The only "optimal" baseline embedding 970 is defined as the vector that generates the best optimization objective value 900 from the baseline embeddings.

[0203] (2.3) In the vector database 901, the distances of the candidate embeddings 930 are compared with the best baseline embedding 970. Sampled candidate embeddings 940 can be obtained, with small, medium, large distances (or any other sampling strategy based on distance sampling) between the candidate embeddings and the best baseline embedding. This step significantly reduces the number of candidate embeddings that need to be input into the optimization Transformer model 950 to evaluate the objective value. Note that one component of the objective value is the optimization objective. Theoretically, the candidate embedding with the smallest distance from the best baseline embedding should generate a similar objective value, while the candidate embedding with a larger distance from the best baseline embedding should generate an objective value different from the best baseline objective value. The embedding that generates the new best objective value among the candidate embeddings sampled by distance will become the new best baseline embedding 960.

[0204] (2.4) Repeat step (2.3), removing candidate embeddings with small distances from the evaluated candidate embeddings or baseline embeddings to determine whether a new best candidate embedding can be found to replace the current best baseline embedding, until the search space is exhausted or the iteration limit is reached. The input corresponding to the best candidate embedding is the best input for the found hardware.

[0205] (2.5) If there are multiple best baseline embeddings, one can choose to loop through each best baseline embedding one by one, or use any parallel computing method to parallelize the search process.

[0206] In terms of control, after finding a set of optimal values for the hardware inputs, there will always be randomness in the input values of each input, making its actual value unable to exactly equal the optimal value. Therefore, a control value is needed to adjust the randomness of the input values to make the input values as close as possible to the optimal values. In the explicit formulas of traditional control theory, it is difficult to optimize the parameters to find the optimal control value, and the final control value is usually a linear combination of multiple control value components. In some embodiments of the present application, if a Transformer-based base model is used (note that the Transformer model is not only an embedding model but can also serve as the base model for all time series models, including sensor models, optimization models, etc.) or other time series models, and the target or output contains a component of the sensor input data y_(si-t), then this problem can be solved by an innovative numerical method. For any randomness (the change Δ in the input value) added to y_(si-t, t<n), the future values of y_(si-t, t≥n) will be accurately predicted by the Transformer model. As a special deep learning model, the Transformer model has the property called universal approximation, which means it can approximate any explicit mathematical formula. Therefore, there is no need to rely on linear combinations to approximate the control value.

[0207] The embodiments of the present application solve a difficult problem faced by traditional control theory when using explicit formulas, that is, the user needs to know the exact parameter values that the explicit formula should adopt to approximate the future values of the input.

[0208] According to some embodiments of the present application, the future values of the input can be accurately predicted without knowing these parameter values. After knowing the future values of y_(si-t, t≥n) at each time step and comparing them with the optimal input value y_si* of this y_(si-t), the control value can be easily calculated using any method of control theory. For example, those skilled in the art can use the slope of the predicted values of the machine learning model changing with the time step to approximate the differential used in control theory, and use the area of the predicted values of the machine learning model changing with the time step to approximate the integral used in control theory, without knowing the exact parameter values or the formula itself in the explicit formula. In addition, this method can simulate multiple y_(si-t), and their interactions are completely processed by the base model.

[0209] This application not only introduces machine learning models that can be used for time series data, but also provides a machine learning architecture system for solving data problems of hardware devices (not only predicting conventional sensor data), and how to stack data and machine learning models together to enable the system to solve hardware device problems. The architecture consists of three layers of models, and the first layer solves three problems. Regarding the hardware optimization problem of the first layer, a new machine learning-based control scheme is also provided. Note that this new control scheme is not limited to the Transformer-based optimization model, but also applicable to any machine learning-based optimization model.

[0210] Section 3.3: Anomaly Event Prediction Based on Transformer Model

[0211] Anomaly event prediction in traditional supervised machine learning (Section 2.6) requires at least some actually occurred anomaly events in the past as labels. Therefore, Section 2.6 focuses on the hardware innovation of the Internet of Things (IoT) for collecting and sharing anomaly event data to build models. While the device monitoring service in some other implementations described in Section 2.8 does not truly predict anomaly events, but uses statistical tests to detect the deviation of sensor data from statistical metrics during normal device operation. The method proposed in this application is completely different from zero-shot, one-shot, and few-shot learning in large language models (LLMs). In large language models, since both the input and output are text, the entire base model learns and organizes the relationships between words, and the learning is based on language input and output. While in this application, the input can be different from the target, and the base model may not have undergone reinforcement learning based on human instructions or human feedback.

[0212] 1(i) When no historical abnormal event (e.g., failure, value is 0) occurs: In the output Y 860 of the basic Transformer model 810, the component y_(ek-t) related to this specific abnormal event will always be 1 (indicating normal operation). However, the other components in the output Y 860 of the basic Transformer model 810 will still change. The basic Transformer model 810 captures how other target values in a normally operating hardware system change according to the input values. When the predicted value of y_(ek-t) (e.g., close to 0) after the changed input vector is much lower than the actual value of y_(ek-t), the actual value is 1 exceeding a certain threshold, or a sample that statistically significantly deviates from the predicted value of y_(ek-t) by the normally operating hardware, this indicates that there is an abnormality at this time step, and it may be that an abnormal event that has never occurred before is about to occur. This zero-negative label learning fundamentally breaks through the limitation of supervised learning, that is, the pre-event used as a label must occur. In the basic model based on Transformer, the clear relationship of the normally operating hardware at least partially indicates a failure that has never occurred, which benefits from the flexibility of the Transformer model objective. In this application, different from large language models, there are no semantic components in the input or output of the basic model. The inventor realizes that the predicted value of y_(ek-t) by Transformer also forms a distribution, the predicted value of y_(ek-t) of the normally operating hardware constitutes a distribution, and the predicted value of y_(ek-t) of the abnormally operating hardware does not belong to this normal distribution. The predicted values of the normally operating hardware and the abnormally operating hardware form two different distributions. Checking whether the predicted value of y_(ek-t) belongs to the distribution of the normally operating hardware through statistical tests becomes the theoretical basis for predicting abnormal events that have never occurred in the embodiments of this application. Once the basic Transformer model outputs the predicted value of y_(ek-t), the method for determining whether an abnormal event will occur can be flexible and is not limited to the threshold or statistical test mentioned here.

[0213] In contrast, the "device monitoring" service in other implementations does not have a model to predict y_(ek-t), does not use a Transformer model to predict y_(ek-t), and does not provide a solution for historical abnormal events that have never occurred as described in this application.

[0214] (ii) When at least one historical abnormal event occurs: At this time, this specific historical abnormal event has been learned and organized in the basic model based on Transformer. (ii.a) First, even if only using the method mentioned in (i), compared with the situation in (i) where no historical abnormal event occurs, it should be able to provide a better indication. (ii.b) Second, due to at least one historical abnormal event, an embedding method can be used. Figure 10Shows single-sample anomaly event prediction based on embedding vectors, relying only on historical anomaly events that occur once. When the input value occurs at historical anomaly event 1010, a new embedding of the input embedding vector 1030 at this time step can be obtained through the Transformer model 1020. When the new input value 1040 generates a new embedding vector 1060 through the Transformer model 1050, and the distance between this embedding vector and the historical anomaly event embedding is very short, it indicates that another new anomaly event may occur. The decision mechanism 1070's determination of how short the distance is to be "short" can be flexible, for example, determined by a threshold, statistical test, or other methods.

[0215] When the number of historical anomaly events is small: This patent provides a method that can make the model better capture the relationship between the input and the anomaly event through data-level enhancement (such as upsampling historical anomaly events) or methods based on Transformer parameter tuning. The above two methods do not limit the scope of this application.

[0216] Section 3.4: Hardware Action Triggering Based on Transformer

[0217] The embodiments of this application provide new applications of multimodal methods and action evaluation methods for the hardware action triggering problem.

[0218] The first method utilizes the multimodal capabilities of the Transformer model. Previously, multimodal techniques have been applied to images and texts. If a specific action is encoded as 0 or 1 to form a vector, this vector can be used as an input or a target. Some embodiments of this application propose to mix the action vector with sensor data and other data to form a multimodal Transformer model. The innovation of EquiFormer lies in that the multimodal Transformer can be applied to manufacturing hardware problems and proposes possible new target vectors.

[0219] The second method is action evaluation. In previous application cases of large language models, Transformer can generate actions, such as calling a calculator or SQL snippet (Fu, Ou, etc., 2022). The action evaluation method uses Python or SQL snippets, or mathematical formulas, where the generated semantic sequences can be passed to a Python interpreter or SQL engine to verify if they run, or passed to a calculator to verify if the calculation is successful. Of course, there is a need for methods to decide when to generate these snippets or formulas in natural language conversations and to decide the start and end positions of the generated text that needs to be evaluated as an action. However, these large language model problems are still very different from the hardware action triggering problem in this patent. Since there is no Python interpreter, SQL engine, or calculator in some embodiments of this application, specific task models are added in the embodiments of this application to replace these evaluators in the large language model literature.

[0220] In common practices of hardware maintenance actions, manufacturers usually propose some schedules (e.g., the oil needs to be changed when the car has traveled x miles) or rules (e.g., a x part needs to be replaced when the x warning light flashes). These schedules are based on time steps, and the warning lights are based on sensor data. Therefore, all this information can be recognized in the Transformer model. A simple rule-based evaluator can be used as in current industry practices, or a technician can build an additional model based on the input of time steps and sensor data, combined with the "should execute (1)" or "should not execute (0)" actions annotated by humans.

[0221] In some embodiments of this application, the actions generated by Transformer can be evaluated by humans and then used for reinforcement learning

[0222] Section 4: Application of EquiFormer in Photonic Programming Processor

[0223] The predecessor of the photonic programming processor was called the Optical Programmed Processor (OPP) and was called the Adaptive Climate Controller (ACC) in some earlier patents, which was the early name of the photonic programming processor.

[0224] In some implementations, the equivalence between the ACC and the photon-programmable processor can be confirmed. The photon-programmable processor is described in the 17 references listed below. This processor uses optical wave control instead of electronic control to output the optimal motor parameters. The OPP directly converts the electromagnet data (light) modulated in real time into an electrical signal (digital or analog signal) without additional conversion, and directly amplifies it into a high-power signal, which can be directly used by an analog motor. This motor converts the analog electrical power into analog mechanical motion. Through this control, the energy efficiency of the motor is improved, usually measured by the percentage ratio of the output mechanical power to the input electrical power. In this previous-generation processor, the optical component parameters (such as the frequency of each light source) used to control the energy efficiency of the motor were determined in the laboratory through physical experiments before the production of the photon-programmable processor. Once determined, these parameters are hard-written into the processor and will not change throughout the life cycle of the processor. In this old processor, the optimization goal is usually a static output variable, such as torque that does not change over time but changes with preset parameters. However, in some practical applications, this optimization goal does change over time. This is also why after adding the photon-programmable processor to the motor, changes in the energy efficiency range can be observed from real-time data.

[0225] In certain embodiments of the present application, the remotely configurable and programmable photon-programmable processor (RCP_OPP) has the following innovations, such as Figure 11 shown:

[0226] (1) The AI / machine learning system 1110 in the general system described in Section 2 or the specialized Transformer-based system in Section 3 will be used to find the optimal hardware parameters for RCP_OPP control, rather than relying solely on laboratory tests. The first-generation photon-programmable processor had no AI components at all. A specific machine learning model stack based on EquiFormer can be described as follows: In the first step, the optical and motor parameters can be modeled as inputs and the torque as the output, without a time component, to accelerate the initial parameter optimization. Then EquiFormer can be used for online optimization. The output of the first model can be used as the input to EquiFormer. The control values are also derived from the AI / machine learning system as described in Section 3.2. In fact, the optimization goal of the hardware controlled by RCP_OPP is not limited to the energy efficiency of the motor. It can be any goal that can be controlled by optical waves.

[0227] (2) RCP_OPP 1120 adds multiple hardware components to enable the hardware to communicate with the AI described in (1). Since the parameters can now be updated, some embodiments of the present application can change the processor name from "programmed" (preset once) to "being programmed" (continuously or intermittently online or offline updated).

[0228] (2.1) A communication component, such as the Internet of Things chip 1122, is used to transmit optical-based controller parameters to the data acquisition service / online optimization service in real-time bidirectional; the specific implementation of the communication component can be flexible and is not limited to the Internet of Things chip.

[0229] (2.2) The remaining hardware components are classified as the RCP_OPP control component 1130.

[0230] (2.2.1) Remote control switch 1131;

[0231] (2.2.2) A rewritable storage component for storing hardware and RCP_OPP parameters 1132;

[0232] (2.2.3) A programmable chip for reading and writing hardware and RCP_OPP parameters 1134;

[0233] (2.2.4) The conversion module 1136 converts the stored digital value of each output hardware parameter into an optical wave parameter. For example, through formulas, experiments, and other methods, the digital value of a continuous variable can be mapped to the parameter value of an optical wave. This mapping is programmable and rewritable, similar to the parameters in 2.2.2 and 2.2.3. From this mapping, those skilled in the art can understand that if the value of the hardware parameter needs to be x, then the corresponding optical parameter y needs to have the value z. For example, if the torque of a single-phase motor needs to be 120 Nm, then the frequency of the light needs to be 80 Hz.

[0234] (2.2.5) RCP_OPP will include an optical component 1138, similar to its predecessor OPP. This component controls the interference of optical waves to generate a new optical wave, and one parameter of this new optical wave represents the hardware parameter mentioned in (2.2.4). Then, this new optical wave is converted into other signals, such as (but not limited to) being converted into an electrical signal through an optical coupler and other devices. In most cases, each hardware parameter in (2.2.4) corresponds to one parameter of the new optical wave. In some cases, a new optical wave can have multiple parameters (such as the frequency and amplitude of light), and these parameters respectively encode the corresponding hardware parameters (such as the frequency and voltage of the input alternating current).

[0235] In (2.1), the Internet of Things is used to assume communication via Internet signals sent through the Earth or satellites. When the Internet signal cannot be received due to excessive distance (e.g., in deep space), some embodiments of the present application provide a novel method for controlling optical components on Earth. Compared with electronic (digital or analog) control, RCP_OPP has unique advantages. RCP_OPP uses light waves instead of analog or digital electronic signals for control. Light has the property of wave-particle duality, which makes it more difficult to observe larger particles. Based on this property, photons have been shown to be able to enter a state of quantum entanglement. Photons in a state of quantum entanglement may be able to communicate with each other over long distances. Based on these facts, some embodiments of the present application propose a photon-controlled optical programming processor control system (Photon-Controlled Optical Programming Processor, hereinafter referred to as PC_OPP) based on photon entanglement: a system capable of unidirectionally controlling remote space devices from Earth, such as Figure 12 as shown. On Earth, the EquiFormer platform 1220 outputs the optimal hardware parameters 1222 for deep space device optimization and control. Another computing platform 1224 takes these hardware parameters as input and outputs the parameters of light waves or photons for controlling the remote device. The photon / light wave control system 1226 generates light waves or photons according to these parameters. The transmitting end Alice 1228 is located on Earth for monitoring or other purposes. The receiving end Bob 1214 is located in deep space and is equipped with its accompanying PC_OPP 1212. The PC_OPP uses photons 2 to generate the light waves required to control the deep space device 1210. The specific implementation of the system can be flexible and scalable. However, it should be emphasized that even without Internet access, unidirectional control of deep space devices can be achieved through the photon entanglement device combined with the PC_OPP in the system.

[0236] (3) The present application proposes three methods for transmitting hardware parameters to deep space for use by the PC_OPP. Although the present application mainly discusses communication between devices located on Earth and devices located in deep space, this arrangement can also be reversed. In other words, placing the devices on Earth in deep space and the deep space devices on Earth can also achieve communication from deep space to Earth. Therefore, there is no need to limit the scope of the present application to only Earth-to-deep space transmission. Currently, the described communication is unidirectional. However, an effective two-way communication link can be established by performing the same unidirectional communication operation twice - once from Earth to deep space and once from deep space to Earth.

[0237] (3.1) The first method uses laser rays. In (2), it is described how certain hardware parameter values correspond to the parameter values of new light waves. These new light waves can be generated by the interference of several light waves. Figure 11The optical component 1138 therein controls the parameters of these light waves, and the solution to this step has been described in the patents related to OPP. The next question is how to transmit the parameters of a light wave (one of the light sources for generating new interference light or the new interference light itself) to deep space if other communication methods cannot be used.

[0238] The specific steps are as follows: First, the hardware parameter values (which may or may not come from EquiFormer) are mapped to the light wave parameter values of the laser beam. Then, using a strong light source with precise aiming ability, the laser beam with the mapped light wave parameters is directly transmitted from the Earth to deep space. To avoid the influence of the atmosphere on the quality of the optical signal, those skilled in the art can place the light source in a location with thin atmosphere, such as a high-altitude area or a device located above the atmosphere. Next, a highly sensitive and capable receiver receives the laser beam signal in deep space and collects photons. After that, the light wave parameter values of the laser beam are converted into other signals, such as electrical signals, for controlling the hardware in PC_OPP. The time series of the light wave signal can be used as a way to indicate the hardware parameters being communicated, where the hardware parameter values are transmitted in sequence. Another way is to use specialized different optical devices to indicate the hardware parameters being communicated, thus allowing the hardware parameter values to be communicated in parallel.

[0239] The following two proposed methods transfer the hardware parameter values from the sending end (on Earth) to the receiving end (in deep space), one is an analog method and the other is a discrete method. In deep space, PC_OPP converts the hardware parameter values into light wave parameter values to control the device. This application also emphasizes that if other control mechanisms are more preferred to be used near the controlled hardware device, PC_OPP can be replaced by other mechanisms in the following two proposed methods. The following two proposed methods are only used as communication mechanisms.

[0240] (3.2) Photon communication based on continuous variable quantum teleportation. This method uses the continuous amplitude and phase quadrature of the optical field as the information carrier. As long as the values to be communicated are continuous, this method can transmit these values. Therefore, both the hardware parameter values and the parameter values of the light wave can be transmitted through this method. This application maintains flexibility regarding the specific values being communicated. This method uses entangled states and classical communication to transfer the unknown quantum state of the sender to the receiver.

[0241] The specific steps are as follows: First, the sender (Alice) and the receiver (Bob) pre-share a pair of continuous-variable entangled states (e.g., two-mode squeezed states). Then, Alice prepares an optical wave containing specific phase and amplitude information to be transmitted. Next, in the joint measurement step, Alice performs a Bell-type joint measurement on the optical wave to be transmitted and a part of her entangled optical field, obtaining the results of continuous variables (i.e., the values of two quadrature components). Alice then sends the measurement results to Bob through a classical communication channel. Subsequently, in the state reconstruction step, after receiving the measurement results, Bob applies corresponding phase shifts and amplitude modulations to his part of the entangled optical field to reconstruct the same quantum state as Alice's original optical wave. For the optical wave frequency value, a technician can encode the frequency information into continuous variables (e.g., phase or amplitude), transmit it through quantum teleportation, and decode it at the receiving end. As for which specific hardware parameters to transmit, a technician can distinguish them through different time series or by using different photon communication devices.

[0242] (3.3) Photon communication based on qubit encoding. In order to transmit a specific hardware parameter value (e.g., 2.453) through photon quantum communication, it is necessary to encode this value into a quantum state to utilize the quantum properties of photons for transmission. If the method of using different time series or different photon communication devices mentioned before is not used to distinguish the transmitted hardware parameters, the ID of the hardware parameter also needs to be encoded as a binary number.

[0243] In (3.3) photon communication based on qubit encoding, the steps are described as follows:

[0244] First, perform binary encoding on the hardware parameter ID and the hardware parameter value. (a) As long as the binary encoding of the parameter ID can be uniquely mapped back to the original parameter ID, it meets the requirements. For the hardware parameter ID, the simplest solution is to sequentially number all hardware parameters, assign each number an ordinal ID, and then represent this ordinal ID in binary form. For example, if at most 16 hardware parameter values need to be transmitted, the hardware parameter ID number 1 will be represented as "0001". (b) The entire hardware parameter value also needs to be represented in binary form. For example, the value 2.453 is converted to 10.0111001111...., where the integer part 2 is converted to the binary form 10, and the fractional part 0.453 is converted to 0.0111001111....

[0245] The calculation steps are: "

[0246] 0.453×2=0.906→0

[0247] 0.906×2=1.812→1

[0248] 0.812 × 2 = 1.624 → 1

[0249] 0.624 × 2 = 1.248 → 1

[0250] 0.248 × 2 = 0.496 → 0

[0251] 0.496 × 2 = 0.992 → 0

[0252] 0.992 × 2 = 1.984 → 1

[0253] 0.984 × 2 = 1.968 → 1

[0254] 0.968 × 2 = 1.936 → 1

[0255] 0.936 × 2 = 1.872 → 1

[0256] ”. Repeat this process to generate the required fractional part of the binary number.

[0257] If the precision is truncated, the value will be 10.0111001111. If this value is represented in floating - point binary form, it will be 1.00111001111 * 2^1. Therefore, combining the hardware parameter ID and the value, along with precision truncation and floating - point representation, the information to be transmitted is: "0001 (hardware parameter ID) 100111001111 (significant digits of the hardware parameter value) 1000 (1, the base - 2 exponent in IEEE754 standard with 4 bits and an offset of 7) 1 (positive sign)". This is just an example. Those skilled in the art can change the specific encoding method of the hardware parameter ID and value information, but essentially all information is encoded in binary form.

[0258] Then, map the binary sequence to qubits, with each bit corresponding to a quantum state. In the quantum state representation, the state |0> represents binary 0 and the state |1> represents binary 1. Those skilled in the art can utilize the polarization direction of photons, such as horizontal polarization (|0>) and vertical polarization (|1>). In some other implementations, those skilled in the art can utilize the phase difference of photons, such as 0 phase (|0>) and π phase (|1>).

[0259] Next is the preparation and transmission of the quantum states. Based on the binary sequence, the sender (Alice) prepares the corresponding sequence of photon quantum states and sends the photons one by one to the receiver through an optical fiber or free space.

[0260] Then is the reception and measurement of the quantum states: The receiver (Bob) receives the sequence of photons sent by Alice. Bob measures the photons using the correct basis (polarization basis or phase basis) and obtains the binary sequence.

[0261] Finally, Bob reconstructs the original information: Based on the agreed-upon parameter ID mapping and the precision of the numerical values, Bob reconstructs the information into the first parameter ID with a value of 2.453.

[0262] In practice, those skilled in the art need to consider error correction and security, and usually additional classical communication methods are required to achieve this. However, the security and error challenges in qubit communication often have many similarities with the problems faced by and attempted to be solved by blockchain technology. This application proposes a blockchain-like mechanism for eliminating the need to use traditional communication methods in qubit communication, called a blockchain-style qubit communication network, as Figure 13 shown. Alice 1302 is similar to the genesis block 1300 with a height of 0. Alice 1302 sends the original qubit-encoded information, not just to one, but possibly to multiple first-level linked Bobs 1312, 1314, and 1316, which have a height of 1 ( Figure 13 of 1310). Each Bob with a height of 1 then sends the qubit-encoded information to zero or more Bobs with a height of 2, and so on. When the end user (such as a device, computer, human, or control device) needs to access the transmitted information, error correction depends on most of the information from the Bobs, rather than verifying a subset of the single Alice-to-Bob information transmitted via classical communication methods. In some embodiments, all the Bobs can be located close to the terminal device, and verifying the information on the Bobs is local. In other embodiments, the Bobs can send their information remotely to the terminal device, whether or not classical communication methods are used. At least between Alice and the Bobs, error correction of the information no longer requires classical communication methods. In terms of security, tampering with at least a small number of Bobs will not affect the correctness of the information. Those skilled in the art can add a distributed quantum key to the information transmitted between Alice and the Bobs, collectively acting as blocks in the blockchain to detect tampering of a certain block. Additionally, in order to compare the information received by each Bob, the Bobs need rewritable storage units to store the received photon information because the quantum state of the photon changes after being received and measured by the Bobs. Similar to the blockchain, the information and numerical values of the previous block are hashed and transmitted to the next block and stored in the next block of the blockchain-style qubit communication network.

[0263] Between Bob and the terminal device, PC_OPP is used to perform continuous and precise control of the device based on light waves. Since the hardware parameter values transmitted in qubit encoding are inherently discrete, PC_OPP is not necessary. Those skilled in the art can retain PC_OPP and regard the reconstructed parameter values as continuous values. If those skilled in the art believe that other device control methods (such as electronic digital control widely used in control theory for discrete values) are more appropriate, the device control method can be replaced from PC_OPP to other methods. The blockchain-based qubit communication network is not limited to communication for hardware control; it is a general qubit communication mechanism.

[0264] Section 5: Innovations related to Transformer but not limited to hardware devices useful for EquiFormer

[0265] Section 5.1: Ring-based global reduction gradient update based on model and data parallelism

[0266] For the training of any deep learning model, including any Transformer model, especially EquiFormer, those skilled in the art have tried various parallel techniques to improve the scalability of training. These techniques include data parallelism, horizontal (tensor) or vertical (pipeline) model parallelism, and gradient averaging. This application proposes a new method that combines a specific gradient averaging technique, namely ring-based global gradient reduction, with data parallelism and model parallelism. When there is data parallelism and model parallelism, gradient averaging is usually performed through a parameter server or other centralized solutions. That is, the gradients calculated by each node in the network need to be sent to a node or centralized device for synchronization and averaging. Of course, this centralized device itself can have multiple computing nodes. This centralized reduction is the reduction in the general sense. If only reduction is mentioned in the literature, it usually refers to this centralized reduction. When ring-based global gradient reduction (a peer-to-peer scheme) is used for gradient averaging, usually only data parallelism is adopted and model parallelism is not involved.

[0267] Section 5.1.1: Ring-based global reduction gradient update based on model and data parallelism applied to the original model

[0268] Figure 14 Shows the ring-based global reduction gradient update based on model and data parallelism. From left to right, the data is first split into batch 0 ( Figure 14 1400 in Figure 14 ), batch 1 ( Figure 14in 1404), and so on. An example model has four layers; the degree of horizontal model parallelism is 2, and the degree of vertical model parallelism is also 2. In the example model, the hidden layers are numbered from 1 to 4 in the order from input (left) to output (right). Pipeline parallelism groups the first and second layers into one group, and the third and fourth layers into another group. Tensor parallelism divides each layer into the upper half, denoted by adding 0 after the layer number. Thus, L10 in GPU0 ( Figure 14 in 1406) represents the upper half of the first hidden layer. The model is divided into four quadrants: L10 and L20, L30 and L40, L11 and L21, and L31 and L41. Thus, the first copy of the model is located in GPU0 ( Figure 14 in 1406), GPU1 ( Figure 14 in 1408), GPU2 ( Figure 14 in 1410), and GPU3 ( Figure 14 in 1412). The second copy of the model is located in GPU4 to 7 ( Figure 14 in 1414 to 1420). The third copy of the model is located in GPU8 to 11 ( Figure 14 in 1422 to 1428). It should be noted that this is just an example of a model structure plus horizontal and vertical model parallelism. The specific model structure and parallel configuration are flexible and do not affect the proposed scheme. If there is more or different types of model parallelism, the ring-based global reduction update will still be applied to GPUs with the same parts of the model but different local gradients due to different data batches.

[0269] When the data of batch 0 (1400) is input to GPUs 0 to 3 (1406 to 1412), the GPUs containing the last output layer of the model will all obtain the gradients corresponding to the upper or lower half partitions of the model. For example, GPU 0 (1406) will not obtain the gradient, GPU 1 will obtain gradient 0_1 (1432), GPU 2 will not obtain it, and GPU 3 will obtain gradient 0_3 (1436). Similarly, batch 1 (1402) will generate gradient 1_1 (1440) and gradient 1_3 (1444); batch 2 (1404) will generate gradient 2_1 (1448) and gradient 2_3 (1452). Thus, gradient 0_1 (1432), gradient 1_1 (1440), and gradient 2_1 (1448) correspond to the upper half partition (L10, L20, L30, and L40) of the same model but come from three data batches (batches 0 to 2, 1400 to 1404). Similarly, gradient 0_3 (1436), gradient 1_3 (1444), and gradient 2_3 (1452) correspond to the lower half partition (L11, L21, L31, and L41) of the model.

[0270] This paper proposes to use the ring allreduce (Patarasuk and Yuan 2009) technique for point-to-point gradient averaging instead of calculating the average gradient of the upper or lower half partitions of the model in a centralized manner. For example, from gradient 0_1 ([[]] Figure 14 [[]]1432 in [[]] Figure 14 [[]]) on GPU1 ([[]] Figure 14 [[]]1408) to gradient 2_1 ([[]] Figure 14 [[]]1448 in [[]] Figure 14 [[]]) on GPU9 ([[]] Figure 14 [[]]1424), from gradient 2_1 ([[]] Figure 14 [[]]1448 in [[]] Figure 14 [[]]) on GPU9 ([[]] Figure 14 [[]]1424) to gradient 1_1 ([[]] Figure 14 [[]]1440 in [[]] Figure 14 [[]]) on GPU5 ([[]] Figure 14 [[]]1416), and then from gradient 1_1 ([[]] Figure 14 [[]]1440 in [[]] Figure 14 [[]]) on GPU5 ([[]] Figure 14 [[]]1416) back to gradient 0_1 ([[]] Figure 14 [[]]1432 in [[]] Figure 14 [[]]) on GPU1 ([[]] Figure 14 [[]]1408), the ring allreduce technique is applied to exchange and average these corresponding gradients. That is, the average gradient 1 of the upper half partition. The direction of data exchange is indicated by the arrows. For better readability, Figure 14 [[]]the arrows indicating the ring allreduce directions for the other lower half partitions are omitted. [[]]

[0271] [[]]When the ring allreduce ends, the average gradients 1 of quadrants L30 and L40 will remain on their original three GPUs. That is, there will be three copies of the same average gradient 1: the average gradient 1 ([[]] Figure 14 [[]]1456 in [[]] Figure 14 [[]]) on GPU1 ([[]] Figure 14 [[]]1432), the average gradient 1 ([[]] Figure 14 [[]]1464 in [[]] Figure 14 [[]]) on GPU5 ([[]] Figure 14 [[]]1416), and the average gradient 1 ([[]] Figure 14 [[]]1472 in [[]] Figure 14 [[]]) on GPU9 ([[]] Figure 14 [[]]1424). Similarly, there will be three copies of the average gradient 3: the average gradient 3 ([[]] Figure 14 [[]]1460 in [[]] Figure 14 [[]]) on GPU3 ([[]] Figure 14 [[]]1412), the average gradient 3 ([[]] Figure 14 [[]]1469 in [[]] Figure 14 [[]]) on GPU7 ([[]] Figure 14 [[]]1420), and the average gradient 3 ([[]] Figure 14 [[]]1476 in [[]] Figure 14 [[]]) on GPU11 ([[]] Figure 14 [[]]1428). [[]]

[0272] [[]]​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​To update the weights in the four quadrants of the model, we need to use both average gradient 1 and average gradient 3. There are at least three ways to achieve this. (1) The first way is a peer-to-peer solution, such as Figure 14 As shown in the figure, three solid arrows are drawn from right to left. Each GPU gets a copy of average gradient 1 and average gradient 3 from its "nearest" node, and the definition of "nearest" is not fixed. For example, GPU 1 ( Figure 14 The average gradient 1( Figure 14 1456) will be copied to GPU 0 ( Figure 14 1406), GPU 2( Figure 14 1410) and GPU 3 ( Figure 14 1412 in); GPU 3( Figure 14 The average gradient 3 (1460) on Figure 14 1460) will be copied to GPU 0 ( Figure 14 1406 in), GPU 1( Figure 14 1408) and GPU 2 ( Figure 14 For GPU 4 ( Figure 14 1414) to GPU11( Figure 14 The situation of 1428 in is similar. (2) The second way is not to move the average gradients 1 and 3, leaving them in place without copying them to other places. When updating the weights in each quadrant, the node responsible for the update calculation reads the average gradient 1 from one GPU and the average gradient 3 from another GPU. (3) The third way is a centralized solution: the average gradient 1 on a GPU and the average gradient 3 on another GPU are sent to a centralized device.

[0273] It should be noted that although the model is parallelized to four quadrants, the weights are updated asynchronously. During the forward pass, GPU 1 ( Figure 14 1408) and GPU 3 ( Figure 14 The calculation of 1412) needs to wait for the GPU 0 ( Figure 14 1406) and GPU 2 ( Figure 14 1410 in , and vice versa in the back propagation. For other data batches, the asynchronous update method is similar. There are many ways to optimize and parallelize back propagation, and optimizer state sharding is one of them, but this is not the focus of this disclosure.

[0274] Section 5.1.2: Model- and data-parallel loop-based global reduction gradient updates for low-rank adaptation

[0275] Those skilled in the art can distribute matrices A and B in Low - Rank Adaptation (LoRA) to different GPUs corresponding to the model - parallel partitions of the original model, and utilize the ring - allreduce technique to update matrices A and B in a data - and - model - parallel manner. This paper presents a method to achieve this goal. For illustration purposes, the mathematical representation of matrices is inevitably used, but this paper focuses on the engineering implementation of how to distribute partitioned matrices on devices in the network for ring - allreduce, rather than any mathematical proof.

[0276] In Low - Rank Adaptation, there are two low - rank matrices A and B. According to the original Low - Rank Adaptation method, the updated model weight matrix is: W_fine_tuned = W_pre_trained+ΔW = W_pre_trained+AB where ΔW is the weight update and the scaling factor is omitted. Let A = {a_i,j} be an m×n matrix, B = {b_i,j} be an n×p matrix, and ΔW = {w_i,j} be an m×p matrix. In this paper, for clarity, a comma is added between the row and column indices of the elements. In the usual nomenclature of linear algebra, there is no comma between the indices. In this paper, a_i,j is equivalent to a_ij in the usual nomenclature. In the normal Low - Rank Adaptation nomenclature, n (the number of columns of matrix A or the number of rows of matrix B) is called the rank. The rank n is usually much smaller than m or p, so it is called low - rank. Although r is often used to represent the index of the dot - product of two matrices, this paper uses the rank n instead of r to assist in explaining the partitioning of the dot - product of matrices A and B. The steps are as follows:

[0277] First, partition W_pre_trained into different devices. Freeze all the weights in W_pre_trained. For the number of horizontal and vertical partitions, those skilled in the art can adjust flexibly. For ease of explanation, W_pre_trained is divided into four quadrants as shown in Figure 14 the following:

[0278] where commas are added to separate the row and column indices of the sub - matrices, and parentheses are added to enclose the indices, and pt refers to pre - trained.

[0279] Second, partition matrices A and B, and place the partitions of A and B on the GPUs corresponding to the same partitions of W_pre_trained. Figure 15This correspondence is shown. This is feasible because in matrix addition, elements with the same row index and column index of W_pre_trained and ΔW are added together to form W_fine_tuned. Therefore, the same model partitions of W_pre_trained and ΔW, or in this description, the same quadrants, should be located on the same GPU. W_fine_tuned and ΔW are partitioned in the same way as W_pre_trained. In Figure 15 , the model is divided into four partitions: model partition 00 (M00) 1500, model partition 01 (M01) 1510, model partition 10 (M10) 1520, and model partition 11 (M11) 1530. Since ΔW is the dot product of A and B, each partition of ΔW can be represented by a_i,j and b_i,j. Figure 15 shows ΔW_[0,0] ( Figure 15 in 1502) and its corresponding a_i,j and b_i,j, as well as ΔW_[0,1] ( Figure 15 in 1512), ΔW_[1,0] (1522), and ΔW_[1,1] ( Figure 15 in 1532). The a_i,j and b_i,j in the same partition should be located on the same GPU.

[0280] The partitioning of A and B should follow the partitioning method corresponding to W_pre_trained. Specifically, according to the matrix partitioning method of W_pre_trained, those skilled in the art partition A and B as follows. A is divided into upper and lower sub-matrices, represented as follows:

[0281] where the colon represents all rows or columns (as in the representation in Python), the comma is used to separate the row index and column index of the sub-matrix, and the parentheses are used to enclose these indices.

[0282] B is divided into left and right sub-matrices, represented as follows:

[0283] B = (B_[:,0] B_[:,1]), where the colon represents all rows or columns (as in the representation in Python), the comma is used to separate the row index and column index of the sub-matrix, and the parentheses are used to enclose these indices.

[0284] Figure 16Shows the correspondence between the partitions of the entire model 1620 and the partitions of A1600 and B1610. Further analysis reveals that model partition 00 (M00) 1622 is only related to the upper half of A, A_[0,:] 1602, and the left half of B, B_[:,0] 1612; model partition 01 (M01) 1624 is related to A_[0,:] 1602 and B_[:,1] 1614; model partition 10 (M10) 1626 is related to A_[1,:] 1604 and B_[:,0] 1612; model partition 11 (M11) 1628 is related to A_[1,:] 1604 and B_[:,1] 1614. Conversely, A_[0,:] 1602 corresponds to M00 1622 and M01 1624; A_[1,:] 1604 corresponds to M10 1626 and M11 1628; B_[:,0] 1612 corresponds to M00 1622 and M10 1626; B_[:,1] 1614 corresponds to M01 1624 and M11 1628. Note that the partitions of W_pre_trained are also located on the model partitions corresponding to the A and B submatrices, but they are frozen. According to the above partitioning method of the low-rank matrices A and B, the low-rank adaptation update formula of the partitioned model still holds. That is, on model partition M00 1622, W_ft_[0,0] = W_pt_[0,0] + A_[0,:]B_[:,0], where ft refers to find_tuned, and W_fine_tuned is partitioned in the same way as W_pre_trained. Similarly, the low-rank adaptation update formulas for model partitions M01 1624, M10 1626, and M11 1628 are also carried out in a similar manner, as Figure 16 shown for each model partition.

[0285] In the third step, apply the ring-style global reduction operation twice to the submatrices of A and B to obtain the average weights of the submatrices of A and B, as Figure 17 shown. From Figure 17 the left to the right, data batch 0 ( Figure 17 1700 in Figure 17 ) to data batch 2 ( Figure 17 1704 in Figure 17 ) are input into the GPU. GPU0 ( Figure 17 1706 in Figure 17 ), GPU1 ( Figure 17 1708 in Figure 17There is a partition M01 on 1708) of, GPU2( Figure 17 There is a partition M10 on 1710) of, GPU3( Figure 17 There is a partition M11 on 1712) of. Data batch 0( Figure 17 There is a partition M00 on 1700) of. Data batch 0( Figure 17 from GPU4( Figure 17 There is another copy of the entire model on 1714) to GPU7( Figure 17 There is a partition M00 on 1702) of; GPU8( Figure 17 from GPU8( Figure 17 There is a third copy of the entire model on 1722) to GPU11( Figure 17 There is a partition M00 on 1704) of. Thus, a specific model partition for a specific data batch on each GPU generates corresponding gradients. For example, for M00 on GPU0( Figure 17 There is a partition M00 on 1706) of, the gradient generated from data batch 0( Figure 17 There is a partition M00 on 1700) of is 0_0_0( Figure 17 There is a partition M00 on 1730) of, where the first two indices represent the model partition and the third index represents the data batch).

[0286] So far, the partitioning of the model weights and the low-rank matrices has been completed, and the model weight partitions and their corresponding A and B sub-matrices are distributed to the distributed computing units in the computing network. It should be noted that the implementation of distributed low-rank adaptation combining data parallelism and model parallelism does not necessarily require relying on a ring-shaped global reduction operation. Those skilled in the art can use a parameter server to obtain the average sub-matrices of A and B without relying on a ring-shaped global reduction operation. In this application, a method using a ring-shaped global reduction operation is further described and can be optionally executed twice, as follows.

[0287] The first ring-shaped global reduction operation is applied to the gradients of each corresponding model partition between data batches to obtain the average gradient of the vertical final partition of the model as described in Section 5.1.1, and obviously also across GPUs.

[0288] For example, gradient 0_1_0(1732), gradient 0_1_1(1740), and gradient 0_1_2(1748) all correspond to Figure 15 the upper halves of the two model partitions M00(1500) and M01(1510) in, but come from batch 0(1700), batch 1(1702), and batch 2(1704) respectively. After these gradients are processed by the ring-shaped global reduction operation, the average gradient 0_1 of the upper half of the M00 model partition is obtained (not shown in Figure 17 ), which isFigure 14 is similar to the average gradient 1 in. GPU 1 (1708), GPU 5 (1716), and GPU 9 (1724) will each have the same copy of this average gradient 0_1.

[0289] Similarly, GPU 3 (1712), GPU 7 (1720), and GPU 11 (1728) will also each have the same copy of the average gradient 1_1 (similar to the average gradient 3 in Figure 14 ), corresponding to the lower half of the model partitions ( Figure 15 M10 and M11 in

[0290] Similar to the update of the original model weights in Section 5.1.1, for each model partition, updating the low-rank adaptation A and B also requires the complete gradient. Similarly, for the "three or more ways" of how to aggregate these average gradients, these average gradients can also be combined together as needed to generate the average gradient _:_1 (omitted for readability in Figure 17 ). Among them, the first index of the average gradient _:_1 uses columns to represent all horizontal partitions; the second index of the average gradient uses 1 to represent the last vertical partition.

[0291] Therefore, there are two submatrices (1756) A_[0,:]_01 and B_[:,1]_01 on GPU1 (1708). The "01" at the end of the submatrices indicates that they are from the model partition M00. Similarly, the two submatrices (1764) A_[0,:]_01 and B_[:,1]_01 on GPU5 (1716), and the two submatrices (1772) A_[0,:]_01 and B_[:,1]_01 on GPU11 (1728) are also updated. At this time, the weights of the submatrices of A and B on each GPU have been updated. Two issues are worth pointing out: (1) Although the submatrices 1756, 1764, and 1772 are from the same M00 model partition and the same average gradient _:_1, due to randomness and other variables, they may not be exactly the same. (2) One submatrix of A or B actually has multiple copies among the model partitions. In this example, each has two copies. For example, the value of the submatrix A_[0,:]_00 (1754) from the M00 partition on GPU0 (1706) is different from the value of the submatrix A_[0,:]_01 (1756) from the M01 partition on GPU1 (1708) because they are from two different model partitions.

[0292] The second ring global reduction operation is applied to the same sub-matrices of A and B across GPUs to obtain the average weights of these sub-matrices. There are two ways to accomplish this operation, both using ring global reduction to obtain the average value. (1) Those skilled in the art can choose to ignore the differences between a sub-matrix across data batches and only take the average of this sub-matrix across model partitions. In many usage scenarios, such as weight quantization, these differences can be ignored. For example, averaging the sub-matrix A_[0,:]_00(1754) on GPU0(1706) and the sub-matrix A_[0,:]_{01(1756) on GPU1(1708) results in obtaining an average A_[0,:] on both GPU0(1706) and GPU1(1708). Another A_[0,:] is obtained on GPU4(1714) and GPU5(1716), and a third A_[0,:] is obtained on GPU8(1722) and GPU9(1724). This type of second ring global reduction is not shown in Figure 17 (2) A more precise but computationally and communicatively more costly method is to use the ring global reduction operation to average the same sub-matrices of all replicas. For example, applying the ring global reduction operation to the sub-matrices A_[0,:]_01(1756), A_[0,:]_00(1754), A_[0,:]_01(1764), A_[0,:]_00(1762), A_[0,:]_01(1772), and A_[0,:]_00(1770) to obtain the same average A_[0,:] on GPU0(1706), GPU1(1708), GPU4(1714), GPU5(1716), GPU8(1722), and GPU9(1724). Then, the sub-matrices of the same model partition will have the same value. For example, the sub-matrices 1778, 1786, and 1794 of the M00 model partition are the same. Similarly, the sub-matrices 1780, 1788, and 1796 of M01 are the same, the sub-matrices 1782, 1790, and 1798 of M10 are the same, and the sub-matrices 1784, 1792, and 1799 of M11 are also the same.

[0293] Step 4, update the model. Whenever the sub-matrices of A and B are calculated, they can be applied to the update formula of low-rank adaptation to obtain W_fine_tuned. In this application, a partitioned form of the low-rank adaptation update formula is used. (1) Those skilled in the art can choose to update the model after the first ring global reduction operation, at which time the values of the A and B sub-matrices may be different. This update is performed by Figure 17is represented by the dashed arrow in. (2) More preferably, a person skilled in the art can update the model after the second ring global reduction operation, at which time each of the sub-matrices of A and B has only one set of values. To ensure that the fine-tuned model remains consistent throughout the computational network, it is recommended to update the model only after the second ring global reduction operation. This update is represented by Figure 17 the solid arrow from the rightmost to the leftmost GPU in.

[0294] However, an obvious obstacle to updating the model only after the second ring global reduction operation is its high computational and communication costs. In data parallelism terms, a "step" describes the process where all model replicas process their data batches once respectively. In Figure 17 , data batch 0 (1700) is input into GPUs 0 (1706) to 3 (1712), data batch 1 (1702) is input into GPUs 4 (1714) to 7 (1720), and data batch 2 (1704) is input into GPUs 8 (1722) to 11 (1728), which together constitute one step. Therefore, the number of data parallel steps is equal to the number of data batches divided by the number of model replicas. In one data parallel step (defined the same in the ring global reduction operation), the number of communications for each sub-matrix is equal to the number of its replicas in all partitions of the model multiplied by the number of model replicas. For example, in Figure 17 , obtaining A_[0,:] requires 6 point-to-point communications because each sub-matrix has 2 replicas and the model has 3 replicas. The total number of communications is equal to the number of communications for each sub-matrix multiplied by the number of sub-matrices. In this example, as shown in Figure 17 , there are 4 sub-matrices, so the total number of communications is 6×4 = 24.

[0295] In contrast, in the first ring global reduction operation, the number of communications for the gradients of the same model partition is equal to the number of model replicas. It is 3 times in Figure 17 . The total number of communications is equal to the number of communications for each partition multiplied by the number of partitions of each model. In the example of Figure 17 , a model has 4 partitions and the model has 3 replicas, so the total number of communications is 4×3 = 12. The number of communications in the first ring global reduction operation is usually much less than that in the second ring global reduction operation.

[0296] A person skilled in the art can synchronously execute the two ring global reduction operations and then perform model update. Synchronously executing the two ring global reduction operations means that in each step of data parallelism, immediately after the first ring global reduction operation, the second ring global reduction operation is executed. This solution will inevitably face a large amount of computation and communication problems in the second ring global reduction operation.

[0297] Alternatively, a person skilled in the art can perform two ring global reduction operations asynchronously and update the model only after the second ring global reduction operation is completed, thereby reducing the communication overhead in the second ring global reduction operation. In a certain step of data parallelism, after the first ring global reduction operation is completed, the generated sub-matrices of A and B can be saved on their respective GPUs. The second ring global reduction operation can then be postponed temporarily. In subsequent data parallelism steps, each subsequent first ring global reduction operation will generate its corresponding copies of the sub-matrices of A and B and save them locally in a similar manner. After a certain number of data parallelism steps, the locally stored sub-matrices of A and B are averaged on their respective GPUs without performing a ring global reduction operation. Once this local averaging operation is completed, the second ring global reduction operation can be executed.

[0298] This solution postpones the second ring global reduction operation in the middle of the intermediate data parallelism steps and only executes it after a series of first ring global reduction operations, thereby significantly reducing the communication cost associated with the second ring global reduction operation. The number of first ring global reduction operations executed before each second ring global reduction operation can be fixed or determined from a random distribution, providing flexibility for optimizing the overall communication cost.

[0299] In practical applications, a person skilled in the art can weigh the pros and cons and choose from the solutions described in this section. If the ring global reduction operation with low-rank adaptation is selected based on data parallelism and model parallelism, it may be that due to the large model, even the low-rank matrices A and B used in low-rank adaptation cannot fully fit into one computing unit, but the communication cost is low.

[0300] The flexibility of low-rank adaptation partitioning is detailed here. In the example partitioning described in detail in this disclosure, the matrix W_pre_trained is divided into four sub-matrices, and the low-rank matrices A and B are correspondingly divided into two sub-matrices according to the partitioning of W_pre_trained. In fact, the partitioning of W_pre_trained can be adjusted according to the requirements of the computing network and arranged in various other configurations. Each corresponding sub-matrix of the A and B dot products (defined by the corresponding rows and columns of W_pre_trained) can always be expressed as the result matrix of performing operations on the sub-matrices of A and B. The partitioning of the low-rank matrices A and B ensures that the generated matrices match the corresponding sub-matrices required for the A and B dot products. A sub-matrix of W_pre_trained and its corresponding sub-matrices of the low-rank matrices A and B are copied to the same computing unit. Such an arrangement allows the distributed LoRA with model parallelism to operate effectively and optionally combine data parallelism and further combine one or more ring global reduction operations in the data parallelism steps.

[0301] The GPU mentioned in this application is the most common computing unit in the network. This application focuses on a circular global reduction gradient update method based on data parallelism and model parallelism. The specific devices used by the computing units in the network are flexible and can use CPUs, quantum computing units, and many other devices.

[0302] Section 5.2 Hybrid large language models represented in general language, graph representation, or a combination of both generate general two-dimensional or three-dimensional visualization charts

[0303] Many web pages, cloud-based, or stand-alone computer applications are used to generate certain types of visualization charts. For example, in a PowerPoint presentation, each slide is similar to a canvas, and each icon, background, or text area is a visualization element. Thus, essentially each PowerPoint slide is a visualization chart. The same applies to charts in Lucidchart, illustrations in BioRender, user interface models in Figma, hardware design applications, floor plan design applications, CAD applications, 3D printing applications, and the visualization interfaces of web pages, especially those created through drag-and-drop applications such as Wix. These different applications are written in their respective computer languages. Web pages or user interfaces are mostly written in JavaScript and HTML, PowerPoint is written in office java script and Office Visual Basic, CAD is written in a hardware description language, and so on. A common approach is to fine-tune the computer code of existing large language models in their respective specific languages to generate visualization charts, with tasks such as predicting tokens of that specific computer language or generating computer language code based on historical prompts or the historical accompanying text of the visualization chart. Examples of these scenarios include generating JavaScript and HTML code for web pages or office_javascript and Visual Basic code for PowerPoint files. This common scenario applies to historical data because existing PowerPoint files or web pages are written in different computer languages. However, this common scenario has certain limitations in the generalization of visualization elements. Those skilled in the art may be more inclined to build independent fine-tuning models for each computer language and its associated type of visualization chart. However, this solution is also limited by language issues: the organization of tokens is usually sequential and either one-way or two-way.

[0304] Various embodiments of this disclosure provide a general solution for generating two-dimensional or three-dimensional visualization charts through general language representations (such as JSON, YAML, etc.).

[0305] The steps are described as follows: First, convert a two-dimensional or three-dimensional visualization chart (whether it is a bitmap or a vector-based chart) into a common language representation form, such as JSON, YAML, etc. For example, a canvas or background can be represented by two-dimensional axis values in JSON, and a spatial background can be represented by three-dimensional axis values. A visualization element can be represented by the values of its attributes such as type, color, position, size, related text, etc. For example, the following JSON is a simple representation that describes a rectangular box containing text and a canvas:

[0306]

[0307] Note that the attributes of some elements represent the graph relationships between different elements, such as "connected_element_id_array". If the historical visualization chart is written in other computer languages and has code, those skilled in the art can write code to output the corresponding JSON file to represent the same visualization chart. If the historical visualization chart has no code but exists in a hand-drawn form or only as a PNG file, first perform visual object detection to identify the visualization elements and their attributes, and then organize them into the above JSON file. These JSON files constitute a natural language corpus and can be used to fine-tune existing large language models.

[0308] Next, in many types of charts, the lowest-level visualization elements are predefined. For example, in Lucidchart, icons belong to the lowest-level elements. Therefore, in addition to defining the tokens in the JSON file as natural language tokens, the lowest-level visualization elements should have their own tokens and corresponding embeddings. For example, if there are only two elements, "rectangle" and "circle", each element should have its own token and embedding. For example, in a one-hot embedding,

[01] represents "rectangle" and

[10] represents "circle". Those skilled in the art can use more appropriate embeddings, such as vertex embeddings from the graph itself. Those skilled in the art can also choose other appropriate token encoding methods. For example, the sequence ["rectangle", "rectangle", "circle"] should be represented as [

[01] ,

[01] ,

[10] ]. If it is treated as a language problem, the token representation of this sequence can be ["rect", "angle", "rect", "angle", "circ", "le"] and its corresponding language-encoded token sequence. How these visualization elements form this sequence will be discussed in the subsequent part of this section.

[0309] Next, in many types of diagrams, the lowest-level visual elements at various levels are usually grouped. These grouped elements are often reused in the diagram and should also be regarded as tokens. To select an appropriate grouping level, those skilled in the art can use techniques similar to those in natural language processing (NLP), such as Byte Pair Encoding. For example, if two connected circles always appear together repeatedly, but not three circles or rectangles, they should form a new element "adjacent_two_circle".

[0310] Next, when the visual elements are not represented in the natural language format of a JSON file, those skilled in the art need to serialize them. The visual elements and their relationships can be described by a graph, where the elements are vertices and the relationships are edges. Those skilled in the art can find all paths on the graph to form a sequence that covers all elements on the graph. Optionally, those skilled in the art can reduce the number of paths by preferentially selecting longer paths to cover as many elements as possible. These paths constitute a new corpus with visual elements as units. This is also the reason why this application is reluctant to call Section 5.2 a multimodal method and prefers to call it a hybrid method. Although the encoding of visual elements is very similar to the multimodal method, serialization is not as simple as simply cutting an image into a grid as tokens in the multimodal method, using convolutional neural network (CNN) embedding for token embedding, and using the order from the upper left to the lower right as the position of the sequence. In this application, this traditional image multimodal method is called the image CNN modality. However, this application uses graph techniques to organize visual element tokens into sequences. In this application, this technique is called the visual graph modality. At the same time, this application still refers to the JSON representation of the diagram as a natural language token sequence as the natural language modality. In this application, the natural language modality and the visual graph modality are two independent modalities in the corpus.

[0311] Next, organize the two-modal input and model structure generated for the visual diagram. Considering that at the time of writing this application, existing multimodal models support the image CNN modality and the natural language modality, but do not support the visual graph modality.

[0312] In the first implementation, if a person skilled in the art wishes to make minimal modifications to an existing basic large language model, including weights, tokens, encodings, and model structures, and only incorporate a general language representation file (such as a JSON file), relevant prompts, and relevant natural language materials into the corpus for continuous training, fine-tuning, etc. of the existing basic large language model without changing the model structure at all. Alternatively, a person skilled in the art can add an adaptation layer to the existing basic large language model, freeze the weights of the existing model layers, but learn new tasks through the learnable weights in the adaptation layer. The task focuses on generating a general language representation file (such as a JSON file). This solution maps visualization charts into natural language.

[0313] In the second implementation, if a person skilled in the art wishes to utilize the visualization graph modality, more modifications to the model structure are required. If a person skilled in the art does not wish to retain the training weights in the existing large language model, they can start from a corpus with both natural language modality and visualization graph modality and train an entirely new model with an entirely new model structure. Attention for different modalities, multi-head attention (for a single modality), position encodings for different modalities, cross-modal attention, and other appropriate techniques can be added. When the edge relationship of the visualization elements is unidirectional, a person skilled in the art can use unidirectional attention. When the edge relationship of the visualization elements is bidirectional, bidirectional attention with absolute position encoding can be used. For example, if the visualization elements A<->B<->C<->D form a bidirectional path on the graph, with the absolute position of element A being 0, B being 1, C being 2, and D being 3 (indexed from 0), the attention can be bidirectional. Optionally, a person skilled in the art can consider this path as two sequences: (1) A->B->C->D, where A is 0, B is 1, C is 2, D is 3 (indexed from 0), and use unidirectional attention; (2) D->C->B->A, where D is 0, C is 1, B is 2, A is 3 (indexed from 0), and also use unidirectional attention. Due to the great flexibility of deep learning models (including transformer models), a person skilled in the art can choose a model structure and implementation method suitable for their specific application scenario.

[0314] Alternatively, in the third implementation, those skilled in the art can retain the pre-trained weights in the existing large language model while still adding the visual graph modality. They can utilize the pre-trained weights for multimodal expansion. The new model structure can modularize the language modality using the structure of the existing large language model. Those skilled in the art can update the model structure and weights using the previous implementations (such as the solution of mapping to JSON mentioned above). These weights can be frozen or at least used as the initial weights for subsequent training. The new model structure also modularizes the visual graph modality by adding a model structure similar to the Transformer because the visual graph modality has been converted from a vertex graph to a token sequence, and this serialized representation of the visual graph modality is more suitable for a model structure similar to the Transformer. The modular model structure of the visual graph modality model can design its own multi-head attention mechanism, positional encoding, etc. for this modality. The modular model structure of the visual graph modality can be trained from scratch. Subsequently, a fusion layer is added. Those skilled in the art can introduce layers in the model to fuse different modalities, such as using a cross-modal attention mechanism to combine the new modality representation with the language representation. These newly added layers can be trained from scratch while keeping the original language model part unchanged or fine-tuning it.

[0315] After generating the general language representation, graph representation, and / or a hybrid representation of language and graph for the visual chart, first, a mechanism is needed to check whether the generated representation actually forms a legal visual chart. This is similar to when a large language model (LLM) generates Python code, first, it is necessary to check whether the generated code can run in a Python executor; or when a large language model generates SQL code, it is necessary to check whether the generated MySQL code can run on an SQL server. Those skilled in the art will build services to map the representation into a visual chart and display these charts, and then verify whether the representation generated by the large language model in this application can display a visual chart in these services. Secondly, a scoring model or a human feedback mechanism is needed to check whether the generated chart meets higher-level requirements, such as the logic of the relationships between elements, the aesthetics of the chart, whether the chart conforms to the prompted instructions, the creativity of the chart, the authenticity of the chart, etc. Further human feedback and / or scoring model feedback reinforcement learning will be applied to further adjust (optionally) the weights of the hybrid large language model.

[0316] Figure 18 An example of the third implementation is shown, in which the first implementation is incorporated as a whole. The second implementation can use a similar Figure 18The model structure or different structures are trained from scratch. For the second implementation, one major change might be not using a modular structure but instead constructing the cross-modal attention mechanism and positional encoding simultaneously. It should be noted that deep learning models (including transformer models) have great flexibility. Those skilled in the art can add more modalities and their corresponding modules, modify the layer structures within the modules, change the types and quantities of attention mechanisms, use unidirectional or bidirectional attention, and adjust all aspects of the attention heads, etc. Therefore, the specific structure of the model is not limited to this.

[0317] Figure 18 Starting from the lower left corner. The chart preprocessing 1800 uses image recognition 1802 and / or other computer languages 1804. On the left is the language modality path, and on the right is the visualization modality path. In the language modality path, the chart preprocessing 1800 maps the chart to a general language expression (such as JSON) at 1822. The natural language prompt, natural language corpus, and JSON corpus constitute the mixed language input 1824, which enters the language module 1826 of the visualization chart generation model. The model includes a language module 1826, a visualization graph modality model 1814, and a fusion layer 1828. The language module 1826 can optionally include an existing large language model 1824, which can optionally use pre-trained weights and an optional hidden adaptation layer. An optional adaptation layer 1828 can be added before the existing large language model, and an adaptation layer 1836 can be added after the existing LLM. The language input follows the token embedding 1830 and positional encoding 1832 that are consistent with the requirements of the large language model 1834.

[0318] In the visualization graph modality path, the visualization elements (including arrows and links that can represent edges, and shapes and icons that can represent vertices) are recognized at 1806. At 1808, the repeated grouping of the visualization elements is processed at an appropriate level. At 1810, a graph structure of edges and vertices is constructed based on graph theory, and the embedding of the vertices is calculated. The vertex paths in the graph structure form the input sequence 1812, which is input into the visualization graph modality module 1814. The token embedding 1816 is based on the graph embedding, and the positional embedding 1818 is based on the sequence of the graph. A dedicated transformer structure module 1820 is used to model the graph structure modality of the visualization chart, with a multi-head attention mechanism (from head 1 1824 to head n 1822). A fusion layer 1838 with a cross-modal attention mechanism is added after the first two modules. The output of the model is multimodal, including natural language, the language representation of the visualization chart, and the graph structure representation of the visualization chart.

[0319] Although this application focuses on chart generation, it does not limit the modality to language and visualization chart modalities. Some charts contain images, so an image modality path can be added in parallel in the model structure, along with other modalities.

[0320] When this hybrid modality chart generation model is used to generate new charts, the prompt input can include not only natural language prompts but also graph-based information. This new method of using graphs for information retrieval will be discussed in detail in Section 5.3. This retrieved information will be used as the prompt input.

[0321] Section 5.3 Enabling Graph Modality Generation

[0322] Traditional retrieval-augmented generation (RAG) of large language models uses the similarity of language embeddings (the opposite of distance) to retrieve the most relevant natural language fragments to a natural language query in the generation step. However, those skilled in the art believe that traditional retrieval-augmented generation lacks connectivity between fragments, so a natural language-based graph retrieval-augmented generation method has evolved from traditional retrieval-augmented generation, a concept popularized by Microsoft and whose pipeline was developed by LlamaIndex. In natural language-based graph retrieval-augmented generation, the similarity retrieval step is graph-based; but in the generation step, it only utilizes the natural language attributes of the retrieved nodes and puts these retrieved natural language fragments into the prompt for generation. For example, if movie IDs are vertices on a graph and each movie ID has natural language attributes such as plot summary, plot, synopsis, director, genre, etc. In natural language-based graph retrieval-augmented generation, in the retrieval phase, a natural language prompt query may be mapped to a movie ID cluster center, or a specific movie ID, to retrieve all movie IDs in that cluster, or the top K movie IDs closest to the specific movie ID mapped by the query. In the generation step, natural language-based graph retrieval-augmented generation can only use the natural language attributes of these retrieved movie IDs to generate a natural language response. For example, if the query is "Summarize the adventure movies of person A", the cluster of person A will be found, all his adventure movie IDs will be found, and the plot or synopsis will be summarized according to specific prompt instructions. However, natural language-based graph retrieval-augmented generation cannot input movie IDs (as vertices) into the prompt and generate a sequence of movie IDs (a sequence of vertices). The fundamental reason for this inability to generate based on the rich connection relationships between vertices in the graph is that the generation model of language-based graph retrieval-augmented generation only has the natural language modality.

[0323] The Transformers4Rec concept proposed by Meta treats item IDs as tokens. However, in Transformers4Rec, these item IDs only form sequences, and graphs are not used to describe the complex relationships between these tokens. In fact, for a recommendation system, the relationships between item IDs are not just the next item in a sequence. Transformers4Rec only switches the method of generating the next item in the sequence from a recurrent neural network to a transformer. For example, the adventure movies of person A can be arranged in chronological order to form a sequence of the next items in Transformers4Rec. However, the movies of person A may share certain similarities with other adventure movies (such as movie B or movie C), and these relationships of person A's movies can form edges on a graph, but it is difficult to capture them in the sequence of the next items. If in the training data, few or no users have watched both person A's movie and movie C, although the graph may discover many other similarities between them, it will be difficult to generate the next recommended item that includes both person A's movie and movie C only using the sequential approach (without layers).

[0324] Various embodiments of the present application provide a new graph modality generation method, which is not limited to the graph specifically used for visualizing diagrams in Section 5.2, but is applicable to graphs with any type of vertices, such as movie IDs, item IDs, academic IDs. Similar to the visual graph modality module 1514, a transformer-based model takes a sequence of vertex IDs as input and outputs a sequence of the next vertex IDs. The difference from the previous graph RAG based on natural language is that the graph modality generation proposed in the present application can generate vertex IDs, which is not available in the previous natural language-based graph RAG. The difference from Transformers4Rec is that the graph modality generation uses graph paths to generate sequences, covering a richer token connectivity than the sequential approach alone.

[0325] The application scope of graph modality generation is very wide and is not limited to the following examples. It can be used for enabling the action sequences of AI agents, such as in areas like price negotiation, law, tax strategies, etc., where actions can be vertices; it can be used for items in a recommendation system, such as products, movies, articles, advertisements, videos, etc. in e-commerce; it can be used for generating visual diagrams or 3D models.

[0326] For the generation of the action sequence of an AI agent, there is already a representation in the figure with actions as vertices. Given an action that has been taken, the goal is to find the next action. In the past, when there was only the figure, the problem was transformed into finding the nearest or next neighbor of the vertex representing the taken action, or finding a path to the target vertex. Now, through the vertex sequence generated by the transformer, given the sequence of past action vertices, the natural language prompt of the target, the prompt of the target vertex, the natural language attributes of the action and the target vertex, and other prompts, the model can generate the sequence of the next action vertices. For the recommendation system, those skilled in the art can borrow the idea of Transformers4Rec and use the next item (vertex) as the recommended item. For chart generation, the relevant methods have been described in detail in Section 5.2.

[0327] Section 5.4 Search and Drop Shipping Enabled by Image Generation

[0328] A user may search for a "backless swing dress", and the items returned may belong to the backless dress group but have a narrow skirt or a tight hem, or be a swing dress or dress that completely covers the back. The reason is that the dresses in the search results do not actually have both of these attributes at the same time, which also constitutes the situation of most of the available dresses on the Internet. The only dress that may match the search query is actually called the "Cinderella" dress on the seller's webpage. Its description does not mention in natural language that the dress is backless or swing, but rather quotes the fairy tale in a romantic way and describes how this dress changes the life of the wearer.

[0329] Given the rich database of dress pictures on the existing Internet, each of which has the attribute of "backless" or "swing", those skilled in the art can first use a convolutional neural network to label these attributes. Then, a large multimodal model with natural language (the natural language attributes of the pictures and other corpora) and the image modality can be constructed to generate dress pictures with both of these attributes for users to select their favorite styles. Subsequently, picture-to-picture search based on the similarity of the convolutional neural network can be used to find matching dresses, even if the existing natural language descriptions do not contain keywords related to these attributes. If no similar dress can be found, a drop shipping service system can be introduced to customize and order the dress.

[0330] Obviously, the proposed solution is not limited to dresses. Any product, such as cars, furniture, etc., can utilize image-enabled search and dropshipping to enhance product discovery and shopping experience. Dropshipping is a retail fulfillment method that does not require inventory. When a customer places an order, the product is produced on demand. The steps are as follows: Image attribute recognition: Segment the product image, annotate the attributes of the segmented image in natural language, and use a convolutional neural network or other models to output the natural language attributes of the image parts; Multimodal model training: Incorporate the image and its natural language attributes and other natural language corpora into the training data for training; Generate product pictures: Generate product pictures based on the user's natural language query; Attach pictures to picture search: Conduct additional picture-to-picture search to enhance search results; Dropshipping service: If the user is not satisfied with the search results, provide a dropshipping service to meet the user's needs.

[0331] Section 5.5 EquiFormer Integrated Multimodal Automotive Base Model

[0332] Those skilled in the art recognize that in the EquiFormer framework, a car is a special example of a device. In Figure 1 a part of the functions of a car driver is to act as a hardware administrator 154, and an autonomous driving system can be combined with the EquiFormer system. In particular, the hierarchical model structure therein can predict the more distant future, providing additional reaction time beyond immediate intervention for the system and the human driver. Those skilled in the art also recognize that the current form of the EquiFormer system in this application mainly focuses on controlling hardware through the car and the model network, but lacks interaction with other human administrators, drivers, pedestrians, passengers, etc. Road conditions, weather conditions, etc. have been integrated as inputs into the EquiFormer system.

[0333] In Section 5.5, this application proposes that a large automotive base model should include a hardware device network module, a human behavior network module, an energy module, etc., and its structure is similar to Figure 18 including a multimodal module and fusing them together, with the EquiFormer model structure as the backbone of the hardware module of this large automotive base model. Figure 19Shows a multi-modal modular base model for automotive concepts. This model includes weather service 1900, traffic control 1902, road conditions 1904, and other input-related services 1906. These services provide input 1908 to the automotive base model, which consists of the following modules: a hardware module 1910 (where a structure similar to EquiFormer can be used), a human module 1912, an energy module 1914, and other modules 1916 (optionally including an autonomous driving module). These modules are fused by a fusion module (1918) to produce an output. Within the hardware module 1910, dedicated sub-modules can be selectively set according to different types of vehicles (such as internal combustion engine vehicles, hybrid vehicles, range-extended electric vehicles, battery electric vehicles, and hydrogen-powered vehicles), especially when non-transformer-based modeling techniques are used in this module. If a "large" model based on transformers is preferred in the hardware module, then these types of vehicles should be regarded as similar to languages in large language models and uniformly built into the transformer model structure.

[0334] References:

[0335] USD962,867S, 2022-09-06, Title: Inductor.

[0336] US10,808,961B2, 2020-10-20, Title: Energy Saving Controller.

[0337] US2019 / 0257539 A1, 2019-08-22, Title: Realtime, Verified and Automated Demand Response Energy Saving Controller.

[0338] US2019 / 0128548 A1, 2019-05-02, Title: Energy Saving Controller.

[0339] US10,174,966B2, 2019-01-08, Title: Energy Saving Controller.

[0340] US10,119,719B2, 2018-11-06, Title: Energy Saving Controller.

[0341] US10,066,849B2, September 4, 2018, Title: Energy Saving Controller.

[0342] US10,047,969B2, August 14, 2018, Title: Energy Saving Controller.

[0343] US2018 / 0038611 A1, February 8, 2018, Title: Energy Saving Controller.

[0344] US2017 / 0051936 A1, February 23, 2017, Title: Energy Saving Controller.

[0345] US 9,419,543 B2, August 16, 2016, Title: Controlled Resonance in Electrical Power Devices.

[0346] US 9,410,713 B2, August 9, 2016, Title: HVAC Fan Controller.

[0347] US2016 / 0223219 A1, August 4, 2016, Title: Energy Saving Controller.

[0348] US2015 / 0159905 A1, June 11, 2015, Title: Energy Saving Controller.

[0349] US2015 / 0060557 A1, March 5, 2015, Title: Energy Saving Apparatus, System and Method.

[0350] US 6,498,546 B1, December 24, 2002, Title: Utilization of Proximity Effect in Ripple Noise Filtering.

[0351] US 6,329,726 B1, December 11, 2001, Title: Proportional Distribution of Power from a Plurality of Power Sources.

[0352] Liukis, A. (2020). Approaching Time-Series with a Tree-based Model from https: / / towardsdatascience.com / approaching-time-series-with-a-tree-based-model-87c6d1fb6603.

[0353] Akiba, T., et al. (2019). Optuna: A next-generation hyperparameter optimization framework. Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining.

[0354] Bergstra, J., et al. (2013). Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures. International conference on machine learning, PMLR.

[0355] Wang, J. (2020). Deep Learning Recommender System, Publishign House of Electronics Industry.

[0356] Pan, J., et al. (2019). "Order matters at fanatics recommending sequentially ordered products by LSTM embedded with Word2Vec", arXiv preprint arXiv:1911.09818.

[0357] Zhang, Z., et al. (2020). "Deep learning on graphs: A survey", IEEE Transactions on Knowledge and Data Engineering 34(1): 249-270.

[0358] Gal, R., et al. (2022). "An image is worth one word: Personalizing text-to-image generation using textual inversion", arXiv preprint arXiv:2208.01618.

[0359] Fu, Y., et al. (2022). "MIGA: A Unified Multi-task Generation Framework for Conversational Text-to-SQL", arXiv preprint arXiv:2212.09278.

[0360] Patarusak, P. and X. Yuan (2009). "Bandwidth optimal all-reduce algorithms for clusters of workstations", Journal of Parallel and Distributed Computing 69(2): 117-124.

[0361] All references cited in this application are hereby incorporated by reference in their entirety into this application.

[0362] For ease of description, the components of the present device can be divided into various modules or units according to their functions and described separately. Of course, when implementing various embodiments of the present application, those skilled in the art can realize that the functions of these modules or units can be implemented using one or more equivalent hardware or software units.

[0363] Various device components, units, modules or parts can have a modular configuration or consist of discrete components, but can still be collectively referred to as "modules". In other words, the "components", "modules" or "units" mentioned in the present application may or may not be in modular form.

[0364] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, various embodiments of the present application can be a completely hardware embodiment, a completely software embodiment, or a hardware-software hybrid embodiment. In addition, various embodiments of the present application can be implemented in the form of a computer program product, which is stored in one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical discs, etc.) and contains computer-executable process code.

[0365] The various embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products of the embodiments of the present application. It should be understood that the computer program instructions implement each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded memory or other programmable data processing device, so that these instructions are executed by the processor of the computer or other programmable data processing device, thereby generating a machine that executes the specified functions in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0366] These computer program instructions can also be stored in a computer-readable memory, such as a non-volatile computer-readable storage medium. These instructions can direct the computer or other programmable data processing device to operate in a specified manner, so that the instructions stored in the computer-readable memory generate a manufactured article, including an instruction device. The instruction device executes the specified functions in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0367] These computer program instructions can also be loaded onto the computer or other programmable data processing device to perform a series of operations and steps on the computer or other programmable data processing device, so that the instructions executed on the computer or other programmable data processing device provide steps for executing the specified functions in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0368] The subject matter and the implementation of the operations of this application can be through digital electronic circuits, computer software, firmware or hardware including the structures disclosed in this application and their structural equivalents, or a combination of one or more of these. The implementation of the subject matter described in this application can be implemented as one or more computer programs, that is, one or more modules of computer program instructions, encoded on one or more computer storage media for execution or control of its operations by a data processing device.

[0369] Optionally, or additionally, the program instructions can be encoded in an artificially generated propagated signal, such as a machine-generated electrical signal, optical signal or electromagnetic signal, for transmitting information to a suitable receiving device for execution by the data processing device. The computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random access or serial access storage array or device, or a combination of one or more of them.

[0370] Furthermore, although a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially generated propagated signal. A computer storage medium can also be, or be included in, one or more separate components or media (such as multiple CDs, disks, drives or other storage devices). Thus, a computer storage medium can be tangible.

[0371] The operations described in this application can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or data received from other sources.

[0372] Processors suitable for executing the above computer program instructions include, for example, general and special microprocessors, and one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory, or both. The components of a computer can include a processor configured to perform operations according to instructions and one or more storage devices for storing instructions and data.

[0373] The processor or processing circuit can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, general processors or other electronic components to perform the above image capture method.

[0374] The subject matter described in this application and its operations can be implemented by digital electronic circuitry, by computer software, firmware, or hardware including the structures disclosed in this application and their structural equivalents, or by a combination of one or more of these. The implementation of the subject matter described in this application can be as one or more computer programs, i.e., one or more portions of computer program instructions, encoded on one or more computer storage media for execution or to control the operation of a data processing apparatus.

[0375] Optionally, or additionally, the program instructions can also be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is used to encode information for transmission to a suitable receiving apparatus and for execution by the data processing apparatus. The computer storage media can be or include a computer-readable storage device, a computer-readable storage substrate, a random access or serial access storage array or device, or a combination of one or more of them.

[0376] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components (such as data servers), middleware components (such as application servers), or frontend components (such as client computers having a graphical user interface or a web browser through which users can interact with embodiments of the subject matter described in this specification), or in any combination including one or more such backend, middleware, or frontend components. The various components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0377] In some embodiments, the model can reside on local processing circuitry and storage devices, and the training of the model can also be performed locally. In other embodiments, the model and the training can be remote or distributed, such as in the cloud.

[0378] Data, such as inputs, outputs, and model predictions, can be presented to a user or operator via a display screen, such as an organic light-emitting diode (OLED) display screen and a liquid crystal display (LCD) in a manufacturing production line and / or a control room.

[0379] Although the preferred embodiments of this application have been described, those skilled in the art can make modifications and variations once they understand the basic innovative concepts. Therefore, the appended claims should be construed to include the preferred embodiments as well as all modifications and variations that fall within the scope of this application.

[0380] This description is only used to help understand some possible methods and concepts. At the same time, those skilled in the art can change the specific implementation methods and application scopes according to the concepts of this application. Therefore, the content of this specification should not be construed as a limitation on this application.

[0381] In the foregoing method embodiments, for the sake of simplicity of description, each step is expressed as a combination of a series of operations. However, those skilled in the art will understand that this application is not limited to the specific step sequences described herein.

[0382] According to other embodiments of this application, certain steps can be executed in other sequences, simultaneously, ignored, or other sequences can be added as needed.

[0383] In addition, although the above features may be described as acting in certain combinations, and even initially claimed in this way, in some cases, one or more features can be deleted from the claimed combination, and the claimed combination can refer to a sub-combination of that combination or its variants.

[0384] Similarly, although the operations are depicted in a specific order in the drawings, this should not be construed as requiring these operations to be performed in the order shown or sequentially, nor that all the operations shown in the drawings must be performed to achieve the desired result. In some cases, multi-tasking and parallel processing may be more advantageous. In addition, the separation of the various system components described in the above implementation should not be construed as required in all implementations. It should be understood that the described program components and systems can generally be integrated into a single software product, or packaged into multiple software products.

[0385] Therefore, specific implementations of the subject matter have been described. Other implementations are within the scope of the following claims. In some cases, the operations recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes shown in the drawings do not necessarily need to be performed in the order shown or sequentially to achieve the desired result. In some implementations, multi-tasking or parallel processing can be utilized.

[0386] In addition, those skilled in the art will also understand that the embodiments described in this specification are only some embodiments, and the operations and parts involved are not all necessary. However, those skilled in the art can judge whether the functions in various embodiments are required for a specific application.

[0387] The various embodiments described in this specification are presented in a progressive manner, where the description of some embodiments focuses on the differences from other embodiments, and the same or similar parts in different embodiments are sometimes only described in one embodiment.

[0388] It should also be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that these entities have such an order or sequence. This does not necessarily require or imply any actual relationship or order between these entities or operations.

[0389] In addition, the term "comprising", "including" or any variant thereof is intended to cover the meaning of non-exclusive inclusion. Thus, a process, method, article or apparatus that comprises a set of elements not only includes those elements, but may also include other elements not expressly listed or other elements inherent to such process, method, article or apparatus.

[0390] Without further limitation, the element defined by the sentence "comprising an..." does not exclude the possibility of the existence of other identical elements in the process, method, article or apparatus that comprises the element.

[0391] In the description, regarding devices, terminals, etc., in some cases the singular form is used, and in some cases the plural form is used, and the same is true when describing various embodiments. However, it should be noted that the singular or plural form does not have a limiting meaning, but is only for illustrative purposes. Unless it is expressly stated to use a single device or terminal, etc., or expressly stated to use multiple devices or terminals, etc., the device or terminal can be singular or plural.

[0392] Based on the various embodiments of this application, the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the above terminal device is only for illustrative purposes, and other types of terminals and devices can also adopt the methods disclosed herein.

[0393] Dividing the terminal or device into different "parts", "regions" or "components" only reflects the various logical functions in some embodiments. The actual implementation can adopt other ways of dividing the "parts", "regions" or "components" to achieve the above similar functions, or no division is required. For example, multiple parts, regions or components can be combined together, or can be integrated into another system. In addition, some features can be omitted, and some method steps can also be skipped.

[0394] Those of ordinary skill in the art will understand that some parts or components, etc. in the devices provided in the above various embodiments can be configured in the above one or more devices. They can also be located in one or more devices different from those described by way of example in the above embodiments or the drawings. For example, the circuits, parts or components, etc. in the above various embodiments can be integrated into a module, or divided into multiple sub-modules.

[0395] The numbers of the above various embodiments are only for illustrative purposes and do not represent the preferred selection of the embodiments.

[0396] Although the specific embodiments have been described in detail above, these descriptions are for illustrative purposes only. Therefore, it should be understood that unless explicitly stated, many of the above aspects are not intended as essential or critical elements.

[0397] Based on the present application, those skilled in the art can make various modifications to the disclosed exemplary embodiments and perform equivalent acts without departing from the spirit and scope of the disclosure defined by the following claims. The scope of the claims should be given the broadest interpretation to cover such modifications and equivalent structures.

Claims

1. A hardware control method, comprising: Acquire device sensor data arranged in a first time series from a plurality of sensors; Acquire equipment optimization target data arranged in a second time series from a plurality of optimization targets; Obtain historical data on equipment abnormal events and intervention actions; as well as Get static device input parameters; Applying a time series model to the acquired device sensor data, the historical data of the abnormal events and intervention actions, and the static device input parameters to obtain predicted device sensor data, predicted optimization target values, and predicted abnormal events; Providing a predicted intervention action for abnormal event intervention based on the acquired static device input parameters, the acquired device sensor data, the acquired device optimization target data, the predicted device sensor data, the predicted optimization target value, and the predicted abnormal event; Wherein, the applying time series model further comprises: applying a machine learning architecture including stacking two or more layers of models to the acquired device sensor data, the historical data of device abnormal events and intervention actions, and the acquired static device input parameters; wherein the two-layer or multi-layer model includes the predicted device sensor data, the predicted optimization target value, the predicted intervention action, and the predicted abnormal event; and Wherein, providing a predicted intervention action for intervening in abnormal events includes: sending a control signal to a control circuit controlling the hardware based on the predicted intervention action, so that the manufacturing process is adjusted toward the predicted optimization target value and intervenes in the predicted abnormal event.

2. The method according to claim 1, further comprising: Iterating at least once between the acquiring of historical data of device abnormal events and intervention events, the acquiring of static device input parameters, and the applying of the time series model; as well as The predicted intervention action is provided to a user based on a result of the iteration.

3. The method according to claim 2, wherein: The providing of the predictive intervention action for the intervention of the abnormal event also includes at least one of the following: displaying the result on a display screen, providing an API interface for controlling the hardware, and sending a mobile phone alert to the user; as well as The method is performed by a hardware control system, which is configured to: (1) generating a hardware control signal from the input of the hierarchical machine learning model and the operating state of the current device to change the operating state of the device; (2) executing an optimization strategy by maintaining the hardware at optimal operating parameters in real time or in a planned manner to achieve the predicted optimization target value; as well as (3) performing intervention actions for the predicted abnormal event by preemptively changing device operation to prevent or mitigate the impact of the predicted abnormal event.

4. The method according to claim 1, further comprising: The device sensor data y_(si-t) obtained from the output of the time series model, where i represents the i-th sensor, and the device sensor data changes with time t; The acquired device sensor data y_(si-t) includes temperature data measured at a specified location; The acquired device sensor data y_(si-t) also includes at least one of the following: amplitude, voltage, current, frequency or force of the motor.

5. The method according to claim 1, further comprising: The equipment optimization target data y_(oj-t) obtained from the output of the time series model, wherein j represents the jth optimization target, and the equipment optimization target data changes with time t; Among them, the acquired equipment optimization target data y_(oj-t) includes at least one of the following: energy output, power, torque or energy efficiency of the motor.

6. The method according to claim 1, further comprising: Generate a sequence of target values ​​beyond a single timestamp and use the generated sequence of target values ​​as input for downstream layer models when real data is not available.

7. The method according to claim 6, further comprising: An abnormal event model is constructed in the second layer of the two-layer or multi-layer model, and the abnormal event model predicts the abnormal events predicted in the long term by: The first layer of the two-layer or multi-layer model is capable of predicting long-term device sensor data, and the first layer includes the predicted device sensor data and the predicted optimization target value; The method of processing historical abnormal event data in behavioral units through the selection of survival classifiers overcomes the problem of fewer historical abnormal event labels in the supervised learning concept.

8. The method according to claim 1, wherein: The time series model includes a Transformer model.

9. The method according to claim 8, wherein: The abnormal event model is based on the Transformer model, and when there are a large number of historical abnormal events, the Transformer model can predict the predicted abnormal events; When there is no previous historical abnormal event, the Transformer model is configured to predict the predicted abnormal event; When there is only one previous historical abnormal event, the Transformer model is configured to predict the predicted abnormal event; When there are multiple previous historical abnormal events, the Transformer model is configured to predict the predicted abnormal event.

10. The method according to claim 8, further comprising: An action recommendation model is constructed, wherein the action recommendation model is configured to output the predicted action in the long term from the Transformer model.

11. The method according to claim 8, further comprising: Hardware device parameter optimization is provided based on the Transformer model, wherein the Transformer model includes a base model configured to learn and organize hardware device data from multiple use cases and multiple types of hardware devices and multiple input and output data types and sources.

12. The method according to claim 3, further comprising: Providing an EquiFormer system design based on a network of interconnected devices, the EquiFormer system design including a remotely configurable and programmable photonic programmable processor having configurable parameters derived from device sensor data and a machine learning architecture; The remotely configurable and programmable photonic programming processor comprises a remote switch, a communication component, a control component and a signal conversion component, wherein the control component has a rewritable storage of new parameters and a programming chip for programming control, and the signal conversion component is configured to convert the model parameters into optical signals; and When the network is available, the communication component is configured to communicate bidirectionally with the machine learning framework to reconfigure the configurable parameters.

13. The method according to claim 12, wherein: When the network is in an offline state, the method further includes: Based on light wave or photon communication, the configurable parameters are sent unidirectionally from the machine learning architecture to a photon-controlled photon-programmable processor; The light wave or photon communication can span the distance between the earth and space.

14. The method of claim 13, further comprising a blockchain-based quantum network for universal qubit communication; in, The blockchain-like mechanism is used to transmit qubit-encoded information from an initial source block to subsequently linked blocks; The quanta represent the smallest discrete values ​​of physical properties, including light.

15. The method according to claim 8, further comprising: Based on model parallelism and data parallelism, ring-all reduce operations are performed on deep learning model training including Transformer and / or low-rank adaptation; in: The low-rank matrices A and B in low-rank adaptation are further partitioned into several sub-matrices and placed in the same computational unit together with the partitions corresponding to the deep learning model; or A parameter server may be used to obtain the means of the sub-matrices of A and B for low-rank adaptation; or One or two ring-type global reduction operations may be used to obtain the means of the sub-matrices of A and B for low-rank adaptation.

16. The method according to claim 8, further comprising: At least through the use of general language representation or visualization graph modalities, combined with large language models to generate 2D or 3D graphs.

17. The method according to claim 16, further comprising: Build a fusion Transformer model, where the language modality comes from an existing large language model and the graph modality comes from the graph structure of a visualization chart; or The Transformer model is used to retrieve and generate graph vertex recognition sequences and applied to AI agents, action sequences, and recommendation scenarios; or Build a multimodal vehicle base model.

18. The method according to claim 8, further comprising: Generate images using large multimodal models of language and images to enable search and dropshipping systems.

19. A computer-readable storage medium, characterized in that: Used to store instructions, when the instructions are executed by one or more processing circuits, implement the method according to any one of claims 1-18.

20. A manufacturing system, characterized in that: The invention comprises one or more processing circuits and a computer-readable storage medium storing instructions, wherein when the instructions are executed by the processing circuits, the method according to any one of claims 1 to 18 is implemented.