Intelligent factory workflow monitoring system, method and apparatus based on intelligent assistant

By building a full-stack monitoring system driven by an intelligent assistant, the problems of low efficiency of manual processing, insufficient cross-system collaboration, and poor data integration in traditional workflows have been solved, enabling rapid fault location and autonomous repair, and improving the production efficiency and responsiveness of smart factories.

CN122453335APending Publication Date: 2026-07-24BEIJING UNITED MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNITED MEDIA TECH CO LTD
Filing Date
2026-02-24
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional workflow adjustments suffer from low efficiency due to manual processing, insufficient cross-system collaboration, and poor data integration, leading to inefficient production line operation and delayed fault diagnosis, making it difficult to achieve rapid response and autonomous repair.

Method used

We will build a full-stack monitoring system driven by an intelligent assistant, including an interaction layer, a semantic understanding layer, a decision-making layer, and an execution layer, to achieve multi-dimensional technological breakthroughs. Through the intelligent assistant, user commands will be transformed into structured semantic features to generate executable fault diagnoses and solutions.

Benefits of technology

It achieves millisecond-level response closed loop, rapid fault location and autonomous repair, effectively avoids the risk of production line paralysis, and improves the real-time monitoring and decision-making capabilities of the production environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453335A_ABST
    Figure CN122453335A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a smart factory workflow monitoring system, method and device based on an intelligent assistant. Applied to the technical field of text big data mining, it comprises an interactive layer, a semantic understanding layer, a decision layer and an execution layer connected in turn; the interactive layer converts user original instructions into structured semantic feature vectors to obtain structured features; the semantic understanding layer converts the structured features into machine-operable semantic instruction trees to generate a four-tuple operation instruction set; the decision layer locates the root cause of faults for the four-tuple operation instruction set and generates a dynamic execution scheme to generate a candidate strategy set; and the execution layer monitors the production environment state in real time by analyzing key elements in each instruction operation in the candidate strategy set. In this way, the present application realizes multi-dimensional technical breakthroughs by constructing an intelligent assistant-driven full-stack monitoring system, thereby solving the technical problems of low manual processing efficiency, insufficient cross-system collaboration and poor data integration in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of text big data mining and processing technology, and in particular to a smart factory workflow monitoring system, method and device based on a smart assistant. Background Technology

[0002] In the field of traditional workflow adjustment, relying on manual experience to determine the cause of anomalies requires a lengthy analysis cycle from problem occurrence to localization. During this period, production lines continue to operate inefficiently, cross-system collaboration suffers from significant delays, and manual transmission of operational instructions across multiple platforms is severely lagging, often missing the optimal intervention window. At the data integration level, various industrial systems form closed data silos, with key indicators scattered across incompatible storage architectures, making global status analysis like the blind men and the elephant. Real-time equipment parameters are disconnected from the process knowledge system, and fault diagnosis lacks multi-dimensional evidence chain support.

[0003] Therefore, how to improve the low efficiency of manual processing, insufficient cross-system collaboration, and poor data integration in the traditional workflow adjustment process has become a technical problem that needs to be solved. Summary of the Invention

[0004] This disclosure provides a smart factory workflow monitoring system, method, and device based on a smart assistant. By constructing a full-stack monitoring system driven by a smart assistant, it achieves multi-dimensional technological breakthroughs, thereby solving the technical problems of low efficiency of manual processing, insufficient cross-system collaboration, and poor data integration in existing technologies.

[0005] According to a first aspect of this disclosure, a smart factory workflow monitoring system based on a smart assistant is provided, comprising an interaction layer, a semantic understanding layer, a decision-making layer, and an execution layer connected in sequence. The interaction layer is used to transform the user's original instructions into structured semantic feature vectors to obtain structured features, providing standardized input for the semantic understanding layer; The semantic understanding layer is used to transform the structured features output by the interaction layer into a machine-operable semantic instruction tree, generating a four-tuple operation instruction set. The decision layer is used to locate the root cause of the fault in the quadruple operation instruction set and generate a dynamic execution plan, and generate a candidate strategy set. The execution layer is used to receive the candidate strategy set generated by the decision layer, and monitor the production environment status in real time by analyzing the key elements in the instruction operation item by item.

[0006] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the interaction layer includes an input preprocessing module, an input type determiner, a modality-specific processor, an intent routing module, and a structured output generator; The input preprocessing module is used to acquire the raw input data, perform denoising and standardization on the raw input, and output a feature sequence. The input type determiner is connected to the input preprocessing module and is used to determine the input data type based on the feature sequence to obtain input features of different modal types. The modality-specific processor is connected to the input type determiner and is used to process the input features of different modality types to form multimodal input; The intent routing module is connected to the modality-specific processor and is used to map multimodal inputs to a predefined intent space to form intent types; The structured output generator is connected to the intent routing module and the modality-specific processor, respectively, to acquire multimodal inputs and intent types, thereby generating and outputting standardized features.

[0007] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the standardized features output by the structured output generator are specifically: in, Y As a standardized feature, M c This is the constraint matrix. E For the entity list, This is an intent type.

[0008] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the input type determiner includes a text purifier, a language feature extractor, and a visual annotation parser. The text purifier is used to extract text signals from the feature sequence and output the segmented text feature sequence. The language feature extractor is used to extract language features from the feature sequence, obtain language features, and construct the Mel spectrum feature matrix; The visual annotation parser is used to extract image features from the feature sequence and annotate them to obtain a set of objects.

[0009] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the modality-specific processor includes a text purifier, a speech feature extractor, and a visual annotation parser; The text purifier is used to purify text data in different modal inputs, eliminate colloquial ambiguity, and form purified text feature input. The speech feature extractor is used to process speech data in different modal types of input, converting speech acoustic features into semantic units to form an output phoneme sequence; The visual annotation parser is used to process image data input from different modal types, extract operable instructions from the images, and output a set of instruction components.

[0010] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the semantic understanding layer includes a domain knowledge enhancement module, a multi-intent decoupler, a spatiotemporal context fusion module, and a semantic instruction set generation module; The domain knowledge enhancement module is used to bind general entities to specific industrial objects and output the triple attribute set of the bound entities. The multi-intent decoupler is connected to the domain knowledge enhancement module and is used to decompose composite intents into atomic parameterized operation sequences based on the entity triplet attribute set. The spatiotemporal context fusion unit is connected to the domain knowledge enhancement module and is used to integrate user constraints and production environment constraints based on the entity triple attribute set to obtain fused constraints. The semantic instruction set generation module is connected to the domain knowledge enhancement module, the multi-intent decoupler, and the spatiotemporal context fusion unit to obtain an executable instruction structure.

[0011] In addition to the aspects and any possible implementations described above, a further implementation is provided in which the decision layer includes a multi-source data fusion engine, a fault diagnosis engine, and a strategy generator; The multi-source data fusion engine is used to execute atomic operation sequences to obtain aggregated device data, standardize the data, and obtain an output matrix; The fault diagnosis engine is connected to the multi-source data fusion engine and is used to locate the root cause of system anomalies based on the output matrix, and to calculate the probability of the anomaly and the contribution of the root cause. The strategy generator is connected to the fault diagnosis engine and is used to generate a set of feasible solutions based on the fault type, thereby obtaining a set of candidate strategies.

[0012] In addition to the aspects and any possible implementations described above, a further implementation is provided, wherein the execution layer includes an execution module and a real-time monitoring module; The execution module is used to parse candidate strategy instructions and allocate resources, generate a complete solution package, and simultaneously monitor the production environment status in real time. The parsing of candidate strategy instructions includes parsing equipment operation instructions, material allocation requirements, and personnel collaboration requirements. The production environment status includes the current operating status of equipment, material inventory status, and personnel availability status. The real-time monitoring module is connected to the execution module and is used to establish a three-dimensional monitoring system in real time based on the state of the winning environment, set dynamic early warning thresholds, and perform real-time monitoring during the execution process.

[0013] According to a second aspect of this disclosure, a smart factory workflow monitoring method based on a smart assistant is provided, comprising the following steps: The user's original commands are transformed into structured semantic feature vectors to obtain structured features; The structured features are transformed into a machine-operable semantic instruction tree, generating a four-tuple operation instruction set; The root cause of the fault is located in the quadruple operation instruction set and a dynamic execution plan is generated, and a candidate strategy set is generated. Receive candidate strategy sets, analyze key elements in instruction operations item by item, and monitor the production environment status in real time.

[0014] According to a third aspect of this disclosure, an electronic device is provided, the electronic device comprising: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to execute modules of the system.

[0015] Compared with the prior art, the present invention has the following technical effects: This invention, through a full-stack monitoring system driven by an intelligent assistant, effectively achieves millisecond-level response closed-loop. The system instantly transforms natural language commands into precise execution plans, constructing an ultra-fast closed loop from problem identification to handling. Fault location is significantly reduced in scale, and major anomalies are autonomously repaired. In typical scenarios, such as sudden malfunctions of critical equipment on the production line, the system completes root cause diagnosis and triggers control mechanisms in a very short time, effectively avoiding the risk of production line paralysis.

[0016] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. The drawings are provided for a better understanding of the invention and are not intended to limit the scope of this disclosure. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure, is shown. Figure 2 This diagram illustrates the structure of the interaction layer 1 of a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure. Figure 3This diagram illustrates the structure of a modal dedicated processor 13 for a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure. Figure 4 This diagram illustrates the semantic understanding layer 2 structure of a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure. Figure 5 This diagram illustrates the structure of the decision layer 3 of a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure. Figure 6 This diagram illustrates the structure of execution layer 4 of a smart factory workflow monitoring system based on a smart assistant, according to an embodiment of the present disclosure. Figure 7 A schematic diagram of a smart factory workflow monitoring method based on a smart assistant, according to an embodiment of the present disclosure, is shown. Figure 8 A schematic diagram of an electronic device structure according to an embodiment of the present disclosure is shown. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Reference Figure 1 As shown, this embodiment provides an intelligent factory workflow monitoring system based on an intelligent assistant, including: an interaction layer 1, a semantic understanding layer 2, a decision layer 3, and an execution layer 4.

[0021] In this embodiment, the interaction layer 1 transforms the user's original instructions into structured semantic feature vectors, providing standardized input for the semantic understanding layer 2. The interaction layer 1 can achieve: multimodal compatibility: supporting unified processing of text, voice, and image (such as device screenshot annotation) inputs; intention routing capability: automatically allocating processing paths according to input type; noise resistance: filtering out industrial environmental noise (such as equipment roaring noise interference); supporting multimodal input (text / voice), receiving user queries (such as "check the fault status of production line A") or operation instructions (such as "optimize the workflow of the welding station").

[0022] Specifically, such as Figure 2As shown, in the interaction layer 1, the original input data is obtained through the input preprocessing module 11, and the original input is denoised and standardized to output the feature sequence.

[0023] Let the original input signal be X, and the output be the standardized signal. : (1) in, An adaptive filtering function (such as Wiener filtering) can suppress environmental noise. M For industrial applications, mask matrix This is the normalization function.

[0024] After standardization, the data type of the feature sequence is determined by the input type judge 12, and the input modality is determined based on the information entropy. If it is a text signal, the word segmentation sequence W={ω1, ω2, ω3, ..., ωk} is output by the text purifier 121. If it is a speech signal, the speech features are extracted by the speech feature extractor 122, and the Mel spectrum feature matrix F is constructed. If it is an image signal, the image features are extracted by the visual annotation parser 123 and the object set O is annotated.

[0025] like Figure 3 As shown, the modal processor 13 performs modal-specific processing on text, speech and image information respectively to form multimodal input.

[0026] The text data is cleaned using a text purifier 131 to eliminate colloquial ambiguity. Domain dictionary enhancement is performed, mapping W to factory terminology vectors. ; (2) in, For factory equipment dictionary, Sim is for word embedding similarity.

[0027] For the term vector V ind Colloquialism disambiguation is performed, removing industrial stop words to obtain the purified text feature input: (S is a set of industrial stop words,) (To clean up text feature input).

[0028] The speech data is processed by the speech feature extractor 132, which converts the speech acoustic features into semantic units to form the output phoneme sequence S.

[0029] (3) Where A is the acoustic model, and the output phoneme sequence S is... This is a connection time-series classification and alignment algorithm.

[0030] The image data is processed by the visual annotation parser 133 to extract operable instructions from the image and output the instruction component set C.

[0031] (4) The output yields a set of instruction components C (e.g., {"object":"robotic arm","action":"calibration","location":[x,y,z]}).

[0032] In the interaction layer 1, the multimodal input is mapped to a predefined intent space through the intent routing module 14 to form intent types.

[0033] (5) (6) in, P The input feature vector is (text word frequency / voice command probability / visual object distribution). α i , β i For industrial scenarios, This is an intent type.

[0034] In interaction layer 1, standardized features are generated by structured output generator 15 and output to semantic understanding layer 2. The specific output structure is as follows: (7) Among them, among them, Y As a standardized feature, M c This is the constraint matrix. E The entity list is as follows: .

[0035] In this embodiment, the interaction layer 1 receives various input information from the user and processes it to generate a structured output that the system understands. For example, it can be used for voice input such as "Check the faults of production line B in the past 2 hours, focusing on the welding robot". The structured output is: {"intent":"monitoring","entities":["production line B","welding robot"],"contraints":{"time_range":["now-2h","now"],"priority":"high"}}.

[0036] In this embodiment, the semantic understanding layer 2 transforms the structured features output by the interaction layer 1 into a machine-operable semantic instruction tree. The semantic understanding layer 2 is the core of the intelligent assistant, employing a pre-trained natural language parsing model from the industrial field (such as the BERT industrial variant) to parse user commands into structured semantic features. User requirement elements: target object (equipment / production line), operation type (monitoring / adjustment / diagnosis), and constraints (time / cost).

[0037] Workflow elements: associated PLC controllers, MES work orders, and SCADA real-time data streams.

[0038] In this embodiment, as Figure 4 As shown, the semantic understanding layer 2 binds general entities to specific industrial objects through the domain knowledge enhancement module 21, thereby mapping them to specific physical devices.

[0039] The specific binding process is as follows: First, the entity is parsed for synonyms, a synonym parsing function is constructed, and the output is as follows: (8) in, e i For input entities, i.e., the name of the original equipment / workstation in the user command, T is the factory thesaurus. Let be the judgment function, and s be the similarity function. Standardized device identifiers.

[0040] Perform device topology expansion to obtain a set of bound entities, specifically: (9) in, For mapping functions, R k For device-associated resource tuples, S k For the associated sensor set, P k This is the system's home path.

[0041] The final domain knowledge enhancement module 21 outputs the attribute set of the bound entity triple. .

[0042] In this embodiment, the semantic understanding layer 2 decomposes the composite intent into a sequence of atomic parameterized operations through the multi-intent decoupler 22.

[0043] Construct the intent-action mapping: (10) Where A is the sequence of atomic operations. For atomic operations, the mapping rule is defined as follows: .

[0044] For example, in this embodiment, if the intent type is monitoring, the corresponding atomic operation sequence is γ1 = data reading, γ2 = threshold detection, and γ3 = status report generation; if the intent type is diagnosis, γ1 = historical data query, γ2 = fault feature matching, and γ3 = root cause analysis.

[0045] Subsequently, the atomic operation parameters were determined through an instance: (11) in, For generating functions, This is an instance of atomic operation parameters.

[0046] For example, when γ j =Data reading, corresponding .

[0047] The final output is the parameterized operation sequence. .

[0048] In this embodiment, the semantic understanding layer 2 uses the spatiotemporal context fusion unit 23 to integrate user constraints and production environment constraints to obtain fused constraints.

[0049] The spatiotemporal context fusion unit 23 obtains the environmental constraints, specifically: (12) in, Extraction function for environmental constraints, This is the environmental constraint matrix, which includes time constraints, process constraints, and safety constraints.

[0050] The final fusion constraints are: (13) in, To constrain the fusion operator, For the user constraint matrix, For the environmental constraint matrix, For user i-th type constraints, Let p be the i-th type of constraint in the environment, and p be the total number of constraints.

[0051] In summary, in this embodiment, the semantic understanding layer 2 constructs an executable instruction structure through the semantic instruction set generation module 24.

[0052] First, the fusion constraints are injected into the operation parameters. Constraints are injected into the parameters of each atomic operation, specifically as follows: (14) For example, time constraints Inject data reading operation.

[0053] Subsequently, based on parameterization, the operation dependencies are obtained, and the transitive dependency relationships between different operations are obtained, specifically: (15) in, It is a set of dependencies between operations. , For atomic operation pairs, Let p be the set of output variables. Let q be the set of input variables for operation q.

[0054] The final output yields the following set of instructions for quadruples operations: .

[0055] In this embodiment, the decision layer 3 processes the quadruple operation instruction set to locate the root cause of the fault and generate a dynamic execution plan.

[0056] like Figure 5 As shown, the decision layer 3 executes an atomic operation sequence through the multi-source data fusion engine 31 to obtain aggregated device data, performs data standardization processing, and obtains an output matrix.

[0057] The specific execution procedure is as follows: (16) in, This is a dependency-based topological sorting algorithm.

[0058] Based on the execution program, atomic operations are performed to obtain the raw output data of the operations. This data is then standardized to produce the output matrix X. (17) in, To manipulate the raw output data, the data matrix , where m is the number of samples and k is the number of features.

[0059] In this embodiment, the decision layer 3 uses the fault diagnosis engine 32 to locate the root cause of system anomalies based on the output matrix, and calculates the probability of anomalies and the contribution of the root cause.

[0060] After obtaining the operational data matrix, the specific fault type is determined through rule matching: (18) in, F For a set of preset fault modes, f i For the i-th fault mode, is the fault feature template, and Sim is the similarity calculation function.

[0061] Based on fault type, neural network prediction is performed in complex scenarios to obtain fault probability and root cause contribution, specifically: (19) MLP stands for Industrial Fine-tuning Neural Network (where the input layer consists of equipment operating condition features). Let W be the feature extraction function, W be the neural network weight matrix, and b be the bias vector.

[0062] The final output yields the failure probability and root cause contribution, for example: bearing wear: 72%, voltage instability: 28%.

[0063] In this embodiment, the decision layer 3 generates a set of feasible solutions based on the fault type through the strategy generator 33, thereby obtaining a set of candidate strategies.

[0064] A set of candidate strategies is generated based on the fault type and the user constraint set: (20) in, For the set of candidate strategies, K is the fault-policy mapping rule base, where K is the number of matching policies.

[0065] In this embodiment, the execution layer 4 receives the candidate strategy set generated by the decision layer 3, and the production environment status is detected in real time by parsing the key elements in the operation instructions one by one.

[0066] In execution layer 4, such as Figure 6 As shown, the execution module 41 performs candidate strategy set instruction parsing and resource allocation. It receives a complete solution package generated by the decision layer (including specific operation steps, resource requirements, expected indicators and rollback plans), and analyzes the key elements in the operation instructions one by one: equipment operation instructions (such as "reduce the welding robot current to 38A"), material allocation requirements (such as "call AGV to transport bearing spare parts to workstation 3"), and personnel collaboration requirements (such as "requires on-site confirmation by a Level 2 qualified technician"), and at the same time, it monitors the production environment status in real time, including: the current working status of the equipment (running status / standby status / fault status), material inventory status (spare parts inventory / transportation status), and personnel dispatchability status (job qualification matching degree / current location).

[0067] Meanwhile, in execution layer 4, a three-dimensional monitoring system is established through real-time monitoring module 42, including: equipment-level monitoring: real-time acquisition of operating parameters of execution equipment; process-level monitoring: tracking the completion progress and quality of each operation step; environmental-level monitoring: monitoring environmental factors such as workshop temperature and humidity, and power supply. Secondly, set dynamic early warning thresholds, including basic thresholds and adaptive thresholds. The basic threshold is the safe range of equipment parameters, while the adaptive threshold needs to be automatically adjusted according to the operating conditions.

[0068] In existing technologies, intelligent factory workflow monitoring has long faced three major technical bottlenecks: 1. Inefficient cross-system collaboration. Traditional systems rely on manual switching between independent platforms such as SCADA, MES, and PLC. Fault diagnosis requires engineers to manually verify data across 5-7 interfaces, with an average response delay exceeding 45 minutes. For example, when an abnormality occurs on the welding production line in an auto parts factory, it is necessary to retrieve equipment logs (SCADA), process parameters (MES), and work order records (ERP) sequentially, with data collection alone taking 32 minutes. 2. Disconnect between decision-making and execution. Existing monitoring systems only provide "dashboard displays" and lack autonomous decision-making capabilities. When motor overheating is detected, the system pops up an alarm but cannot generate an executable solution. Maintenance personnel must rely on experience to judge whether to reduce the load (affecting production capacity) or replace the cooling module (2-hour downtime), resulting in a decision error rate of up to 35%. 3. Difficulty in knowledge reuse. The fault-handling experience accumulated in the factory is stored in paper work orders or scattered electronic documents. The 17 methods for handling bearing jamming on a certain home appliance production line are scattered in the maintenance foreman's manual, equipment manufacturer's emails, and senior technicians' notes. New employees need an average of 6.8 trials and errors to find the optimal solution.

[0069] This embodiment utilizes a full-stack monitoring system driven by an intelligent assistant to effectively achieve a millisecond-level response closed loop. The system instantly transforms natural language commands into precise execution plans, constructing a rapid closed loop from problem identification to handling. Fault location is significantly reduced in scale, and major anomalies are autonomously repaired. In typical scenarios, such as sudden malfunctions of critical equipment on the production line, the system completes root cause diagnosis and triggers control mechanisms in a very short time, effectively mitigating the risk of production line shutdown.

[0070] like Figure 7 As shown, this embodiment also provides a smart factory workflow monitoring method based on a smart assistant, including the following steps: S101. Transform the user's original instructions into a structured semantic feature vector to obtain structured features; S102. The structured features are transformed into a machine-operable semantic instruction tree to generate a four-tuple operation instruction set. S103. Perform fault root cause localization on the quadruple operation instruction set and generate a dynamic execution plan, and generate a candidate strategy set; S104. Receive the candidate strategy set and monitor the production environment status in real time by analyzing the key elements in the instruction operation item by item.

[0071] Furthermore, some embodiments of this application also provide an electronic device. The electronic device can be various forms of digital computer, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0072] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the methods provided in any one or more of the above embodiments. Figure 8 An exemplary structural diagram of the electronic device is disclosed. For example... Figure 8 As shown, the electronic device includes one or more processors 1101, a memory 1102, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0073] The electronic device may further include an input device 1103 and an output device 1104. The processor 1101, memory 1102, input device 1103, and output device 1104 may be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.

[0074] Input device 1103 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include a display device, auxiliary lighting device (e.g., LED), and haptic feedback device (e.g., vibration motor). The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0075] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0076] In this embodiment, a computer-readable medium stores a computer program / instructions that, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. This computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into that device. The aforementioned computer-readable medium carries one or more computer-readable instructions.

[0077] The memory 1102 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, thereby implementing the program instructions / modules corresponding to the methods provided in any one or more of the embodiments described above in this application.

[0078] The memory 1102 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1102 may optionally include memory remotely located relative to the processor 1101, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0079] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0080] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0081] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0082] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, magnetic or optical drives, floppy disks, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.

[0083] The computer program product provided in this application includes one or more computer programs / instructions. When executed by a processor, these computer programs / instructions generate, in whole or in part, the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0084] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0085] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0086] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A smart factory workflow monitoring system based on a smart assistant, characterized in that, It includes an interaction layer (1), a semantic understanding layer (2), a decision layer (3), and an execution layer (4) connected in sequence. The interaction layer (1) is used to transform the user's original instructions into a structured semantic feature vector to obtain structured features, which provide standardized input for the semantic understanding layer (2). The semantic understanding layer (2) is used to transform the structured features output by the interaction layer (1) into a machine-operable semantic instruction tree, and generate a set of four-tuple operation instructions. The decision layer (3) is used to locate the root cause of the fault in the quadruple operation instruction set and generate a dynamic execution plan and a candidate strategy set. The execution layer (4) is used to receive the candidate strategy set generated by the decision layer (3), and monitor the production environment status in real time by analyzing the key elements in the instruction operation item by item.

2. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 1, characterized in that, The interaction layer (1) includes an input preprocessing module (11), an input type determiner (12), a modality-specific processor (13), an intent routing module (14), and a structured output generator (15). The input preprocessing module (11) is used to acquire the original input data, and to perform noise reduction and standardization on the original input to output the feature sequence. The input type determiner (12) is connected to the input preprocessing module (11) and is used to determine the input data type based on the feature sequence to obtain input features of different modal types; The modality-specific processor (13) is connected to the input type determiner (12) and is used to process the input features of different modality types to form multimodal input; The intent routing module (14) is connected to the modality-specific processor (13) and is used to map multimodal inputs to a predefined intent space to form intent types; The structured output generator (15) is connected to the intent routing module (14) and the modality dedicated processor (13) respectively, and is used to acquire multimodal input and intent type respectively, thereby generating and outputting standardized features.

3. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 2, characterized in that, The standardized features output by the structured output generator (15) are specifically as follows: in, Y As a standardized feature, M c This is the constraint matrix. E For the entity list, This is an intent type.

4. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 2, characterized in that, The input type determiner (12) includes a text purifier (121), a language feature extractor (122), and a visual annotation parser (123). The text purifier (121) is used to extract text signals from the feature sequence and output the segmented text feature sequence. The language feature extractor (122) is used to extract language features from the feature sequence to obtain language features and construct the Mel spectrum feature matrix; The visual annotation parser (123) is used to extract image features from the feature sequence and annotate them to obtain an object set.

5. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 2, characterized in that, The modality-specific processor (13) includes a text purifier (131), a speech feature extractor (132), and a visual annotation parser (133). The text purifier (131) is used to purify the text data in different modal type inputs, eliminate colloquial ambiguity, and form purified text feature input; The speech feature extractor (132) is used to process speech data in different modal types of input, convert speech acoustic features into semantic units, and form an output phoneme sequence; The visual annotation parser (133) is used to process image data in different modal types of input, extract operable instructions from the image, and output a set of instruction components.

6. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 3, characterized in that, The semantic understanding layer (2) includes a domain knowledge enhancement module (21), a multi-intent decoupler (22), a spatiotemporal context fusion module (23), and a semantic instruction set generation module (24). The domain knowledge enhancement module (21) is used to bind general entities to specific industrial objects and output the bound entity triple attribute set; The multi-intent decoupler (22) is connected to the domain knowledge enhancement module (21) and is used to decompose the composite intent into an atomic parameterized operation sequence based on the entity triplet attribute set; The spatiotemporal context fusion unit (23) is connected to the domain knowledge enhancement module (21) and is used to integrate user constraints and production environment constraints based on the entity triple attribute set to obtain fused constraints; The semantic instruction set generation module (24) is connected to the domain knowledge enhancement module (21), the multi-intent decoupler (22), and the spatiotemporal context fusion unit (23) respectively to obtain an executable instruction structure.

7. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 6, characterized in that, The decision layer (3) includes a multi-source data fusion engine (31), a fault diagnosis engine (32), and a strategy generator (33). The multi-source data fusion engine (31) is used to execute atomic operation sequences to obtain aggregated device data, standardize the data, and obtain an output matrix; The fault diagnosis engine (32) is connected to the multi-source data fusion engine (31) and is used to locate the root cause of system anomalies based on the output matrix, and to calculate the probability of anomalies and the contribution of the root cause. The strategy generator (33) is connected to the fault diagnosis engine (32) and is used to generate a set of feasible solutions based on the fault type to obtain a set of candidate strategies.

8. The intelligent factory workflow monitoring system based on an intelligent assistant according to claim 1, characterized in that, The execution layer (4) includes an execution module (41) and a real-time monitoring module (42). The execution module (41) is used to parse and allocate resources for candidate strategy instructions, generate a complete solution package, and simultaneously detect the production environment status in real time. The parsing of candidate strategy instructions includes parsing equipment operation instructions, material allocation requirements and personnel collaboration requirements. The production environment status includes the current working condition of the equipment, the material inventory status and the personnel schedulable status. The real-time monitoring module (42) is connected to the execution module (41) and is used to establish a three-dimensional monitoring system in real time based on the state of the winning environment, and set dynamic early warning thresholds to perform real-time monitoring during the execution process.

9. A smart factory workflow monitoring method based on a smart assistant, characterized in that, The intelligent factory workflow monitoring system based on an intelligent assistant, as described in any one of claims 1-8, is characterized by comprising the following steps: The user's original commands are transformed into structured semantic feature vectors to obtain structured features; The structured features are transformed into a machine-operable semantic instruction tree, generating a four-tuple operation instruction set; The root cause of the fault is located in the quadruple operation instruction set and a dynamic execution plan is generated, and a candidate strategy set is generated. Receive candidate strategy sets, analyze key elements in instruction operations item by item, and monitor the production environment status in real time.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform modules of the system as described in any one of claims 1 to 8.