An industrial data acquisition system based on AI visual recognition and edge computing
Patent Information
- Application Number
- CN202610729959.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-01
AI Technical Summary
当企业需要将这些系统与其他应用进行数据对接时,面临巨大困难
[0044] 1. This invention achieves automatic learning of user operation processes through an operation sequence learning framework, eliminating the need for professional technicians to develop RPA scripts, shortening the implementation cycle from the traditional 2-4 weeks to 2 hours, and reducing implementation costs by more than 80%.
Abstract
Description
Technical Field
[0001] This invention relates to an industrial data acquisition system based on AI visual recognition and edge computing, belonging to the field of industrial automation and data acquisition technology. Background Technology
[0002] Industrial systems (including Manufacturing Execution System (MES), Enterprise Resource Planning (ERP), Equipment Asset Management (EAM), Customer Relationship Management (CRM), and Quality Management System (QMS)) are the core support platforms for the digital operation of industrial enterprises, storing critical data related to their production and operations. However, due to historical reasons and limitations in their technical architecture, many industrial systems suffer from the following problems:
[0003] 1. Closed Systems and Lack of Open Interfaces: Many industrial systems were built 10-20 years ago, employing closed technical architectures and lacking standardized API interfaces. This creates significant difficulties for enterprises when they need to interface these systems with other applications. While some system vendors offer interface development services, the fees are exorbitant (typically ranging from 500,000 to 2 million RMB per system) and the development cycle can take 3-6 months.
[0004] 2. Severe data silos: Enterprises often have multiple heterogeneous industrial systems that cannot share data, creating data silos. Business personnel need to spend a lot of time manually exporting, organizing, and summarizing data, which seriously affects the efficiency of data-driven decision-making.
[0005] 3. High implementation cost of traditional RPA: Although Robotic Process Automation (RPA) technology can simulate manual operation for data collection, traditional RPA solutions have the following shortcomings: long script development cycle, with a single process configuration requiring 2-4 weeks; requiring professional technical personnel for development and maintenance, making it difficult for business personnel to operate independently; minor changes in the system interface can cause script failures, resulting in high maintenance costs; and limited ability to recognize complex tables and dynamic content. Summary of the Invention
[0006] The purpose of this invention is to provide an industrial data acquisition system based on AI visual recognition and edge computing. By using an operation sequence learning framework, the system enables automatic learning of user operation processes, eliminating the need for professional technicians to develop RPA scripts. This reduces the implementation cycle from the traditional 2-4 weeks to 2 hours and lowers implementation costs by more than 80%.
[0007] Current industrial data acquisition methods primarily rely on system interface integration or manual export, lacking effective technical means for systems with closed APIs. Some solutions employ traditional RPA technology, but its high implementation cost and technical barriers hinder large-scale application. Existing OCR technology is mainly designed for scanned documents, exhibiting low accuracy in recognizing complex tables, dynamic content, and multilingual scenarios in industrial system interfaces, and lacks deep integration with RPA technology. Existing solutions mostly employ cloud processing models, requiring data to be uploaded to the cloud for processing, posing data security risks and privacy breaches, and failing to meet the data localization requirements of industrial enterprises. Existing RPA solutions lack AI capabilities, cannot automatically learn and generate operation processes, requiring significant manual configuration work, resulting in low implementation efficiency.
[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0009] An industrial data acquisition system based on AI visual recognition and edge computing includes:
[0010] Edge Computing Box Module: Serving as the local execution environment, it is responsible for RPA script execution, AI inference, temporary data storage, and secure transmission, connecting to the target PC via plug-and-play. AI Intelligent Recording Module: Based on an operation sequence learning framework, it uses a temporal behavior modeling subsystem, a UI element perception subsystem, and an intent understanding subsystem to automatically learn user operation flows and generate RPA scripts, then sends the generated RPA scripts to the edge computing box module. RPA Execution Engine Module: Receives instructions from the edge computing box module and executes automated operation flows, including desktop application control and automated browser browsing. AI Multimodal Data Structured Processing Module: Based on a vision-language joint understanding architecture, it uses a screen semantic parsing subsystem, an information unit extraction subsystem, and a data paradigm conversion subsystem to achieve intelligent recognition and structured processing of screen data. Demand Understanding Module: Converts user natural language demands into structured data acquisition tasks and sends these tasks to the AI Intelligent Recording Module.
[0011] Operation Sequence Learning Framework: This refers to a machine learning framework based on time-series data modeling, used to learn operation flow patterns from user operation sequences and generate executable automated scripts. The framework comprises three core subsystems: time-series behavior modeling, UI element awareness, and intent understanding.
[0012] Temporal behavior modeling: refers to the technique of modeling and analyzing user operation behavior in the time dimension, including the capture, alignment, segmentation and feature extraction of operation events, which is used to identify behavioral fragments with independent semantics.
[0013] UI element awareness: refers to the technology of detecting, recognizing and tracking user interface elements on the screen through computer vision technology, used to establish the relationship between operation events and interface elements.
[0014] Intent understanding: refers to the technology of recognizing and understanding the business intent of user operations based on natural language processing technology, which is used to map operation sequences into business semantics and extract variable parameters.
[0015] Visual-Language Joint Understanding Architecture: This refers to a multimodal deep learning architecture that integrates visual and linguistic information. It extracts image features through a visual encoder, extracts semantic features through a text encoder, and establishes visual-language associations through a cross-modal fusion layer, thereby achieving a comprehensive understanding of screen content.
[0016] Screen semantic parsing: refers to the technology of visually analyzing screen images, identifying the functional area divisions, UI component types, and visual-language correspondences of the screen, and constructing a semantic representation graph of the screen.
[0017] Information unit extraction: refers to the technology of detecting, identifying and extracting information elements with business meaning from screen content, and assembling them into structured data records according to business logic.
[0018] Data paradigm shift refers to the techniques used to infer data type, standardize format, verify quality, and convert target format of extracted information units, in order to convert raw screen data into standard structured data.
[0019] Multimodal pre-trained models: These are deep learning models pre-trained on large-scale visual-language data. They have the ability to extract visual features, extract text features, and align across modalities. They can be used for downstream tasks such as visual understanding, text recognition, and semantic understanding.
[0020] In a preferred embodiment, the AI intelligent recording module is based on an operation sequence learning framework and includes three subsystems: temporal behavior modeling, UI element perception, and intent understanding.
[0021] The temporal behavior modeling subsystem includes an input capture unit, a temporal alignment unit, a behavior segment extraction unit, and a temporal feature encoding unit. The input capture unit acquires screen frame sequences and input device event sequences at a preset sampling frequency; the temporal alignment unit establishes a frame-event mapping relationship; the behavior segment extraction unit segments the continuous event sequence into behavior segments with independent semantics based on a time window and event density; and the temporal feature encoding unit extracts temporal features from each behavior segment.
[0022] The UI element perception subsystem includes a visual perception unit, a DOM parsing unit, an element association unit, and an element state tracking unit. The visual perception unit detects and identifies UI elements on the screen based on a deep learning object detection algorithm; the DOM parsing unit parses the DOM tree of the web application page; the element association unit establishes associations between input events and UI elements; and the element state tracking unit tracks changes in the state of UI elements.
[0023] The intent understanding subsystem includes an operation semantic recognition unit, a parameter extraction unit, a process template matching unit, and a script generation unit. The operation semantic recognition unit identifies the business intent of user operations based on a pre-trained language model; the parameter extraction unit extracts variable parameters from the operation sequence; the process template matching unit matches behavioral fragments with a predefined process template library; and the script generation unit generates parameterized RPA scripts.
[0024] The AI multimodal data structuring module is based on a vision-language joint understanding architecture, comprising three subsystems: screen semantic parsing, information unit extraction, and data paradigm transformation. This vision-language joint understanding architecture is based on a multimodal pre-trained model, including a visual encoder, a text encoder, a cross-modal fusion layer, and a task decoder.
[0025] The screen semantic parsing subsystem includes a layout analysis unit, a component identification unit, a visual-language alignment unit, and a screen semantic graph construction unit. The layout analysis unit identifies the functional area divisions of the screen; the component identification unit identifies the types of UI components on the screen; the visual-language alignment unit establishes the correspondence between visual elements and text content; and the screen semantic graph construction unit constructs a semantic representation graph of the screen.
[0026] The information unit extraction subsystem includes a text detection unit, a text recognition unit, a table parsing unit, a field semantic understanding unit, and an information unit assembly unit. The text detection unit detects text regions based on a deep learning object detection algorithm; the text recognition unit performs character recognition based on an end-to-end sequence recognition model; the table parsing unit performs structured parsing of table components; the field semantic understanding unit identifies the business semantics of fields based on named entity recognition and a domain knowledge base; and the information unit assembly unit assembles fields into information units.
[0027] The data paradigm conversion subsystem includes a data type inference unit, a data standardization unit, a data quality verification unit, and a target format conversion unit. The data type inference unit infers data types based on regular expression matching and machine learning classification; the data standardization unit performs format standardization processing; the data quality verification unit performs integrity verification, consistency verification, and reasonableness verification; and the target format conversion unit converts the data into the target output format.
[0028] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the temporal behavior modeling subsystem of the AI intelligent recording module includes: an input capture unit for capturing screen frame sequences and input device event sequences, wherein the screen frame sequences are acquired at a preset sampling frequency, and the input device event sequences include mouse events and keyboard events, each carrying a timestamp and coordinate information; a temporal alignment unit for aligning the screen frame sequences with the input device event sequences in time, establishing a frame-event mapping relationship, which is used to determine the screen state corresponding to each input event; a behavior segment extraction unit for segmenting continuous input event sequences into behavior segments with independent semantics based on time windows and event density, wherein each behavior segment corresponds to a complete intent unit of user operation; and a temporal feature encoding unit for extracting temporal features from each behavior segment, wherein the temporal features include operation type sequences, time interval distributions, and operation trajectory features.
[0029] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the UI element perception subsystem of the AI intelligent recording module includes: a visual perception unit: used for visual analysis of screen frames, detecting and recognizing UI elements on the screen, including buttons, input boxes, dropdown boxes, tables, labels, and icons; the visual perception unit outputs the category label, bounding box coordinates, and visual feature vector of each UI element; a DOM parsing unit: used for DOM tree parsing of the web application page, obtaining the hierarchical structure, attribute information, and text content of page elements; the DOM parsing serves as a supplement to visual perception to improve the accuracy of element recognition; an element association unit: used to establish the association between detected UI elements and input events, determining the target element operated by each input event; the association is based on dual verification of coordinate matching and temporal matching; and an element state tracking unit: used to track the state changes of UI elements during the execution of the operation sequence, including element appearance, disappearance, position movement, and content update.
[0030] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the intent understanding subsystem of the AI intelligent recording module includes: an operation semantic recognition unit, used to perform semantic understanding of behavioral fragments based on a pre-trained language model, and identify the business intent of user operations, including data query, report export, information entry, and system navigation; a parameter extraction unit, used to extract variable parameters from the operation sequence, including date range, query conditions, and selection options, the parameter extraction being based on a combination of rule matching and semantic understanding; a process template matching unit, used to match the identified behavioral fragments with a predefined process template library, the process template library containing common process patterns for common business operations; and a script generation unit, used to generate parameterized RPA scripts based on the identified business intents and extracted parameters, the RPA scripts supporting variable substitution, conditional branching, and loop control.
[0031] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the workflow of the operation sequence learning framework includes: S1: Recording begins, and the input capture unit starts acquiring screen frame sequences and input device event sequences; S2: The timing alignment unit establishes a frame-event mapping relationship to determine the screen state corresponding to each input event; S3: The visual perception unit performs UI element detection on key frames to obtain element categories, positions, and features; S4: The element association unit establishes the association relationship between input events and UI elements; S5: The behavior fragment extraction unit divides the continuous event sequence into behavior fragments based on event density and operation type; S6: The operation semantic recognition unit performs semantic understanding on each behavior fragment to identify business intent; S7: The parameter extraction unit extracts variable parameters from the operation sequence; S8: The script generation unit generates RPA scripts based on business intent and parameters.
[0032] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the screen semantic parsing subsystem of the AI multimodal data structuring processing module includes: a layout analysis unit for performing visual layout analysis on screen images and identifying the functional areas of the screen, including navigation areas, content areas, operation areas, and status areas; a component recognition unit for identifying the types of UI components on the screen, including forms, tables, lists, cards, and charts, wherein the component recognition is based on joint judgment of visual features and semantic features; a visual-language alignment unit for establishing the correspondence between visual elements and text content, wherein the visual-language alignment is based on the cross-modal alignment capability of a multimodal pre-trained model; and a screen semantic graph construction unit for constructing a semantic representation graph of the screen, wherein the semantic representation graph includes the functional structure, element relationships, and interaction logic of the screen.
[0033] The information unit extraction subsystem of the AI multimodal data structuring processing module includes: a text detection unit for detecting text regions in screen images, the text detection being based on a deep learning object detection algorithm, outputting the bounding box and confidence score of the text region; a text recognition unit for recognizing characters in the detected text regions, the text recognition being based on an end-to-end sequence recognition model, supporting mixed Chinese and English recognition; a table parsing unit for structured parsing of table components, the table parsing including table boundary detection, cell segmentation, row and column relationship recognition, and table header understanding; a field semantic understanding unit for recognizing the business semantics of the extracted text fields, the business semantics including field name, field type, and field value, the field semantic understanding being based on named entity recognition and a domain knowledge base; and an information unit assembly unit for assembling the recognized fields into information units according to business logic, the information unit being a data record with complete business meaning.
[0034] The data paradigm conversion subsystem of the AI multimodal data structuring processing module includes: a data type inference unit, used to infer the data type of the extracted information units, including numeric, date, text, and enumeration types, based on regular expression matching and machine learning classification; a data standardization unit, used to standardize the format of the identified data, including date format unification, numerical precision standardization, and encoding format conversion; a data quality verification unit, used to verify the quality of the structured data, including integrity verification, consistency verification, and rationality verification; and a target format conversion unit, used to convert the standardized data into a target output format, including JSON, XML, CSV, and database records.
[0035] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the vision-language joint understanding architecture is based on a multimodal pre-trained model. This multimodal pre-trained model includes: a visual encoder for extracting visual features from screen images, including global layout features and local element features; a text encoder for extracting semantic features from screen text, including word-level features and sentence-level features; a cross-modal fusion layer for fusing visual and text features to establish semantic associations between visual elements and text content; and a task decoder for performing specific tasks based on the fused multimodal features, including element detection, text recognition, and semantic understanding.
[0036] The workflow of the AI multimodal data structuring processing module includes: S1: Receiving a screen capture image, the layout analysis unit performs visual layout analysis and identifies functional areas; S2: The component identification unit identifies the types of UI components on the screen; S3: The text detection unit detects text areas in the image; S4: The text recognition unit performs text recognition on the text areas; S5: The visual-language alignment unit establishes the correspondence between visual elements and text content; S6: The table parsing unit performs structured parsing on table components; S7: The field semantic understanding unit identifies the business semantics of fields; S8: The information unit assembly unit assembles fields into information units; S9: The data type inference unit infers the data type of the information units; S10: The data standardization unit performs format standardization processing; S11: The data quality verification unit performs quality verification; S12: The target format conversion unit converts the data to the target output format.
[0037] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, this system employs an anomaly recovery mechanism, specifically including the following steps: S1: Monitor the RPA execution process and detect anomalies, including element not found, timeout, and recognition failure; S2: Classify the detected anomalies and determine their levels, including element-level, page-level, and process-level anomalies; S3: Execute corresponding recovery strategies based on the anomaly level, including retry positioning, page refresh, and process rollback; S4: When automatic recovery fails, record the anomaly information and notify manual intervention; S5: Continuously optimize the AI model and recovery strategies based on anomaly cases.
[0038] The edge computing box module adopts a multi-level security mechanism, including the following: local data processing: sensitive data is processed locally at the edge and is not uploaded to the cloud; encrypted transmission: communication with the cloud uses TLS encryption to ensure data transmission security; device authentication: supports hardware-level device authentication to prevent unauthorized access; operation auditing: all operation logs are fully recorded and audit traceability is supported.
[0039] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the system adopts the following workflow: S1: The user inputs data acquisition requirements or initiates intelligent recording mode via natural language; S2: The AI intelligent recording module, based on an operation sequence learning framework, captures user screen operations and automatically generates RPA scripts through three subsystems: temporal behavior modeling, UI element perception, and intent understanding; S3: The RPA execution engine module loads the RPA scripts and controls the target system to perform corresponding operations; S4: The screen capture module captures page images of the target system; S5: The AI multimodal data structuring processing module, based on a vision-language joint understanding architecture, intelligently recognizes and structures page images through three subsystems: screen semantic parsing, information unit extraction, and data paradigm transformation; S6: The structured data is imported into a large model for analysis or output to a specified target system.
[0040] In the aforementioned industrial data acquisition system based on AI visual recognition and edge computing, the workflow of the AI intelligent recording module in step S2 includes: S21: Recording is started, and the input capture unit acquires screen frame sequences and input device event sequences at a preset sampling frequency; S22: The timing alignment unit establishes a frame-event mapping relationship and determines the screen state corresponding to each input event; S23: The visual perception unit performs UI element detection on key frames and obtains element categories, positions, and features; S24: The element association unit establishes the association relationship between input events and UI elements; S25: The behavior fragment extraction unit divides the continuous event sequence into behavior fragments based on event density and operation type; S26: The operation semantic recognition unit performs semantic understanding on each behavior fragment based on a pre-trained language model and identifies business intent; S27: The parameter extraction unit extracts variable parameters from the operation sequence; S28: The script generation unit generates parameterized RPA scripts based on business intent and parameters.
[0041] The workflow of the AI multimodal data structuring module described in step S5 includes: S51: Receiving a screen capture image, the layout analysis unit performs visual layout analysis and identifies functional areas; S52: The component recognition unit identifies the types of UI components on the screen; S53: The text detection unit detects text regions in the image based on a deep learning object detection algorithm; S54: The text recognition unit performs text recognition on the text regions based on an end-to-end sequence recognition model; S55: The visual-language alignment unit establishes the correspondence between visual elements and text content based on the cross-modal alignment capability of the multimodal pre-trained model; S56: The table parsing unit processes the table components... Row structure parsing includes table boundary detection, cell segmentation, and row-column relationship recognition; S57: Field semantic understanding unit identifies the business semantics of fields based on named entity recognition and domain knowledge base; S58: Information unit assembles the identified fields into information units according to business logic; S59: Data type inference unit infers the data type of the information units based on regular expression matching and machine learning classification; S510: Data standardization unit performs format standardization processing; S511: Data quality verification unit performs integrity verification, consistency verification, and rationality verification; S512: Target format conversion unit converts the data into the target output format;
[0042] The vision-language joint understanding architecture is based on a multimodal pre-trained model, which includes: a visual encoder for extracting visual features from screen images, including global layout features and local element features; a text encoder for extracting semantic features from screen text, including word-level features and sentence-level features; a cross-modal fusion layer for fusing visual and text features to establish semantic associations between visual elements and text content; and a task decoder for performing element detection, text recognition, and semantic understanding tasks based on the fused multimodal features.
[0043] Compared with the prior art, the technical solution of the present invention has the following advantages:
[0044] 1. This invention achieves automatic learning of user operation processes through an operation sequence learning framework, eliminating the need for professional technicians to develop RPA scripts, shortening the implementation cycle from the traditional 2-4 weeks to 2 hours, and reducing implementation costs by more than 80%.
[0045] 2. This invention adopts a vision-language joint understanding architecture, which achieves a comprehensive understanding of screen content through a multimodal pre-trained model. Compared with traditional OCR technology, it can better handle complex tables, dynamic content and multilingual mixed scenarios, and improve the recognition accuracy to over 99%.
[0046] 3. This invention adopts an edge computing architecture, where data processing is completed locally at the edge, ensuring that sensitive data does not leave the factory and meeting the data security and compliance requirements of industrial enterprises. Simultaneously, the edge-cloud collaborative architecture guarantees low latency for local processing while leveraging large cloud models to handle complex scenarios.
[0047] 4. This invention supports natural language input of requirements. Business personnel do not need to learn professional technical knowledge. They can generate data collection tasks by describing them in natural language, which lowers the threshold for use and enables business personnel to complete the data collection configuration by themselves.
[0048] 5. This invention adopts an adaptive learning mechanism, which enables the system to continuously learn from the execution log, optimize the operation process and recognition model, and achieve the effect of "becoming smarter with use", with the accuracy continuously improving over long-term use. Detailed Implementation
[0049] Example 1: An industrial data acquisition system based on AI visual recognition and edge computing, comprising the following modules:
[0050] Edge computing box module: Serving as the local execution environment, it is responsible for RPA script execution, AI inference, temporary data storage, and secure transmission, connecting to the target PC via plug-and-play. The edge computing box uses an RK3576 processor as its main control chip, integrating an 8-core ARM CPU (4×Cortex-A72@2.2GHz + 4×Cortex-A53@1.8GHz) and 6 TOPS+26TOPS NPU computing power (INT4 / INT8), sufficient for local AI inference needs. The box is equipped with 8GB LPDDR4X memory, 64GB eMMC system storage, and 256GB NVMe SSD data cache, supporting dual gigabit Ethernet ports, WiFi 6 wireless connectivity, and an optional 5G module. Screen capture uses an HDMI loop-out solution, supporting 1080p@60Hz video capture without affecting the normal use of the existing monitor.
[0051] AI Intelligent Recording Module: Based on the operation sequence learning framework, it realizes automatic learning of user operation flow and RPA script generation through three subsystems: temporal behavior modeling, UI element perception, and intent understanding.
[0052] The AI multimodal data structuring module, based on a vision-language joint understanding architecture, achieves intelligent recognition and structured processing of screen data through three subsystems: screen semantic parsing, information unit extraction, and data paradigm transformation. The vision-language joint understanding architecture is based on a multimodal pre-trained model, including a visual encoder, a text encoder, a cross-modal fusion layer, and a task decoder.
[0053] Example 2: Operation Sequence Learning Framework
[0054] The workflow of the operation sequence learning framework described in this embodiment includes:
[0055] S1: Start recording. The input capture unit collects screen frame sequences and input device event sequences at a preset sampling frequency (e.g., 30 frames per second). The input device event sequences include mouse events (movement, click, scroll) and keyboard events (key press, input), each event carrying a timestamp and coordinate information.
[0056] S2: Timing Alignment. The timing alignment unit establishes a frame-event mapping relationship to determine the screen state corresponding to each input event. The mapping relationship is based on timestamp matching, associating each input event with the nearest screen frame.
[0057] S3: UI Element Detection. The visual perception unit performs UI element detection on keyframes based on the YOLOv8 object detection model, obtaining element categories, bounding box coordinates, and visual feature vectors. The YOLOv8 model runs on the edge box's NPU, with an inference latency of less than 100ms, and can detect common UI elements such as buttons, input boxes, dropdown lists, tables, and labels in real time.
[0058] S4: Element Association. The element association unit establishes the association between input events and UI elements. For mouse click events, the clicked target element is determined based on coordinate matching; for keyboard input events, the input target element is determined by combining focus tracking.
[0059] S5: Behavior fragment extraction. The behavior fragment extraction unit divides a continuous sequence of input events into behavior fragments with independent semantics based on a time window and event density. Each behavior fragment corresponds to a complete intent unit of a user operation, such as "login to the system," "query a report," or "export data."
[0060] S6: Intent Recognition. The operation semantic recognition unit performs semantic understanding on each behavioral segment based on a pre-trained large language model (such as DeepSeek-V3 or Qwen2.5) to identify the user's business intent. The business intent includes types such as data query, report export, information entry, and system navigation.
[0061] S7: Parameter extraction. The parameter extraction unit extracts variable parameters from the operation sequence, such as date ranges, query conditions, and selection options. The parameter extraction is based on a combination of rule matching and semantic understanding to identify the variable parts of the operation.
[0062] S8: Script Generation. The script generation unit generates a parameterized RPA script based on the identified business intent and extracted parameters. The RPA script is written in Python and uses the PyAutoGUI or Playwright library to simulate operations, supporting variable substitution, conditional branching, and loop control.
[0063] Example 3: Vision-Language Joint Understanding Architecture
[0064] The vision-language joint understanding architecture described in this embodiment is based on a multimodal pre-trained model, including:
[0065] Visual encoder: Used to extract visual features from screen images. Based on the VisionTransformer architecture, the visual encoder divides the screen image into a sequence of image blocks and extracts global layout features and local element features through a self-attention mechanism. These visual features include visual attributes such as element position, shape, color, and texture.
[0066] Text Encoder: Used to extract semantic features from screen text. The text encoder is based on the Transformer architecture and extracts word-level and sentence-level features through a self-attention mechanism. These semantic features include the text's lexical semantics, syntactic structure, and contextual relationships.
[0067] Cross-modal fusion layer: Used to fuse visual and textual features to establish semantic associations between visual elements and textual content. Based on a cross-attention mechanism, this layer enables bidirectional interaction between visual and textual features, outputting a fused multimodal representation vector.
[0068] Task decoder: Used to perform specific tasks based on the fused multimodal features. The task decoder includes an element detection head, a text recognition head, and a semantic understanding head, which are used to perform UI element detection, text content recognition, and field semantic understanding tasks, respectively.
[0069] Example 4: AI Multimodal Data Structured Process
[0070] The workflow of the AI multimodal data structuring processing module described in this embodiment includes:
[0071] S1: Layout Analysis. The layout analysis unit performs visual layout analysis on the screen image to identify the functional areas of the screen, including navigation area, content area, operation area, status area, etc. The layout analysis is based on joint judgment of visual and semantic features, and determines the boundaries of functional areas by analyzing the spatial distribution and semantic relationships of elements.
[0072] S2: Component Recognition. The component recognition unit identifies the types of UI components on the screen, including forms, tables, lists, cards, charts, etc. The component recognition is based on the classification capabilities of a multimodal pre-trained model, combining visual and textual features to determine the component type.
[0073] S3: Text detection. The text detection unit detects text regions in the screen image based on the DBNet (Differentiable Binarization Network) algorithm, and outputs the bounding box and polygon outline of the text region. The DBNet algorithm can detect text regions of arbitrary shapes, including horizontal text, slanted text, and curved text.
[0074] S4: Text Recognition. The text recognition unit uses the CRNN (Convolutional Recurrent Neural Network) algorithm to recognize characters in the detected text regions and extract the text content. The CRNN algorithm combines CNN feature extraction and RNN sequence modeling, supports mixed Chinese and English recognition, and achieves an accuracy rate of over 99%.
[0075] S5: Visual-Language Alignment. The visual-language alignment unit establishes a correspondence between visual elements and text content based on the cross-modal alignment capability of a multimodal pre-trained model. The cross-modal alignment determines the UI component to which the text element belongs by calculating the similarity between visual features and text features.
[0076] S6: Table parsing. The table parsing unit performs structured parsing of table components, including table boundary detection, cell segmentation, row and column relationship recognition, and table header understanding. The table parsing is based on line detection and cell alignment algorithms to construct the row and column structure of the table.
[0077] S7: Field Semantic Understanding. The field semantic understanding unit, based on Named Entity Recognition (NER) technology and a domain knowledge base, identifies the business semantics of extracted text fields, including field name, field type, and field value. The business semantics include industrial data fields such as order number, production date, quantity, and amount.
[0078] S8: Information Unit Assembly. The information unit assembly unit assembles the identified fields into information units according to business logic. Each information unit is a data record with complete business meaning; for example, a production record may contain fields such as product number, production date, output, and quality status.
[0079] S9: Data type inference. The data type inference unit performs data type inference on the extracted information unit based on regular expression matching and machine learning classification algorithms. The data types include numeric, date, text, and enumeration types.
[0080] S10: Data standardization, wherein the data standardization unit performs format standardization processing on the identified data, including date format unification (e.g., unification to YYYY-MM-DD format), numerical precision standardization, encoding format conversion, etc.
[0081] S11: Data quality verification. The data quality verification unit performs quality verification on structured data, including integrity verification (whether required fields are empty), consistency verification (whether cross-field logical relationships are correct), and reasonableness verification (whether the values are within a reasonable range).
[0082] S12: Target format conversion. The target format conversion unit converts the standardized data into a target output format, which includes JSON, XML, CSV, database records, etc., and outputs it to the specified target system.
[0083] Example 5: Data Acquisition Application of MES System
[0084] A car parts manufacturer is using a Manufacturing Execution System (MES) purchased 10 years ago. This system lacks an open API, and the vendor's interface quote is 800,000 yuan with a development cycle of 6 months. The company needs to obtain daily production output, quality data, and equipment status data from the MES system for production analysis and decision-making.
[0085] The deployment process of the industrial data acquisition system of the present invention is as follows:
[0086] 1. Hardware Deployment: Connect the edge computing box to the graphics card output of the MES client computer via an HDMI cable, loop the box out via HDMI to the monitor, connect the network cable to the enterprise intranet, and power on.
[0087] 2. Process Recording: The IT administrator logs into the edge box management interface, enables intelligent recording mode, and demonstrates a complete operation process: Log in to the MES system → Select production line → Query production report → Set date range → Export Excel report. The AI intelligent recording module's operation sequence learning framework captures the entire operation process and automatically generates parameterized RPA scripts through three subsystems: temporal behavior modeling, UI element perception, and intent understanding.
[0088] 3. Task Configuration: Administrators configure scheduled tasks via natural language: "Automatically export production reports at 2 AM every day, with a date range of the previous day, and save the data to the enterprise data warehouse." The requirement understanding module parses the configuration and generates structured task configurations.
[0089] 4. Task Execution: Every day at 2:00 AM, the RPA execution engine automatically logs into the MES system and performs report export operations. The AI multimodal data structuring module, based on a vision-language joint understanding architecture, intelligently identifies and structures the exported report data. Through three subsystems—screen semantic parsing, information unit extraction, and data paradigm conversion—the data is converted into JSON format and pushed to the enterprise data warehouse.
[0090] 5. Anomaly Handling: When minor changes occur to the MES system interface (such as button repositioning), the RPA execution engine's multimodal element location strategy automatically adapts to ensure normal task execution. In the event of system maintenance or network anomalies, the anomaly recovery mechanism automatically retryes or notifies the administrator.
[0091] 6. The RPA execution engine module supports three execution modes: Preset script mode: executes pre-recorded and optimized RPA scripts, suitable for high-frequency fixed processes; AI-guided mode: AI analyzes the current screen status in real time and dynamically guides RPA execution operations, suitable for changing processes; Hybrid mode: mainly uses preset scripts, and switches to AI-guided mode when anomalies or interface changes are detected.
[0092] Implementation results: The deployment cycle was shortened from 6 months to 3 days, the implementation cost was reduced from 800,000 yuan to 150,000 yuan (including hardware), the data collection frequency was increased from once a day to once every 30 minutes, and the manual operation time was reduced from 2 hours a day to 0.
Claims
1. An industrial data acquisition system based on AI visual recognition and edge computing, characterized in that, include: Edge computing box module: As a local execution environment, it is responsible for RPA script execution, AI inference, temporary data storage and secure transmission, and connects to the target PC in a plug-and-play manner; AI intelligent recording module: Based on the operation sequence learning framework, it realizes automatic learning of user operation flow and RPA script generation through the temporal behavior modeling subsystem, UI element perception subsystem and intent understanding subsystem, and sends the generated RPA script to the edge computing box module; The RPA execution engine module receives instructions from the edge computing box module and executes automated operation processes, including desktop application control and automated browser browsing. The AI multimodal data structuring module, based on a vision-language joint understanding architecture, uses a screen semantic parsing subsystem, an information unit extraction subsystem, and a data paradigm conversion subsystem to achieve intelligent recognition and structured processing of screen data. The demand understanding module converts user natural language demands into structured data acquisition tasks and sends these tasks to the AI intelligent recording module.
2. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The AI intelligent recording module's temporal behavior modeling subsystem includes: an input capture unit for capturing screen frame sequences and input device event sequences, wherein the screen frame sequences are acquired at a preset sampling frequency, and the input device event sequences include mouse events and keyboard events, each carrying a timestamp and coordinate information; a temporal alignment unit for aligning the screen frame sequences with the input device event sequences in time, establishing a frame-event mapping relationship, which is used to determine the screen state corresponding to each input event; a behavior segment extraction unit for segmenting continuous input event sequences into behavior segments with independent semantics based on time windows and event density, wherein each behavior segment corresponds to a complete intent unit of user operation; and a temporal feature encoding unit for extracting temporal features from each behavior segment, wherein the temporal features include operation type sequences, time interval distributions, and operation trajectory features.
3. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The UI element perception subsystem of the AI intelligent recording module includes: a visual perception unit, used for visual analysis of screen frames, detecting and recognizing UI elements on the screen, including buttons, input boxes, dropdown boxes, tables, labels, and icons; the visual perception unit outputs the category label, bounding box coordinates, and visual feature vector of each UI element; a DOM parsing unit, used for DOM tree parsing of the web application page, obtaining the hierarchical structure, attribute information, and text content of page elements; the DOM parsing serves as a supplement to visual perception to improve the accuracy of element recognition; an element association unit, used to establish the association between detected UI elements and input events, determining the target element operated by each input event; the association is based on dual verification of coordinate matching and temporal matching; and an element state tracking unit, used to track the state changes of UI elements during the execution of the operation sequence, including element appearance, disappearance, position movement, and content update.
4. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The intent understanding subsystem of the AI intelligent recording module includes: an operation semantic recognition unit, used to perform semantic understanding of behavioral fragments based on a pre-trained language model, and identify the business intent of user operations, including data query, report export, information entry, and system navigation; a parameter extraction unit, used to extract variable parameters from the operation sequence, including date range, query conditions, and selection options, and the parameter extraction is based on a combination of rule matching and semantic understanding; a process template matching unit, used to match the identified behavioral fragments with a predefined process template library, which contains common process patterns for common business operations; and a script generation unit, used to generate parameterized RPA scripts based on the identified business intent and extracted parameters, the RPA scripts supporting variable substitution, conditional branching, and loop control.
5. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The workflow of the operation sequence learning framework includes: S1: Recording begins, and the input capture unit starts collecting screen frame sequences and input device event sequences; S2: The timing alignment unit establishes a frame-event mapping relationship and determines the screen state corresponding to each input event; S3: The visual perception unit performs UI element detection on keyframes and obtains element categories, positions, and features; S4: The element association unit establishes the association relationship between input events and UI elements; S5: The behavior fragment extraction unit divides the continuous event sequence into behavior fragments based on event density and operation type; S6: The operation semantic recognition unit performs semantic understanding on each behavior fragment and identifies the business intent; S7: The parameter extraction unit extracts variable parameters from the operation sequence; S8: The script generation unit generates RPA scripts based on business intent and parameters.
6. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The screen semantic parsing subsystem of the AI multimodal data structuring processing module includes: a layout analysis unit for visual layout analysis of screen images and identification of functional areas of the screen, including navigation area, content area, operation area, and status area; a component identification unit for identifying UI component types on the screen, including forms, tables, lists, cards, and charts, with component identification based on joint judgment of visual and semantic features; a visual-language alignment unit for establishing the correspondence between visual elements and text content, with visual-language alignment based on the cross-modal alignment capability of a multimodal pre-trained model; and a screen semantic graph construction unit for constructing a semantic representation graph of the screen, which includes the screen's functional structure, element relationships, and interaction logic. The information unit extraction subsystem of the AI multimodal data structuring processing module includes: a text detection unit for detecting text regions in screen images, the text detection being based on a deep learning object detection algorithm, outputting the bounding box and confidence score of the text region; a text recognition unit for recognizing characters in the detected text regions, the text recognition being based on an end-to-end sequence recognition model, supporting mixed Chinese and English recognition; a table parsing unit for structured parsing of table components, the table parsing including table boundary detection, cell segmentation, row and column relationship recognition, and table header understanding; a field semantic understanding unit for recognizing the business semantics of the extracted text fields, the business semantics including field name, field type, and field value, the field semantic understanding being based on named entity recognition and a domain knowledge base; and an information unit assembly unit for assembling the recognized fields into information units according to business logic, the information unit being a data record with complete business meaning. The data paradigm conversion subsystem of the AI multimodal data structuring processing module includes: a data type inference unit, used to infer the data type of the extracted information units, including numeric, date, text, and enumeration types, based on regular expression matching and machine learning classification; a data standardization unit, used to standardize the format of the identified data, including date format unification, numerical precision standardization, and encoding format conversion; a data quality verification unit, used to verify the quality of the structured data, including integrity verification, consistency verification, and rationality verification; and a target format conversion unit, used to convert the standardized data into a target output format, including JSON, XML, CSV, and database records.
7. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The vision-language joint understanding architecture is based on a multimodal pre-trained model, which includes: a visual encoder for extracting visual features from screen images, including global layout features and local element features; a text encoder for extracting semantic features from screen text, including word-level features and sentence-level features; a cross-modal fusion layer for fusing visual and text features to establish semantic associations between visual elements and text content; and a task decoder for performing specific tasks based on the fused multimodal features, including element detection, text recognition, and semantic understanding. The workflow of the AI multimodal data structuring processing module includes: S1: Receiving a screen capture image, the layout analysis unit performs visual layout analysis and identifies functional areas; S2: The component identification unit identifies the types of UI components on the screen; S3: The text detection unit detects text areas in the image; S4: The text recognition unit performs text recognition on the text areas; S5: The visual-language alignment unit establishes the correspondence between visual elements and text content; S6: The table parsing unit performs structured parsing on table components; S7: The field semantic understanding unit identifies the business semantics of fields; S8: The information unit assembly unit assembles fields into information units; S9: The data type inference unit infers the data type of the information units; S10: The data standardization unit performs format standardization processing; S11: The data quality verification unit performs quality verification; S12: The target format conversion unit converts the data to the target output format.
8. The industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, This system employs an anomaly recovery mechanism, specifically... Includes the following steps: S1: Monitor the RPA execution process and detect abnormal situations, including element not found, timeout, and recognition failure; S2: Classify the detected abnormalities and determine their levels, including element-level, page-level, and process-level; S3: Execute corresponding recovery strategies based on the abnormality level, including retry positioning, page refresh, and process rollback; S4: When automatic recovery fails, record the abnormal information and notify manual intervention; S5: Continuously optimize AI models and recovery strategies based on abnormal cases; The edge computing box module adopts a multi-level security mechanism, including the following: local data processing: sensitive data is processed locally at the edge and is not uploaded to the cloud; Encrypted transmission: TLS encryption is used for communication with the cloud to ensure secure data transmission; Device authentication: Hardware-level device authentication is supported to prevent unauthorized access; Operation auditing: All operation logs are fully recorded, supporting audit traceability.
9. An industrial data acquisition system based on AI visual recognition and edge computing according to claim 1, characterized in that, The system adopts the following workflow: S1: The user inputs data collection requirements through natural language or starts the intelligent recording mode; S2: The AI intelligent recording module, based on the operation sequence learning framework, captures user screen operations and automatically generates RPA scripts through three subsystems: temporal behavior modeling, UI element perception, and intent understanding; S3: The RPA execution engine module loads the RPA scripts and controls the target system to perform corresponding operations; S4: The screen capture module captures page images of the target system; S5: The AI multimodal data structuring processing module, based on the vision-language joint understanding architecture, performs intelligent recognition and structuring processing of page images through three subsystems: screen semantic parsing, information unit extraction, and data paradigm transformation; S6: Import structured data into a large model for analysis, or output it to a specified target system.
10. An industrial data acquisition system based on AI visual recognition and edge computing according to claim 9, characterized in that, The workflow of the AI intelligent recording module in step S2 includes: S21: Start recording; the input capture unit collects screen frame sequences and input device event sequences at a preset sampling frequency; S22: The timing alignment unit establishes a frame-event mapping relationship and determines the screen state corresponding to each input event; S23: The visual perception unit performs UI element detection on key frames and obtains element categories, positions, and features; S24: The element association unit establishes the association relationship between input events and UI elements; S25: The behavior segment extraction unit divides the continuous event sequence into behavior segments based on event density and operation type; S26: The operation semantic recognition unit performs semantic understanding on each behavior segment based on a pre-trained language model and identifies business intent; S27: The parameter extraction unit extracts variable parameters from the operation sequence; S28: The script generation unit generates parameterized RPA scripts based on business intent and parameters. The workflow of the AI multimodal data structuring module described in step S5 includes: S51: Receiving a screen capture image, the layout analysis unit performs visual layout analysis and identifies functional areas; S52: The component recognition unit identifies the types of UI components on the screen; S53: The text detection unit detects text regions in the image based on a deep learning object detection algorithm; S54: The text recognition unit performs text recognition on the text regions based on an end-to-end sequence recognition model; S55: The visual-language alignment unit establishes the correspondence between visual elements and text content based on the cross-modal alignment capability of the multimodal pre-trained model; S56: The table parsing unit processes the table components... Row structure parsing includes table boundary detection, cell segmentation, and row-column relationship recognition; S57: Field semantic understanding unit identifies the business semantics of fields based on named entity recognition and domain knowledge base; S58: Information unit assembles the identified fields into information units according to business logic; S59: Data type inference unit infers the data type of the information units based on regular expression matching and machine learning classification; S510: Data standardization unit performs format standardization processing; S511: Data quality verification unit performs integrity verification, consistency verification, and rationality verification; S512: Target format conversion unit converts the data into the target output format; The vision-language joint understanding architecture is based on a multimodal pre-trained model, which includes: a visual encoder for extracting visual features from screen images, including global layout features and local element features; a text encoder for extracting semantic features from screen text, including word-level features and sentence-level features; a cross-modal fusion layer for fusing visual and text features to establish semantic associations between visual elements and text content; and a task decoder for performing element detection, text recognition, and semantic understanding tasks based on the fused multimodal features.