Data acquisition method and device, computer equipment, storage medium and program product
By obtaining control handles and generating data acquisition strategies, monitoring and capturing control interaction behavior data in production equipment in real time, the problem of lack of data interface is solved and efficient data acquisition of various controls is achieved.
Patent Information
- Application Number
- CN202510748104.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
In the prior art, the data of certain control interaction behaviors of production equipment cannot be collected, mainly due to the lack of data interfaces or limited data types provided by the interface, resulting in insufficient data collection.
By obtaining the control handle of the target application, determining the control type, and generating corresponding data acquisition policies, monitoring the interactive behavior data of the control in real time, including controls such as lists, buttons and input boxes, data capture is used by event listeners or message hooks, and transmitting it to the receiver through encryption protocols.
It realizes the acquisition of interactive data without relying on the data interface provided by production equipment, improves the applicability and integrity of data acquisition, and is suitable for a variety of control types in industrial scenarios.
Smart Images

Figure CN120256253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet of Things technology, and particularly to a data acquisition method, device, computer device, storage medium, and program product. Background Art
[0002] With the rapid development of industrial automation, the demand for data acquisition and analysis of production equipment is increasing day by day.
[0003] The existing technology mainly relies on data interfaces provided by production equipment for remote communication acquisition, or uses industrial gateways for connection for acquisition. However, some production equipment does not provide data interfaces, or the types of collectible data provided by the data interfaces are limited, resulting in the inability to collect some interaction behavior data of controls in industrial scenarios. Summary of the Invention
[0004] The purpose of this application is to provide a data acquisition method, device, computer device, storage medium, and program product.
[0005] To achieve the above purpose, this application provides the following solutions: In the first aspect, this application provides a data acquisition method, including: Obtain the control handle in the current interface of the target application program; the control handle is at least used to determine the control type; among them, the control whose control type belongs to the target type in the current interface is the target control; the target type includes lists, buttons, and input boxes; the target controls in the current interface include at least one of list controls, button controls, and input box controls; Input the feature data of the target control into the acquisition policy generation model to generate a data acquisition policy corresponding to the control type of the target control; the feature data at least includes: control type; Based on the data acquisition policy, perform real-time monitoring on the target control to capture the interaction behavior data of the target control when a state change is detected.
[0006] In an embodiment, the control handle of the target control is the target control handle; the process of obtaining the feature data includes: Call a preset API or function to read the text content and storage location of the target control handle; Obtain the status-related information and control type of the target control according to the target control handle; Among them, the feature data includes the handle value, the text content, the storage location, the control type of the target control, and the status-related information.
[0007] In one embodiment, the step of generating a data collection strategy corresponding to the control type of the target control includes: Determine a basic collection rule that matches the control type of the target control from a preset policy template library; Based on the control listening configuration corresponding to the target control, adjust the matching basic collection rule to obtain a data collection strategy.
[0008] In one embodiment, the step of adjusting the matching basic collection rule includes: Based on the control listening configuration corresponding to the list control, adjust the chunk loading rule for the list control. The chunk loading rule includes a scroll trigger condition and a data chunk size, and collect the data of the currently visible list items when a scroll event for the list control is detected, and collect the data of the clicked list item when a click event for the list control is detected; Based on the control listening configuration corresponding to the button control, adjust the listening rule for the button control. The listening rule includes a baseline listening interval and a minimum listening interval; Based on the control listening configuration corresponding to the input box control, adjust the difference comparison rule for the input box control. The difference comparison rule includes a comparison period and a change content extraction rule.
[0009] In one embodiment, after the step of capturing the interaction behavior data of the target control, it further includes: Encode the interaction behavior data into transmission data in a preset format and transmit it to the receiving end through an encryption protocol; Optimize the data collection strategy corresponding to the target control based on the feedback result of the receiving end.
[0010] In one embodiment, the step of optimizing the data collection strategy corresponding to the target control based on the feedback result of the receiving end includes: Calculate the policy optimization weight corresponding to each target type according to the data integrity score and system load index in the feedback result; When the policy optimization weight corresponding to the list control reaches the list control weight threshold, adjust the data chunk size in the chunk loading rule; When the policy optimization weight corresponding to the button control reaches the button control weight threshold, switch the listening mode in the listening rule to a hybrid listening mode. The hybrid listening mode includes an event-driven mode and a polling mode.
[0011] In one embodiment, after the step of optimizing the data collection strategy corresponding to the target control, it further includes: Synchronize the optimized control listening configuration to the running environment of the target application through the hot loading mechanism; Verify the compatibility between the optimized control listening configuration and the control handle; If the verification passes, re-enter the step of performing real-time monitoring on the target control based on the data collection strategy.
[0012] In one embodiment, the step of encoding the interaction behavior data into transmission data in a preset format includes: Encapsulate the interaction behavior data into a JSON structure; Perform Base64 encoding on the JSON structure to generate intermediate encoded data; Encrypt the intermediate encoded data through a symmetric encryption algorithm to generate an encrypted data packet; Add a transmission protocol header to the encrypted data packet to generate the transmission data that conforms to the preset communication specification.
[0013] In one embodiment, the step of performing real-time monitoring on the target control based on the data collection strategy includes: Listen for target events of the list control through an event listener or a message hook; wherein, the target events of the list control include a scroll event and a selection event; Listen for target events of the button control according to the listening mode in the control listening configuration; wherein, the listening mode in the control listening configuration is an event-driven mode, a polling mode, or a hybrid listening mode; the hybrid listening mode includes: an event-driven mode and a polling mode; the target events of the button control include a click event and a long-press event; Listen for target events of the input box control through a polling mode; wherein, the target events of the input box control include a text input event and a focus switching event; The interaction behavior data includes the monitored target events.
[0014] In one embodiment, after the step of capturing the interaction behavior data of the target control, it further includes: Output a data collection report of the target application, where the data collection report includes content obtained through at least one of the following methods: Statistical operation frequency distribution and timestamps of the interaction behavior data; Identify abnormal operations in the interaction behavior data through a clustering algorithm and generate warning events; Associate the interaction behavior data with the device operation log based on the timestamp to obtain an association table; Render the analysis result of the interaction behavior data into a visual chart and export it as a document.
[0015] In one embodiment, after the step of obtaining the control handle in the current interface of the target application, the method further includes: If a control fingerprint matching the control handle of the target control is found in the pre-loaded static control fingerprint library, the data collection policy associated with the matching control fingerprint is determined as the data collection policy of the target control; the static control fingerprint library includes associated control fingerprints and data collection policies; The step of inputting the feature data of the target control into the collection policy generation model is executed after no matching control fingerprint is found in the static control fingerprint library.
[0016] In a second aspect, the present application provides a data collection device, including: An acquisition module, configured to obtain the control handle in the current interface of the target application; the control handle is at least used to determine the control type; among them, the control whose control type belongs to the target type in the current interface is the target control; the target type includes lists, buttons, and input boxes; the target controls in the current interface include at least one of list controls, button controls, and input box controls; A processing module, configured to input the feature data of the target control into a collection policy generation model to generate a data collection policy corresponding to the control type of the target control; the feature data at least includes: control type; A collection module, configured to perform real-time monitoring on the target control based on the data collection policy, so as to capture the interaction behavior data of the target control when a state change is monitored.
[0017] In one embodiment, the control handle of the target control is the target control handle; the acquisition module is further configured to: Call a preset API or function to read the text content and storage location of the target control handle; Obtain the status-related information and control type of the target control according to the target control handle; Among them, the feature data includes the handle value, the text content, the storage location, the control type of the target control, and status-related information.
[0018] In one embodiment, in terms of generating a data collection policy corresponding to the control type of the target control, the processing module is configured to: Determine a basic collection rule matching the control type of the target control from a preset policy template library; Adjust the matching basic collection rule based on the control monitoring configuration corresponding to the target control to obtain a data collection policy.
[0019] In one embodiment, in terms of adjusting the base acquisition rules that match, the processing module is configured to: Based on the control listening configuration corresponding to the list control, adjust the chunk loading rules for the list control, where the chunk loading rules include a scroll trigger condition and a data chunk size, and collect the currently visible list item data when a scroll event for the list control is detected, and collect the clicked list item data when a click event for the list control is detected; Based on the control listening configuration corresponding to the button control, adjust the listening rules for the button control, where the listening rules include a baseline listening interval and a minimum listening interval; Based on the control listening configuration corresponding to the input box control, adjust the difference comparison rules for the input box control, where the difference comparison rules include a comparison period and a change content extraction rule.
[0020] In one embodiment, the processing module is further configured to: Encode the interaction behavior data into transmission data in a preset format and transmit it to the receiving end through an encryption protocol; Optimize the data acquisition strategy corresponding to the target control based on the feedback result of the receiving end.
[0021] In one embodiment, in terms of optimizing the data acquisition strategy corresponding to the target control based on the feedback result of the receiving end, the processing module is configured to: Calculate the strategy optimization weights corresponding to each target type according to the data integrity score and system load index in the feedback result; When the strategy optimization weight corresponding to the list control reaches the list control weight threshold, adjust the data chunk size in the chunk loading rules; When the strategy optimization weight corresponding to the button control reaches the button control weight threshold, switch the listening mode in the listening rules to a hybrid listening mode, where the hybrid listening mode includes an event-driven mode and a polling mode.
[0022] In one embodiment, the processing module is further configured to: Synchronize the optimized data acquisition strategy to the running environment of the target application through a hot loading mechanism; Verify the compatibility between the optimized control listening configuration and the control handle; If the verification passes, re-enter the step of performing real-time listening on the target control based on the data acquisition strategy.
[0023] In one embodiment, in terms of encoding the interaction behavior data into transmission data in a preset format, the processing module is configured to: Encapsulate the interaction behavior data into a JSON structure; Perform Base64 encoding on the JSON structure to generate intermediate encoded data; Encrypt the intermediate encoded data through a symmetric encryption algorithm to generate an encrypted data packet; Add a transport protocol header to the encrypted data packet to generate the transport data that conforms to the preset communication specification.
[0024] In one embodiment, in terms of the real-time monitoring of the target control based on the data collection strategy, the collection module is used for: Monitor the target events of the list control through an event listener or a message hook; wherein, the target events of the list control include a scroll event and a selection event; Monitor the target events of the button control according to the monitoring mode in the control monitoring configuration; wherein, the monitoring mode in the control monitoring configuration is an event-driven mode, a polling mode, or a hybrid monitoring mode; the hybrid monitoring mode includes: an event-driven mode and a polling mode; the target events of the button control include a click event and a long-press event; Monitor the target events of the input box control through a polling mode; wherein, the target events of the input box control include a text input event and a focus switching event; The interaction behavior data includes the monitored target events.
[0025] In one embodiment, the collection module is further used for: Output a data collection report of the target application, and the data collection report includes the content obtained through at least one of the following methods: Statistically analyze the operation frequency distribution and time stamps of the interaction behavior data; Identify abnormal operations in the interaction behavior data through a clustering algorithm and generate warning events; Associate the interaction behavior data with the device operation log based on the time stamp to obtain an association table; Render the analysis result of the interaction behavior data into a visual chart and export it as a document.
[0026] In one embodiment, the collection module is further used for: If a control fingerprint matching the control handle of the target control is found in the pre-loaded static control fingerprint library, determine the data collection strategy associated with the matching control fingerprint as the data collection strategy of the target control; the static control fingerprint library includes associated control fingerprints and data collection strategies; The step of inputting the feature data of the target control into the acquisition strategy generation model is performed after no matching control fingerprint is found in the static control fingerprint library.
[0027] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the data acquisition method described in any one of the above.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the data acquisition method described in any one of the above.
[0029] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the data acquisition method described in any one of the above.
[0030] According to the specific embodiments provided by the present application, the following technical effects are disclosed: Through the data acquisition technology based on control handle capture, the present application can obtain control handles and generate corresponding data acquisition strategies for different target controls according to the types of control handles, so as to monitor the target controls according to the data acquisition strategies to implement the data acquisition of various interaction data, without relying on the data interfaces provided by production devices and being restricted by the types of collectable data provided by the data interfaces, improving the applicability of the interaction data acquisition in production devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic flowchart of a data acquisition method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a data acquisition method applied to a data acquisition system provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the functional modules of a data acquisition device provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Technical term explanations: Windows API: Windows API, also known as WinAPI, is the core programming interface of Microsoft's Windows operating system. It is a collection of C language functions that allow developers to interact with the Windows operating system. These APIs cover a large number of functions, including window management, system services, device input and output, file operations, network communication, memory management, etc. Due to their low-level nature, Windows API functions talk directly to the operating system kernel, which means that they run efficiently but may be more complicated to use than modern programming frameworks. Over time, Microsoft introduced higher-level programming interfaces and frameworks such as NET Framework and Windows Runtime (WinRT), but Windows API remains the cornerstone behind these technologies.
[0033] Base64: Base64 is a method of representing binary data based on 64 printable characters. Since , every 6 bits is a unit, corresponding to a printable character. 3 bytes have 24 bits, corresponding to 4 Base64 units, that is, 3 bytes can be represented by 4 printable characters. It can be used as a transmission code for emails. The printable characters in Base64 include letters AZ, az, and numbers 0-9, so there are 62 characters in total. In addition, the two printable symbols are different in different systems. Base64 is often used in situations where text data is usually processed to represent, transmit, and store some binary data, including MIME emails and some complex data in XML.
[0034] Java: Java is an object-oriented high-level programming language. It absorbs the advantages of C++ and discards the difficult-to-understand concepts of multiple inheritance and pointers in C++. Therefore, it has two significant features: powerful functions and simple and easy to use. Java language implements object-oriented theory very well, allowing programmers to use rigorous thinking to perform complex programming. The main features of Java include simplicity, object-oriented, distributed, robust, security, platform independence and portability, multi-threading, and dynamism.
[0035] Style ID: A style ID is a constant or flag used to define and specify the appearance and behavior of a window, control, or other object. These IDs are usually a set of predefined values that can be used in combination to set a specific style or property of an object.
[0036] Monitoring Filter (Event Filter): An event filter is a powerful mechanism that allows an object to monitor and intercept events received by other objects. Through an event filter, events can be processed, modified, or intercepted before they reach the target object. Its features include: Event filters can centrally process events from multiple controls without having to write separate event handling logic in each control (for example, multiple button click events can be uniformly monitored through an event filter without having to rewrite the event handling function for each button); Global event monitoring can be performed, for example, it can be used to globally monitor certain types of events, such as keyboard shortcuts or mouse events; Through event filters, custom behaviors can be added to specific controls while retaining the default event handling logic; It can be used for debugging and optimizing event handling logic; It can monitor events between multiple components to implement complex interaction logics.
[0037] Event Queue: The event queue is used to store and manage events to be processed. Its main functions include: 1. Event storage: When an event occurs, the event is first placed in the event queue and waits to be processed; 2. Event distribution: The event loop takes the event out of the event queue and distributes it to the target object; 3. Event filtering: Before event distribution, the event filter has the opportunity to process or intercept the event.
[0038] Bloom Filter: In the fingerprint database, the Bloom filter can be used to quickly determine whether a fingerprint is likely to exist in the database. If the Bloom filter returns "possibly in the set", the database can be further queried; if it returns "not in the set", it can be directly excluded.
[0039] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.
[0040] The present application claims to protect a data acquisition method, a data acquisition device, a computer device, a storage medium, and a program product.
[0041] The above-mentioned data acquisition device and computer device can be used to execute the above-mentioned data acquisition method. The data acquisition device and computer device can be collectively referred to as a system, or a device / system including the data acquisition device and computer device of the present application can be referred to as a system.
[0042] As Figure 1 shown, some embodiments of the present application provide a data acquisition method. In the embodiments of the present application, it includes: Step 101: Obtain the control handle in the current interface of the target application; the above control handle is at least used to determine the control type; among them, the control whose control type belongs to the target type in the current interface is the target control; the above target type refers to the control categories for which data needs to be collected, such as lists, buttons, and input boxes by way of example; these control types may correspond to different interaction behaviors and data collection requirements.
[0043] The "current interface" may refer to the screen or window with which the user is currently interacting, which contains all the elements that the user can see and operate, such as buttons, text boxes, menus, icons, etc., or may refer to the window or view that the program is currently displaying; or, in the operating system, the "current interface" may refer to the window or desktop environment that is currently in the foreground.
[0044] Each interface may become the current interface. It should be noted that in different interfaces, there may be only one type of list control, button control, and input box control, or there may be multiple types. Therefore, the target controls in the current interface include at least one of the list control, button control, and input box control.
[0045] In the embodiments of the present application, the control handle (referred to as the handle for short) is the unique identifier assigned by the operating system to interface elements (such as windows, buttons, input boxes, etc.), and is used to identify and operate controls during program operation.
[0046] By way of example, the control handle can be obtained through an application programming interface (API) or services provided by the operating system or by using an automated testing tool such as the pywinauto library.
[0047] Taking obtaining through the API as an example, in the Windows system, the handle is obtained through the Win32 API (such as FindWindow, FindWindowEx). Specifically, the API of the operating system (such as the UI Automation API of Windows or the Spy++ tool) can be called to traverse the current interface elements of the target application to obtain the handles of all controls in the current interface. The system filters out the controls of the target type (such as lists, buttons, input boxes) according to the class name, window name, or other attributes of the controls. For example, use FindWindowEx to recursively find the child window handle and verify the control type through GetClassName. Finally, the system stores the handles of the target controls in memory for use in subsequent steps.
[0048] Step 102: Input the feature data of the above target control into the acquisition policy generation model to generate a data acquisition policy corresponding to the control type of the above target control; the above feature data at least includes: control type.
[0049] Exemplarily, the feature data is structured information describing the properties of the control, including the control type, size, location, hierarchical relationship, etc., and is used to guide the system to generate an adapted data collection strategy. The collection strategy generation model is composed of rules or machine learning algorithms and is used to dynamically generate a data collection strategy according to the control feature data.
[0050] In one example, the system can serialize the obtained control feature data into the JSON format and input it into the collection strategy generation model. The model first determines the control type through a classifier, and then selects the corresponding basic collection rules according to the type. For example: for a list control, a chunk loading rule based on a scroll event is generated; for a button control, a listening rule is generated; for an input box control, a difference comparison rule is used. The data collection strategy output by the model includes parameters such as a listening mode and a data processing method, and the system can compile these strategies into an executable monitoring instruction set.
[0051] Specifically, for the Linux system, the parsing of the control tree in the application can be implemented based on the XQueryTree / XGetWindowProperty interfaces of the X Window System, and then the events corresponding to the controls can be captured by combining the D-Bus interfaces of the GTK+ / Qt frameworks.
[0052] For the macOS system, the NSAccessibility protocol can be called through Objective-C runtime reflection, and the Core Graphics API is used to implement visual feature matching between the screenshot of the control area and a pre-prepared standard control image to obtain the change situation of the control.
[0053] Furthermore, when the control handle cannot be successfully obtained, the OpenCV template matching algorithm can be used to locate the control position. Specifically, visual matching is performed between the OpenCV templates corresponding to different pre-set controls and the interface screenshot of the application to determine the positions of each control. Then, the text recognition function of the OCR engine is used to parse the screenshot at the position where the control is located to obtain the text content in the control.
[0054] Specifically, the Tesseract OCR engine can be integrated into the control handle capture program to parse the control text content (supporting multi-language recognition).
[0055] In one example, the collection strategy generation model can include a connected intelligent optimization layer and an adaptive adjustment layer, and the collection strategy generation model can be trained in the following way: 1) Data collection and preprocessing: Collect control operation logs from multi-platform devices such as Windows, Android, and iOS. Each log includes but is not limited to: control type, historical state sequence, system parameters, clock deviation, and business tags.
[0056] 2) Data cleaning and standardization: Remove invalid records (such as missing key fields, abnormal timestamps). Normalize numerical features (CPU load, network latency) (Min-Max Scaling) to distribute them in the [0,1] interval. One-hot encode the control type (One-Hot Encoding), for example: button → 1,0,01,0,0, list → 0,1,00,1,0. Split the time series data into fixed window lengths to form LSTM input samples.
[0057] 3) Intelligent optimization layer training.
[0058] The model architecture includes: input layer: receiving historical state sequences of length 10 (including control type, state value, system load and other features); LSTM layer: 2 layers of stacked LSTM units, each layer with 64 hidden nodes, used to capture long-term dependencies of time series; fully connected layer: mapping LSTM output to 3D vectors, corresponding to the three probability distributions of "high / medium / low"; output layer: Softmax activation function, outputting the probability of state change within the next 3 seconds.
[0059] Model training process: The loss function can use cross-entropy loss to measure the difference between the predicted probability and the true label (the change state of the manual annotation). The optimizer can use the Adam optimizer. During training, the training set and the validation set are divided according to the preset ratio to prevent overfitting. The model is trained with the training set and validated with the validation set.
[0060] 4) Adaptive adjustment layer training.
[0061] The Q-Learning algorithm is used to learn the optimal strategy in an unknown environment and output the predicted action.
[0062] In this embodiment, the state space of the Q-Learning algorithm includes the LSTM prediction results (high / medium / low probability) output by the intelligent optimization layer, real-time system parameters (CPU load, network delay) and business priorities, and its action space may exemplarily include polling mode (checking the control status at fixed intervals), event-driven mode (triggering collection only when the state changes), and hybrid mode (combining polling and event-driven in proportion).
[0063] The reward value takes into account data integrity, system load and energy consumption. The reward function can be seen in the following formula (1): reward = (data_completeness * 0.4) + (1 / system_load * 0.3) -(energy_consumption * 0.3) (1) Among them, data_completeness represents data integrity, which is calculated by comparing the actually collected data with the true state to obtain the coverage rate. System_load represents the system load, which is the reciprocal of the CPU and memory occupancy, encouraging low-load strategies. Energy_consumption represents the energy consumption, which is the energy consumption cost of the acquisition action estimated according to the device power consumption model.
[0064] The training process is as follows: First, assign an initial Q value (usually 0 or a small random value) to each state-action pair, and then adopt the ε-greedy strategy (initial ε = 0.9, gradually decaying to 0.1) to balance random exploration and optimal policy execution. After each action is executed, update the current Q value according to the immediate reward value and the maximum Q value of the next state. When the change range of the Q value is less than the threshold (such as 1e-5) or the cumulative reward tends to be stable, terminate the training. Finally, convert the optimized Q table into a policy function to select the optimal acquisition action in real time. Support online incremental learning: When the environment changes (such as a new control type is added), trigger a local Q value update.
[0065] After training is completed, the optimal Q value of each state-action pair is stored in the Q table. According to the Q value, the optimal action in each state can be determined, thus obtaining the optimal policy.
[0066] Step 103, based on the above data acquisition strategy, perform real-time monitoring on the above target control to capture the interaction behavior data of the above target control when a state change is detected.
[0067] The states of a control usually include but are not limited to: visibility state (visible / hidden / collapsed), enable state (enabled / disabled), focus state (has focus / loses focus), selected state (selected / unselected), active state (active / inactive), etc.
[0068] Taking a list control as an example, its states can exemplarily include one or more of the visibility state, enable state, focus state, selected state, detailed state of the selected item (single selection / multi-selection / no selected item), scroll state (scroll position / scroll to the bottom / scroll to the top), etc.
[0069] Taking a button control as an example, its states can exemplarily include one or more of the visibility state, enable state, focus state, selected state, button state (normal state / pressed state / hover state / selected state / unselected state).
[0070] Taking the input box control as an example, its status may include one or more of visibility status, enabled status, focus status, selected status, input status (normal input / read-only / password input), data status (data complete / data missing / data validation), selected status (selected content / no selected content), and data binding status (bound / unbound).
[0071] Changes in control state can be caused by a variety of factors, including event triggering, program logic, data binding, timers, system state changes, and window messages.
[0072] In the embodiment of the present application, real-time monitoring is a process in which the system continuously monitors changes in the control state through an event hook, an event listener, or a polling mechanism, such as using a Windows message hook (WH_CALLWNDPROC) to capture click or input events of a control.
[0073] The above-mentioned interactive behavior data is the information generated when the user interacts with the control, including the trigger time (which can be represented by a timestamp), the trigger method or the operation type (click, input), the trigger result or the operation result (such as input text, list selection), etc.
[0074] Event hooks, event listeners, functions, etc. that perform real-time monitoring can be collectively referred to as monitoring modules. In an example, for different control types, the system can configure and deploy monitoring modules according to the corresponding data collection strategy: for list controls, register the WM_VSCROLL message hook, call GetScrollInfo to obtain the scroll position when the scroll event is triggered, and use ListView_GetItemText to collect visible item data; for button controls, set the WH_CALLWNDPROC hook to capture the BN_CLICKED message and record the click time and context information; for input box controls, start the timer to periodically call the GetWindowText or SendMessage function to obtain text content, and filter the unchanged content through the difference algorithm.
[0075] When a state change is detected, the monitoring module encapsulates the interaction behavior data into a predefined format (such as Protocol Buffers), adds a timestamp and session ID, and transmits it to the receiving end through a secure channel. The receiving end can use the interaction behavior data for analysis, so the receiving end can also be called a data analysis module.
[0076] Further, in some embodiments of the present application, the three functions of identifying control handles, collecting data, and listening for events can be implemented in the form of plugins by defining a standardized plugin interface (IControlCollectorPlugin) at the software design level. Specifically, the hot-plug function of plugin loading can be implemented through the Java SPI mechanism or dynamic link library, so that the functions of identifying control handles, collecting data, and listening for events in some embodiments of the present application can be updated in real time during the listening process without restarting the listening process.
[0077] The system can also collect operation events of physical buttons on the production device through USB HID during the process of collecting interaction behavior data of the control, and associate the interaction behavior data with the operation events of the physical buttons in the time dimension, so that the operation events of the physical buttons can be used to reflect the user operations when the interaction behavior data is generated for later analysis.
[0078] It can be seen that through the data collection technology based on control handle capture in the embodiments of the present application, control handles can be obtained, and corresponding data collection strategies can be generated for different target controls according to the types of control handles, so as to listen to the target controls according to the data collection strategies to implement the data collection of various interaction data, without relying on the data interface provided by the production device and being restricted by the types of collectable data provided by the data interface, improving the applicability of the interaction data collection in the production device.
[0079] In one embodiment, for the sake of convenience, the control handle of the above-mentioned target control can be referred to as the target control handle, and the process of obtaining the above-mentioned characteristic data includes: Step 201, call a preset API or function to read the text content and storage location of the above-mentioned target control handle.
[0080] Exemplarily, the preset API or function can refer to a set of standard interface functions provided by the operating system or programming framework for interacting with graphical user interface (GUI) elements. Taking Windows as an example, on the Windows platform, it mainly includes the exported functions in User32.dll and UIAutomationCore.dll. The text content refers to the readable text information currently displayed by the control, such as the label text on the button, the input text in the input box, the display content of the list item, etc. The storage location refers to the physical address or logical location identifier of the control in the memory, including information such as the window handle, the hierarchical path of the control in the window tree, and the screen coordinate position.
[0081] Specifically, the system can call native API functions (such as the GetWindowTextW function and the GetWindowRect function) through JNI (Java Native Interface) or JNA (Java Native Access) technology to obtain the text content and storage location. Among them: for standard Windows controls, the GetWindowTextW function is used to read the text content, and this function needs to pass in the control handle and the character buffer; the GetWindowRect function is used to obtain the screen coordinates of the control and calculate the storage location. For complex controls such as ListView (list), additional macros such as ListView_GetItemText can be called to extract the sub-item text. The system converts the obtained text content to UTF-8 encoding and converts the storage location information to a rectangular structure containing left, top, right, and bottom. After all the data is verified, it can be stored in the temporary buffer for the next step of processing.
[0082] For mobile applications, data in the mobile device can be collected by using reflection to call the internal methods or private fields of AccessibilityNodeInfo.
[0083] Step 202, obtain the status-related information and control type of the target control according to the above target control handle; among them, in this embodiment, the foregoing characteristic data includes the above handle value, the above text content, the above storage location, the control type of the above target control, and status-related information.
[0084] In this embodiment, the status-related information may include the current interaction state and display attributes of the control, including but not limited to dynamic attributes such as enabled / disabled state, visibility, focus state, selected state (indicating whether it is selected), scroll position, value range, etc.
[0085] The control type refers to the functional classification identifier of the control, such as Win32 standard control types such as BUTTON, EDIT, LISTBOX, COMBOBOX, or custom control types defined by a third-party UI framework.
[0086] The following introduces the specific methods for obtaining status-related information and control type: The system can query the control status by sending messages such as WM_GETTEXT and LB_GETCURSEL through the SendMessage function. Or, the system can use GetWindowLong to obtain the control style identifier to determine whether there is an available style for the control, and combine functions such as IsWindowEnabled (used to determine whether the specified window is enabled) and IsWindowVisible (used to check whether the window is visible) to detect status-related information. It can also confirm whether the control is selected by checking the BM_GETCHECK property value of the button control, confirm the selection status of the detected item by the LB_GETSEL property value of the list box, and obtain the modification flag by the EM_GETMODIFY property value of the edit box (the modification flag is used to confirm the modification of the edited content). Of course, the above-described ways of obtaining status-related information of the control are only exemplary descriptions and can be specifically set according to actual needs, which are not limited herein.
[0087] The control type can obtain the standard class name through GetClassName and be converted into an enumeration type in combination with a predefined mapping table. The system combines this information with the data obtained in step 201 to construct a complete feature data set, including the handle value (hexadecimal), text content (string), storage location (coordinate structure), control type (enumeration value), and status information (set of key-value pairs).
[0088] It can be seen that in this embodiment, the text content, physical location, type identifier, and real-time status of the control are obtained through a standardized interface, constructing a complete control feature data set. The collected feature data provides an accurate data basis for subsequent applications such as automated testing, user behavior analysis, and interface monitoring, enabling the system to determine an accurate data collection strategy based on the feature data and thus accurately identify changes in the status of interface elements.
[0089] In one embodiment, step 102 above includes: Step 1021: Determine the basic collection rule that matches the control type of the above target control from a preset policy template library.
[0090] In the embodiments of the present application, the policy template library is a set of policies predefined by the system. Exemplarily, it can be stored in a key-value pair data structure, where the key is the enumeration value of the control type and the value is the corresponding basic collection rule. Each basic collection rule can be regarded as a template, which includes standard fields such as a listening mode, a collection frequency, and a data processing method.
[0091] Among them, the collection frequency field can store the collection frequency value or the collection interval (i.e., the reciprocal of the collection frequency value).
[0092] In an example, if the specific value of the monitoring mode field is the event-driven mode, the corresponding acquisition frequency field may be empty.
[0093] The value stored in the listening mode field can be used to represent an event-driven mode, a polling mode, or a mixed listening mode. For example, 1, 2, 3 or other characters can be used to represent the event-driven mode, the polling mode, and the mixed listening mode, respectively, or the listening mode field can also directly store the text form of "event-driven mode, polling mode, or mixed listening mode".
[0094] The basic collection rule is the smallest strategy unit designed for a specific control type. In its standard field, the common collection parameters and default configuration values of the corresponding control type are stored. For example, the basic collection rule of the list control contains fixed parameters such as scroll trigger conditions and data block size.
[0095] In one example, the system can retrieve matching items in the policy template library through a hash search algorithm: first, the type identifier of the target control (such as "LIST_VIEW", "BUTTON", etc.) is used as the query key, and the hash value is calculated using the MurmurHash3 algorithm to locate the corresponding control type within O(1) time complexity. The system loads the basic collection rules corresponding to the control type, such as differential comparison rules, monitoring rules, and block loading rules, and also loads additional configuration options applicable to the control type. Additional configuration options can be, for example, the range of loaded list items for list controls, the limit on the amount of loaded data for button controls, the data collection limit for input controls, and so on. All loaded rules (basic collection rules, or basic collection rules + additional configuration options) are stored in the policy buffer, waiting for the next step of optimization and adjustment.
[0096] Step 1022, based on the control monitoring configuration corresponding to the above target control, the matching basic collection rules are adjusted to obtain a data collection strategy (ie, a dynamic adjustment mechanism is used to obtain a final data collection strategy).
[0097] In an embodiment of the present application, the control monitoring configuration may include a dynamically generated set of optimization parameters, including collection constraints, current operating parameters and optimization targets for the target control under the current operating environment. The collection constraints may be real-time factors such as the upper limit of CPU occupancy, network delay tolerance, and service priority weights. The optimization target may be, for example, reducing CPU occupancy, shortening network delays, etc. The current operating parameters may include the current CPU usage rate. In addition, the control monitoring configuration may also include statistical data on the interactive behavior of the target control or target type, the characteristics of the target control, etc. The data collection strategy is the final generated, immediately executable monitoring collection plan that integrates the core logic of the basic rules and the runtime optimization parameters to form a complete control instruction sequence.
[0098] Exemplarily, to adjust the basic acquisition rule, the system may create a policy adjustment engine and input the basic acquisition rule and real-time monitoring configuration. The engine first performs resource evaluation and adjusts the acquisition frequency according to the current CPU utilization rate (obtained through GetSystemTimes) and memory pressure (GlobalMemoryStatusEx): for example, when the system load > CPU occupancy upper limit (70%), the polling interval is automatically extended by a set step length. For example, the step length can be set to 20% (of course, for the event-driven mode, the acquisition frequency can be set to empty subsequently). Then, the monitoring mode is adjusted by applying the business priority weight. For example, the event-driven mode is enabled for high-priority controls (business weight ≥ weight threshold 0.8), and the polling mode is maintained for low-priority controls, etc. (of course, there can be other adjustment methods, which will be introduced later in this article); then, environmental parameters are injected, and the data transmission timeout threshold is set according to the network latency (ping test result). The system compiles the data acquisition policy into a binary instruction set and deploys it to the kernel-level monitoring module for execution.
[0099] Those skilled in the art can flexibly design the specific values of the CPU occupancy upper limit, weight threshold, step length, etc., which will not be elaborated here.
[0100] Furthermore, the data acquisition policy can also be set by determining the usage scenario of the production equipment. The usage scenario of the production equipment can be set by the user, such as high-frequency data interaction scenario, industrial control scenario, low-end device scenario, etc. Exemplarily, for the high-frequency data interaction scenario, a millisecond-level polling monitoring mechanism and a memory write-only method can be adopted to ensure the accuracy and integrity of the collected interaction behavior data; or when the production equipment belongs to the industrial control scenario with high safety requirements, the stability of data transmission can be ensured through Modbus protocol conversion, and an abnormal fuse mechanism is introduced to ensure automatic fusing in case of data anomalies, avoiding affecting industrial production safety; or for the scenario where the production equipment is a mobile device, considering the limited performance of the mobile device, multi-sensor data fusion can be adopted: let the data be sent together after aggregation to reduce the transmission load of the mobile device and ensure the smooth completion of the data acquisition process. The embodiments of the present application provide a basic acquisition rule and configure and adjust the basic acquisition rule through a dynamic adjustment mechanism based on the control type and scenario to obtain a data acquisition policy, ensuring that the policy adapts to the real-time operating environment.
[0101] The following introduces exemplary operations for adjusting the basic acquisition rule for different control types.
[0102] In one embodiment, for the list control, step 1022 described above includes: Step 10221: Based on the control listening configuration corresponding to the above list control, adjust the chunk loading rule for the above list control.
[0103] Among them, the above chunk loading rule includes a scroll trigger condition, a data chunk size, a scroll listening operation, and a click event listening operation. The scroll listening operation is used to collect the data of the currently visible list items when a scroll event for the above list control is detected. The click event listening operation is used to collect the data of the clicked list item when a click event for the above list control is detected. The chunk loading rule can reduce system resource consumption by dividing the list data items into multiple small chunks (chunks) and gradually loading these data chunks when the user needs them, instead of loading all data items at once.
[0104] This chunk loading rule includes two core parameters: a scroll trigger condition and a data chunk size. The scroll trigger condition is one of the criteria for triggering data collection for the list control, including physical quantity indicators such as a scroll distance threshold and a dwell time threshold. Among them, the scroll distance threshold is at least used to determine whether the list is in a scroll scenario. If it is in a scroll scenario, chunk loading will be performed.
[0105] The data chunk size is the amount of data contained in each data chunk. For example, each chunk contains 10 or 50 list item data.
[0106] The dwell time threshold is used to trigger full - volume collection. Taking the dwell time threshold equal to 200 ms as an example, if the dwell time reaches 200 ms, full - set collection can be triggered to collect the data of all currently visible list items.
[0107] It should be noted that the scroll event itself is not determined by the scroll distance threshold. The scroll event is usually automatically triggered when the user performs a scroll operation. The scroll distance threshold can be used to judge the degree or condition of the scroll to achieve specific functions or interactions, such as triggering certain operations when the scroll distance reaches the threshold, rather than determining whether the scroll event occurs.
[0108] More specifically, to achieve data collection for the list control in the scrolling scenario, the system can first analyze the structural characteristics of the list control, obtain the scroll bar parameters through GetScrollInfo, and calculate the ratio of the visible area height to the list item height. According to the current memory usage (obtained through GlobalMemoryStatusEx), the data block size is dynamically set. Exemplarily, when the available memory > 2GB, it is set to 50 items, otherwise it is set to 30 items. And the scroll trigger condition is adjusted according to the control response speed. For example: use GetMessageTime to measure the scroll event interval. When the average interval < 200ms, set the scroll trigger threshold to 5 pixels; otherwise, set it to 10 pixels. The system encapsulates these parameters into the SCROLL_PARAMS structure and injects them into the configuration register of the list monitoring thread. The monitoring thread implements intelligent loading based on these parameters: when it detects that the scroll distance exceeds the threshold and the residence time reaches the residence time threshold, it triggers full-scale collection.
[0109] For the button control, step 1022 further includes: Step 10222, based on the corresponding control listening configuration of the button control, adjust the listening rules for the button control.
[0110] Exemplarily, the listening rules can include two parameters: the baseline listening interval and the minimum listening interval.
[0111] In the embodiment of the present application, the listening rules belong to the basic collection rules and can also be regarded as the event capture strategy of the button control, which define the frequency and method for the system to check the button state change, including control parameters such as the time interval (exemplarily, it can be stored in the aforementioned collection frequency field or other fields) and the mode switching condition (exemplarily, it can be stored in the aforementioned listening mode field or other fields). Among them, the time interval includes the baseline listening interval and the minimum listening interval. The baseline listening interval is the default polling time interval for button monitoring and serves as the reference value for the basic sampling frequency. The minimum listening interval is the shortest polling period allowed by the system. This parameter ensures that all valid events can still be captured in high-frequency operation scenarios while avoiding excessive consumption of CPU resources. Generally, the baseline listening interval is adjusted, and the adjusted baseline listening interval is not less than the minimum listening interval.
[0112] In this embodiment, the corresponding control listening configuration of the button control can include the interaction behavior statistical data of the button control, such as click frequency distribution, recent click frequency, click interval, etc. In addition, it can also include the aforementioned optimization parameter set.
[0113] The system may include a historical data analysis module, which can count interaction behavior statistics data such as the click frequency distribution of buttons through the historical data analysis module. The exponential weighted moving average algorithm can be used to perform exponential weighted calculations on each point in the click frequency distribution to obtain the recent click frequency.
[0114] In one example, when the recent click frequency is higher than the click frequency threshold, step 10222 can be executed. For example, when the recent click frequency > 5 times / second, the dynamic interval adjustment mode (i.e., step 10222) is started.
[0115] The initial value of the above baseline listening interval can be set to 100ms. There are various ways to dynamically adjust the baseline listening interval. For example, it can be dynamically adjusted according to the CPU usage rate (obtained through GetSystemTimes): When the CPU usage rate is in the first range (e.g., < 60%), the current or default baseline listening interval is maintained; when the CPU usage rate is in the second range (e.g., between 60% - 80%), it is extended by 20% based on the current or default baseline listening interval; when the CPU usage rate is in the third range (e.g., > 80%), the minimum listening interval is adopted.
[0116] Alternatively, the baseline listening interval can be dynamically adjusted based on the click interval: for example, when it is detected that the click interval is < 50ms for N consecutive times (N can be 3 or any other natural number), the baseline listening interval is automatically set to the minimum listening interval and monitored continuously for 300ms. Within these 300ms, if it is still detected that the click interval is < 50ms for N consecutive times, the minimum listening interval is maintained; otherwise, it can be restored to the initial value or the value of the previous adjustment.
[0117] The minimum listening interval is determined through hardware performance detection: the CPUID instruction is called to obtain the processor model. For Intel Core i5 and above, it is set to 10ms, and for others, it is set to 20ms. The system writes these parameters into the BUTTON_MONITOR_CONFIG configuration object, and the event monitoring service reads and adjusts the actual listening frequency in real time.
[0118] For the input box control, the above step 1022 further includes: Step 10223, based on the control listening configuration corresponding to the above input box control, adjusts the difference comparison rule for the above input box control. The above difference comparison rule includes the comparison period and the change content extraction rule.
[0119] In the embodiments of the present application, the difference comparison rule is the content change detection strategy of the input box control. By comparing the text differences at different time points, only the changed parts are collected to optimize the data transmission efficiency. The comparison period is the time interval at which the system performs text difference detection, and this parameter affects the real-time performance of content changes and the system overhead. The changed content extraction rule defines how to identify and record text differences.
[0120] Exemplarily, the control listening configuration corresponding to the input box control may include the content characteristics (text length) of the input box and the interactive behavior statistical data (e.g., input speed, input frequency).
[0121] Specifically, the system first detects the content characteristics of the input box: obtains the text length through GetWindowTextLength, and selects the comparison algorithm according to the length: for <100 characters, an exact comparison based on LCS is used, and for ≥100 characters, a fast comparison based on hashing is used. The comparison period is dynamically adjusted according to the input speed: uses a keyboard hook to count the input frequency. When the input interval < the interval threshold (e.g., 500 ms), the comparison period duration is set to the first duration (e.g., 300 ms); otherwise, it is set to the second duration (e.g., 500 ms), and the first duration should be less than the second duration. The changed content extraction rule adopts a three-level processing: locates the change position in the text through the Levenshtein distance; uses regular expressions to identify special formats (such as phone numbers, email addresses); merges consecutive deletion operations. The system compiles the parameters in the difference comparison rule into DIFF_RULE bytecode and is interpreted and executed by a dedicated text difference engine. After each comparison, a DELTA structure is generated, which contains fields such as the change position, old value, and new value, and is transmitted after being Base64 encoded.
[0122] The embodiments of the present application ensure that the strategy adapts to the real-time operating environment by providing a basic acquisition rule and a dynamic adjustment mechanism based on the control type and scenario to configure and adjust the basic acquisition rule to obtain a data acquisition strategy.
[0123] In other embodiments of the present application, after step 103, all the above embodiments further include: Step 301, encodes the above interactive behavior data into transmission data in a preset format and transmits it to the receiving end through an encryption protocol.
[0124] In the embodiments of the present application, the transmission data in the preset format refers to a standardized data structure defined by the system. For example, the Protocol Buffers serialization format can be used, and the message body structure containing fixed fields can be used to ensure the consistency of data parsing across platforms. The format contains necessary fields such as timestamp, session ID, control path, operation type, data content, etc. The encryption protocol refers to a transport layer security protocol that complies with the TLS 1.3 standard, using the AES-256-GCM encryption algorithm and the ECDHE key exchange mechanism to provide end-to-end data encryption protection. The protocol configuration includes parameters such as the encryption suite list, certificate verification rules, and session recovery strategy.
[0125] The system first encapsulates the interactive behavior data into a PROTOBUF_MSG structure, uses Varint encoding to compress integer fields, and applies UTF-8 encoding and Huffman compression to string fields. After data serialization, the system calls the secure transmission module to initialize the TLS session: generates an elliptic curve key pair through BCryptGenerateKeyPair; uses CertOpenSystemStore to obtain the preset CA certificate; and completes the TLS handshake after establishing a TCP connection. After the encrypted channel is established, the system transmits the data in blocks, each block size is fixed at 16KB, and an HMAC-SHA256 checksum is attached. The transmission process adopts a double buffering mechanism. When data in the foreground buffer is being sent, the background buffer continues to receive new monitoring data to ensure transmission continuity. The system monitors the network status in real time and automatically switches to the UDP+QUIC protocol when the delay exceeds 500ms.
[0126] Step 302: Based on the feedback result of the receiving end, optimize the data collection strategy (basic collection rule) corresponding to the target control.
[0127] In the embodiment of the present application, the feedback result refers to the data processing status report returned by the receiving end, which can be exemplarily in JSON format or XML. The data processing status report includes performance indicators such as data reception integrity indicators (packet loss rate, disorder rate), processing delay statistics, resource usage, and business-level data validity verification results.
[0128] Exemplarily, the data collection strategy may be stored in a shared memory configuration area.
[0129] For example, the system can establish a feedback analysis engine to process the receiving end data: parse the JSON feedback message and extract key performance indicators; use the sliding window algorithm to calculate the moving average of the packet loss rate in the feedback results received nearly 10 times; establish a correlation model between network indicators and collection frequency through regression analysis to confirm the correlation coefficient between network delay and collection frequency. Then optimize the data collection strategy (basic collection rules) based on the correlation coefficient.
[0130] Specifically, it can include three methods: 1) Modify the polling interval according to network metrics (for example, extend the polling interval by 5 ms for every 1% increase in the packet loss rate); 2) Reconstruct the event listener based on the data validity verification result (for example, increase the data collection frequency of the event listener when the proportion of invalid data > the proportion threshold, and those skilled in the art can flexibly design the proportion threshold, such as 10%, 20%, etc.); 3) Adjust the event trigger condition (for example, the aforementioned rolling trigger condition, and how to adjust it can refer to the previous description, which will not be elaborated here).
[0131] The optimized data collection strategy is synchronized to all monitoring threads through a memory-mapped file, and the CAS (Compare-And-Swap) operation is used to ensure the updateability of the strategy. The system records the strategy change history and rolls back to the version before optimization when the performance has not been improved after three consecutive optimizations.
[0132] The embodiments of the present application ensure data security and integrity through efficient encoding and encrypted transmission, and the dynamic optimization process introduces machine learning algorithms, enabling the system to automatically identify the optimal configuration combination, reducing the need for manual intervention, and improving the operation and maintenance efficiency.
[0133] The above data validity verification can ensure the accuracy and availability of data.
[0134] In an embodiment of the present application, data validity verification can be achieved through methods such as data quality inspection and outlier detection.
[0135] Among them, data quality inspection refers to the process of evaluating and verifying various aspects of data, such as accuracy, integrity, consistency, timeliness, reliability, and availability. Data quality is the basis for data analysis, data mining, machine learning, and any data-based decision-making.
[0136] Common methods of data quality inspection include: One, data cleaning: Handling missing values: Handle missing data through methods such as filling (such as mean, median, mode), deletion, or interpolation.
[0137] Correcting incorrect values: Identify and correct errors or outliers in the data.
[0138] Standardizing data formats: Unify data formats, such as date formats, currency units, etc.
[0139] Two, data validation: Range check: Ensure that the data value is within a reasonable range. For example, the output should not be negative.
[0140] Format check: Verify whether the data conforms to the predetermined format.
[0141] Uniqueness check: Ensures that the values of certain fields (such as primary keys) are unique.
[0142] Completeness Check: Make sure all required fields are filled.
[0143] 3. Data consistency check: Cross-table consistency: Check whether the data between different data tables is consistent.
[0144] Time series consistency: Checks whether the time series data is continuous and has no duplication.
[0145] 4. Data Deduplication: Check and delete duplicate records to ensure the uniqueness of the data.
[0146] 5. Data integrity check: Foreign key integrity: Ensure that the value of the foreign key field exists in the related table.
[0147] Field Completeness: Make sure all required fields are filled.
[0148] 6. Data Quality Monitoring: Use data quality monitoring tools (such as Informatica, Talend, etc.) to regularly check data quality; Set data quality indicators (such as missing value ratio, error value ratio, etc.) and monitor them.
[0149] Exemplary data quality checking tools may include: Excel: Suitable for preliminary inspection of small-scale data, discovering outliers through functions such as sorting, filtering, and conditional formatting.
[0150] SQL: Check the integrity, consistency and accuracy of data through SQL query statements.
[0151] Python: Use libraries like Pandas, NumPy for data cleaning and validation.
[0152] R language: Use dplyr, tidyverse and other packages for data processing and quality checking.
[0153] Professional data quality tools: such as Informatica Data Quality and Talend Data Quality, which provide powerful data quality inspection and repair functions.
[0154] For outlier detection, it can be implemented based on statistical methods (such as the standard deviation method), machine learning-based methods (such as clustering), etc.
[0155] In another embodiment of the present application, the data validity verification at the receiving end may include: For the interaction behavior data of the list view (LIST_VIEW) control, both adaptive numerical range verification (dynamically adjusting the threshold of 0 - 1000) and cross-field association verification (checking the logical consistency between the index and the content) are applied.
[0156] For the interaction behavior data of the button control (BUTTON), considering that the response of the button control has a certain response time, thus enabling timing behavior verification: screening the interaction behavior data according to a preset click interval, eliminating continuous invalid operations or preventing invalid and redundant interaction behavior data caused by misoperations. Exemplarily, if the response time interval after each press of the button control corresponding to the interaction behavior data is 100 ms, the preset click interval can be designed as 100 ms. If the user repeatedly clicks the button control at a frequency of 10 ms for up to 100 ms, it will result in collecting the interaction behavior data 10 times repeatedly within 100 ms. Then, according to the preset click interval of 100 ms, only the interaction behavior data collected when the button control is clicked between 90 ms - 100 ms is retained, and the interaction behavior data collected when the button control is clicked between 0 ms - 90 ms is eliminated to avoid generating invalid and redundant interaction behavior data.
[0157] The verification result triggers a three-level response mechanism: single-dimensional anomalies (such as values exceeding the preset range) mark the problem data and generate a warning log, keep the data in the database but trigger manual review. Two-dimensional conflicts (such as violating both the numerical range and the field association rules simultaneously) immediately terminate the current data stream, start an automatic re-collection program and lock the problem data source. Complex anomalies (failure of multiple rules such as numerical, association, timing, etc.) activate an emergency isolation protocol, transfer the abnormal data to the sandbox environment, and synchronously push an alarm notification containing a fault location map to the operation and maintenance terminal.
[0158] Then, by tracing the abnormal link, potential problems such as hardware drift, logical conflicts, or interface rendering anomalies are identified. The characteristic data such as the link identifier, data type, data metrics, etc. related to the interaction behavior data corresponding to the potential problems are hot-deployed to the verification engine in the form of a binary package, that is, the characteristics of the unknown potential problems are extracted and added to the verification engine, so that the verification engine can quickly identify abnormal problems / abnormal characteristics in the subsequent interaction behavior data. The updated policy supports a hybrid execution mode: event-driven verification (data change immediately triggers the verification process for the interaction behavior data) and periodic polling verification (trigger the verification process for the interaction behavior data at periodic time intervals) operate in parallel.
[0159] Finally, the system adopts a versioned incremental update protocol, carrying an incremental version number for each policy update, and maintaining zero service interruption during the smooth transition between the old and new rule sets (or old and new data collection policies). All verification operation records are encrypted and stored as evidence, and the exception handling process generates a timestamped audit trail to support policy effect traceability and compliance proof. Through the self-evolving cycle of "collection - verification - analysis - optimization", this closed-loop system can continuously improve data quality while reducing the cost of manual intervention, effectively coping with complex and changing data risk scenarios in a dynamic business environment.
[0160] Furthermore, when using the optimized data collection strategy to collect interaction behavior data, if a system failure occurs, it can be processed hierarchically.
[0161] Specifically, fault analysis can be performed according to the occurrence frequency of the same fault: when the fault frequency is lower than the fault frequency threshold, it can be determined as a level 1 error, and at this time, the current acquisition channel can be automatically switched to the backup acquisition channel. When the fault frequency is higher than the fault frequency threshold, it can be confirmed as a level 2 error. Here, the system snapshot stored by using the data collection strategy of the previous stable version can be used to restore the current system, so that the system can resume data collection using the data collection strategy of the previous stable version. Among them, the data collection strategy of the total stable version can be set by the user himself.
[0162] In one embodiment, step 302 includes: Step 3021, calculate the policy optimization weight corresponding to each target type according to the data integrity score and system load index in the above feedback result.
[0163] In the embodiments of the present application, the data integrity score is an index for quantitatively evaluating the data transmission quality, with a value range of 0 - 100%, calculated through the verification result at the receiving end, and reflecting the comprehensive influence of the packet loss rate, out-of-order rate, and verification failure rate.
[0164] The scoring algorithm uses weighted average. Exemplarily, the weight of the verification failure rate accounts for 60%, the packet loss rate accounts for 30%, and the out-of-order rate accounts for 10%.
[0165] The system load index is a set of metric values reflecting the usage of computer resources, including four core dimensions: CPU utilization rate (percentage), memory occupancy rate (MB), disk I / O waiting time (ms), and network bandwidth occupancy rate (Mbps). Data for each dimension is collected in real time through the operating system performance counter. The system load index can be obtained by normalizing the CPU utilization rate, memory occupancy rate, disk I / O waiting time, and network bandwidth occupancy rate to parameter values of 0 - 1, and then summing the parameter values corresponding to each core dimension.
[0166] The policy optimization weight is a policy adjustment priority value calculated for different control types, ranging from 0 to 1. The larger the value, the more urgently the collection policy of this type of control needs to be adjusted. The weight calculation comprehensively considers two dimensions: the importance of business optimization and the system impact degree.
[0167] The system establishes a multi-dimensional evaluation model to process feedback data in order to calculate the policy optimization weight: extract the data integrity scores of the last 5 transmissions from the feedback message and calculate the moving average; obtain the instantaneous value of the current system load indicator through PerformanceCounter; normalize each indicator to a value in the range of 0 - 1. The weight calculation uses a linear weighting formula: Policy optimization weight = α × (1 - Data integrity score) + β × System load indicator, where α and β are type-related coefficients (for list controls, α = 0.6, β = 0.4; for button controls, α = 0.4, β = 0.6). The calculation process uses SIMD instructions for parallel processing and updates the policy optimization weight every 100 ms. The system maintains a weight history queue (containing the historical policy optimization weights calculated over a period of time). When a weight mutation (change rate > 15%) is detected, an anomaly detection mechanism is triggered to perform anomaly detection on the interaction data collected by the control, determine the cause of the anomaly, and prompt the staff to handle the anomaly in a timely manner.
[0168] Step 3022, when the policy optimization weight corresponding to the list control reaches the list control weight threshold, adjust the data block size in the above-mentioned block loading rule.
[0169] In the embodiment of the present application, the list control weight threshold is a preset policy adjustment trigger value. Exemplarily, the default setting is 0.75. When the policy optimization weight exceeds this value, the system starts the list control collection policy optimization process. The threshold can be dynamically adjusted through the management interface according to business requirements. For the data block size, please refer to the foregoing description and will not be elaborated here.
[0170] More specifically, the system continuously monitors the policy optimization weight of the list control. When the weight value exceeds the threshold, in one example, the following adjustments can be made: analyze the correlation curve between the current data block size and the system CPU occupancy rate and the memory occupancy rate (obtained through GlobalMemoryStatusEx) to calculate the allowable value of the data block size; consider the network latency (obtained through ping test) to determine the recommended value of the data block size from the allowable value to complete the adjustment.
[0171] Exemplarily, the CPU occupancy rate at different data block sizes can be measured through experimental and performance monitoring tools (such as top, iostat, etc.), and then the correlation curve between the two can be obtained. After that, based on the current CPU occupancy rate and the upper limit of the CPU occupancy rate, the corresponding data block size value can be found in this correlation curve, so as to determine the value range of the data block size (which can be called the first value range).
[0172] Meanwhile, the value (which can be called the second value) or value range (which can be called the second value range) of the data block size can be determined according to the memory occupancy rate.
[0173] The intersection of the second value / second value range and the first value range can be taken as the allowed value of the data block size. If there is no intersection, the range with a smaller upper limit value of the interval can be taken as the allowed value. For example, assuming the second value range is [64, 128] and the first value range is [256, 512], [64, 128] can be taken as the allowed value. Another example, assuming the second value is 64 and the first value range is [256, 512], 64 can be taken as the allowed value.
[0174] Exemplarily, the method for determining the data block size according to the memory occupancy rate includes: According to the memory occupancy rate, calculate the currently available memory amount, and the formula is: available memory = total memory - used memory; where, used memory = memory occupancy rate * total memory; The following formula can be used to calculate the size of the data block: Data block size = available memory / number of data blocks.
[0175] For example, assuming that the available memory needs to be divided into 5 data blocks, then the size of each data block is: The maximum value of the data block size = available memory / 5. The lower limit of the data block size can be 0, or default settings, etc.
[0176] Exemplarily, the recommended value of the data block size can be determined from the allowed values in the following way: set a latency threshold. When the network latency is greater than the threshold, preferably the minimum value in the allowed values is taken as the recommended value; while when it is not greater than the threshold, the maximum value in the allowed values can be taken as the recommended value.
[0177] In another example, the adjustment algorithm for the data block size can adopt a PID control model: new block number = current value + Kp × (1 - data integrity score) + Ki × Σ (load deviation) + Kd × (weight change rate), where Kp = 8, Ki = 0.5, Kd = 2 are tuning parameters, and the load deviation is the deviation between the system load of the current device and the preset standard load. The adjusted value is written into the configuration register after being clamped by the upper and lower limits, and at the same time, all monitoring threads are notified to synchronize the new configuration through a memory barrier.
[0178] Step 3023: When the policy optimization weight corresponding to the button control reaches the button control weight threshold, switch the listening mode in the above listening rule to the hybrid listening mode, where the hybrid listening mode includes: the event-driven mode and the polling mode.
[0179] In the embodiment of the present application, the button control weight threshold is the activation condition value of the hybrid listening mode. Exemplarily, it is default set to 0.8. Further, when the policy optimization weight exceeds this threshold continuously for 3 times, the mode switch is triggered. This threshold setting takes into account the performance overhead brought by the mode switch. The hybrid listening mode is a composite monitoring strategy that combines event-driven and timed polling. Among them, the event-driven mode processes high-frequency operations, and the polling mode ensures the capture of low-frequency operations. The switching conditions of the two modes are dynamically adjusted based on the operation frequency, and the switching delay is controlled within 10 ms. High frequency and low frequency can be defined based on the frequency threshold, and those skilled in the art can flexibly design its value, which will not be elaborated here.
[0180] When the button control weight continuously exceeds the threshold, the system performs mode switching: save the state snapshot of the current event listener; initialize the polling thread. Exemplarily, the initial interval can be set to (100 - weight value × 80) ms. When running in the hybrid mode: in the high-frequency stage (for example, the operation frequency > 5 times / second), the event-driven mode is mainly used, and only the polling is enabled to fill in the gaps for the uncaught events. At this time, the baseline listening interval of the polling mode is 10 times / second; in the low-frequency stage, the polling frequency is automatically reduced to the default baseline listening interval. The system decides the current dominant mode based on three dimensions: operation frequency, CPU usage rate, and event loss rate. Each time a switch occurs, the system will compare the data capture rates in the previous and subsequent 3 seconds. When the improvement effect of the hybrid mode is < 5%, it will fallback to the original dominant mode.
[0181] The embodiment of the present application realizes the best balance between data acquisition quality and resource consumption through a quantitative evaluation mechanism, providing high-reliability data acquisition guarantee for the intelligent manufacturing scenario. All optimization operations record detailed logs, support post-event analysis and algorithm optimization, and form a closed-loop system for continuous improvement.
[0182] In one implementation, after step 302, it further includes: Step 303: Synchronize the optimized data acquisition policy to the running environment of the above target application through the hot loading mechanism.
[0183] In the embodiments of the present application, the hot loading mechanism is a technical solution for the system to dynamically update component configurations during operation, and seamless switching is achieved by using memory-mapped files and atomic operations. This mechanism includes three core modules: a configuration parser, a version manager, and a memory synchronization controller. The operating environment refers to the execution context in which the target application and its associated monitoring service are located, including runtime elements such as the process memory space, the thread pool state, and the loaded dynamic link libraries.
[0184] In one example, the system implements hot loading through the following process: serialize the optimized data acquisition policy into a binary format and write it to a temporary memory-mapped file; use the InterlockedCompareExchange atomic operation to replace the pointer to the data acquisition policy to ensure thread safety; notify the configuration manager of the control listening process (exemplarily used to execute step 103) through IPC to reload the data acquisition policy. In a specific implementation, the system creates a double-buffer structure, where the foreground buffer serves the current request and the background buffer loads the new data acquisition policy. During the switch, first freeze the control monitoring process (up to 50 ms), update the configuration references in all thread local storage (TLS), and finally release the thread to continue execution. The entire process maintains semaphore synchronization to ensure the atomicity and visibility of policy changes. The monitoring data generated during hot loading will be marked with special status bits for subsequent processing identification.
[0185] Step 304, verify the compatibility between the optimized data acquisition policy and the above control handle.
[0186] In the embodiments of the present application, the compatibility verification is an inspection process to ensure that the new configuration matches the characteristics of the target control, including three verification dimensions: handle validity verification, control type matching degree detection, and operation permission verification. The verification results are divided into three levels: fully compatible, partially compatible, and incompatible. The control handle is an interface element identifier assigned by the operating system, and during the verification process, it is necessary to check the survival status (IsWindow), access permission (GetSecurityInfo), and associated attributes (GetWindowLong) of the handle.
[0187] In one example, multi - layer collaboration can be adopted to achieve verification: the basic verification layer calls IsWindow to verify the validity of the handle and checks cross - process permissions through GetWindowThreadProcessId; the type verification layer uses GetClassName to match the control type and compares the control types corresponding to the old and new configurations; the function verification layer simulates sending messages such as WM_GETTEXT to test the actual operability. A timeout mechanism (default 300ms) is introduced in the verification process, and a compatibility score (0 - 100 points) is generated for each control handle. The system maintains a verification state machine. When the score ≥ 80 points, it is determined to pass (corresponding to the full compatibility mentioned above); when the score is 60 - 79 points, a downgraded adaptation is triggered (corresponding to the partial compatibility mentioned above); when the score is less than 60 points, a rollback strategy is adopted (corresponding to the incompatibility mentioned above). All verification results are recorded in the compatibility matrix as the basis for subsequent optimization.
[0188] Exemplarily, the compatibility matrix is composed of multiple row vectors. Each row vector corresponds to one verification and includes the verification results of three verification dimensions: the verification of handle validity, the detection of control type matching degree, and the verification of operation permissions in this verification, as well as the final verification result. Each row vector can be expressed as {a1, a2, a3, A}, where a1 represents the result of the handle validity verification, and different characters can be used to represent whether the handle validity verification passes or fails; a2 represents the result of the control type matching degree detection, and different characters can be used to represent matching or non - matching; a3 represents the result of the operation permission verification (actual operability), and different characters can be used to represent passing or failing. A is calculated based on a1' - a3'. a1' - a3' are obtained by normalizing a1 - a3. Taking a1 as an example, its value can be 1 or 0. A = (ɑa1' + ßa2' + γa3') / 3 * 100%, where ɑ, ß, γ represent the importance degrees of each dimension, which can all be equal to 1, or the values of ɑ, ß, γ can also be flexibly set, which will not be elaborated here.
[0189] Step 305, if the verification passes, then re - enter the step of real - time monitoring of the target control based on the data collection strategy (that is, the monitoring process restarts).
[0190] In the embodiment of the present application, the restart of the monitoring process is a state reset operation of the monitoring service, including three standard stages: releasing old resources, initializing new instances, and restoring data streams. The restart process ensures business continuity, and the lost windows are controlled within 3 events. The version number of the strategy is a 64 - bit incrementing identifier, which includes two parts: a time stamp (high 32 bits) and a serial number (low 32 bits). The version number is also used as the metadata of data records and supports traceability analysis according to the strategy.
[0191] Exemplarily, the monitoring process can be restarted in the following manner: The system sends a pause signal to the monitoring thread and waits for the events being processed to complete (up to 100 ms); calls FreeLibrary to unload the old monitoring module and loads the new implementation via LoadLibrary; initializes core components such as the monitoring filter and event queue with the new policy parameters. The version management service generates a new version number: obtains a timestamp based on GetSystemTimeAsFileTime and generates a sequence number through an atomic increment operation. The system writes the version number to the registry hive and the log file header, and simultaneously updates the version index table in memory. After the restart is complete, the monitoring service sends a ready event, and the data processor starts receiving monitoring data in the new format and embeds a version marker in each record.
[0192] In the embodiments of this application, zero-downtime deployment of policy changes is achieved through a hot reload mechanism, and the average effective time is shortened to within 200 ms. The compatibility verification mechanism ensures the security of policy updates and reduces the monitoring interruption rate caused by policy errors to less than 0.1%. Versioning management and atomic operations guarantee data consistency and support policy backtracking accurate to the millisecond level. This solution enables the system to achieve minute-level policy iteration and optimization without affecting business continuity, and the overall operation and maintenance efficiency is increased by 70%. The adaptive ability of the monitoring service is significantly enhanced, and it can automatically handle more than 95% of environmental change scenarios, providing continuous and stable monitoring guarantee for industrial data collection. All operation records are detailed in the audit log, meeting the security audit requirements of ISO27001.
[0193] In one embodiment, step 301 above includes: Step 3011, encapsulate the above interaction behavior data into a JSON structure; In the embodiments of this application, the JSON (JavaScript Object Notation) structure is a lightweight data interchange format that uses a text format completely independent of programming languages to store and transmit data. It consists of key-value pairs and includes two basic structures: objects (represented by curly braces {}) and arrays (represented by square brackets []), and supports basic data types such as strings, numbers, and booleans.
[0194] The system first parses the interaction behavior data into structured objects, including core fields such as timestamp (in ISO 8601 format), control handle (hexadecimal string), operation type (enumeration value), operation content, etc. The system organizes the data using a tree structure: the root node contains metadata (version number, session ID), and the child nodes store detailed information classified by operation type. For complex operations (such as list scrolling), the system treats additional parameters (scroll position, visible item index, etc.) as nested objects. The generated JSON strictly follows the RFC 8259 standard, with strings encoded in UTF-8 and numerical values in IEEE 754 double-precision format. The system manages the JSON construction process through a memory pool to avoid frequent memory allocation. After construction, it performs a syntax check (verified by JSON.parse) to ensure the validity of the data structure.
[0195] Step 3012: Perform Base64 encoding on the above JSON structure to generate intermediate encoded data. In the embodiments of this application, Base64 encoding is a method of encoding binary data based on 64 printable characters, converting every 3 bytes (24 bits) of data into 4 Base64 characters. The encoding character set includes 62 alphanumeric characters from A-Z, a-z, 0-9, as well as two symbols, "+", " / ", and usually uses "=" as the padding character.
[0196] The system takes the UTF-8 encoded JSON string as input and processes it in chunks of 118 bytes (meeting the MIME specification line length limit). Each chunk is converted by a Base64 encoder: grouped by 3 bytes, padded with zeros for less than 3 bytes; every 6 bits are converted into the corresponding Base64 character; the positions corresponding to the padded zero bytes are filled with "=". The encoding process is accelerated using SIMD instructions (such as the Intel AVX2 instruction set), with a 40% performance improvement. The system inserts a CRLF line break after every 76 characters (in line with RFC 2045), and the final output is in the standard Base64 MIME format. After encoding, the system verifies whether the output length is a multiple of 4 and ensures data losslessness through a decoding loopback test.
[0197] Step 3013: Encrypt the above intermediate encoded data using a symmetric encryption algorithm to generate an encrypted data packet.
[0198] In the embodiments of this application, a symmetric encryption algorithm refers to a cryptographic algorithm that uses the same key for encryption and decryption. This system uses the AES-256 (Advanced Encryption Standard) algorithm, with a 256-bit key, operating in the GCM (Galois / Counter Mode) mode, providing data encryption and integrity verification functions.
[0199] The system generates a 256-bit key through CryptGenRandom, creates a key object using BCryptGenerateSymmetricKey, and generates a 12-byte random IV (Initialization Vector). The encryption process is executed in three steps: perform PKCS7 padding on the Base64 data according to the AES block size (16 bytes), call the BCryptEncrypt interface for GCM mode encryption, and generate a 128-bit authentication tag at the same time. Concatenate the IV, ciphertext, and authentication tag in the format of [IV(12B)][Ciphertext(NB)][Tag(16B)]. The system uses hardware acceleration (such as the Intel AES-NI instruction set), and the encryption throughput reaches 5GB / s. Each data packet carries a key version number, and the receiving end obtains the corresponding decryption key from the key management service according to the version number.
[0200] Step 3014, add a transport protocol header to the above encrypted data packet to generate the above transport data that conforms to the preset communication specification.
[0201] In the embodiment of the present application, the transport protocol header is control information attached to the front of the data packet, including fields such as protocol version (2 bytes), data packet length (4 bytes), timestamp (8 bytes), checksum (4-byte CRC32), etc., and is encoded in network byte order (big endian).
[0202] The system constructs a protocol header structure: the protocol version field is fixed at 0x0102, calculates the length of the encrypted data packet (including the IV and authentication tag), obtains a timestamp accurate to 100ns through GetSystemTimeAsFileTime, and calculates the CRC32 checksum for the complete data packet. The protocol header is stored in a structure with memory alignment (4-byte boundary), and byte order conversion is processed through ntohs / htonl series functions. The final transport data is assembled in the format of [Protocol Header(18B)][Encrypted Data Packet], and a 2-byte END marker (0x0D0A) is appended at the end. The system manages data packets through a pre-allocated send buffer, supports zero-copy transmission, and each data packet is attached with a sequence number for receiving end reassembly and packet loss detection.
[0203] In the embodiments of the present application, standard JSON format is used to ensure data readability, Base64 encoding provides binary-safe transmission capabilities, AES-256-GCM encryption safeguards data confidentiality, and a structured protocol header enables reliable transmission. This solution controls the data encapsulation time within 2 ms, achieves a transmission efficiency of up to 98% of the theoretical bandwidth limit, and a packet integrity rate of 99.999%. The encryption strength meets the requirements of FIPS 140-2 Level 3 and supports tens of thousands of encryption operations per second. The design of the protocol header enables the receiving end to quickly verify data integrity and identify and repair transmission errors below 0.01%. The entire mechanism provides a highly secure and reliable transmission guarantee for industrial data collection, meeting the security requirements of the third level of information security protection while ensuring data real-time performance.
[0204] In one embodiment, step 103 includes: Step 1031, listening for target events of the above list control through an event listener or a message hook; wherein, the target events of the above list control include a scroll event and a selection event.
[0205] In the embodiments of the present application, an event listener is a set of callback functions registered by the system, implemented through the Windows message mechanism (such as WM_NOTIFY) or the UI Automation event interface (such as UIA_ScrollPattern_ScrollEvent), and is used to respond to asynchronous notifications of specific control events. A message hook is a system-level message interception mechanism, installing a WH_CALLWNDPROC type hook through the SetWindowsHookEx function to capture window messages sent to the target control (such as WM_VSCROLL, LBN_SELCHANGE). A scroll event is an operation notification triggered when the position of the vertical or horizontal scroll bar of the list control changes, including parameters such as the scroll direction (SB_LINEUP / SB_LINEDOWN) and the current position (the return value of GetScrollInfo). A selection event is a status change notification generated when a list item is selected by the user, including information such as the selected item index (the wParam of LBN_SELCHANGE) and the selection status (the return value of LB_GETSEL).
[0206] The system establishes a two - layer listening system for the list control: register ScrollPattern and SelectionPattern event listeners through UI Automation, and set event filters (such as UIA_ScrollHorizontallyScrollablePropertyId) to limit the listening scope; install a thread - level message hook using SetWindowsHookEx to capture WM_VSCROLL and LBN_SELCHANGE messages. The event handlers are managed by a priority queue: UI Automation events are processed with high priority (delay < 10ms), and message hook events are processed with low priority (delay < 50ms).
[0207] Among them, for the scroll event, the system records the starting position (GetScrollPos), the ending position, and the duration; for the selection event, it captures the text of the currently selected item (LB_GETTEXT) and the selection method (mouse / keyboard). All event data is attached with a timestamp accurate to the microsecond level (obtained by QueryPerformanceCounter) and is passed to the data - processing thread through a lock - free queue.
[0208] In addition, after capturing the scroll event, the list control can be listened to in real - time according to the aforementioned chunk - loading rules for chunk - loading, and when the stay time reaches the stay - time threshold, full - set collection is triggered - collecting data for all currently visible list items.
[0209] Step 1032, listen to the target events of the above - mentioned button control according to the listening mode in the corresponding data - collection strategy; among them, the above - mentioned listening mode is specifically the event - driven mode, the polling mode, or the hybrid listening mode; the above - mentioned hybrid listening mode includes: the event - driven mode and the polling mode; the target events of the above - mentioned button control include click events and long - press events.
[0210] In the embodiments of the present application, the event-driven mode is a passive listening method implemented through the Windows message pump (such as WM_COMMAND / BN_CLICKED) or the UI Automation event interface (UIA_InvokePattern_InvokedEvent). It triggers callbacks only when events occur, with a low CPU occupancy rate but may miss high-frequency events. The polling mode is an active listening method that periodically checks the button state (such as calling GetWindowLong every 50 ms to detect style changes), consuming more resources but ensuring the event capture rate. The hybrid listening mode is a composite strategy that dynamically combines event-driven and polling. By default, it uses the event-driven mode and automatically enables polling to fill in the gaps when detecting event loss (by counting the difference between BN_CLICKED and actual clicks). The click event is a standard operation triggered by a left mouse button click or a keyboard enter on the button, including metadata such as click coordinates (GET_X_LPARAM) and trigger time. The long-press event is a special operation where the button is continuously pressed for more than a threshold (default 800 ms), calculated by the time difference between WM_LBUTTONDOWN and WM_LBUTTONUP, and includes a parameter for the press duration.
[0211] Specifically, in the event-driven mode, a BN_CLICKED message processor and a UIA Invoke event listener can be registered; in the polling mode, a high-precision timer (timeSetEvent) can be started to periodically detect the button state (through BM_GETSTATE); in the hybrid listening mode, a basic event listener can be combined with a watchdog timer or a timer (default 100 ms interval). When the interval between two consecutive events is less than 50 ms, the polling mode is started for assistance. Any of the two events here can be a click event or a long-press event.
[0212] For long-press event detection, the system records the mouse press time (WM_LBUTTONDOWN) and starts a long-press detection timer (default 800 ms). When the timer times out, a long-press event is triggered. If WM_LBUTTONUP is received in advance, the timer is cancelled. All event data is attached with a hardware input timestamp (GetMessageTime) and transmitted through a thread-safe queue to ensure that the event order is consistent with the actual operation.
[0213] Step 1033, listen for the target events of the above input box control through the polling mode; wherein, the target events of the above input box control include text input events and focus switching events; the above interaction behavior data includes the monitored target events.
[0214] In the embodiments of the present application, the polling mode is a timing check policy for input box controls. The detection period is set by SetTimer (default 300 ms), and APIs such as GetWindowText and GetFocus are called regularly to obtain changes in the control state. The polling interval is dynamically adjusted according to the CPU load (calculated by GetSystemTimes). When the load > CPU occupancy upper limit (e.g., 70%), the polling interval can be automatically extended by a set step length. For example, the step length can be set to 20%, 50%, etc.
[0215] A text input event is an operation record of a change in the content of the input box, including details such as the text before the change (obtained through a difference comparison algorithm), the new text, and the change position. A focus switch event is a state transition of the input box gaining or losing the keyboard focus, including context information such as the focus state (WM_SETFOCUS / WM_KILLFOCUS) and the switching direction (Tab key / mouse click).
[0216] An exemplary implementation is as follows: Obtain the initial text (GetWindowText) and focus state (GetFocus) during initialization; start a timer to check regularly (configurable 300 - 1000 ms): detect changes in text length through GetWindowTextLength, and compare text content differences using memcmp; capture WM_SETFOCUS / WM_KILLFOCUS messages through the WH_CALLWNDPROC hook (WM_SETFOCUS and WM_KILLFOCUS are messages in the Windows system used to notify a window of gaining or losing input focus. When a window gains input focus, the system sends the WM_SETFOCUS message to the window, and when a window loses input focus, the system sends the WM_KILLFOCUS message). Both WM_SETFOCUS and WM_KILLFOCUS belong to WM_ messages.
[0217] Among them, three - level optimization is adopted for text change detection: directly determine a change when the lengths are different; compare hash values (MurmurHash3) when the lengths are equal; perform character - by - character comparison when the hashes are inconsistent.
[0218] The focus switch event is accompanied by the operation type that caused the switch (detect the Tab key / Shift - Tab key combination through GetAsyncKeyState). All focus switch event data records the complete window message sequence (the last 5 WM_ messages) for analyzing the operation context.
[0219] Embodiments of the present application achieve optimal resource utilization through a differential listening strategy: The list control adopts high-precision event listening to improve the capture rate; the button control supports dynamic mode switching, controlling the CPU occupancy to a relatively low level while ensuring the event capture rate; the input box polling mechanism reduces the data transmission volume through intelligent optimization. All event data is attached with nanosecond-level timestamps and complete operation contexts, providing a high-fidelity data source for behavior analysis. The solution supports processing a large number of control events per second, reducing the end-to-end latency and having a low memory occupancy rate. The dynamic switching mechanism of multi-mode listening enables the system to adapt to different load scenarios and maintain operational stability in the complex industrial field environment, providing reliable data support for device monitoring and operation analysis.
[0220] In one embodiment, after step 103 above, it further includes: outputting a data collection report of the target application, where the data collection report includes content obtained through at least one of the following methods: Statistical operation frequency distribution and timestamps of the interaction behavior data.
[0221] Identifying abnormal operations in the interaction behavior data through a clustering algorithm and generating warning events.
[0222] Associating the interaction behavior data with the device operation log based on timestamps to obtain an association table.
[0223] Rendering the analysis results of the interaction behavior data into visual charts (such as heat maps, line charts, scatter plots, etc.) and exporting them as documents.
[0224] In embodiments of the present application, the operation frequency distribution is the statistical result of the occurrence frequencies of various interaction behaviors in the target application in the time dimension, including a histogram of the number of operations per minute / hour / day, as well as derivative indicators such as peak frequency and average frequency. The timestamp is a record of the event occurrence moment accurate to the millisecond level. Exemplarily, it can be stored in the ISO 8601 extended format (YYYY-MM-DDThh:mm:ss.sssZ) and uniformly converted to the UTC time zone to eliminate the influence of time zone differences.
[0225] After obtaining the interactive data, the receiving end can clean the data. The data cleaning methods include but are not limited to deduplication, error correction, and missing value processing. Deduplication can be based on specific key fields (such as user ID, order number) to delete duplicate data, retain the latest or most complete records, and use similarity algorithms (such as Levenshtein distance, Jaccard similarity) to process records with similar but not identical spellings. Error correction can be detected according to business rules (such as negative age, wrong date format), and text errors can be corrected using natural language processing (NLP) tools (such as TextBlob, SymSpell). Missing value processing can be directly deleting records or fields with a high missing rate (applicable to cases with fewer missing values).
[0226] The interaction data may also be normalized by using Min-Max standardization or Z-score standardization. The format unification may be performed by unifying the date and time format, text size, encoding type, etc. in the data.
[0227] The interactive behavior data of the same control type or the same target control in the same time period can be clustered using a clustering algorithm to obtain multiple clusters, each of which includes interactive behavior data corresponding to at least one interactive behavior.
[0228] Exemplarily, the above clustering algorithm adopts the improved DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm, with parameters set to Eps=0.5 (normalized distance) and MinPts=10, which supports processing density anomalies in high-dimensional time series data.
[0229] The steps of the DBSCAN algorithm exemplarily include: Initialization: mark all points as unvisited; in this embodiment, the interaction behavior data corresponding to one interaction behavior corresponds to one point.
[0230] Select core points: Randomly select an unvisited point and check the number of points in its neighborhood.
[0231] Expand the cluster: If the point is a core point, add the points in its neighborhood to the cluster and expand it recursively.
[0232] Label Noise: Points that are not assigned to any cluster are labeled as noise.
[0233] Repeat the above steps until all points have been visited.
[0234] The DBSCAN algorithm can discover clusters of any shape, does not require the number of clusters to be specified in advance, and can handle noisy data.
[0235] In one example, the above abnormal operation may refer to the interaction behavior data or operation corresponding to a noise point; In another example, an abnormal operation refers to an interaction behavior that deviates from historical interaction data by more than 3σ (standard deviation), including predefined patterns such as an unconventional operation sequence (e.g., 10 consecutive rapid clicks), no operation for a timeout (interval > 30 minutes), etc.
[0236] An alarm event is a structured abnormal record, including fields such as the abnormal ID of the abnormal operation, the trigger time, the abnormal type (enumeration value), the associated control, the confidence level (0 - 100%), etc., and conforms to the Common Event Format standard. Exemplarily, the confidence level may refer to the confidence level of the abnormal operation. Among them, the abnormal ID can be used to uniquely identify an abnormal operation, and the abnormal type can be classified according to the statistically abnormal operations to obtain different abnormal categories. For example, the continuous rapid click type, the no operation for a timeout type, etc.
[0237] The device operation log is a standardized status record generated by industrial devices, including structured data such as the device ID (MAC address), the timestamp (synchronized to the NTP server), the operation parameters (e.g., temperature, rotation speed, etc., more than 200 indicators), and the alarm code. The association table is a relational database table structure, with the primary key being a composite key (device ID + event timestamp), and the fields include association results such as interaction behavior data, device status parameters, and association confidence scores (0 - 1).
[0238] A heat map is a two-dimensional density visualization chart. Exemplarily, the x-axis can be the time (24-hour system), the y-axis can be the control type, the color saturation represents the operation frequency, and the HSL color space (H = 120° green to 0° red gradient) is used to encode the frequency intensity. It can be seen that the heat map can be drawn based on the operation frequency distribution and timestamp of the interaction behavior data. That is to say, the operation frequency distribution and timestamp of the interaction behavior data belong to the analysis results.
[0239] A line chart is a time series trend chart, with the x-axis being the time (supporting the granularity from seconds to months), the y-axis being the statistical indicator (frequency / number of anomalies, etc.), including a confidence interval band and key point annotations (such as peaks, inflection points). The exported document is a PDF / A-3 file that conforms to the ISO 32000 standard, including vector graphics, structured bookmarks, accessibility tags, and digital signature blocks, and supports archiving for long-term preservation. Exemplarily, the line chart can be drawn based on the operation frequency distribution and timestamp of the interaction behavior data and the abnormal operation. That is to say, the abnormal operation also belongs to the analysis results.
[0240] In addition, the x-axis represents time (supporting granularity from seconds to months), and the y-axis represents statistical metrics (such as frequency / number of exceptions), and it can also be plotted as a scatter plot.
[0241] Exemplarily, the following analysis can be performed to obtain the analysis results.
[0242] The system processes the raw data through a distributed computing framework: 1) Use a time window aggregator (TumblingWindow) to count the number of various operations at a 1-minute granularity; 2) Generate P50 / P90 / P99 frequency metrics through a percentile calculator (Percentile Calculator); 3) Apply Fourier transform to identify periodic operation patterns. The statistical process adopts columnar storage optimization, builds a B+ tree index for the timestamp field, and controls the frequency query response time within 100ms. The analysis results are stored in a three-columnar structure (operation type, time bucket, count value), supporting dynamic range queries and sliding window analysis. The system automatically generates a frequency distribution report, including visualization components such as a time series histogram classified by operation type and a frequency comparison radar chart.
[0243] The system executes an anomaly detection pipeline: 1) Extract features such as operation interval, duration, and sequence pattern in the feature engineering stage; 2) Use a normalizing flow to map the features to a Gaussian space; 3) Run a clustering algorithm to identify outliers and calculate the Mahalanobis Distance for each cluster. The alarm generation module performs the following for confirmed abnormal operations (confidence > 85%): 1) Associate context data (the previous 5 operations); 2) Match a predefined rule library (such as industrial operation procedures); 3) Generate an alarm event containing repair suggestions. The alarm events are distributed through a message queue, supporting multiple notification methods such as email, SMS, and Syslog, and are recorded in the audit database for post-event analysis.
[0244] The system implements precise time correlation: 1) Unify the time reference (PTP protocol synchronization, error < 1ms) for two types of data sources (interaction behavior data and device operation logs); 2) Establish a time alignment window (a default ±500ms sliding window); 3) Perform correlation: For each target event, find all device logs within the time window of the time alignment window and calculate the correlation probability through a random forest model. Among them, the device log with the highest correlation probability will be associated with this target event.
[0245] The association table is stored in columnar format (Parquet), sharded by device ID, and a time range index is established for each shard. The system automatically maintains an association quality dashboard, which displays real-time metrics such as matching rate (>95% compliance) and latency distribution (P99 < 200ms). When the matching rate drops, a self-optimization process (adjusting the time window or model) is triggered. Among them, the displayed matching rate is the degree of matching between the successfully displayed data and the actually required displayed data, and the latency distribution refers to the deviation between the real-time data generation time and the real-time data display time.
[0246] The system generates reports through a visualization pipeline: 1) Use a heatmap renderer accelerated by WebGL to process millions of data points and dynamically adjust the color scale (0 - max frequency); 2) Smooth the line chart with Bezier curves and add a trend line (LOESS local regression); 3) The document assembly engine combines charts, tables, and text analysis according to a template and applies the enterprise VI style (font, color scheme). The export process includes: 1) PDF / A compliance check (verified by veraPDF); 2) Add a digital signature (RSA-PSS algorithm); 3) Generate attached machine-readable data (embed the original data summary in JSON-LD format). The final document is distributed through CDN, supporting version control and differential download (only update the modified parts).
[0247] In addition to being returned to the industrial device side to provide data collection support for subsequent device production and operation, the above analysis results can also be used in other application scenarios, such as customer segmentation, product recommendation, risk assessment, market trend prediction, etc.
[0248] In the embodiment of this application, through an automated analysis process, the system can process tens of millions of interactive data per hour and generate a comprehensive report containing multiple analysis dimensions. The report document meets the GMP data integrity ALCOA+ principle (attributable, legible, contemporaneous, original, accurate), supports audit trail and electronic signature verification. The entire mechanism improves the industrial operation analysis efficiency by 80% and shortens the problem location time by 90%, providing data support for equipment optimization, personnel training, and quality traceability. The system resource consumption is stable, with the peak memory occupancy of a single node <8GB and the CPU utilization maintained below 30%, suitable for long-term operation in the industrial field.
[0249] In one embodiment, after the above step 101, it further includes: If a control fingerprint matching the control handle of the target control is found in the pre-loaded static control fingerprint library, determine the data collection policy associated with the matching control fingerprint as the data collection policy of the above target control; the above static control fingerprint library includes associated control fingerprints and data collection policies; the step of inputting the feature data of the above target control into the collection policy generation model (step 102) is executed after no matching control fingerprint is found in the above static control fingerprint library.
[0250] The static control fingerprint library is a predefined control feature database loaded when the system starts. Exemplarily, it can be stored in a hash table structure. The hash table structure accesses records by mapping a key to a location in the table through a hash function, thus enabling fast data insertion, lookup, and deletion operations. The key is the control fingerprint (a 64-bit hash value), and the value is the associated data collection strategy. The fingerprint library exists in memory in a read-only manner and enables fast access through memory-mapped files. The control fingerprint is a digest value that uniquely identifies the control features and is generated by calculating the control attributes (including more than 20 dimensions of features such as window class name, style bits, location hash, parent window ID, etc.) using the SHA-256 algorithm, and it has the anti-collision property. The data collection strategy is a predefined set of monitoring scheme configurations, encoded in the Protocol Buffers format, and includes a listening mode (event-driven mode, polling mode, or hybrid listening mode), event listening configurations (such as message hook type, callback function pointer), data collection parameters (sampling rate, collection interval, data format), transmission settings (compression algorithm, encryption method), and other complete instruction sets. The feature data is a structured data set that describes the dynamic characteristics of the control, containing more than 30 dimensions of metrics collected in real time (handle value, text content, storage location, control type, status-related information, message response latency, number of child controls, drawing frequency, etc.), organized in a columnar storage format (Apache Arrow) to support fast vectorized calculations. The collection strategy generation model is a decision tree ensemble model based on XGBoost, with a 200-dimensional feature vector as the input and the optimal strategy configuration as the output.
[0251] An exemplary process for the system to perform fingerprint matching: 1) Extract the complete attribute set of the target control (obtained through APIs such as GetClassName and GetWindowLong); 2) Calculate the fingerprint value using the SIMD-accelerated SHA-256 algorithm, with a processing throughput of up to 1 GB / s; 3) Perform a fast pre-check in the Bloom filter (false positive rate < 0.1%) of the fingerprint library to confirm whether there may be a matching fingerprint; 4) If a possible match is determined, perform an exact lookup in the hash table (time complexity O(1)). The lookup process uses optimistic concurrency control, allowing 100,000 query requests per second simultaneously. When the match is successful, a read-only view of the strategy configuration is returned, containing more than 50 parameter items such as the listening mode and collection frequency.
[0252] When the data acquisition policy found in the static control fingerprint library, the system can further execute the policy loading process: 1) Verify the version compatibility of the data acquisition policy (check the match of the major version number of the data acquisition policy); 2) Parse the policy configuration into the memory structure, including the function pointer table (more than 20 callback functions) and the parameter block (more than 200 bytes of configuration data); 3) Initialize the policy execution context (allocate thread local storage and create an event queue); 4) Replace the policy pointer of the current control through an atomic operation, and the switching process takes less than 1 ms. After the policy takes effect, the system records the policy binding log (including the control handle, fingerprint hash, and loading timestamp), and updates the runtime policy cache (managed by the LRU algorithm, with a capacity of 1000 entries).
[0253] In addition, in other embodiments of the present application, the static control fingerprint library can also be updated using a dynamic data acquisition policy: 1) Start the feature acquisition pipeline, and collect the feature data of the control in real time through the performance counter (QueryPerformanceCounter) and the API hook (Detours library); 2) Perform feature engineering (standardization, missing value filling, PCA dimensionality reduction); 3) Call the model inference engine (ONNX Runtime) to generate the data acquisition policy configuration, which takes less than 5 ms; 4) Write the new data acquisition policy and its fingerprint into the dynamic policy cache (ConcurrentDictionary), and start a background thread to perform policy verification (simulate 100 operations). The data acquisition policies that pass the verification will be batch merged into the static fingerprint library regularly (for example, every 24 hours) to complete knowledge accumulation.
[0254] In the embodiments of the present application, sub-millisecond policy matching for 95% of common controls is achieved through the static fingerprint library, and 5% of the long-tail controls are processed by the dynamic model. The average policy decision time of the system is reduced from 50 ms to 1.2 ms, the memory occupancy is reduced by 40% (through fingerprint sharing), and the CPU utilization is reduced by 35%. The dynamic policy generation mechanism enables the system to adapt to new types of controls, and the policy accuracy rate increases by 0.5% per month. The hot update function of the fingerprint library supports updating more than 1000 policies per hour without restarting the service, providing continuous optimization of data acquisition capabilities for industrial sites. The entire mechanism ensures that 100,000 policy queries per second are supported on an 8-core CPU, and the false decision rate is less than 0.01%, meeting the industrial-level reliability requirements.
[0255] In a feasible embodiment of the present application, referring to Figure 2 , the data acquisition system may include: production equipment, edge computing, and cloud services.
[0256] Production equipment refers to physical equipment in industrial automation scenarios (such as PLCs, numerically controlled machine tools, etc.). These devices interact with application programs through control handles to generate operation data and production logs. Edge computing is a distributed computing node close to the data source, responsible for real-time processing of the raw data generated by production equipment, reducing the transmission load to the cloud. Cloud services refer to the storage, computing, and analysis capabilities provided by remote servers, which receive the data processed by edge computing for in-depth analysis and long-term storage.
[0257] Among them, after the target application program is started, the control handle capture program in the production equipment will start to obtain the control handles in the interface of the target application program, and then will monitor the interface elements corresponding to the control handles, analyze the monitored interaction data, and send the interaction data to edge computing.
[0258] Edge computing will clean, assemble, and then encrypt the received interaction data and send it to cloud services.
[0259] Cloud services will decode and parse the encrypted data and insert it into the corresponding business table according to business processing.
[0260] Of course, the above is an exemplary description. In actual applications, edge computing and cloud services can be the same device or different devices. Edge computing can also be the same device as the production equipment, which can be specifically set according to actual needs and is not limited here.
[0261] Based on the same inventive concept, the embodiments of the present application also provide a data collection device for implementing the data collection method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the following data collection devices can refer to the limitations on the data collection method in the above text and will not be repeated here.
[0262] In an exemplary embodiment, as Figure 3 shown, a data collection device 50 is provided, including: An acquisition module 501, configured to acquire the control handles in the current interface of the target application program; the above control handles are at least used to determine the control type; among them, in the current interface, the control whose control type belongs to the target type is the target control; the above target type includes lists, buttons, and input boxes; the target controls in the above current interface include at least one of list controls, button controls, and input box controls.
[0263] A processing module 502, configured to input the feature data of the above target control into a collection policy generation model to generate a data collection policy corresponding to the control type of the above target control; the above feature data at least includes: control type.
[0264] The acquisition module 503 is configured to perform real-time monitoring on the target control based on the above data acquisition strategy, so as to capture the interaction behavior data of the target control when a state change is detected.
[0265] Reference may be made to the foregoing description, and details are not repeated here.
[0266] In one embodiment, the control handle of the target control is the target control handle; the acquisition module 501 is further configured to: Call a preset API or function to read the text content and storage location of the target control handle.
[0267] Obtain the status-related information and control type of the target control according to the target control handle.
[0268] Wherein, the characteristic data includes the handle value, the text content, the storage location, the control type and status-related information of the target control. Reference may be made to the foregoing description, and details are not repeated here.
[0269] In terms of generating a data acquisition strategy corresponding to the control type of the target control, the processing module 501 is configured to: Determine a basic acquisition rule matching the control type of the target control from a preset policy template library.
[0270] Adjust the matching basic acquisition rule based on the control listening configuration corresponding to the target control to obtain a data acquisition strategy. Reference may be made to the foregoing description, and details are not repeated here.
[0271] In one embodiment, in terms of adjusting the matching basic acquisition rule, the processing module 502 is configured to: Based on the control listening configuration corresponding to the list control, adjust the block loading rule for the list control, where the block loading rule includes a scrolling trigger condition and a data block size, and collect the currently visible list item data when a scrolling event for the list control is detected, and collect the clicked list item data when a click event for the list control is detected.
[0272] Based on the control listening configuration corresponding to the button control, adjust the listening rule for the button control, where the listening rule includes a baseline listening interval and a minimum listening interval.
[0273] Based on the control listening configuration corresponding to the input box control, adjust the difference comparison rule for the input box control, where the difference comparison rule includes a comparison period and a change content extraction rule. Reference may be made to the foregoing description, and details are not repeated here.
[0274] In one embodiment, the processing module 502 is further configured to: Encode the above interactive behavior data into transmission data in a preset format, and transmit it to the receiving end through an encryption protocol.
[0275] Optimize the data collection strategy corresponding to the above target control based on the feedback result of the above receiving end. Refer to the foregoing description, and details are not repeated here.
[0276] In one embodiment, in terms of optimizing the data collection strategy corresponding to the target control based on the feedback result of the receiving end, the processing module 502 is used for: Calculate the policy optimization weight corresponding to each target type according to the data integrity score and system load index in the above feedback result.
[0277] When the policy optimization weight corresponding to the list control reaches the list control weight threshold, adjust the data block size in the above block loading rule.
[0278] When the policy optimization weight corresponding to the button control reaches the button control weight threshold, switch the listening mode in the above listening rule to the hybrid listening mode, and the above hybrid listening mode includes: event-driven mode and polling mode. Refer to the foregoing description, and details are not repeated here.
[0279] In one embodiment, the above processing module 502 is further used for: Synchronize the optimized data collection strategy to the running environment of the above target application through the hot loading mechanism; Verify the compatibility between the optimized control listening configuration and the above control handle.
[0280] If the verification passes, re-enter the step of performing real-time listening on the target control based on the data collection strategy. Refer to the foregoing description, and details are not repeated here.
[0281] In one embodiment, in terms of encoding the interactive behavior data into transmission data in a preset format, the processing module 502 is used for: Encapsulate the above interactive behavior data into a JSON structure.
[0282] Perform Base64 encoding on the above JSON structure to generate intermediate encoded data.
[0283] Encrypt the above intermediate encoded data through a symmetric encryption algorithm to generate an encrypted data packet.
[0284] Add a transmission protocol header to the above encrypted data packet to generate the above transmission data that conforms to the preset communication specification. Refer to the foregoing description, and details are not repeated here.
[0285] In one embodiment, in terms of real-time monitoring of a target control based on a data collection strategy, the acquisition module 503 is configured to: Monitor the target events of the above list control through an event listener or a message hook; wherein, the target events of the above list control include a scroll event and a selection event.
[0286] Monitor the target events of the above button control according to the monitoring mode in the control monitoring configuration; wherein, the monitoring mode in the above control monitoring configuration is an event-driven mode, a polling mode, or a hybrid monitoring mode; the above hybrid monitoring mode includes: an event-driven mode and a polling mode; the target events of the above button control include a click event and a long-press event.
[0287] Monitor the target events of the above input box control through a polling mode; wherein, the target events of the above input box control include a text input event and a focus switching event.
[0288] The above interaction behavior data includes the monitored target events. Refer to the foregoing description for details and no further elaboration will be provided here.
[0289] In one embodiment, the above acquisition module 503 is further configured to: Output a data collection report of the above target application, and the above data collection report includes the content obtained through at least one of the following methods: Statistically analyze the operation frequency distribution and timestamps of the above interaction behavior data; Identify abnormal operations in the above interaction behavior data through a clustering algorithm and generate an alarm event; Associate the above interaction behavior data with the device operation log based on the timestamp to obtain an association table; Render the analysis result of the above interaction behavior data into a visualization chart (such as a heat map, a line chart, a scatter chart, etc.) and export it as a document. Refer to the foregoing description for details and no further elaboration will be provided here.
[0290] In one embodiment, the above acquisition module 503 is further configured to: If a control fingerprint matching the control handle of the target control is found in the pre-loaded static control fingerprint library, determine the data collection strategy associated with the matching control fingerprint as the data collection strategy of the above target control; the above static control fingerprint library includes associated control fingerprints and data collection strategies.
[0291] The step of inputting the feature data of the above target control into the acquisition strategy generation model is executed after no matching control fingerprint is found in the above static control fingerprint library. Refer to the foregoing description for details and no further elaboration will be provided here.
[0292] In the embodiments of the present application, through the data acquisition technology based on control handle capture, the control handle can be obtained, and corresponding data acquisition strategies can be generated for different target controls according to the type of the control handle, so as to monitor the target controls according to the data acquisition strategies to implement the data acquisition of various interaction data, without relying on the data interface provided by the production equipment and being restricted by the types of collectable data provided by the data interface, thereby improving the applicability of the interaction data acquisition in the production equipment.
[0293] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0294] Exemplarily, the above computer device may be a server or a terminal, and its internal structure diagram may be exemplarily as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data acquisition data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the data acquisition method in the foregoing embodiments is implemented.
[0295] Those skilled in the art can understand that Figure 4 the structure shown in
[0296] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0297] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0298] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0299] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0300] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0301] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0302] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. A data acquisition method, characterized in that, The data acquisition method includes: Obtaining the control handle in the current interface of the target application; the control handle is at least used to determine the control type; among them, the control whose control type belongs to the target type in the current interface is the target control; the target type includes list, button, and input box; the target controls in the current interface include at least one of list controls, button controls, and input box controls; Inputting the feature data of the target control into the acquisition policy generation model to generate a data acquisition policy corresponding to the control type of the target control; the feature data at least includes: control type; Based on the data acquisition policy, performing real-time monitoring on the target control to capture the interaction behavior data of the target control when a state change is detected.
2. The data acquisition method according to claim 1, wherein The control handle of the target control is the target control handle; the process of obtaining the feature data includes: Calling a preset API or function to read the text content and storage location of the target control handle; Obtaining the status-related information and control type of the target control according to the target control handle; Among them, the feature data includes the handle value, the text content, the storage location, the control type of the target control, and status-related information.
3. The data acquisition method according to claim 1, wherein The step of generating a data acquisition policy corresponding to the control type of the target control includes: Determining a basic acquisition rule that matches the control type of the target control from a preset policy template library; Adjusting the matched basic acquisition rule based on the control monitoring configuration corresponding to the target control to obtain a data acquisition policy.
4. The data acquisition method according to claim 3, characterized in that The step of adjusting the matched basic acquisition rule includes: Based on the control monitoring configuration corresponding to the list control, adjusting the chunk loading rule for the list control, where the chunk loading rule includes a scroll trigger condition and a data chunk size, and collecting the currently visible list item data when a scroll event for the list control is detected, and collecting the clicked list item data when a click event for the list control is detected; Based on the control monitoring configuration corresponding to the button control, adjusting the monitoring rule for the button control, where the monitoring rule includes a baseline monitoring interval and a minimum monitoring interval; Based on the control monitoring configuration corresponding to the input box control, adjusting the difference comparison rule for the input box control, where the difference comparison rule includes a comparison period and a change content extraction rule.
5. The data acquisition method according to claim 4, wherein After the step of capturing the interaction behavior data of the target control, it further includes: Encoding the interaction behavior data into transmission data in a preset format and transmitting it to the receiving end through an encryption protocol; Optimizing the data acquisition policy corresponding to the target control based on the feedback result of the receiving end.
6. The data acquisition method according to claim 5, characterized in that, The step of optimizing the data acquisition policy corresponding to the target control based on the feedback result of the receiving end includes: Calculating the policy optimization weights corresponding to each target type according to the data integrity score and system load index in the feedback result; When the policy optimization weight corresponding to the list control reaches the list control weight threshold, adjusting the data chunk size in the chunk loading rule; When the policy optimization weight corresponding to the button control reaches the button control weight threshold, switch the listening mode in the listening rule to the hybrid listening mode, where the hybrid listening mode includes: an event-driven mode and a polling mode.
7. The data acquisition method according to claim 5, wherein After the step of optimizing the data collection policy corresponding to the target control, the following steps are further included: Synchronize the optimized data collection policy to the running environment of the target application through a hot loading mechanism; Verify the compatibility between the optimized data collection policy and the control handle; If the verification passes, re-enter the step of performing real-time monitoring on the target control based on the data collection policy.
8. The data acquisition method according to claim 5, characterized in that, The step of encoding the interaction behavior data into transmission data in a preset format includes: Encapsulate the interaction behavior data into a JSON structure; Perform Base64 encoding on the JSON structure to generate intermediate encoded data; Encrypt the intermediate encoded data through a symmetric encryption algorithm to generate an encrypted data packet; Add a transmission protocol header to the encrypted data packet to generate the transmission data that conforms to the preset communication specification.
9. The data acquisition method according to claim 1, wherein The step of performing real-time monitoring on the target control based on the data collection policy includes: Monitor the target events of the list control through an event listener or a message hook; where the target events of the list control include a scroll event and a selection event; Monitor the target events of the button control according to the listening mode in the control listening configuration; where the listening mode in the control listening configuration is an event-driven mode, a polling mode, or a hybrid listening mode; the hybrid listening mode includes: an event-driven mode and a polling mode; the target events of the button control include a click event and a long press event; Monitor the target events of the input box control through a polling mode; where the target events of the input box control include a text input event and a focus switching event; The interaction behavior data includes the monitored target events.
10. The data acquisition method according to claim 1, wherein After the step of capturing the interaction behavior data of the target control, the following steps are further included: Output a data collection report of the target application, where the data collection report includes the content obtained through at least one of the following methods: Statistical operation frequency distribution and timestamps of the interaction behavior data; Identify abnormal operations in the interaction behavior data through a clustering algorithm and generate alarm events; Associate the interaction behavior data with the device operation log based on the timestamp to obtain an association table; Render the analysis result of the interaction behavior data into a visualization chart and export it as a document.
11. The data acquisition method according to claim 1, wherein After the step of obtaining the control handle of the current interface of the target application, the following steps are further included: If a control fingerprint matching the control handle of the target control is found in the pre-loaded static control fingerprint library, determine the data collection policy associated with the matching control fingerprint as the data collection policy of the target control; the static control fingerprint library includes associated control fingerprints and data collection policies; The step of inputting the characteristic data of the target control into the collection policy generation model is executed after no matching control fingerprint is found in the static control fingerprint library.
12. A data acquisition device, characterized in that, The data collection device includes: An acquisition module, configured to acquire the control handle in the current interface of the target application; the control handle is at least used to determine the control type; wherein, in the current interface, the control whose control type belongs to the target type is the target control; the target type includes lists, buttons, and input boxes; the target controls in the current interface include at least one of list controls, button controls, and input box controls; A processing module, configured to input the feature data of the target control into a collection policy generation model to generate a data collection policy corresponding to the control type of the target control; the feature data at least includes: control type; A collection module, configured to perform real-time monitoring on the target control based on the data collection policy to capture the interaction behavior data of the target control when a state change is monitored.
13. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data collection method according to any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the data collection method according to any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes instructions, which when run, cause the data collection method according to any one of claims 1 to 11 to be executed.
Citation Information
Patent Citations
Page data acquisition method and device and server
CN109739717A
Graphical interface control object generation method and system
CN115129311A
Multi-platform data acquisition method
CN119377043A
Profile based capture component
US20050246588A1
Data-processing device, data-processing method, program, and computer-readable medium
US20120185760A1
Cited By
Message analysis method and device, storage medium and electronic equipment
CN115034624A
Heterogeneous acceleration method of database, electronic equipment, storage medium and program product
CN120561153A
Robot control strategy migration method, system and device, medium and program product
CN120755885A
Polling strategy determination method and device, equipment, storage medium and product
CN120880951A
Smart home voice recognition method based on Internet of Things
CN121262027A