Heterogeneous terminal multi-dimensional trace data identification method and system

By developing a multi-dimensional trace data recognition method for heterogeneous terminals in terminal equipment, the problems of low data processing efficiency and insufficient adaptability in the prior art are solved, and more efficient and accurate network security threat identification and early warning are achieved.

CN119945705APending Publication Date: 2025-05-06INFORMATION CENT OF YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411729482.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing multi-dimensional trace data recognition methods for terminals have shortcomings in terms of processing efficiency and adaptability, making it difficult to effectively identify and early warning of network security threats.

Method used

A multi-dimensional trace data recognition method for heterogeneous terminals is proposed. By obtaining the target terminal equipment data, preprocessing, training the data to identify the model and optimizing the model, the real-time data is finally input to the model for identification and preprocessing to obtain the recognition results.

Benefits of technology

It improves the processing efficiency and accuracy of terminal equipment data, can adapt to different types of terminal equipment, provide accurate trace data identification results, and helps network personnel better understand and respond to network security incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945705A_ABST
    Figure CN119945705A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous terminal multi-dimensional trace data identification method and system. The method comprises the following steps: acquiring first target terminal equipment data; performing first preprocessing on the first target terminal equipment data to obtain second target terminal equipment data; taking the second target terminal equipment data as a training sample, and training a first data identification model; and performing second preprocessing on the first data identification model to obtain a second data identification model. The data processing efficiency and accuracy of the terminal equipment can be effectively improved, so that the performance of the whole system is improved. The method can adapt to different types of terminal devices, and has good universality and expansibility. By providing an accurate trace data identification result, network personnel are helped to better understand and cope with network security events. And an identification result is converted into a uniform format, so that subsequent data analysis and processing are facilitated, and powerful data support is provided for decision support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-source heterogeneous data recognition and detection of terminal devices, and in particular to a method and system for recognizing multi-dimensional trace data of heterogeneous terminals. Background Art

[0002] The terminal multi-dimensional trace data recognition method is a method for identifying and detecting multi-source heterogeneous data of terminal devices. The current multi-source heterogeneous data recognition and detection method for terminal devices includes: collecting multi-source heterogeneous data; using annotation tools to process multi-source heterogeneous data to generate a multi-source heterogeneous data training set; training the yolov5 learning model through the multi-source heterogeneous data training set to obtain a target detection model and a model weight file; inputting multi-source heterogeneous data into the target detection model, outputting detection information, and obtaining the ancillary information of multi-source heterogeneous data at the same time. With the continuous development of science and technology, people's requirements for terminal multi-dimensional trace data recognition methods are getting higher and higher.

[0003] The existing terminal multi-dimensional trace data identification method has certain drawbacks when used. Network security risks are inherently complex and uncertain. Without an early warning mechanism, it is difficult to detect and prevent problems in a timely manner. Compared with traditional defense systems, network security early warning models are more inclined to prevent and predict risks. Its importance in network security maintenance is becoming increasingly apparent. Modern network attack threats have become increasingly complex and advanced, requiring continuous updating and upgrading of security systems and continuous improvement of the skills of network personnel. Summary of the invention

[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the present invention provides a method and system for identifying multi-dimensional trace data of heterogeneous terminals, which can solve the problems mentioned in the background technology.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, the present invention provides a method for identifying multi-dimensional trace data of heterogeneous terminals, comprising:

[0009] Acquiring first target terminal device data;

[0010] Performing a first preprocessing on the first target terminal device data to obtain second target terminal device data;

[0011] Using the second target terminal device data as a training sample to train a first data recognition model;

[0012] Perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

[0013] As a preferred solution of the heterogeneous terminal multi-dimensional trace data recognition method of the present invention, wherein: the first data recognition model includes:

[0014] The first data recognition model is an arbitrary model that inputs terminal device data of different types, different manufacturers and different operating systems, and outputs at least information to be detected and auxiliary information corresponding to the information to be detected.

[0015] As a preferred solution of the heterogeneous terminal multi-dimensional trace data recognition method of the present invention, the second preprocessing includes:

[0016] Determining whether the first data recognition model meets a first preset standard;

[0017] The first data recognition model is updated according to the judgment result to obtain a second data recognition model.

[0018] As a preferred solution of the heterogeneous terminal multi-dimensional trace data recognition method of the present invention, the first preprocessing includes:

[0019] Establish pre-processing strategies for different types, manufacturers and operating systems;

[0020] The preprocessing strategy at least includes preprocessing strategies for domestic operating systems, iOS operating systems, Android operating systems, LINUX operating systems, and Windows operating systems;

[0021] The preprocessing strategies all include extracting features from the data corresponding to the first target terminal device.

[0022] As a preferred solution of the method for identifying multi-dimensional trace data of heterogeneous terminals described in the present invention, the feature extraction includes extracting feature data including at least timestamp, geographic location, device information, network information and user behavior.

[0023] As a preferred solution of the heterogeneous terminal multi-dimensional trace data identification method of the present invention, it also includes:

[0024] Obtain real-time target terminal device data;

[0025] inputting the real-time target terminal device data into the second data recognition model;

[0026] Perform a third preprocessing on the output result of the second data recognition model.

[0027] As a preferred solution of the heterogeneous terminal multi-dimensional trace data recognition method of the present invention, the third preprocessing includes:

[0028] The output result of the second data recognition model is subjected to field mapping, data format conversion, and data cleaning to obtain a multi-dimensional trace data recognition result of heterogeneous terminals with the same structure and format.

[0029] In a second aspect, the present invention provides a heterogeneous terminal multi-dimensional trace data recognition system, comprising:

[0030] A data acquisition module, used to acquire the first target terminal device data;

[0031] A preprocessing module, configured to perform a first preprocessing on the first target terminal device data to obtain second target terminal device data;

[0032] A first model acquisition module, configured to use the second target terminal device data as a training sample to train a first data recognition model;

[0033] The second model acquisition module is used to perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

[0034] In a third aspect, the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method when executing the computer program.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the method described above when executed by a processor.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention proposes a method and system for multi-dimensional trace data identification of heterogeneous terminals, which obtains the data of a first target terminal device; performs a first preprocessing on the data of the first target terminal device to obtain the data of a second target terminal device; uses the data of the second target terminal device as a training sample to train a first data identification model; performs a second preprocessing on the first data identification model to obtain a second data identification model. It can effectively improve the processing efficiency and accuracy of terminal device data, thereby improving the performance of the entire system. It can adapt to different types of terminal devices, including but not limited to personal computers, smart phones, tablet computers, etc., and has good versatility and extensibility. By providing accurate trace data identification results, it helps network personnel better understand and respond to network security incidents. The identification results are converted into a unified format to facilitate subsequent data analysis and processing, providing strong data support for decision support. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:

[0038] Figure 1 A method flow chart of a method and system for multi-dimensional trace data recognition of heterogeneous terminals provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0040] Example 1

[0041] Reference Figure 1 , is the first embodiment of the present invention, which provides a method and system for identifying multi-dimensional trace data of heterogeneous terminals, such as Figure 1 As shown, including:

[0042] There are some problems in the existing related technologies, such as low data processing efficiency, difficulty in adapting to various types of terminal devices, low accuracy of trace data recognition, etc. This application provides a method that can effectively solve the above-mentioned problems. Next, we will combine multiple embodiments to explain in detail how to implement the heterogeneous terminal multi-dimensional trace data recognition method;

[0043] S100: Acquire first target terminal device data;

[0044] In an optional embodiment, the first target terminal device may be a multi-source heterogeneous data acquisition device, including devices of different types, different manufacturers, and different operating systems, such as a personal computer, a smart phone, a tablet computer, etc.

[0045] In an optional embodiment, obtaining the first target terminal device data may require in-depth understanding of the characteristics of various devices and operating systems, and using appropriate technologies and tools to achieve the goal of data collection. At the same time, it is also necessary to comply with relevant laws and regulations, respect user privacy, and ensure the legality and security of the data collection process.

[0046] In an optional embodiment, device type and manufacturer: First, it is necessary to clarify the device types and manufacturers to be supported. This includes smartphones, tablets, laptops, IoT devices, etc. According to the characteristics of the target device, corresponding data collection methods and technologies need to be developed.

[0047] In an optional embodiment, heterogeneous terminals may run various operating systems, such as iOS, Android, Windows, Linux, etc. In order to successfully collect data, it is necessary to understand the data access and collection methods of each operating system and develop applicable collection tools or applications accordingly.

[0048] In an optional embodiment, appropriate methods are required for data collection depending on the device type, manufacturer, and operating system. The following technologies are included: a. API integration: Use the API interface provided by the operating system to directly obtain device data. This can be implemented through the operating system's development kit (SDK) or other supported interfaces. b. Application: Develop a dedicated application and install it on the target device to collect specific data. The application can be released through the app store and requires the user to install and authorize access to device data.

[0049] In an optional embodiment, for some devices, specific types of data can be obtained through sensors (such as accelerometers, gyroscopes, position sensors, etc.) These data can be collected through the API interface of the device itself or related libraries.

[0050] In the embodiments of the present application, the data acquisition operation is not limited. Any method randomly selected from the above should be within the protection scope of the present application, and any method randomly selected from the above is suitable for subsequent operations of the present application.

[0051] It should be noted that obtaining the first target terminal device data can provide rich and diverse original information for subsequent data processing and analysis. Through these data, a more comprehensive and accurate device usage and behavior model can be constructed, providing a solid foundation for trace data identification. In addition, the acquired data can also be used for equipment performance monitoring, fault diagnosis, user behavior analysis and other aspects, thereby providing support for the management and maintenance of terminal devices. In the process of data collection, it is crucial to ensure the integrity and accuracy of the data.

[0052] S200: performing a first preprocessing on the first target terminal device data to obtain second target terminal device data;

[0053] In the embodiment of the present application, the first preprocessing includes:

[0054] Establish pre-processing strategies for different types, manufacturers and operating systems;

[0055] The preprocessing strategy at least includes preprocessing strategies for domestic operating systems, iOS operating systems, Android operating systems, LINUX operating systems, and Windows operating systems;

[0056] The preprocessing strategies all include extracting features from the data corresponding to the first target terminal device.

[0057] In an embodiment of the present application, feature extraction includes extracting feature data including at least timestamp, geographic location, device information, network information, and user behavior.

[0058] In an optional embodiment, the preprocessing strategy can be specifically planned according to different operating systems. For example, for domestic operating systems, log analysis technology can be used to extract key information from system logs, such as application usage frequency, system startup time, etc.;

[0059] In an optional embodiment, for the iOS operating system, developer tools provided by Apple, such as Instruments in Xcode, may be used to monitor and record device usage;

[0060] In an optional embodiment, for an Android operating system, the Android Debug Bridge (ADB) tool may be used to obtain device status and application usage data;

[0061] In an optional embodiment, for a LINUX operating system, user activity and system configuration information may be obtained by reading system log files and configuration files;

[0062] In an optional embodiment, for Windows operating system, Windows Management Instrumentation (WMI) can be used to collect system and application operation data. Through these pre-processing strategies, standardization and normalization of data can be ensured, laying a foundation for subsequent data analysis and processing.

[0063] In an optional embodiment, the step of extracting feature data of timestamp, geographic location, device information, network information and user behavior for the features of the domestic operating system may include the following steps:

[0064] Step 1: Extract timestamp information from system logs to determine the specific time when the event occurred;

[0065] Step 2: Use the location service API to obtain the device's geographic location information;

[0066] Step 3: Collect the device's hardware and software information, such as device model, operating system version, etc.

[0067] Step 4: Obtain network information such as network connection status and IP address through the network interface;

[0068] Step 5: Analyze application usage records and user interaction data to identify user behavior patterns.

[0069] In an optional embodiment, the step of extracting feature data of timestamp, geographic location, device information, network information, and user behavior for the features of the iOS operating system may include the following steps:

[0070] Step 1: Use Xcode's Instruments tool to monitor device usage and record application startup and shutdown events;

[0071] Step 2: Get the real-time location information of the device through the location service API;

[0072] Step 3: Get the detailed hardware and software configuration of the iOS device through the device information interface;

[0073] Step 4: Use the network interface to obtain the network connection status and the device's Wi-Fi or cellular network information;

[0074] Step 5: Analyze application usage data and user interaction logs to identify user behavior habits.

[0075] In an optional embodiment, the step of extracting feature data of timestamp, geographic location, device information, network information, and user behavior for the features of the Android operating system may include the following steps:

[0076] Step 1: Use the ADB tool to obtain the device's operating status and application usage;

[0077] Step 2: Use the location service API to obtain the real-time location information of the device;

[0078] Step 3: Obtain the hardware and software configuration of the Android device through the device information interface;

[0079] Step 4: Obtain network information such as network connection status and device IP address through the network interface;

[0080] Step 5: Analyze application usage records and user interaction data to identify user behavior patterns.

[0081] In an optional embodiment, the step of extracting feature data of timestamp, geographic location, device information, network information and user behavior for the features of the LINUX operating system may include the following steps:

[0082] Step 1: Extract key system activity logs through system log analysis tools to determine timestamp information;

[0083] Step 2: Use the GPS module or network positioning service to obtain the device's geographic location;

[0084] Step 3: Collect detailed information of the device by reading the system configuration file and hardware information file;

[0085] Step 4: Use network management tools to obtain network information such as network connection status and IP address;

[0086] Step 5: Identify user behavior patterns by analyzing system logs and user activity records.

[0087] In an optional embodiment, the step of extracting feature data of timestamp, geographic location, device information, network information, and user behavior for the features of the Windows operating system may include the following steps:

[0088] Step 1: Use the Windows Management Instrumentation (WMI) interface to collect system and application operation data;

[0089] Step 2: Obtain the real-time location information of the device through the GPS module or network positioning service;

[0090] Step 3: Obtain the hardware and software configuration details of the device through the system information interface;

[0091] Step 4: Obtain network information such as network connection status and device IP address through the network interface;

[0092] Step 5: Identify user behavior patterns by analyzing system logs and application usage. Through the above steps, it is possible to ensure that representative and consistent feature data is extracted from different operating systems, providing high-quality input for subsequent data analysis and processing.

[0093] It should be noted that performing a first preprocessing on the first target terminal device data to obtain the second target terminal device data can significantly improve the efficiency and accuracy of data processing. Through preprocessing, the original data is converted into a structured and standardized format, which not only simplifies the data storage and management process, but also facilitates subsequent data analysis. For example, the extraction of timestamp information makes it possible to track and analyze the changing trends of user behavior over time; the acquisition of geographic location information helps to analyze the geographical distribution of user activities; the collection of device information and network information provides a detailed record of device usage and network behavior. The extraction of user behavior feature data provides an important basis for understanding user habits and preferences.

[0094] In an optional embodiment, the implementation of the preprocessing strategy can also be adjusted and optimized according to different application scenarios to meet specific data analysis requirements. In short, the first preprocessing step is an indispensable part of the entire heterogeneous terminal multi-dimensional trace data identification method and system, which provides a solid foundation for subsequent data analysis and processing.

[0095] S300: Using the second target terminal device data as a training sample to train the first data recognition model;

[0096] In the embodiment of the present application, the first data recognition model includes:

[0097] The first data recognition model is an arbitrary model that inputs terminal device data of different types, different manufacturers and different operating systems, and outputs at least information to be detected and auxiliary information corresponding to the information to be detected.

[0098] In an optional embodiment, the first data recognition model can be a classifier based on machine learning, such as a support vector machine (SVM), a random forest, a neural network, etc. These models can automatically identify and classify new data samples by learning a large amount of labeled training data. For example, the SVM model distinguishes different categories of data by finding the best hyperplane in a high-dimensional space, while the random forest improves the accuracy of classification by building multiple decision trees and voting. The neural network model can handle complex nonlinear relationships by simulating the structure and function of neurons in the human brain, and is suitable for processing large-scale and high-dimensional data sets.

[0099] In the embodiment of the present application, the first data recognition model is designed using yolov5, the input is terminal device data of different types, different manufacturers and different operating systems, and the output is at least information to be detected and auxiliary information corresponding to the information to be detected;

[0100] In the embodiment of the present application, multi-source heterogeneous data is the pattern data of classified devices such as servers, routers, switches, disks, u ports, cabinets, cables, serial ports, etc. collected by different devices. The multi-source heterogeneous data is processed using a labeling tool, including labelling the multi-source heterogeneous data using a picture labeling tool labellmg, and the labeling information mainly includes the category information and coordinate position information of the multi-source heterogeneous data.

[0101] In an optional embodiment, the designed yolov5 learning model network includes an input end, a backbone end, a neck end and a prediction end.

[0102] In the embodiment of the present application, the input end adopts a mosaic data enhancement method, including adaptive anchor frame calculation and adaptive image scaling;

[0103] In the embodiment of the present application, the backbone end includes a focus structure and two CSP (cross stage partial) structures, which are used to aggregate multi-source heterogeneous data and form image features;

[0104] In an embodiment of the present application, the neck end adopts the structure of FPN (feature pyramid networks) and PAN (path aggregation network) to extract image features of multi-source heterogeneous data and pass the image features to the prediction layer.

[0105] In an embodiment of the present application, the prediction end uses giou_loss (generalized intersection over union loss) as a loss function of a bounding box, which is used to predict multi-source heterogeneous data to be detected based on image features.

[0106] It should be noted that training the yolov5 learning model includes using a multi-scale sliding window with different anchors (anchor frames) to train a weight file that meets the requirements. The steps of inputting multi-source heterogeneous data into the target detection model and then outputting detection information and simultaneously obtaining the ancillary information of the multi-source heterogeneous data include: the mobile terminal obtains multi-source heterogeneous data in the real scene through the camera, and inputs the processed multi-source heterogeneous data into the mobile terminal containing the target detection model for recognition and detection; the mobile terminal outputs recognition and detection information, draws the detection frame in the multi-source heterogeneous data of the target asset, and obtains the ancillary information of the multi-source heterogeneous data through the communication unit.

[0107] In an optional embodiment, when training the first data recognition model, a cross-validation method can be used to evaluate the performance of the model to ensure that the model has good generalization ability. In addition, feature selection technology can be used to reduce the complexity of the model, improve training efficiency, and avoid overfitting. Through these methods, the first data recognition model can effectively extract valuable information from heterogeneous terminal devices to provide support for subsequent data analysis and decision-making.

[0108] It should be noted that the benefit of using the second target terminal device data as training samples to train the first data recognition model is that it can significantly improve the recognition accuracy and processing speed of the model for heterogeneous terminal data. In this way, the model can learn the characteristics of different device data and adapt to data changes in various complex scenarios.

[0109] S400: Perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

[0110] In the embodiment of the present application, the second preprocessing includes:

[0111] Determining whether the first data recognition model meets a first preset standard;

[0112] The first data recognition model is updated according to the judgment result to obtain a second data recognition model.

[0113] In an optional embodiment, the first preset standard may be that performance indicators such as accuracy, recall rate, F1 score, etc. of the model reach a predetermined threshold.

[0114] In an optional embodiment, if the first data recognition model fails to meet the first preset criterion, the following steps are performed: re-adjusting model parameters, such as learning rate, batch size, optimizer type, etc.;

[0115] In an optional embodiment, different data enhancement techniques, such as rotation, scaling, cropping, etc., may also be used to increase the generalization ability of the model.

[0116] In an optional embodiment, a regularization technique, such as L1 or L2 regularization, may also be introduced to prevent overfitting.

[0117] In the embodiment of the present application, the first preset standard is specifically to set the threshold value as the accuracy of the model must reach more than 95%, the recall rate must reach more than 90%, and the F1 score must reach more than 92%. These indicators ensure that the model has high accuracy and reliability when identifying heterogeneous terminal data.

[0118] In the embodiment of the present application, if the first data recognition model fails to meet these preset standards, it will be adjusted and optimized accordingly based on the performance feedback of the model. For example, if the accuracy is low, it may be necessary to add more training data or adjust the model structure to improve its recognition ability.

[0119] In the embodiment of the present application, if the recall rate is insufficient, it may be necessary to optimize the data preprocessing steps to ensure that all relevant data are fully learned by the model.

[0120] In the embodiment of the present application, the optimization of the F1 score may require balancing the accuracy and recall rate, which can be achieved by adjusting the classification threshold or improving the decision logic of the model. Through these careful adjustments, it can be ensured that the performance of the second data recognition model is significantly improved to better meet the needs of practical applications.

[0121] S105, obtaining real-time target terminal device data;

[0122] In an embodiment of the present application, the real-time target terminal device data is input into the second data recognition model;

[0123] The output result of the second data recognition model is subjected to a third preprocessing.

[0124] In the embodiment of the present application, the third preprocessing includes:

[0125] The output results of the second data recognition model are subjected to field mapping, data format conversion, and data cleaning to obtain multi-dimensional trace data recognition results of heterogeneous terminals with the same structure and format.

[0126] It should be noted that the third processing is to ensure the accuracy and consistency of the data to facilitate subsequent data analysis and processing. Field mapping is to match the data fields in the recognition results with the predefined field templates to ensure that each data item can be correctly classified into the corresponding field. Data format conversion is to unify data from different sources into a standard format, such as converting all timestamps into a unified time format, or converting values ​​in different units into standard units. Data cleaning includes removing duplicate data, correcting errors and outliers, and filling in missing values ​​to improve data quality. Through these preprocessing steps, it can be ensured that the final multi-dimensional trace data recognition results are both accurate and easy to analyze, providing reliable data support for decision makers.

[0127] In an optional embodiment, on the Windows operating system, the Minifilter of the file system is used to record traces of file operations. The Minifilter driver is a driver for file system filtering and interception. It inserts custom code on the I / O path of the file system, allowing developers to monitor, modify and intercept file system operations. The Minifilter driver can intercept and filter file system operations, such as file creation, opening, reading, writing and closing. Developers can selectively intercept or record these operations as needed and perform custom processing. The Minifilter driver can also monitor the activities of the file system in real time. It can provide real-time notifications of access, modification and deletion of files and directories, and record these events for security audits, troubleshooting or other purposes. Use kernel callback functions to record behavioral traces such as creation, opening, and termination of processes and threads. Use the WFP framework to record network connection operations. WFP (Windows Filtering Platform) is a network packet processing framework in the Windows operating system. It provides a flexible way to allow developers to filter, inspect and modify traffic during network packet transmission. WFP allows developers to monitor the transmission and processing of network packets. It can record information about application-initiated network connection quintuples, packet flows, and events to aid in network traffic analysis, troubleshooting, and security auditing.

[0128] This embodiment also provides a heterogeneous terminal multi-dimensional trace data recognition system, including:

[0129] A data acquisition module, used to acquire the first target terminal device data;

[0130] A preprocessing module, used for performing a first preprocessing on the first target terminal device data to obtain second target terminal device data;

[0131] A first model acquisition module, used to train a first data recognition model using the second target terminal device data as a training sample;

[0132] The second model acquisition module is used to perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

[0133] The above-mentioned unit modules may be embedded in or independent of the processor in the computer device in the form of hardware, or may be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0134] This embodiment also provides a computer device, which may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for identifying multi-dimensional trace data of a heterogeneous terminal is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0135] This embodiment further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0136] Acquiring first target terminal device data;

[0137] Performing a first preprocessing on the first target terminal device data to obtain second target terminal device data;

[0138] Using the second target terminal device data as a training sample to train the first data recognition model;

[0139] The first data recognition model is subjected to a second preprocessing to obtain a second data recognition model.

[0140] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0141] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0142] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0143] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0145] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0146] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for identifying multi-dimensional trace data of heterogeneous terminals, characterized in that: include: Acquiring first target terminal device data; Performing a first preprocessing on the first target terminal device data to obtain second target terminal device data; Using the second target terminal device data as a training sample to train a first data recognition model; Perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

2. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 1, characterized in that: The first data recognition model includes: The first data recognition model is an arbitrary model that inputs terminal device data of different types, different manufacturers and different operating systems, and outputs at least information to be detected and auxiliary information corresponding to the information to be detected.

3. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 2, characterized in that: The second preprocessing comprises: Determining whether the first data recognition model meets a first preset standard; The first data recognition model is updated according to the judgment result to obtain a second data recognition model.

4. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 3, characterized in that: The first preprocessing comprises: Establish pre-processing strategies for different types, manufacturers and operating systems; The preprocessing strategy at least includes preprocessing strategies for domestic operating systems, iOS operating systems, Android operating systems, LINUX operating systems, and Windows operating systems; The preprocessing strategies all include extracting features from the data corresponding to the first target terminal device.

5. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 4, characterized in that: The feature extraction includes extracting feature data including at least timestamp, geographic location, device information, network information and user behavior.

6. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 5, characterized in that: Also includes: Obtain real-time target terminal device data; inputting the real-time target terminal device data into the second data recognition model; Perform a third preprocessing on the output result of the second data recognition model.

7. The method for identifying multi-dimensional trace data of heterogeneous terminals according to claim 6, characterized in that: The third preprocessing comprises: The output result of the second data recognition model is subjected to field mapping, data format conversion, and data cleaning to obtain a multi-dimensional trace data recognition result of heterogeneous terminals with the same structure and format.

8. A heterogeneous terminal multi-dimensional trace data recognition system, characterized in that: include: A data acquisition module, used to acquire the first target terminal device data; A preprocessing module, configured to perform a first preprocessing on the first target terminal device data to obtain second target terminal device data; A first model acquisition module, configured to use the second target terminal device data as a training sample to train a first data recognition model; The second model acquisition module is used to perform a second preprocessing on the first data recognition model to obtain a second data recognition model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.