Data processing method and device and electronic equipment
By receiving data on the identity, location and time information of the carrying device, using the data risk prediction model to conduct risk assessment and determine the processing method, the problem of insufficient data risk analysis in the prior art is solved, and rapid and accurate data processing and storage pressure are achieved.
Patent Information
- Application Number
- CN202510039804.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art fails to effectively perform data risk analysis when submitting data to the server, resulting in increased pressure on invalid data storage and storage.
By receiving data on the identity, location and time information of the carrying device, using the data risk prediction model for risk assessment, and determining the data processing method based on the evaluation results, and performing verification in advance to avoid excessive risk operations.
It realizes rapid data risk analysis, reduces the storage of invalid data, reduces storage pressure, and improves the accuracy and efficiency of data processing.
Smart Images

Figure CN120105448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data risk identification, and in particular to a data processing method, device and electronic equipment. Background Art
[0002] In the current environment, every time data is submitted to the server and saved in the system, there is a risk. Invalid data and disguised operations have a huge impact on the validity of the data saved by the system for analysis and display. Invalid data cannot correctly help system users make effective decisions. Most systems basically ignore the verification of data validity. A few systems involving finance and payment will pay attention to the validity of operations, and avoid identification risks through face recognition and SMS verification codes. Post-processing is better than pre-prevention, resulting in increased cleaning costs for later data and reduced reliability of early statistics.
[0003] In related technologies, most systems store data directly into the system without processing its validity, or perform manual inspection or spot checks. Regardless of whether the data is stored without processing or manually inspected, the data has been stored in the system database, which undoubtedly increases the pressure on database storage. There are two types of manual inspections: one is to manually call some tools to calculate and judge the data; the other is to use no tools, purely manual calculation and judgment methods. Post-verification will increase the database storage pressure and the complexity of query processing.
[0004] How to design a data processing method that can quickly implement data risk analysis when submitting data to a server is a technical problem that needs to be solved urgently. However, there is currently no technical solution that can solve the above technical problem, and there is no data processing method, device or electronic device. Summary of the invention
[0005] The present invention provides a data processing method, device and electronic device, which perform verification in advance instead of making judgments during query processing, thereby avoiding operations with excessive risks, reducing the storage of invalid data and reducing storage pressure.
[0006] In a first aspect, the present invention provides a data processing method, comprising:
[0007] Receive current data sent from the current device, the current data carrying original data, current device identity, current location information and current time information, the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment;
[0008] Inputting the current device identity, the current location information, and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model;
[0009] Determining a processing method for the original data according to the data risk value;
[0010] The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data;
[0011] The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
[0012] According to the data processing method provided by the present invention, before receiving current data sent from the current device, the method includes:
[0013] The current device recombines the system version number, the underlying software program number, and the core hardware unique identifier corresponding to the current operating environment according to a preset rule to determine recombined data, wherein the recombined data includes, arranged from left to right, the first 3 digits of the system version number, the first 3 digits of the underlying software program number, the first 2 digits of the core hardware unique identifier, the length digit of the core hardware unique identifier, the remaining string of the core hardware unique identifier, the length digit of the underlying software program number, the remaining digits of the underlying software program number, the length digit of the system version number, and the remaining digits of the system version number;
[0014] The current device processes the reorganized data using a HASH algorithm to generate the current device identity.
[0015] According to the data processing method provided by the present invention, after generating the current device identity, the method further includes:
[0016] The current device performs public key encryption on the original data, the current device identity, the current location information, and the current time information to obtain encrypted current data;
[0017] The current device sends the encrypted current data to the server.
[0018] According to the data processing method provided by the present invention, before inputting the current device identity, the current location information and the current time information into the data risk prediction model, the method further includes:
[0019] The current data is decrypted using the private key of the server to obtain the original data, the current device identity, the current location information and the current time information.
[0020] According to the data processing method provided by the present invention, before inputting the current device identity, the current location information and the current time information into the data risk prediction model, the method further includes:
[0021] For any sample device, perform public key encryption on the sample data, the sample device identification, the sample location information, the sample time information, the verification mark and the parity check code to obtain the encrypted sample data, wherein the verification mark is generated according to the parity check code;
[0022] The sample device sends the encrypted sample data to the server;
[0023] Decrypting the encrypted sample data using the private key of the server to obtain sample data, sample device identification, sample location information, sample time information, verification mark, and parity check code;
[0024] The parity check code is used to perform a parity check on the verification mark. If it is determined that the parity check fails, the check weight is determined to be 0; otherwise, the check weight is determined to be 1.
[0025] According to the data processing method provided by the present invention, after performing a parity check on the verification mark using the parity check code, the method further includes:
[0026] If it is determined that the sample device identity exists, the sample device identity weight is determined to be 1; otherwise, the sample device identity weight is determined to be 0;
[0027] When the sample time information is within the first preset time period, the sample time weight is determined to be 0.4; when the sample time information exceeds the first preset time length and the second preset time length, the sample time weight is determined to be 0.2; when the sample time information exceeds the second preset time length, the sample time weight is determined to be 0, and the first preset time length is less than the second preset time length;
[0028] When the sample position information is within the first range, the sample position weight is determined to be 0.6; when the sample position information is greater than the first range and less than the second range, the sample position weight is determined to be 0.4; when the sample position information is greater than the second range and less than the third range, the sample position weight is determined to be 0.2; when the sample position information is greater than the third range, the sample position weight is determined to be 0; the first range is smaller than the second range, and the second range is smaller than the third range.
[0029] According to the data processing method provided by the present invention, after determining the sample device identification weight, the sample location weight and the sample time weight, the sample risk value is determined according to the sample device identification weight, the sample location weight and the sample time weight;
[0030] Determining the sample risk value according to the sample device identification weight, the sample location weight, and the sample time weight includes:
[0031] Y=(1-v×d×(t+p))×100%
[0032] Among them, Y is the sample risk value, v is the verification weight, d is the sample device identification weight, t is the sample time weight, and p is the sample location weight.
[0033] According to the data processing method provided by the present invention, the method of determining the processing mode of the original data according to the data risk value includes:
[0034] When the data risk value is greater than or equal to 60%, a first processing instruction is generated, wherein the first processing instruction is used to instruct recording the original data, terminating the business operation, and sending a warning email to the system administrator;
[0035] When the data risk value is greater than or equal to 30% and less than 60%, a second processing instruction is generated, wherein the second processing instruction is used to instruct to execute a business operation of the original data and mark the original data so as to remove the mark after manual review;
[0036] When the data risk value is less than 30% and greater than 0, a third processing instruction is generated, wherein the third processing instruction is used to instruct execution of a business operation of the original data and sending an in-site reminder message and a reminder email;
[0037] When the data risk value is 0, a fourth processing instruction is generated, where the fourth processing instruction is used to instruct execution of a business operation on the original data.
[0038] In a second aspect, a data processing device is provided, comprising:
[0039] A receiving unit, the receiving unit is used to receive current data sent from the current device, the current data carries original data, current device identity, current location information and current time information, the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment;
[0040] An input unit, the input unit is used to input the current device identity, the current location information and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model;
[0041] A determination unit, the determination unit is used to determine a processing method of the original data according to the data risk value;
[0042] The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data;
[0043] The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
[0044] According to a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the data processing method is implemented when the processor executes the program.
[0045] The present invention provides comprehensive input for data risk prediction by receiving current data carrying rich information (including device identity, location, and time), and uses a data risk prediction model to conduct risk assessment on data, thereby improving the accuracy and efficiency of data processing, and flexibly determining the processing method of the original data according to the data risk value, thereby achieving effective management and control of the data;
[0046] The present invention also recombines device information through preset rules, increases the confidentiality and cracking difficulty of device identity, generates device identity through HASH algorithm, ensures the uniqueness and security of device identity, ensures the security of data during transmission through public key encryption, prevents data from being stolen or tampered with, and the server decrypts data through private key to ensure the accuracy and reliability of the decryption process. By comprehensively considering device identity, time and location information, multi-dimensional weight information is provided for data risk prediction. By combining multiple weight information to determine the sample risk value, a comprehensive evaluation index is provided for data risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 It is a flow chart of the data processing method provided by the present invention;
[0049] Figure 2 It is a structural schematic diagram of a data processing device provided by the present invention;
[0050] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] Figure 1 : is a flow chart of a data processing method provided by the present invention, wherein the data processing method comprises:
[0053] Step 101: receiving current data sent from a current device, wherein the current data carries original data, a current device identity, current location information, and current time information, wherein the current device identity is generated according to an operating system version number, an underlying software program number, and a core hardware unique identifier corresponding to a current operating environment;
[0054] Step 102: input the current device identity, the current location information, and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model;
[0055] Step 103: Determine a processing method for the original data according to the data risk value;
[0056] The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data;
[0057] The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
[0058] In step 101, the current device sends data to the data processing center through a preset communication protocol. The current device may be a smart phone, an IoT device, etc. The preset communication protocol may be HTTP, HTTPS, MQTT, etc. The received data packet contains original data, current device identity, current location information, and current time information. The original data may be sensor readings and user input information. The current location information may be GPS coordinates and base station information. The current time information may be a timestamp. The current device identity is a unique identifier generated by comprehensively considering the current operating environment, the underlying software program number, and the core hardware unique identifier. The current operating environment may be an operating system version number such as Android 10 and iOS14, the underlying software program number may be a specific firmware version, and the core hardware unique identifier may be a MAC address or an IMEI number.
[0059] Optionally, before receiving current data sent from the current device, the method includes:
[0060] The current device recombines the system version number, the underlying software program number, and the core hardware unique identifier corresponding to the current operating environment according to a preset rule to determine recombined data, wherein the recombined data includes, arranged from left to right, the first 3 digits of the system version number, the first 3 digits of the underlying software program number, the first 2 digits of the core hardware unique identifier, the length digit of the core hardware unique identifier, the remaining string of the core hardware unique identifier, the length digit of the underlying software program number, the remaining digits of the underlying software program number, the length digit of the system version number, and the remaining digits of the system version number;
[0061] The current device processes the reorganized data using a HASH algorithm to generate the current device identity.
[0062] Optionally, before sending the current data, the current device first recombines the system version number, underlying software program number, and core hardware unique identifier corresponding to the current operating environment according to preset rules, and uses the HASH algorithm to process the recombined data to generate the current device identity. The introduction of this step aims to improve the complexity and security of the device identity and prevent the identity from being easily cracked or forged. In an optional embodiment, assume that the operating environment of the current device is as follows: system version number: Android 12.0.1, underlying software program number: FW1003, core hardware unique identifier: MAC address 00-1A-2B-3C-4D-5E, and recombines the current operating environment information according to preset rules: extract the first 3 digits of the system version number: And, extract the first 3 digits of the underlying software program number: FW1, extract the first 2 digits of the core hardware unique identifier: 00, core hardware unique identifier length: 6, the remaining string of the core hardware unique identifier: 1A-2B-3C-4D-5E, underlying software program number length: 4, underlying The remaining digits of the software program number: 003, the length of the system version number: 6, the remaining digits of the system version number: 2.0.1, and in another optional embodiment, it can be simplified, for example, the system version number takes the major version number and the minor version number, such as 120, the bottom software program number takes the first 4 digits, such as FW10, and the core hardware unique identifier takes the complete MAC address, but for security reasons, it can be hashed before taking only part or deforming it, and the recombined data is: 120FW1000-1A-2B-3C-4D-5E64. Then, the HASH algorithm (such as SHA-256) is used to process the above reorganized data to generate the current device identity.
[0063] The present invention generates a device identity with high complexity and uniqueness through recombination and HASH algorithm processing, which is difficult to crack or forge, thereby improving data security. The device identity contains key information such as the system version number, the underlying software program number and the core hardware unique identifier, so that the source of the data can be traced, which is helpful for subsequent data management and risk control. The technical solution can be applied to different types of devices and operating system environments, and only needs to be adjusted according to the operating environment information of the specific device.
[0064] Optionally, after generating the current device identity, the method further includes:
[0065] The current device performs public key encryption on the original data, the current device identity, the current location information, and the current time information to obtain encrypted current data;
[0066] The current device sends the encrypted current data to the server.
[0067] Optionally, after generating the device identity, the current device does not directly send the original data and related information, but first performs public key encryption on the information, and then sends the encrypted data to the server, in order to enhance the security of the data transmission process and prevent the data from being intercepted or tampered with during the transmission process. Specifically, the current device uses the pre-acquired server public key to encrypt the original data, the current device identity, the current location information and the current time information. Public key encryption is an asymmetric encryption method, in which the public key is used to encrypt data and the private key is used to decrypt data. Since only the server holds the corresponding private key, only the server can decrypt and read the encrypted data. After the encryption process is completed, the current device sends the encrypted data to the server through a secure communication protocol. Since the data has been encrypted with the public key, even if it is intercepted during the transmission process, the attacker cannot easily decrypt and obtain the original data. Through public key encryption, the confidentiality and integrity of the data during the transmission process are ensured. Even if the data is intercepted during the transmission process, the attacker cannot decrypt and obtain the original data, thereby effectively preventing the risk of data leakage.
[0068] Optionally, before inputting the current device identity, the current location information, and the current time information into the data risk prediction model, the method further includes:
[0069] The current data is decrypted using the private key of the server to obtain the original data, the current device identity, the current location information and the current time information.
[0070] Optionally, after receiving the encrypted current data sent by the current device, the server does not directly input it into the data risk prediction model, but first uses the server's private key to decrypt the encrypted data to restore the original data, the current device identity, the current location information and the current time information, to ensure that before the data risk prediction, the server can accurately and securely obtain the complete data of the current device. The server first receives the encrypted current data sent by the current device through a secure communication protocol, and the server uses the pre-held private key to decrypt the received encrypted data. Since the encrypted data is encrypted using the server's public key, only the server can decrypt it using the corresponding private key. After the decryption process is completed, the server will restore the original data, the current device identity, the current location information and the current time information. This information is necessary for subsequent data risk prediction, and the restored information will be input into the data risk prediction model to assess the risk level of the current data. Through private key decryption processing, the server can securely obtain the complete data of the current device without worrying about the risk of data being intercepted or tampered with during transmission. This helps to protect the confidentiality and integrity of the data. Combining public key encryption and private key decryption processing, this technical solution significantly enhances the security of the entire data processing system. From data generation, transmission to processing, corresponding security measures are taken in each link to ensure the security and reliability of data.
[0071] In step 102, the current device identity, current location information and current time information received in step 101 are used as input features and passed to a pre-trained data risk prediction model. The data risk prediction model is constructed based on a machine learning or deep learning algorithm, and can automatically learn and identify data risk patterns under different input feature combinations. Through the data risk prediction model, the risk level of the current data can be evaluated in real time and dynamically, thereby improving the efficiency and accuracy of data processing. The data risk prediction model is trained based on historical sample data. Each sample data contains a sample device identity, sample location information, sample time information and a corresponding sample risk value. The sample risk value is determined by comprehensively evaluating the sample device identity weight, sample location weight and sample time weight. These weights reflect the importance of different features in data risk assessment. The sample data is trained using a machine learning algorithm (such as logistic regression, random forest, neural network, etc.) to obtain a data risk prediction model.
[0072] Optionally, before inputting the current device identity, the current location information, and the current time information into the data risk prediction model, the method further includes:
[0073] For any sample device, perform public key encryption on sample data, sample device identification, sample location information, sample time information, verification mark and parity check code to obtain encrypted sample data, wherein the verification mark is generated according to the parity check code;
[0074] The sample device sends the encrypted sample data to the server;
[0075] Decrypting the encrypted sample data using the private key of the server to obtain sample data, sample device identification, sample location information, sample time information, verification mark, and parity check code;
[0076] The parity check code is used to perform a parity check on the verification mark. If it is determined that the parity check fails, the check weight is determined to be 0; otherwise, the check weight is determined to be 1.
[0077] Optionally, when the current device submits data to the server, it attaches a unique identifier, location information, and time information, and generates a check code after combining them in order to obtain a new tag. After obtaining the tag, a code is generated through the parity check method, and then replaced with the corresponding string through the corresponding dictionary information, appended to the end of the tag, and the data after the splicing tags is encrypted through the server's public key, and then submitted to the server. After the server receives the data, it decrypts the submitted data through the server's private key. After decryption, obtain the spliced tags in the data, first perform a parity check, and obtain the corresponding string through the dictionary information after the check code generated after the check. Compare the check codes at the end of the tag to see if they are consistent. If they are inconsistent, the credibility of the identity of the data sender is extremely low, and the identity verification weight is 0.
[0078] Optionally, different from the input data when the model is applied, in an optional embodiment, the sample risk value is determined after the sample device identification weight, sample location weight, verification weight and sample time weight are determined based on the sample device identification, sample location information, sample time information, verification mark and parity check code, and the sample credibility is determined based on the sample device identification weight, sample location weight, verification weight and sample time weight. The sample device collects and prepares sample data, including the sample data itself, the sample device identification, sample location information, sample time information, etc., generates a parity check code, which is calculated based on a specific algorithm of the sample data (which may include other information, such as device identification, location information, etc.) and is used for subsequent verification of data integrity. A verification mark is generated based on the parity check code, which can be a direct application of the parity check code or the result of some transformation. The sample device uses the public key of the server to check the sample data, sample device identification, sample location information, sample time information, verification mark and parity check The sample data is encrypted with a code to obtain the encrypted sample data. The sample device sends the encrypted sample data to the server through a secure communication protocol. After receiving the encrypted sample data, the server uses its own private key to decrypt it and recover the original sample data, sample device identity, sample location information, sample time information, verification mark and parity check code. The server uses the recovered parity check code to perform a parity check on the verification mark. If the check passes (that is, the verification mark matches the parity check code), the check weight is determined to be 1, indicating that the sample data is complete and accurate; if the check fails (that is, the verification mark does not match the parity check code), the check weight is determined to be 0, indicating that the sample data may be erroneous or has been tampered with. By introducing the parity check code and the verification mark, the server can perform an integrity check on the received sample data to ensure that the data has not been tampered with during transmission. At the same time, in actual application scenarios, by introducing the check weight, the data risk prediction model can more accurately assess the risk level of the data, thereby making more reasonable decisions.
[0079] Optionally, after performing a parity check on the verification mark using the parity check code, the method further includes:
[0080] If it is determined that the sample device identity exists, the sample device identity weight is determined to be 1; otherwise, the sample device identity weight is determined to be 0;
[0081] When the sample time information is within the first preset time period, the sample time weight is determined to be 0.4; when the sample time information exceeds the first preset time length and the second preset time length, the sample time weight is determined to be 0.2; when the sample time information exceeds the second preset time length, the sample time weight is determined to be 0, and the first preset time length is less than the second preset time length;
[0082] When the sample position information is within the first range, the sample position weight is determined to be 0.6; when the sample position information is greater than the first range and less than the second range, the sample position weight is determined to be 0.4; when the sample position information is greater than the second range and less than the third range, the sample position weight is determined to be 0.2; when the sample position information is greater than the third range, the sample position weight is determined to be 0; the first range is smaller than the second range, and the second range is smaller than the third range.
[0083] Optionally, after the parity check is completed, obtain the unique identifier, longitude and latitude, and time information in the tag. Check whether there is a unique identifier for the device in the system. If it exists, it is 1. If not, the identity weight is 0. After the time information is obtained, compare the valid time range of the corresponding device in the system. If there is no separate time range, obtain the public valid time range. Within the valid time range, the weight is 0.4; within 1 hour above or below the valid time, the weight is 0.2; if it exceeds one hour, the weight is 0. The longitude and latitude information is compared with the longitude and latitude recorded in the system. If the difference does not exceed 3 meters, the value is 0.6; if it is greater than 3 meters and less than or equal to 5 meters, the value is 0.4; if it is greater than 5 meters and less than or equal to 10 meters, the value is 0.2; if it is greater than 10 meters, the value is 0; the larger the difference, the lower the weight value.
[0084] Optionally, the present invention further introduces the concepts of sample device identification weight, sample time weight and sample location weight, aiming to provide a more sophisticated and accurate weight allocation mechanism for the data risk prediction model by comprehensively considering information in multiple dimensions (device identification, time, location). This weight allocation mechanism helps the model to more accurately evaluate the importance and reliability of sample data, thereby improving the accuracy of the prediction. If there is a valid sample device identity, the sample device identification weight is determined to be 1, indicating that the sample data comes from a known and trusted device. If there is no valid sample device identity, the sample device identification weight is determined to be 0, indicating that the source of the sample data is unknown or unreliable.
[0085] Optionally, based on the sample time information, it is compared with a preset time period. If the sample time information is within a first preset time period (for example, within the last hour), the sample time weight is determined to be 0.4, indicating that the sample data is newer and has a higher timeliness. If the sample time information exceeds the first preset time length but is within a second preset time length (for example, within the last day), the sample time weight is determined to be 0.2, indicating that although the sample data is slightly older, it still has a certain reference value. If the sample time information exceeds the second preset time length (for example, more than one day), the sample time weight is determined to be 0, indicating that the sample data is too old and may no longer have a reference value.
[0086] Optionally, based on the sample location information, it is compared with a preset range. If the sample location information is within a first range (for example, inside a specific area), the sample location weight is determined to be 0.6, indicating that the sample data comes from a key or important geographical location. If the sample location information is larger than the first range but smaller than the second range (for example, the surrounding area of a specific area), the sample location weight is determined to be 0.4, indicating that although the sample data is not in the core area, it still has a certain geographical relevance. If the sample location information is larger than the second range but smaller than the third range (for example, a more distant surrounding area or adjacent area), the sample location weight is determined to be 0.2, indicating that the geographical relevance of the sample data is weak. If the sample location information is larger than the third range, the sample location weight is determined to be 0, indicating that the sample data has no direct association with the core geographical area.
[0087] By comprehensively considering the sample device identification, time and location information, a more comprehensive and accurate weight allocation mechanism is provided for the data risk prediction model, which helps to improve the accuracy of the prediction. The weight allocation mechanism takes into account information in multiple dimensions, enabling the model to make more reasonable judgments when faced with sample data from different sources, at different times and at different locations, thereby enhancing the robustness of the model.
[0088] Optionally, after determining the sample device identification weight, the sample location weight, and the sample time weight, the sample risk value is determined according to the sample device identification weight, the sample location weight, and the sample time weight;
[0089] Determining the sample risk value according to the sample device identification weight, the sample location weight, and the sample time weight includes:
[0090] Y=(1-v×d×(t+p))×100%
[0091] Among them, Y is the sample risk value, v is the verification weight, d is the sample device identification weight, t is the sample time weight, and p is the sample location weight.
[0092] Optionally, the data is verified based on the set threshold and device information to determine the validity of the identity of the party sending the data, and the risk coefficient of this operation is generated. By calculating the sample risk value, a quantitative indicator is provided for evaluating the potential risk of the sample data, making the risk assessment more objective and accurate. It comprehensively considers multiple dimensions of sample data (integrity, source reliability, timeliness and geographical relevance) to make the risk assessment more comprehensive and detailed. The present invention makes the calculation of the sample risk value more accurate by "determining the sample risk value according to the sample device identification weight, the sample location weight and the sample time weight", and the sample risk value will be used for subsequent model training, thereby improving the prediction accuracy of the model.
[0093] Optionally, the present invention trains a risk assessment model through machine learning while processing sample data, and optimizes the model according to each processing result. The corresponding data is obtained according to the key information such as the identity, longitude and latitude and time point when the data is uploaded, and analyzed by the random forest algorithm, and then the credibility is predicted by linear regression, and the feasibility is compared with the algorithm processing result. The evaluation model uses the above steps and algorithms to obtain the credibility of the data more quickly through the model according to the current time period, reduce the time consumption caused by verification and manual review, and improve the timeliness of the data; the data verified by the model is added to the data adjusted by the model algorithm, and the verification model is optimized so that the algorithm can dynamically adjust the threshold and weight settings in the model, calculate the success rate of the model verification comparison, and the success rate is based on the configuration information (number of successes and percentages). The verification mode is switched to the evaluation model, which effectively improves the verification speed and reduces the frequency of manual review processing.
[0094] In step 103, the data risk value output by the data risk prediction model is used to measure the potential risk level of the current data. The data risk value is classified according to a preset risk threshold or strategy. Different processing methods are adopted for different risk classifications, such as direct storage, encrypted storage, isolated review, discarding, etc.
[0095] Optionally, determining a processing method of the original data according to the data risk value includes:
[0096] When the data risk value is greater than or equal to 60%, a first processing instruction is generated, wherein the first processing instruction is used to instruct recording the original data, terminating the business operation, and sending a warning email to the system administrator;
[0097] When the data risk value is greater than or equal to 30% and less than 60%, a second processing instruction is generated, wherein the second processing instruction is used to instruct to execute a business operation of the original data and mark the original data so as to remove the mark after manual review;
[0098] When the data risk value is less than 30% and greater than 0, a third processing instruction is generated, wherein the third processing instruction is used to instruct execution of a business operation of the original data and sending an in-site reminder message and a reminder email;
[0099] When the data risk value is 0, a fourth processing instruction is generated, where the fourth processing instruction is used to instruct execution of a business operation on the original data.
[0100] Optionally, the present invention can set four data risk value thresholds: 0%, 30%, 60% and 100%. According to the size of the data risk value, it is divided into four intervals: 0%-30%, 30%-60%, 60%-100% and 0%. By setting different data risk value thresholds and processing methods, it can ensure that high-risk data is processed and investigated in a timely manner, and avoid the further dissemination and use of erroneous data, thereby improving the accuracy of data processing. For high-risk data, through measures such as recording, terminating operations and sending warning emails, potential security threats can be discovered and responded to in a timely manner, thereby enhancing the security of the system; for low-risk data, business operations can be directly executed or only reminder information can be sent to avoid unnecessary review and waiting time, thereby improving data processing efficiency. Through reasonable processing method settings, while ensuring data security and accuracy, the impact on user business operations can be minimized to enhance user experience.
[0101] Optionally, the present invention collects set operating environment information, such as the operating system, the underlying software version that the program depends on, the longitude and latitude of the device, the operation time, and the core hardware identification of the device. The operating system, the underlying software version that the program depends on, the core hardware identification of the device, and other information that is not easy to be modified are generated into a unique tag through the HASH algorithm, and then the tag, the operation time, and the longitude and latitude are combined and a check code is added to generate a new tag information, which is encrypted by the public key of the server to generate a final tag and transmitted to the server; after the server receives the tag attached to the data, it decrypts and obtains the complete data. After verifying the data, the unique tag, operation time, longitude and latitude and other information are obtained, and these information are compared and judged with the relevant information recorded in the system to obtain the reliability score of the current operation. For operations below the threshold, the system can process them according to the settings, such as rejecting the operation, ignoring the operation, or notifying the administrator; while processing the data, the processing results are compared with the model processing results to train and tune the evaluation model. When the number and proportion of the validity of the evaluation model processing results reach the set threshold, the data verification is frequently switched to the model evaluation process.
[0102] The present invention can perform verification in advance instead of making judgments during query processing, so that operations with excessive risks may be avoided, the storage of invalid data can be reduced, and storage pressure can be reduced; the validity of data can be verified to avoid contamination of system data by invalid or false data, and query processing time can be reduced, and verification of data can be avoided every time data is acquired, thereby reducing repeated processing of data; the verification evaluation speed can be improved through an algorithm model, the number of manual processing times can be reduced, and the verification cost can be reduced.
[0103] The present invention provides comprehensive input for data risk prediction by receiving current data carrying rich information (including device identity, location, and time), and uses a data risk prediction model to conduct risk assessment on data, thereby improving the accuracy and efficiency of data processing, and flexibly determining the processing method of the original data according to the data risk value, thereby achieving effective management and control of the data;
[0104] The present invention also recombines device information through preset rules, increases the confidentiality and cracking difficulty of device identity, generates device identity through HASH algorithm, ensures the uniqueness and security of device identity, ensures the security of data during transmission through public key encryption, prevents data from being stolen or tampered with, and the server decrypts data through private key to ensure the accuracy and reliability of the decryption process. By comprehensively considering device identity, time and location information, multi-dimensional weight information is provided for data risk prediction. By combining multiple weight information to determine the sample risk value, a comprehensive evaluation index is provided for data risk prediction.
[0105] Figure 2 It is a structural diagram of the data processing device provided by the present invention, the data processing device includes a receiving unit 1, the receiving unit is used to receive current data sent from the current device, the current data carries original data, current device identity, current location information and current time information, the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment, the working principle of the receiving unit 1 can refer to the aforementioned step 101, which will not be repeated here.
[0106] The data processing device also includes an input unit 2, which is used to input the current device identity, the current location information and the current time information into the data risk prediction model to obtain the data risk value output by the data risk prediction model. The working principle of the input unit 2 can refer to the aforementioned step 102 and will not be repeated here.
[0107] The data processing device further includes a determination unit 3, which is used to determine a processing method for the original data according to the data risk value. The working principle of the determination unit 3 can be referred to the aforementioned step 103 and will not be described in detail here.
[0108] The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data;
[0109] The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
[0110] The present invention provides comprehensive input for data risk prediction by receiving current data carrying rich information (including device identity, location, and time), and uses a data risk prediction model to conduct risk assessment on data, thereby improving the accuracy and efficiency of data processing, and flexibly determining the processing method of the original data according to the data risk value, thereby achieving effective management and control of the data;
[0111] The present invention also recombines device information through preset rules, increases the confidentiality and cracking difficulty of device identity, generates device identity through HASH algorithm, ensures the uniqueness and security of device identity, ensures the security of data during transmission through public key encryption, prevents data from being stolen or tampered with, and the server decrypts data through private key to ensure the accuracy and reliability of the decryption process. By comprehensively considering device identity, time and location information, multi-dimensional weight information is provided for data risk prediction. By combining multiple weight information to determine the sample risk value, a comprehensive evaluation index is provided for data risk prediction.
[0112] Figure 3 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 3As shown, the electronic device may include: a processor (processor) 310, a communication interface (Communications Interface) 320, a memory (memory) 330 and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logic instructions in the memory 330 to execute the data processing method, which includes: receiving current data sent from the current device, the current data carrying original data, current device identity, current location information and current time information, and the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment; inputting the current device identity, the current location information and the current time information into the data risk prediction model to obtain the data risk value output by the data risk prediction model; determining the processing method of the original data according to the data risk value; the data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data; the sample risk value is determined after determining the sample device identity weight, sample location weight and sample time weight according to the sample device identity, sample location information and sample time information, and determining the sample credibility according to the sample device identity weight, sample location weight and sample time weight.
[0113] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0114] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a data processing method provided by the above methods, which method includes: receiving current data sent from a current device, the current data carrying original data, a current device identity, current location information and current time information, and the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment; inputting the current device identity, the current location information and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model; determining a processing method for the original data according to the data risk value; the data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data; the sample risk value is determined after determining the sample device identity weight, sample location weight and sample time weight according to the sample device identity, sample location information and sample time information, and determining the sample credibility according to the sample device identity weight, sample location weight and sample time weight.
[0115] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the data processing method provided by the above-mentioned methods, the method comprising: receiving current data sent from a current device, the current data carrying original data, a current device identity, current location information and current time information, the current device identity being generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment; inputting the current device identity, the current location information and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model; determining a processing method for the original data according to the data risk value; the data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data; the sample risk value is determined after determining the sample device identity weight, sample location weight and sample time weight according to the sample device identity, sample location information and sample time information, and then determining the sample credibility according to the sample device identity weight, sample location weight and sample time weight.
[0116] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0117] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that: include: Receive current data sent from the current device, the current data carrying original data, current device identity, current location information and current time information, the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment; Inputting the current device identity, the current location information, and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model; Determining a processing method for the original data according to the data risk value; The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data; The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
2. The data processing method according to claim 1, characterized in that: Before receiving current data sent from the current device, the method includes: The current device recombines the system version number, the underlying software program number, and the core hardware unique identifier corresponding to the current operating environment according to a preset rule to determine recombined data, wherein the recombined data includes, arranged from left to right, the first 3 digits of the system version number, the first 3 digits of the underlying software program number, the first 2 digits of the core hardware unique identifier, the length digit of the core hardware unique identifier, the remaining string of the core hardware unique identifier, the length digit of the underlying software program number, the remaining digits of the underlying software program number, the length digit of the system version number, and the remaining digits of the system version number; The current device processes the reorganized data using a HASH algorithm to generate the current device identity.
3. The data processing method according to claim 2, characterized in that: After generating the current device identity, the method further includes: The current device performs public key encryption on the original data, the current device identity, the current location information, and the current time information to obtain encrypted current data; The current device sends the encrypted current data to the server.
4. The data processing method according to claim 3, characterized in that: Before inputting the current device identity, the current location information, and the current time information into the data risk prediction model, the method further includes: The current data is decrypted using the private key of the server to obtain the original data, the current device identity, the current location information and the current time information.
5. The data processing method according to claim 1, characterized in that: Before inputting the current device identity, the current location information, and the current time information into the data risk prediction model, the method further includes: For any sample device, perform public key encryption on the sample data, the sample device identification, the sample location information, the sample time information, the verification mark and the parity check code to obtain the encrypted sample data, wherein the verification mark is generated according to the parity check code; The sample device sends the encrypted sample data to the server; Decrypting the encrypted sample data using the private key of the server to obtain sample data, sample device identification, sample location information, sample time information, verification mark, and parity check code; The parity check code is used to perform a parity check on the verification mark. If it is determined that the parity check fails, the check weight is determined to be 0; otherwise, the check weight is determined to be 1.
6. The data processing method according to claim 5, characterized in that: After performing a parity check on the verification mark using the parity check code, the method further includes: If it is determined that the sample device identity exists, the sample device identity weight is determined to be 1; otherwise, the sample device identity weight is determined to be 0; When the sample time information is within the first preset time period, the sample time weight is determined to be 0.4; when the sample time information exceeds the first preset time length and the second preset time length, the sample time weight is determined to be 0.2; when the sample time information exceeds the second preset time length, the sample time weight is determined to be 0, and the first preset time length is less than the second preset time length; When the sample position information is within the first range, the sample position weight is determined to be 0.6; when the sample position information is greater than the first range and less than the second range, the sample position weight is determined to be 0.4; when the sample position information is greater than the second range and less than the third range, the sample position weight is determined to be 0.2; when the sample position information is greater than the third range, the sample position weight is determined to be 0; the first range is smaller than the second range, and the second range is smaller than the third range.
7. The data processing method according to claim 6, characterized in that: After determining the sample device identification weight, the sample location weight, and the sample time weight, determine the sample risk value according to the sample device identification weight, the sample location weight, and the sample time weight; Determining the sample risk value according to the sample device identification weight, the sample location weight, and the sample time weight includes: Y=(1-v×d×(t+p))×100% Among them, Y is the sample risk value, v is the verification weight, d is the sample device identification weight, t is the sample time weight, and p is the sample location weight.
8. The data processing method according to claim 1, characterized in that: The determining of a processing method of the original data according to the data risk value includes: When the data risk value is greater than or equal to 60%, a first processing instruction is generated, wherein the first processing instruction is used to instruct recording the original data, terminating the business operation, and sending a warning email to the system administrator; When the data risk value is greater than or equal to 30% and less than 60%, a second processing instruction is generated, wherein the second processing instruction is used to instruct to execute a business operation of the original data and mark the original data so as to remove the mark after manual review; When the data risk value is less than 30% and greater than 0, a third processing instruction is generated, wherein the third processing instruction is used to instruct execution of a business operation of the original data and sending an in-site reminder message and a reminder email; When the data risk value is 0, a fourth processing instruction is generated, where the fourth processing instruction is used to instruct execution of a business operation on the original data.
9. A data processing device, characterized in that: include: A receiving unit, the receiving unit is used to receive current data sent from the current device, the current data carries original data, current device identity, current location information and current time information, the current device identity is generated according to the operating system version number, the underlying software program number and the core hardware unique identifier corresponding to the current operating environment; An input unit, the input unit is used to input the current device identity, the current location information and the current time information into a data risk prediction model to obtain a data risk value output by the data risk prediction model; A determination unit, the determination unit is used to determine a processing method of the original data according to the data risk value; The data risk prediction model is determined after training based on the sample device identity, sample location information and sample time information corresponding to all sample data, and the sample risk value corresponding to each sample data; The sample risk value is determined after determining the sample device identification weight, sample location weight and sample time weight based on the sample device identity, sample location information and sample time information, and determining the sample credibility based on the sample device identification weight, sample location weight and sample time weight.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the data processing method according to any one of claims 1 to 8 is implemented.