Data identification method and device, electronic equipment and storage medium

By extracting features related to encryption protocols from the data to be identified and using benchmark features to perform similarity matching with the features to be identified, the problem of encryption protocol data being tampered with in network communication is solved, thus improving the security of network communication.

CN115879166BActive Publication Date: 2026-08-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2022-08-31
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify and prevent the tampering of encryption protocol-related data during network communications, thus affecting network security.

Method used

By extracting features related to the encryption protocol from the data to be identified, and using benchmark features to perform similarity matching with the features to be identified, the identification result of the data to be identified is determined, and it is determined whether the data has been tampered with.

Benefits of technology

It improves the security of network communication, enabling rapid and effective identification and prevention of tampering with encrypted protocol data, thus enhancing the security of network communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115879166B_ABST
    Figure CN115879166B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data identification method and device, electronic equipment, storage medium and program product, relates to the technical field of computers, in particular to the technical field of network security. The specific implementation scheme is: determining a to-be-identified feature from to-be-identified data, the to-be-identified feature including a feature related to an encryption protocol; and determining an identification result of the to-be-identified data based on the to-be-identified feature and a reference feature, the identification result being used to represent whether the to-be-identified data is tampered with, and the reference feature being a feature related to the encryption protocol and extracted from first historical data with an identification result. The present disclosure uses the feature related to the encryption protocol to judge the to-be-identified data, which can improve the security of network communication, and effectively counter the behavior of maliciously tampering with data related to the encryption protocol.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more particularly to the field of network security technology, specifically to data identification methods, devices, electronic devices, storage media, and program products. Background Technology

[0002] With the development of information technology, the openness and interconnectivity of internet information have become increasingly prominent, making cybersecurity particularly important. In terms of scope, cybersecurity can include the security of hardware architecture and operating procedures. In terms of characteristics, cybersecurity needs to ensure the confidentiality, integrity, availability, and resistance to attacks of information. How to improve cybersecurity has become a key research focus. Summary of the Invention

[0003] This disclosure provides a data identification method, apparatus, electronic device, storage medium, and program product.

[0004] According to one aspect of this disclosure, a data identification method is provided, comprising: determining a feature to be identified from data to be identified, wherein the feature to be identified includes a feature related to an encryption protocol; and determining an identification result of the data to be identified based on the feature to be identified and a reference feature, wherein the identification result is used to characterize whether the data to be identified has been tampered with, and the reference feature is a feature extracted from first historical data with known identification results and related to the encryption protocol.

[0005] According to another aspect of this disclosure, a data identification device is provided, comprising: a first determining module for determining a feature to be identified from data to be identified, wherein the feature to be identified includes a feature related to an encryption protocol; and a second determining module for determining an identification result of the data to be identified based on the feature to be identified and a reference feature, wherein the identification result is used to characterize whether the data to be identified has been tampered with, and the reference feature is a feature extracted from first historical data with known identification results and related to the encryption protocol.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as disclosed herein.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods as disclosed herein.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as disclosed herein.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This illustration schematically shows an exemplary system architecture to which data identification methods and apparatus can be applied according to embodiments of the present disclosure;

[0012] Figure 2 A flowchart illustrating a data identification method according to an embodiment of the present disclosure is shown schematically.

[0013] Figure 3 A signaling diagram of a data identification method according to an embodiment of the present disclosure is illustrated schematically;

[0014] Figure 4 The schematic diagram illustrates a flowchart of determining the identification result according to an embodiment of the present disclosure;

[0015] Figure 5 A flowchart illustrating a data identification method according to another embodiment of the present disclosure is shown schematically;

[0016] Figure 6 A block diagram of a data identification device according to an embodiment of the present disclosure is schematically shown; and

[0017] Figure 7 A block diagram of an electronic device suitable for implementing a data identification method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0019] This disclosure provides a data identification method, apparatus, electronic device, storage medium, and program product.

[0020] According to embodiments of this disclosure, a data identification method is provided, comprising: determining a feature to be identified from data to be identified, wherein the feature to be identified includes features related to an encryption protocol; and determining an identification result of the data to be identified based on the feature to be identified and a reference feature, wherein the identification result is used to characterize whether the data to be identified has been tampered with, and the reference feature is a feature extracted from first historical data with known identification results and related to an encryption protocol.

[0021] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0022] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0023] Figure 1 The illustration schematically depicts an exemplary system architecture to which data identification methods and apparatus can be applied according to embodiments of the present disclosure.

[0024] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the data identification method and apparatus can be applied may include a terminal device, but the terminal device can implement the data identification method and apparatus provided by the embodiments of this disclosure without interacting with the server.

[0025] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a terminal device 101, an application server 102, and a detection server 103. Networks may be provided between the terminal device 101 and the application server 102, and between the application server 102 and the detection server 103. The network, as the medium for communication links, may include various connection types, such as wired and / or wireless communication links, etc.

[0026] Users can use terminal device 101 to interact with application server 102 via a network to receive or send messages, etc. Messages sent by terminal device 101, such as network requests, can be requests to establish an encrypted communication connection. This allows application server 102 to respond to the network request and establish an encrypted communication connection with terminal device 101. Various communication client applications can be installed on terminal device 101, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software (for example only).

[0027] Terminal device 101 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0028] Application server 102 may be a server that provides various services, such as a back-end management server that supports the content browsed by the user using terminal device 101 (for example only).

[0029] The detection server 103 can be a server providing network security detection. The detection server 103 can receive network requests sent by the application server 102 and determine whether the network request is abnormal, such as whether the data to be identified in the network request has been tampered with. If it is determined that the data to be identified has been tampered with, the network request is determined to be abnormal; if it is determined that the data to be identified has not been tampered with, the network request is determined to be normal. The detection server 103 can send the identification results to the application server 102 so that the application server 102 can provide feedback to the terminal device 101 based on the identification results. For example, if the identification result indicates that the data to be identified is normal data that has not been tampered with, the application server 102 can establish an encrypted communication connection with the terminal device 101.

[0030] It should be noted that the data identification method provided in this embodiment can generally also be executed by the detection server 103. Correspondingly, the data identification device provided in this embodiment can generally be located in the server 103. The data identification method provided in this embodiment can also be executed by a server or server cluster that is different from the server 103 but capable of communicating with the server 103. Correspondingly, the data identification device provided in this embodiment can also be located in a server or server cluster that is different from the server 103 but capable of communicating with the server 103.

[0031] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0032] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0033] Figure 2 A flowchart illustrating a data identification method according to an embodiment of the present disclosure is shown schematically.

[0034] like Figure 2 As shown, the method includes operations S210 to S220.

[0035] In operation S210, the features to be identified are determined from the data to be identified. The features to be identified include features related to the encryption protocol.

[0036] In operation S220, based on the features to be identified and the baseline features, the identification result of the data to be identified is determined. The identification result is used to characterize whether the data to be identified has been tampered with. The baseline features are features extracted from the first historical data with known identification results and are related to the encryption protocol.

[0037] According to embodiments of this disclosure, the data identification method can be applied to scenarios where a terminal device and an application server establish a secure communication connection over a network. The data to be identified may include data related to an encryption protocol, which is a protocol used to provide security support for network communication. However, it is not limited to this; the data to be identified may also include authentication data and other data used to provide security support for network communication.

[0038] According to embodiments of this disclosure, features to be identified can be determined from data to be identified. These features include characteristics related to encryption protocols. The features to be identified can be used to reflect the characteristics of the data to be identified.

[0039] According to embodiments of this disclosure, first historical data can be collected. The first historical data is historical data with known identification results. The identification results of the first historical data can be used to characterize the first historical data as tampered abnormal data, but are not limited thereto; the identification results of the first historical data can also be used to characterize the first historical data as untampered normal data.

[0040] According to embodiments of this disclosure, features related to the encryption protocol are extracted from first historical data as baseline features. The baseline features are features extracted from the first historical data and are used to reflect the characteristics of the first historical data. The relationship between the baseline features and the features to be identified can be determined by comparing them, thereby determining the relationship between the data to be identified and the first historical data. The amount of first historical data is not limited, but the larger the amount of data, the more representative and accurate the extracted baseline features will be.

[0041] According to embodiments of this disclosure, determining the recognition result of the data to be recognized based on the feature to be recognized and a reference feature may include: performing similarity matching on the reference feature and the feature to be recognized to obtain a similarity matching result. If the similarity matching result is greater than a predetermined similarity threshold, then the recognition result of the data to be recognized is determined to be the same as the recognition result of the first historical data. If the similarity matching result is less than or equal to the predetermined similarity threshold, then the recognition result of the data to be recognized is determined to be different from the recognition result of the first historical data. The method for determining the recognition result of the data to be recognized based on the feature to be recognized and the reference feature is not limited to this; any method that can determine whether the recognition result of the data to be recognized is consistent with the recognition result of the first historical data based on the feature to be recognized and the reference feature is acceptable.

[0042] According to embodiments of this disclosure, by using a reference feature as a comparison standard for the feature to be identified, the identification result of the data to be identified can be determined simply and effectively. At the same time, using features related to encryption protocols to evaluate the data to be identified can improve the security of network communication, thereby effectively combating malicious tampering with data related to encryption protocols.

[0043] According to embodiments of this disclosure, when performing such Figure 2 Before determining the features to be identified from the data to be identified, the data identification method may also include the following operations in operation S210 shown.

[0044] For example, receiving a network request. Preprocessing the initial data to be identified to obtain the data to be identified.

[0045] According to embodiments of this disclosure, the network request includes initial data to be identified, which is data related to an encryption protocol.

[0046] According to embodiments of this disclosure, the encryption protocol can be a network security protocol used to provide secure data transmission between a terminal device and an application server. For example, the encryption protocol includes one or more of TLS (Transport Layer Security) and SSL (Secure Sockets Layer) encryption protocols.

[0047] According to embodiments of this disclosure, a network request can be a request sent by the terminal device to the application server during a negotiation session between the terminal device and the application server to establish an encrypted communication connection based on an encryption protocol. For example, during the TLS handshake phase between the terminal device and the application server, a network request sent by the terminal device to the application server to complete the encrypted communication connection may include a Client Hello request packet. This request packet can serve as initial identification data.

[0048] According to optional embodiments of this disclosure, data related to the encryption protocol may include, but is not limited to, fingerprint data within the encryption protocol. Data related to the encryption protocol may also include data related to the device environment. Furthermore, data related to the encryption protocol may include business-related data.

[0049] According to embodiments of this disclosure, the fingerprint data in the encryption protocol may include one or more of the following: TLS Version protocol version, Ciphers encryption suite, Extensions extension list, Elliptic Curves elliptic algorithm field, and EllipticCurve Point Formats elliptic standard.

[0050] Taking the JA3 TLS fingerprint data of a terminal device as an example, the fingerprint data in the encryption protocol includes the TLS Version protocol version, Ciphers encryption suite, Extensions extension list, Elliptic Curves elliptic algorithm fields, and Elliptic Curve Point Formats elliptic standard.

[0051] According to embodiments of this disclosure, data related to the device environment may include at least one of the following: the IP address of the device terminal, the IP address of the application server, user identification data, device identification data of the device terminal, and device identification data of the application server.

[0052] According to embodiments of this disclosure, business-related data may include at least one of the following: verification code, historical records of whether the verification code has been approved, and verification code input duration.

[0053] According to embodiments of this disclosure, preprocessing the initial data to be identified may include: cleaning the initial data to be identified, such as removing dirty data that is different in format from the initial data to be identified or removing empty data.

[0054] According to embodiments of this disclosure, preprocessing the initial data to be identified to obtain the data to be identified can remove noisy data from the initial data to be identified, simplify the processing of the data to be identified, improve processing efficiency, and improve the accuracy of determining the identification result.

[0055] Figure 3 A signaling diagram of a data identification method according to an embodiment of the present disclosure is illustrated schematically.

[0056] like Figure 3 As shown, the data recognition method may include operations S310 to S360.

[0057] When operating S310, the terminal device sends a network request to the application server.

[0058] When operating the S320, the application server sends network requests to the detection server.

[0059] When operating S330, the detection server performs a data recognition method on network requests to determine the recognition result of the data to be recognized in the network request.

[0060] When operating the S340, the detection server sends the identification results to the application server.

[0061] In operation S350, determine whether the identification result indicates that the data to be identified has been tampered with. If it is determined that the identification result indicates that the data to be identified has been tampered with, stop the operation. If it is determined that the identification result indicates that the data to be identified has not been tampered with, execute operation S360.

[0062] When operating the S360, a feedback message is sent to the terminal device to establish an encrypted communication connection between the terminal device and the application server.

[0063] According to the embodiments of this disclosure, it should be noted that a detection server can be set up to execute the data recognition method, but it is not limited to this. A detection module can also be set up on the application server to execute the data recognition method.

[0064] According to embodiments of this disclosure, applying the data identification method to the handshake phase between a terminal device and an application server can improve the security of encrypted communication connections. It can effectively and quickly identify fingerprint data in encryption protocols from the data to be identified, thereby improving the security of network communication.

[0065] According to embodiments of this disclosure, the feature to be identified includes multiple sub-features to be identified. A reference feature may include multiple reference sub-features. The number of sub-features to be identified can be the same as the number of reference sub-features, or the number of reference sub-features can be greater than the number of sub-features to be identified.

[0066] According to embodiments of this disclosure, when the feature to be identified includes multiple sub-features to be identified, for example... Figure 2 The operation S220 shown, which determines the recognition result of the data to be recognized based on the features to be recognized and the reference features, may include the following operations.

[0067] For example, for each of multiple sub-features to be identified, a target reference sub-feature is determined from multiple reference sub-features that matches the feature type of the sub-feature to be identified. Based on the sub-feature to be identified and the target reference sub-feature, the identification sub-result of the data to be identified is determined. Based on the multiple identification sub-results, the identification result of the data to be identified is determined. Each identification sub-result corresponds one-to-one with one of the multiple sub-features to be identified.

[0068] According to embodiments of this disclosure, for each sub-feature to be identified, a target reference sub-feature matching the feature type of the sub-feature to be identified can be determined from multiple reference sub-features. Determining the identification sub-result of the data to be identified based on the sub-feature to be identified and the target reference sub-feature makes the identification sub-result accurate and effective. The more sub-features to be identified among the features to be identified, the more accurate the final identification result of the data to be identified. The more different feature types among the multiple sub-features to be identified, the wider the feature range involved, and the more factors considered, the more accurate the identification result of the data to be identified.

[0069] According to embodiments of this disclosure, determining the identification result of the data to be identified based on multiple identification sub-results may include: weighted summation of multiple identification sub-results to obtain the identification result of the data to be identified. For example, the multiple identification sub-results include identification sub-result A, identification sub-result B, and identification sub-result C. Weights are pre-configured for each identification sub-result, such as weight Wa for identification sub-result A, weight Wb for identification sub-result B, and weight Wc for identification sub-result C. The identification result = identification sub-result A * weight Wa + identification sub-result B * weight Wb + identification sub-result C * weight Wc.

[0070] According to embodiments of this disclosure, a confidence level can be used to characterize the identification sub-result, and the confidence level can be a numerical value. For example, a value between 0 and 1 or between 0 and 100 can be used as the confidence level. It can be predetermined that the higher the confidence level, the greater the likelihood that the identification sub-result is used to characterize the data to be identified as having been tampered with.

[0071] According to embodiments of this disclosure, using confidence level to characterize the identification sub-results of the sub-features to be identified can quantify the identification sub-results, thereby facilitating the determination of the identification result of the data to be identified based on multiple identification sub-results.

[0072] According to other embodiments of this disclosure, determining the identification result of the data to be identified based on multiple identification sub-results may further include: sorting the multiple identification sub-results in descending order to obtain a sorting result. Based on the sorting result, determining the target identification sub-result from the multiple identification sub-results, and using the target identification sub-result as the identification result of the data to be identified. The target identification sub-result can be the identification sub-result with the highest value, but is not limited to this; the target identification sub-result can also be the identification sub-result with the lowest value. It can be determined according to the actual situation.

[0073] Figure 4 The schematic diagram illustrates a flowchart of determining the identification result according to an embodiment of the present disclosure.

[0074] like Figure 4 As shown, taking an example where there are two sub-features to be identified. Multiple sub-features to be identified are determined from the data to be identified 410, such as a first sub-feature 420 and a second sub-feature 430. For each of the multiple sub-features to be identified, a target benchmark sub-feature matching the feature type of the sub-feature to be identified is determined from a benchmark feature database storing multiple benchmark sub-features. The benchmark features in the benchmark feature database can be features extracted from first historical data and pre-stored in the benchmark feature database. A first target benchmark sub-feature 440 matching the first sub-feature to be identified 420 and a second target benchmark sub-feature 450 matching the second sub-feature to be identified 430 can be determined from the benchmark feature database. Based on the first sub-feature to be identified 420 and the first target benchmark sub-feature 440, a first identification result 460 of the data to be identified 410 is determined. Based on the second sub-feature to be identified 430 and the second target benchmark sub-feature 450, a second identification result 470 of the data to be identified 410 is determined. Based on the first identification sub-result 460 and the second identification sub-result 470, the identification result 480 of the data to be identified 410 is determined.

[0075] According to embodiments of this disclosure, the sub-features to be identified can be divided into source data and summary data based on the generation category of the sub-features. For example, source data is a portion of data related to encryption protocols, and summary data is common data determined from the data related to encryption protocols.

[0076] According to embodiments of this disclosure, the sub-features to be identified can be divided into features related to fingerprint data in the encryption protocol, features related to device environment information, and features related to business information based on the feature type of the sub-features to be identified.

[0077] According to embodiments of this disclosure, the sub-feature to be identified may include features related to fingerprint data in the encryption protocol. For example, the sub-feature to be identified may include at least one of the following: the protocol version of the encryption protocol, the target field sequence of the encryption protocol, and the string length of the encryption protocol.

[0078] According to embodiments of this disclosure, the target field sequence of the encryption protocol includes a string sequence of the cipher suite field and / or a string sequence of the extended list field. The string length of the encryption protocol may include, but is not limited to, the string lengths of the protocol version field, the cipher suite field, and the extended list field; it may also refer to the string length of all original fields in the fingerprint data. For example, the string lengths of the protocol version field, the cipher suite field, the elliptic algorithm field, the elliptic standard field, and the extended list field.

[0079] According to embodiments of this disclosure, determining the sub-features to be identified from the fingerprint data in the encryption protocol enables the identification of whether the fingerprint data of the encryption protocol in the data to be identified has been attacked or tampered with. Furthermore, by determining multiple sub-features to be identified based on multiple fields in the fingerprint data, it is possible to utilize multiple sub-features of different categories to identify whether the fingerprint data has been tampered with or attacked from different degrees and perspectives.

[0080] According to embodiments of this disclosure, the sub-features to be identified related to device environment information may include at least one of the following: the IP address of the device terminal, the IP address of the application server, user identity information, device identification information of the device terminal, and device identification information of the application server.

[0081] According to embodiments of this disclosure, the sub-features to be identified related to business information may include at least one of the following: verification code, historical records of whether the verification code has been passed, and verification code input duration.

[0082] According to embodiments of this disclosure, in addition to features related to fingerprint data in the encryption protocol and features related to address information in the encryption protocol, the sub-features to be identified may also include features related to business information or features related to device environment information. The more types of sub-features to be identified, the more accurate the determined identification sub-result, and the more valuable the identification sub-result.

[0083] According to embodiments of this disclosure, the multiple sub-features to be identified each have different feature types, and the methods for determining the identification sub-results of the data to be identified based on the sub-features to be identified and the target reference sub-features are also different. Target recognition rules can be pre-configured for the sub-features to be identified based on their feature types, so that the determination of the identification sub-results is accurate and effective.

[0084] According to embodiments of this disclosure, determining the identification sub-result of the data to be identified based on the sub-feature to be identified and the target reference sub-feature may include the following operations.

[0085] For example, a target recognition rule is determined from a set of predefined recognition rules that matches the feature type of the sub-feature to be recognized. Based on the sub-feature to be recognized and the target reference sub-feature, it is determined whether the sub-feature to be recognized satisfies the target recognition rule. If the sub-feature to be recognized is determined to satisfy the target recognition rule, the recognition sub-result of the data to be recognized is determined to indicate that the data to be recognized has been tampered with.

[0086] According to embodiments of this disclosure, the feature types of the multiple sub-features to be identified are different from each other, and the identification rules for the multiple sub-features to be identified are also different from each other. A mapping relationship between identification rules and feature types can be configured, and the multiple feature types and multiple identification rules can have a one-to-one correspondence. Based on the feature type of the sub-feature to be identified, a target identification rule matching the sub-feature to be identified is determined from the mapping relationship between identification rules and feature types.

[0087] According to embodiments of this disclosure, if it is determined that the sub-feature to be identified satisfies the target identification rules, the identification sub-result of the data to be identified indicates that the data to be identified has been tampered with. If it is determined that the sub-feature to be identified does not satisfy the target identification rules, the identification sub-result of the data to be identified indicates that the data to be identified has not been tampered with.

[0088] According to embodiments of this disclosure, a target recognition rule matching the sub-feature to be identified is determined based on the feature type of the sub-feature to be identified. By determining whether the sub-feature to be identified and the reference sub-feature satisfy the target recognition rule, the identification sub-result of the data to be identified is determined, which makes the method of determining the identification sub-result both benchmark and evaluable.

[0089] According to embodiments of this disclosure, the identification rule can be a pre-set rule. Determining whether the sub-feature to be identified satisfies the target identification rule based on the sub-feature to be identified and the target reference sub-feature can include the following operations.

[0090] For example, determine the target association relationship between the sub-feature to be identified and the target baseline sub-feature. Based on the target association relationship and the predetermined association relationship, determine whether the sub-feature to be identified satisfies the target identification rules.

[0091] According to embodiments of this disclosure, the target association relationship can refer to whether the sub-feature to be identified and the target reference sub-feature are the same. However, it is not limited to this. The target association relationship can also refer to whether the similarity result between the sub-feature to be identified and the target reference sub-feature is greater than a predetermined similarity threshold. The target association relationship can also refer to whether the sub-feature to be identified is part of the target reference sub-feature, or whether the sub-feature to be identified includes the target reference sub-feature. The target association relationship only needs to reflect the association relationship between the sub-feature to be identified and the target reference sub-feature.

[0092] According to embodiments of this disclosure, the predetermined association relationship is an association relationship consistent with the target association relationship type.

[0093] According to embodiments of this disclosure, taking the protocol version of an encryption protocol as an example, the target identification rule matching the protocol version of the encryption protocol includes: if the target association relationship is determined to be different between the sub-feature to be identified and the target baseline sub-feature, and the predetermined association relationship is that the sub-feature to be identified is different from the target baseline sub-feature, then the sub-feature to be identified satisfies the target identification rule. Otherwise, the sub-feature to be identified is determined not to satisfy the target identification rule.

[0094] For example, a target baseline sub-feature with protocol version 1.2 corresponds to the identification result of the first historical data, indicating that the first historical data is unaltered. If the sub-feature to be identified has protocol version 1.1, and the target association relationship between the sub-feature to be identified and the target baseline sub-feature is determined to be different, while the target association relationship is the same as the predetermined association relationship, then the sub-feature to be identified is determined to satisfy the target identification rules. The identification result is then used to indicate that the data to be identified has been tampered with.

[0095] According to embodiments of this disclosure, taking a target field of an encryption protocol as an example, the target identification rule matching the target field of the encryption protocol includes: if it is determined that the target association relationship is the same as the target reference sub-feature, and the predetermined association relationship is the same as the target reference sub-feature, then the target sub-feature is determined to satisfy the target identification rule. Otherwise, it is determined that the target sub-feature does not satisfy the target identification rule.

[0096] For example, the target field of the encryption protocol, such as the target field sequence AAAA of the encryption suite or the target field sequence BBBB of the extended list, is a target baseline sub-feature. The identification result of the corresponding first historical data indicates that the first historical data has been tampered with. When the sub-feature to be identified is the target field sequence AAAA of the encryption suite or the target field sequence BBBB of the extended list, it is determined that the target association relationship between the sub-feature to be identified and the target baseline sub-feature is the same, and the target association relationship is the same as the predetermined association relationship. Therefore, it is determined that the sub-feature to be identified satisfies the target identification rule. The identification result is used to characterize the data to be identified as tampered with.

[0097] According to embodiments of this disclosure, taking the string length of fingerprint data in an encryption protocol as an example, the target identification rule that matches the string length of the fingerprint data in the encryption protocol includes: if the target association relationship is determined to be the same as the target reference sub-feature, and the predetermined association relationship is the same as the target reference sub-feature, then the target sub-feature is determined to satisfy the target identification rule; otherwise, the target sub-feature is determined not to satisfy the target identification rule.

[0098] For example, a target baseline sub-feature with a string length of 300 in the fingerprint data of an encryption protocol corresponds to the identification result of the first historical data, indicating that the first historical data has been tampered with. When the sub-feature to be identified is a string length of 300 in the fingerprint data of the encryption protocol, it is determined that the target association relationship between the sub-feature to be identified and the target baseline sub-feature is the same, and the target association relationship is the same as the predetermined association relationship. Therefore, it is determined that the sub-feature to be identified satisfies the target identification rule. The identification result is used to characterize the data to be identified as tampered with.

[0099] According to embodiments of this disclosure, the target association relationship and the predetermined association relationship in the target recognition rule can be matched based on the target reference sub-features, and then set as a rule that is more in line with reality and easier to judge, so that the method of determining the recognition sub-result is flexible and easy to implement.

[0100] According to embodiments of this disclosure, massive amounts of network traffic can be collected, and the data with determined identification results can be used as first historical data. The larger the amount of first historical data, the more representative the extracted sub-features to be identified will be. The feature type of the benchmark sub-feature can be the same as the feature type of the sub-features to be identified. The method for determining the benchmark sub-features can be similar to the method for determining the sub-features to be identified.

[0101] For example, a portion of the encryption protocol-related data from the first historical data can be directly used as a baseline sub-feature. For instance, data related to the encryption protocol, such as target fields within the encryption protocol, can be extracted from the first historical data and used as a baseline sub-feature. Alternatively, encryption protocol-related data can be extracted from multiple first historical data sets, processed, and used as the baseline sub-feature. Common data can also be identified from the encryption protocol-related data in the first historical data and used as a baseline sub-feature. For example, indicator data can be identified from data related to device environment information within the encryption protocol. Based on the indicator data, data related to fingerprint data within the encryption protocol can be identified as a baseline sub-feature.

[0102] According to embodiments of this disclosure, taking the string length of fingerprint data in the encryption protocol as an example, the data identification method may further include the following method for determining the reference sub-features.

[0103] For example, first historical data within a predetermined historical period is acquired. Data related to device environment information is determined from the first historical data. Based on the data related to device environment information, indicator data is determined. Based on the indicator data and a predetermined indicator threshold, the target string length of the fingerprint data in the encryption protocol is determined from the first historical data, serving as a baseline sub-feature.

[0104] According to embodiments of this disclosure, the identification result of the first historical data can be used to characterize that the first historical data has been tampered with. Data related to device environment information can be determined from the first historical data. Data related to device environment information may include at least one of the following: the IP address of the device terminal, the IP address of the application server, user identification information, device identification information of the device terminal, and device identification information of the application server.

[0105] According to embodiments of this disclosure, when the amount of first historical data is sufficiently large, indicator data can be determined based on data related to device environment information within the first historical data. This indicator data can be used to indirectly indicate the mapping relationship between the identification result and the benchmark features.

[0106] According to embodiments of this disclosure, the indicator data may include one or more of the following: IP address portability rate, identification information validity rate, login status, login duration, and pageviews.

[0107] According to embodiments of this disclosure, the IP address portability rate may include one or more of the IP address portability rate of the device terminal and the IP address portability rate of the application server. The validity rate of the identification information may include one or more of the validity rate of the user's identity identification information, the validity rate of the device's identification information, and the validity rate of the application server's device identification information. Whether or not login is enabled may include whether or not login is enabled using the user's identity identification information. Login duration may refer to the login duration when login is enabled using the user's identity identification information. Page views may refer to the page views of the application server's IP address.

[0108] According to other embodiments of this disclosure, the indicator data includes multiple indicator sub-data, each of which may include, for example, one of the following: IP address portability rate, identification information validity rate, login status, login duration, and pageviews. For each indicator sub-data, a target predetermined indicator sub-threshold matching the indicator sub-data can be determined. Based on the indicator sub-data and the target predetermined sub-threshold, a matching result is determined, resulting in multiple matching results, each corresponding one-to-one with multiple indicator sub-data. The multiple matching results are weighted and summed to determine the target matching result. Based on the target matching result, the target string length of the fingerprint data in the encryption protocol is determined from the first historical data as a baseline sub-feature. However, this is not limited to this. Multiple indicator sub-data can also be weighted and summed to obtain indicator data. This allows the target string length of the fingerprint data in the encryption protocol to be determined from the first historical data based on the indicator data and the predetermined indicator threshold, serving as a baseline sub-feature.

[0109] According to embodiments of this disclosure, determining the target string length of fingerprint data in the encryption protocol from first historical data as a benchmark sub-feature may include: comparing indicator data with a predetermined indicator threshold; if it is determined that the indicator data is less than the predetermined indicator threshold, then determining the target string length of fingerprint data in the encryption protocol from the first historical data as a benchmark sub-feature.

[0110] For example, if the indicator data is determined to be greater than or equal to a predetermined indicator threshold, it indicates that the fingerprint data of the encryption protocol in the first historical data is tamper-proof, and the target string length is 200, so it can be excluded from the baseline sub-feature. If the indicator data is determined to be less than the predetermined indicator threshold, it indicates that the fingerprint data of the encryption protocol in the first historical data is tamper-proof, and the target string length is 300, so data with a target string length of 300 can be used as the baseline sub-feature.

[0111] According to embodiments of this disclosure, it should be noted that the first historical data is time-sensitive, and when collecting the first historical data, a predetermined historical time period can be used for filtering. The predetermined historical time period can be historical data within one month prior to the current time. Using a predetermined historical time period not only makes the collected first historical data representative, but also allows the first historical data to be updated based on the predetermined historical time period, so that the benchmark sub-features are time-sensitive features.

[0112] According to embodiments of this disclosure, by determining benchmark sub-features related to the identification result through indicator data in the first historical data, the determined benchmark sub-features can be made representative, providing a basis for identifying whether the data to be identified has been tampered with.

[0113] According to embodiments of this disclosure, the data identification method may further include: acquiring a set of historical data collected within a predetermined historical period; and, if it is determined that there is target second historical data in the historical data set that matches the data to be identified, updating the identification result of the data to be identified based on the identification result of the target second historical data.

[0114] According to embodiments of this disclosure, each second historical data in the historical data set is data with known identification results. The second historical data includes features that are the same as the baseline features, and the identification results of the second historical data are different from the identification results of the first historical data.

[0115] According to embodiments of this disclosure, the encryption protocol-related features present in the second historical data are the same as the baseline features, but the identification result of the second historical data differs from the identification result of the first historical data. For example, encryption protocol-related features are extracted from the second historical data, and these features are the same as the baseline features. Referring to the baseline features, it can be determined that the identification result of the second historical data is the same as the identification result of the first historical data. The identification result of the first historical data is used to characterize the first historical data as tampered abnormal data. However, in reality, the second historical data is normal data, that is, normal data that has not been tampered with.

[0116] According to embodiments of this disclosure, when the identification result of the data to be identified has been determined based on benchmark data, the data to be identified can be compared with data in a historical data set to determine a target second historical data that matches the data to be identified. If it is determined that a target second historical data matching the data to be identified exists in the historical data set, the identification result of the target second historical data can be used to update the identification result of the data to be identified. For example, based on benchmark features, the identification result of the data to be identified is determined as a first identification result, used to characterize the data to be identified as tampered data. Based on the target second historical data, the identification result of the data to be identified is determined as a second identification result, used to characterize the data to be identified as untampered data. The second identification result can be used to update the first identification result. For example, the second identification result can be determined from the first and second identification results as the final target identification result.

[0117] According to embodiments of this disclosure, by utilizing second historical data from a pre-collected historical data set, data to be identified that does not conform to predetermined identification rules can be reclassified based on actual conditions. This avoids identification errors caused by using benchmark features to determine the identification results of the data to be identified, and ultimately achieves accurate identification of the untampered data to be identified.

[0118] According to embodiments of this disclosure, before performing the operation of updating the identification result of the data to be identified based on the identification result of the target second historical data when it is determined that there is target second historical data in the historical data set that matches the data to be identified, the data identification method may include the following operations.

[0119] For example, the data to be identified is transformed using a predetermined algorithm to obtain transformed data to be identified. Based on the transformed data to be identified and multiple transformed second historical data in the historical data set, it is determined whether there is target second historical data in the historical data set that matches the data to be identified. The multiple transformed second historical data are data transformed by the predetermined algorithm.

[0120] According to embodiments of this disclosure, the predetermined algorithm may refer to a hash algorithm. A hash algorithm can be a one-way mathematical function, and may include at least one of the following functions: MD2, MD4, MD5.

[0121] According to embodiments of this disclosure, a predetermined algorithm can be used to transform the data to be identified and multiple second historical data sets, such that the data to be identified and the multiple second historical data sets are in the same format. Furthermore, after the predetermined algorithm transformation, the resulting transformed data to be identified and the transformed multiple second historical data sets have a small data volume, making them easy to store and compute. This reduces the storage space required for the second historical data sets while improving the processing efficiency for determining the target second historical data.

[0122] Figure 5 A schematic flowchart of a data identification method according to another embodiment of the present disclosure is shown.

[0123] like Figure 5 As shown, initial data to be identified 520 can be determined from network request 510. The initial data to be identified 520 is preprocessed to obtain data to be identified 530. Identification features 540 are determined from the data to be identified 530. Based on the identification features 540 and the baseline features 550, a first identification result 560 is determined for the data to be identified. The data to be identified 530 is matched with multiple second historical data in historical data set 570 to determine whether there is a target second historical data 571 in historical data set 570 that matches the data to be identified 530. If the target second historical data 571 is determined to exist in historical data set 570, the second identification result 580 of the target second historical data 571 is used as the identification result 590 of the data to be identified 530. If the target second historical data is determined not to exist in historical data set, the first identification result 560 is used as the identification result 590 of the data to be identified 530.

[0124] According to embodiments of this disclosure, a network request can be a request sent by a terminal device to an application server to establish an encrypted communication connection based on an encryption protocol. The initial data to be identified can be data from the client hello packet during the TLS handshake between the terminal device and the application server. Based on the TLS protocol, when exchanging information between the terminal device and the application server, the application server records data such as the terminal device's protocol version, supported encryption suites, and extension list, forming fingerprint data, such as JA3 fingerprint data, to describe device environment information or browser information. The initial data to be identified can be preprocessed to obtain the data to be identified. Identification features related to TLS, specifically the JA3 fingerprint dimension, are extracted from the data to be identified. Based on the baseline features and the identification features, a first identification result for the data to be identified is determined. Combined with a collection of historical data with specific rules, it is determined whether the data to be identified is special data. If a second target historical data is determined from the historical data set, then the data to be identified is special data. For example, new normal fingerprint data arising from a device upgrade. For this new normal fingerprint data, the baseline features are used for evaluation, and it is determined to be tampered data. In this scenario, new, normal fingerprint data can be pre-classified into a historical dataset. When identical fingerprint data to be identified appears, the second identification result can be used to correct the first identification result, ultimately determining the final identification result. This avoids the problem of using a single baseline feature to determine an incorrect first identification result and then accepting that incorrect first identification result as the final identification result.

[0125] The data identification method provided in this disclosure can effectively identify tampered data in abnormal device environments, such as when the encryption suite or extended list supported by JA3 fingerprint is tampered with, thereby improving the security of network communication connections.

[0126] Figure 6 A block diagram of a data identification device according to an embodiment of the present disclosure is shown schematically.

[0127] like Figure 6 As shown, the data identification device 600 includes: a first determining module 610 and a second determining module 620.

[0128] The first determining module 610 is used to determine the features to be identified from the data to be identified, wherein the features to be identified include features related to the encryption protocol.

[0129] The second determining module 620 is used to determine the identification result of the data to be identified based on the features to be identified and the reference features. The identification result is used to characterize whether the data to be identified has been tampered with, and the reference features are features extracted from the first historical data with known identification results and are related to the encryption protocol.

[0130] According to embodiments of this disclosure, the feature to be identified includes multiple sub-features to be identified, and the reference feature includes multiple reference sub-features.

[0131] According to embodiments of this disclosure, the second determining module includes: a first determining submodule, a second determining submodule, and a third determining submodule.

[0132] The first determination submodule is used to determine, for each of the multiple sub-features to be identified, a target reference sub-feature that matches the feature type of the sub-feature to be identified from multiple reference sub-features.

[0133] The second determination submodule is used to determine the identification sub-result of the data to be identified based on the sub-features to be identified and the target reference sub-features.

[0134] The third determination submodule is used to determine the recognition result of the data to be recognized based on multiple recognition sub-results, wherein the multiple recognition sub-results correspond one-to-one with multiple sub-features to be recognized.

[0135] According to embodiments of this disclosure, the data identification device further includes a first acquisition module and an update module.

[0136] The first acquisition module is used to acquire a set of historical data collected within a predetermined historical period. Each second historical data in the historical data set is data with a known recognition result. The second historical data includes features that are the same as the baseline features, and the recognition result of the second historical data is different from the recognition result of the first historical data.

[0137] The update module is used to update the identification result of the data to be identified based on the identification result of the target second historical data when it is determined that there is target second historical data in the historical data set that matches the data to be identified.

[0138] According to embodiments of this disclosure, the second determining submodule includes: a first determining unit, a second determining unit, and a third determining unit.

[0139] The first determining unit is used to determine the target recognition rule that matches the feature type of the sub-feature to be recognized from a plurality of predetermined recognition rules.

[0140] The second determining unit is used to determine whether the sub-feature to be identified satisfies the target identification rule based on the sub-feature to be identified and the target reference sub-feature.

[0141] The third determining unit is used to determine, when it is determined that the sub-features to be identified satisfy the target identification rules, that the identification sub-results of the data to be identified represent that the data to be identified has been tampered with.

[0142] According to embodiments of this disclosure, the second determining unit includes: a first determining subunit and a second determining subunit.

[0143] The first determining sub-unit is used to determine the target association relationship between the sub-feature to be identified and the target reference sub-feature.

[0144] The second determining subunit is used to determine whether the sub-feature to be identified satisfies the target identification rules based on the target association relationship and the predetermined association relationship.

[0145] According to embodiments of this disclosure, the reference sub-feature includes the string length of the fingerprint data in the encryption protocol.

[0146] According to embodiments of this disclosure, the data identification device further includes: a second acquisition module, a third determination module, a fourth determination module, and a fifth determination module.

[0147] The second acquisition module is used to acquire the first historical data within a predetermined historical period.

[0148] The third determination module is used to determine data related to the equipment environment information from the first historical data.

[0149] The fourth determination module is used to determine indicator data based on data related to equipment environmental information.

[0150] The fifth determination module is used to determine the target string length of the fingerprint data in the encryption protocol from the first historical data based on the indicator data and the predetermined indicator threshold, as a benchmark sub-feature.

[0151] According to embodiments of this disclosure, the data identification device further includes, prior to the update module: a conversion module and a sixth determination module.

[0152] The conversion module is used to convert the data to be identified using a predetermined algorithm to obtain the converted data to be identified.

[0153] The sixth determining module is used to determine whether there is a target second historical data in the historical data set that matches the data to be identified, based on the transformed data to be identified and multiple transformed second historical data in the historical data set. The multiple transformed second historical data are data transformed by a predetermined algorithm.

[0154] According to embodiments of this disclosure, the data identification device further includes, prior to the first determining module: a receiving module and a preprocessing module.

[0155] The receiving module is used to receive network requests, which include initial data to be identified, which is data related to the encryption protocol.

[0156] The preprocessing module is used to preprocess the initial data to be identified to obtain the data to be identified.

[0157] According to embodiments of this disclosure, the sub-feature to be identified includes at least one of the following:

[0158] The encryption protocol version, the target field sequence of the encryption protocol, and the string length of the encryption protocol.

[0159] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0160] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in the embodiments of the present disclosure.

[0161] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform methods as described in embodiments of the present disclosure.

[0162] According to embodiments of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the methods as described in embodiments of the present disclosure.

[0163] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0164] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0165] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0166] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as data recognition methods. For example, in some embodiments, the data recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the data recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform data recognition methods by any other suitable means (e.g., by means of firmware).

[0167] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0168] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0169] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0171] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0172] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0173] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0174] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A data identification method, comprising: Determine the features to be identified from the data to be identified, wherein the features to be identified include features related to the encryption protocol; Based on the features to be identified and the baseline features, an identification result is determined for the data to be identified, wherein the identification result is used to characterize whether the data to be identified has been tampered with, the baseline features include multiple baseline sub-features, and common data is used as the baseline sub-features, the common data being determined from data related to the encryption protocol in the first historical data; and If it is determined that there is a target second historical data in the historical data set that matches the data to be identified, the identification result of the data to be identified is updated based on the identification result of the target second historical data. Each second historical data in the historical data set is data with a known identification result. The second historical data includes features that are the same as the benchmark features, and the identification result of the second historical data is different from the identification result of the first historical data.

2. The method of claim 1, wherein, The feature to be identified includes multiple sub-features to be identified; The step of determining the recognition result of the data to be recognized based on the features to be recognized and the reference features includes: For each of the plurality of sub-features to be identified, a target reference sub-feature that matches the feature type of the sub-feature to be identified is determined from the plurality of reference sub-features; Based on the sub-features to be identified and the target reference sub-features, the identification sub-results of the data to be identified are determined; and Based on the multiple identification sub-results, the identification result of the data to be identified is determined, wherein the multiple identification sub-results correspond one-to-one with the multiple sub-features to be identified.

3. The method according to claim 1 or 2, further comprising: Obtain the set of historical data collected within a predetermined historical period.

4. The method of claim 2, wherein, The step of determining the identification sub-result of the data to be identified based on the sub-feature to be identified and the target reference sub-feature includes: Determine a target recognition rule from a set of predefined recognition rules that matches the feature type of the sub-feature to be recognized; Based on the sub-feature to be identified and the target reference sub-feature, determine whether the sub-feature to be identified satisfies the target identification rule; and If the sub-feature to be identified satisfies the target identification rule, the identification sub-result of the data to be identified indicates that the data to be identified has been tampered with.

5. The method of claim 4, wherein, The step of determining whether the sub-feature to be identified satisfies the target identification rule based on the sub-feature to be identified and the target reference sub-feature includes: Determine the target association relationship between the sub-feature to be identified and the target reference sub-feature; and Based on the target association relationship and the predetermined association relationship, determine whether the sub-feature to be identified satisfies the target identification rule.

6. The method of claim 2, wherein, The baseline sub-feature includes the string length of the fingerprint data in the encryption protocol; The method further includes: Obtain the first historical data within the predetermined historical time period; Identify data related to the equipment environment information from the first historical data; Based on the data related to the equipment environment information, the indicator data are determined; and Based on the indicator data and the predetermined indicator threshold, the target string length of the fingerprint data in the encryption protocol is determined from the first historical data and used as the baseline sub-feature.

7. The method according to claim 3, further comprising, before updating the identification result of the data to be identified based on the identification result of the target second historical data when it is determined that there is target second historical data in the historical data set that matches the data to be identified: The data to be identified is transformed using a predetermined algorithm to obtain the transformed data to be identified; and determining whether there is target second historical data matching the to-be-recognized data in the historical data set based on the converted to-be-recognized data and the converted plurality of second historical data in the historical data set, wherein, The multiple second historical data after conversion are data converted by the predetermined algorithm.

8. The method of claim 2, further comprising, before determining the feature to be identified from the data to be identified: receiving a network request, wherein, The network request includes initial data to be identified, which is data related to the encryption protocol; and The initial data to be identified is preprocessed to obtain the data to be identified.

9. The method of claim 8, wherein, The sub-feature to be identified includes at least one of the following: The encryption protocol version, the target field sequence of the encryption protocol, and the string length of the encryption protocol.

10. A data identification device, comprising: The first determining module is used to determine the features to be identified from the data to be identified, wherein the features to be identified include features related to the encryption protocol; The second determining module is used to determine the identification result of the data to be identified based on the features to be identified and the benchmark features. The identification result is used to characterize whether the data to be identified has been tampered with. The benchmark features include multiple benchmark sub-features, and common data is used as the benchmark sub-features. The common data is determined from data related to the encryption protocol in the first historical data. An update module is used to update the identification result of the data to be identified based on the identification result of the target second historical data when it is determined that there is a target second historical data in the historical data set that matches the data to be identified. Each second historical data in the historical data set is data with a known identification result. The second historical data includes features that are the same as the benchmark features, and the identification result of the second historical data is different from the identification result of the first historical data.

11. The apparatus of claim 10, wherein, The feature to be identified includes multiple sub-features to be identified; The second determining module includes: The first determining submodule is used to determine, for each of the plurality of sub-features to be identified, a target reference sub-feature that matches the feature type of the sub-feature to be identified from the plurality of reference sub-features; The second determining submodule is used to determine the identification sub-result of the data to be identified based on the sub-feature to be identified and the target reference sub-feature; and The third determining submodule is used to determine the recognition result of the data to be recognized based on the multiple recognition sub-results, wherein the multiple recognition sub-results correspond one-to-one with the multiple sub-features to be recognized.

12. The apparatus according to claim 10 or 11, further comprising: The first acquisition module is used to acquire the set of historical data collected within a predetermined historical period.

13. The apparatus of claim 11, wherein, The second determining submodule includes: The first determining unit is used to determine a target recognition rule that matches the feature type of the sub-feature to be identified from a plurality of predetermined recognition rules; The second determining unit is configured to determine, based on the sub-feature to be identified and the target reference sub-feature, whether the sub-feature to be identified satisfies the target identification rule; and The third determining unit is used to determine, when it is determined that the sub-feature to be identified satisfies the target identification rule, that the identification sub-result of the data to be identified represents that the data to be identified is tampered with.

14. The apparatus of claim 13, wherein, The second determining unit includes: A first determining subunit is configured to determine the target association relationship between the sub-feature to be identified and the target reference sub-feature; and The second determining subunit is used to determine whether the sub-feature to be identified satisfies the target identification rule based on the target association relationship and the predetermined association relationship.

15. The apparatus of claim 11, wherein, The baseline sub-feature includes the string length of the fingerprint data in the encryption protocol; The device further includes: The second acquisition module is used to acquire the first historical data within a predetermined historical time period; The third determining module is used to determine data related to the equipment environment information from the first historical data; The fourth determining module is used to determine indicator data based on the data related to the equipment environment information; and The fifth determining module is used to determine the target string length of the fingerprint data in the encryption protocol from the first historical data based on the indicator data and the predetermined indicator threshold, as the benchmark sub-feature.

16. The apparatus of claim 12, further comprising, prior to the updating module: A conversion module is used to perform a predetermined algorithm conversion on the data to be identified to obtain the converted data to be identified; and The sixth determining module is used to determine, based on the transformed data to be identified and multiple transformed second historical data in the historical data set, whether there is target second historical data in the historical data set that matches the data to be identified. The multiple second historical data after conversion are data converted by the predetermined algorithm.

17. The apparatus of claim 11, further comprising, prior to the first determining module: The receiving module is used to receive network requests, wherein... The network request includes initial data to be identified, which is data related to the encryption protocol; as well as The preprocessing module is used to preprocess the initial data to be identified to obtain the data to be identified.

18. The apparatus of claim 17, wherein, The sub-feature to be identified includes at least one of the following: The encryption protocol version, the target field sequence of the encryption protocol, and the string length of the encryption protocol.

19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 9.

20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 9.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 9.