A method and related apparatus for processing preset identifiers in data.

By querying the characteristic attributes of the target preset identifier in the collected data and performing a one-to-one unidirectional mapping, the reliability and security issues of diverting online stored data to offline testing are solved, achieving wider applicability and higher quality offline testing.

CN116992170BActive Publication Date: 2025-11-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211228036.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-11-14
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

In existing technologies, when data is collected and transferred from online storage to offline testing, there are problems such as weak reliability, low security, and poor applicability, which affect the effectiveness of subsequent offline testing.

Method used

By extracting the target preset identifier from the collected data stored online, querying the associated feature attributes that cannot identify the target subject, and using a one-to-one unidirectional mapping algorithm to map it to a target mapping identifier that cannot identify any subject, the target mapping identifier is finally replaced with a target test identifier that has the target feature attributes, thus forming data for offline testing.

Benefits of technology

This improved the reliability, security, and applicability of data collection from online storage to offline testing, thereby enhancing the effectiveness and quality of offline testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992170B_ABST
    Figure CN116992170B_ABST
Patent Text Reader

Abstract

This application discloses a method and related apparatus for processing preset identifiers in data. The method involves extracting a target preset identifier that identifies a target entity from first collected data stored online; querying target feature attributes associated with the target preset identifier that cannot identify the target entity based on the business scenario of the first collected data; mapping the target preset identifier in the first collected data to obtain second collected data using a one-to-one unidirectional mapping preset algorithm, the second collected data including target mapping identifiers that cannot identify any entity in the set of entities to which the target entity belongs; and replacing the target mapping identifier in the second collected data with a target test identifier that is different from the target preset identifier but has target feature attributes, using the correspondence between the target mapping identifier and the target test identifier that has target feature attributes, to obtain third collected data for offline testing. This method improves the reliability, security, and applicability of transferring collected data from online storage to offline testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular to a method and related apparatus for processing preset identifiers in data. Background Technology

[0002] With the rapid development of data processing technology, online-stored collected data can be redirected to offline testing. In this process, preset identifiers within the collected data identify the subject; that is, the preset identifiers represent the identification information of the subject. These preset identifiers need to be modified to enable offline testing.

[0003] In related technologies, when transforming preset identifiers in collected data, methods such as masking, randomization, data replacement, symmetric encryption, averaging, or offset rounding are commonly used to change the preset identifiers in the collected data, thereby transforming the preset identifiers.

[0004] However, some of the aforementioned related technologies alter the uniqueness of the preset identifier, some methods are not reversible enough and are not secure enough, and some methods have too narrow an application scenario. This results in weak reliability, low security and poor applicability of the data collected from online storage to offline testing, thus affecting the subsequent offline testing results. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a method and related apparatus for processing preset identifiers in data, thereby improving the reliability, security, and applicability of data collected from online storage to offline testing, and thus enhancing the effectiveness of subsequent offline testing.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] On the one hand, this application provides a method for processing preset identifiers in data, the method comprising:

[0008] Extract a target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject.

[0009] Based on the business scenario of the first collected data, query the target feature attributes associated with the target preset identifier; the target feature attributes cannot identify the target subject;

[0010] According to a preset mapping algorithm for one-to-one unidirectional mapping, the target preset identifier in the first collected data is mapped to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs;

[0011] Based on the correspondence between the target mapping identifier and the target test identifier, which is different from the target preset identifier but has the target characteristic attribute, the target mapping identifier in the second collected data is replaced to obtain the third collected data for offline testing.

[0012] On the other hand, this application provides a processing apparatus for preset identifiers in data, the apparatus comprising: an extraction unit, a query unit, a mapping unit, and a replacement unit;

[0013] The extraction unit is used to extract a target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject;

[0014] The query unit is used to query the target feature attribute associated with the target preset identifier based on the business scenario of the first collected data; the target feature attribute cannot identify the target subject;

[0015] The mapping unit is used to map the target preset identifier in the first collected data according to a preset mapping algorithm of one-to-one unidirectional mapping to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs;

[0016] The replacement unit is used to replace the target mapping identifier in the second collected data according to the correspondence between the target mapping identifier and the target test identifier which is different from the target preset identifier but has the target characteristic attribute, so as to obtain the third collected data for offline testing.

[0017] On the other hand, this application provides a computer device, which includes a processor and a memory:

[0018] The memory is used to store program code and transmit the program code to the processor;

[0019] The processor is used to execute the processing method for the preset identifier in the data described above, according to the instructions in the program code.

[0020] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program, which, when executed by a processor, performs the processing method for a preset identifier in the data described above.

[0021] On the other hand, embodiments of this application provide a computer program product, which includes a computer program or instructions; when the computer program or instructions are executed by a processor, the processing method for preset identifiers in the data described above is executed.

[0022] As can be seen from the above technical solution, firstly, a target preset identifier for identifying the target subject is extracted from the first collection data stored online; secondly, based on the business scenario of the first collection data, the target feature attributes associated with the target preset identifier that cannot identify the target subject are queried; then, through a preset mapping algorithm of one-to-one unidirectional mapping, the target preset identifier in the first collection data is mapped to obtain the second collection data, which includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs; finally, the target mapping identifier in the second collection data is replaced with the target test identifier that is different from the target preset identifier but has target feature attributes, based on the correspondence between the target mapping identifier and the target test identifier that has target feature attributes, to obtain the third collection data for offline testing.

[0023] As can be seen, for the first collection data stored online, additional querying is performed on the target feature attributes associated with the target preset identifier that cannot identify the target subject. After obtaining the second collection data by using a preset mapping algorithm with one-to-one unidirectional mapping, the target preset identifier is mapped to a target mapping identifier that cannot identify any subject. Then, the target mapping identifier is replaced with a target test identifier that is different from the target preset identifier but has target feature attributes, resulting in the third collection data for offline testing. Based on this, the method ensures that the target test identifier has both the target feature attributes of the target preset identifier and cannot be traced back to the target subject identified by the target preset identifier. It also has a wider range of applicable scenarios, thereby improving the reliability, security, and applicability of the collection data from online storage to offline testing, thus improving the subsequent offline testing results. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A schematic diagram illustrating common variations of related technologies provided in the embodiments of this application;

[0026] Figure 2 A schematic diagram illustrating an application scenario of a method for processing preset identifiers in data, provided in an embodiment of this application;

[0027] Figure 3 A flowchart illustrating a method for processing preset identifiers in data, provided as an embodiment of this application;

[0028] Figure 4 A schematic diagram illustrating the specific process stages of processing a preset identifier in data, provided in an embodiment of this application;

[0029] Figure 5 A schematic diagram of a data processing device for a given preset identifier provided in an embodiment of this application;

[0030] Figure 6 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application.

[0031] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0032] The embodiments of this application will now be described with reference to the accompanying drawings.

[0033] At present, in the scenario of diverting online-stored collected data to offline testing, considering that the preset identifier in the collected data identifies the subject, that is, the preset identifier represents the identification information of the subject, common transformation methods such as masking, randomization, data replacement, symmetric encryption, averaging or offset rounding are usually used to transform the preset identifier, so that the preset identifier in the collected data changes, realizing the transformation of the preset identifier, so as to conduct offline testing later.

[0034] See Figure 1 The diagram illustrates common variations in related technologies. For example, "preset identifier A" can be modified by masking, randomizing, data replacement, or symmetric encryption; "preset identifier B" can be modified by averaging; and "preset identifier C" can be modified by offsetting and rounding.

[0035] Masking refers to using mask characters to cover up part of the characters, that is, using mask characters to replace part of the characters. The part that needs to be masked can be adjusted according to actual needs.

[0036] Randomization refers to replacing existing characters with random characters, such as replacing existing letters with random letters, existing numbers with random numbers, and existing text with random text. The advantage of this method is that it can preserve the original character format to a certain extent, making it less likely to be noticed.

[0037] Data replacement refers to replacing the original characters with set characters, which is similar to the masking method mentioned above. The difference is that it does not replace the original characters with mask characters.

[0038] Symmetric encryption refers to encrypting original characters using an encryption key and an encryption algorithm. The encrypted ciphertext characters are consistent with the original characters in terms of logical rules, and the original characters can be recovered using the decryption key. Since this method is reversible, the security of the key must be taken into account.

[0039] Averaging refers to calculating the mean of the original numerical characters, so that the transformed original characters are randomly distributed around the mean, thus keeping the sum of the original characters unchanged. It is suitable for statistical scenarios.

[0040] Offset rounding refers to changing the original character by random shifting. While maintaining the security of the original character, it ensures that the range of the original character after transformation is basically true. That is, the original character after transformation is closer to the original character. It is suitable for big data analysis scenarios.

[0041] However, research has revealed that the three methods mentioned above—masking, randomization, and data replacement—all alter the uniqueness of the preset identifier. Symmetric encryption methods are not secure enough due to their reversibility, and averaging and offset rounding methods have too narrow an application scenario. This results in weak reliability, low security, and poor applicability of the collected data when it is transferred from online storage to offline testing, thus affecting the subsequent offline testing results.

[0042] In view of this, this application proposes a method and related apparatus for processing preset identifiers in data. For the first collection data stored online, additional query is performed to identify target feature attributes associated with the target preset identifier that cannot identify the target subject. Then, using a one-to-one unidirectional preset mapping algorithm, the target preset identifier is mapped to a target mapping identifier that cannot identify any subject, resulting in the second collection data. Finally, the target mapping identifier is replaced with a target test identifier that is different from the target preset identifier but possesses target feature attributes, resulting in the third collection data for offline testing. Based on this, the method ensures that the target test identifier possesses both the target feature attributes of the target preset identifier and cannot be traced back to the target subject identified by the target preset identifier. Furthermore, it has a wider range of applicable scenarios, thereby improving the reliability, security, and applicability of data transferred from online storage to offline testing, thus enhancing the subsequent offline testing results.

[0043] It is understood that in the specific implementation of this application, if the preset identifier and other related data involve user information, when the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain separate permission or consent from the user, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0044] To facilitate understanding of the technical solution of this application, the following describes the method for processing preset identifiers in the data provided in the embodiments of this application, in conjunction with actual application scenarios.

[0045] See Figure 2 , Figure 2 This is a schematic diagram illustrating an application scenario of a method for processing preset identifiers in data, provided in an embodiment of this application. Figure 2 The application scenario shown includes database 201 and server 202.

[0046] Database 201 stores the first collected data, which includes a target preset identifier that identifies the target subject. When the first collected data is redirected to offline testing, after server 202 retrieves the first collected data stored online from database 201, it extracts the target preset identifier from the first collected data stored online. As an example, the first collected data is "first file X", and "first file X" includes a target preset identifier "preset identifier Y" that identifies the target subject "subject a". Server 202 can extract "preset identifier Y" from the "first file X" stored online.

[0047] Based on the business scenario of the first collected data, server 202 queries the target feature attributes associated with the target preset identifier; the target feature attributes cannot identify the target subject. As an example, based on the above example, server 202 can query the target feature attributes associated with "preset identifier Y" that cannot identify "subject a" as "feature attribute 1, feature attribute 2, ..., feature attribute N" based on the business scenario of "first file X".

[0048] Server 202 maps the target preset identifier in the first collected data according to a preset mapping algorithm of one-to-one unidirectional mapping to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs. As an example, based on the above example, the preset mapping algorithm of one-to-one unidirectional mapping is a "hash salting algorithm". Server 202 can use the "hash salting algorithm" to map the "preset identifier Y" in "first file X" to obtain "second file X", which is the second collected data. "Second file X" includes the target mapping identifier "mapping identifier Y" that cannot identify any subject in the subject set to which "subject a" belongs.

[0049] Server 202 replaces the target mapping identifier in the second collected data according to the correspondence between the target mapping identifier and the target test identifier, which is different from the target preset identifier but has target characteristic attributes, to obtain the third collected data for offline testing. As an example, based on the above example, the target test identifier is a "test identifier Y" that is different from "identifier Y" and has "characteristic attribute 1, characteristic attribute 2, ..., characteristic attribute N". Server 202 can replace the "mapping identifier Y" in "second file X" through the correspondence between "mapping identifier Y" and "test identifier Y" to obtain "third file X", i.e., the third collected data, for offline testing. "Third file X" includes "test identifier Y".

[0050] As can be seen, for the first collection data stored online, additional querying is performed on the target feature attributes associated with the target preset identifier that cannot identify the target subject. After obtaining the second collection data by using a preset mapping algorithm with one-to-one unidirectional mapping, the target preset identifier is mapped to a target mapping identifier that cannot identify any subject. Then, the target mapping identifier is replaced with a target test identifier that is different from the target preset identifier but has target feature attributes, resulting in the third collection data for offline testing. Based on this, the method ensures that the target test identifier has both the target feature attributes of the target preset identifier and cannot be traced back to the target subject identified by the target preset identifier. It also has a wider range of applicable scenarios, thereby improving the reliability, security, and applicability of the collection data from online storage to offline testing, thus improving the subsequent offline testing results.

[0051] The method for processing preset identifiers in data provided in this application can be applied to devices capable of processing preset identifiers in data, such as servers and terminal devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, but is not limited thereto. Terminal devices include, but are not limited to, mobile phones, tablets, computers, smart cameras, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft, but are not limited thereto. Terminal devices and servers can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this connection.

[0052] The method for processing the preset identifiers in the data provided in this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, vehicle scenarios, intelligent transportation, and assisted driving.

[0053] Next, the processing method for preset identifiers in data provided in the embodiments of this application will be specifically described using a server or terminal device as the processing device for preset identifiers in data.

[0054] See Figure 3 This figure is a flowchart of a method for processing preset identifiers in data according to an embodiment of this application. Figure 3 As shown, the method for processing the preset identifier in this data includes the following steps:

[0055] S301: Extract the target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject.

[0056] In this embodiment of the application, for the application scenario of diverting the first collected data stored online to offline testing, the first collected data includes a target preset identifier that identifies the target subject. That is, the target preset identifier in the first collected data involves the identification information of the target subject. Based on this, it is first necessary to extract the target preset identifier from the first collected data so that the target preset identifier can be transformed to achieve offline testing.

[0057] In a specific implementation of S301, the location of the target preset identifier in the first collected data can be located by using the data format of the first collected data and the preset delimiter in the first collected data, and then the target preset identifier can be extracted from the first collected data. Therefore, this application provides a possible implementation method, and S301 can be specifically as follows: extract the target preset identifier from the first collected data according to the data format of the first collected data and the preset delimiter in the first collected data.

[0058] As an example, the first collected data is “2601357921|1|1”. “2601357921|1|1” includes the target preset identifier “2601357921” that identifies the target subject “subject b”. Based on the data format of “2601357921|1|1” and the preset separator “|” in “2601357921|1|1”, “2601357921” is extracted from “2601357921|1|1”.

[0059] S302: Based on the business scenario of the first collected data, query the target feature attributes associated with the target preset identifier; the target feature attributes cannot identify the target subject.

[0060] In related technologies, methods such as masking, randomization, data replacement, symmetric encryption, averaging, or offset rounding are commonly used to transform the target preset identifier in the first collected data, thereby changing the target preset identifier for subsequent offline testing. However, research has found that the masking, randomization, and data replacement methods all alter the uniqueness of the target preset identifier; symmetric encryption is reversible and not secure enough; and averaging and offset rounding methods have too narrow an application scenario. This results in weak reliability, low security, and poor applicability of the first collected data transferred from online storage to offline testing, thus affecting the effectiveness of subsequent offline testing.

[0061] Therefore, in this embodiment of the application, in order to ensure that the modified target preset identifier still has the inherent characteristic attributes associated with the target preset identifier, but cannot be traced back to the target subject identified by the target preset identifier, so as to avoid the impact on subsequent offline testing; based on the inherent characteristic attributes associated with the first collected data determined by the business scenario of the first collected data, especially the inherent characteristic attributes associated with the target preset identifier, it is necessary to query the inherent characteristic attributes associated with the target preset identifier that cannot identify the target subject through the business scenario of the first collected data, and use them as target characteristic attributes, so that the modification of the target preset identifier can be realized based on the target characteristic attributes in the future.

[0062] It should be noted that whether the inherent characteristic attributes associated with the target preset identifier identify the target subject needs to be determined specifically according to the specific protection strategy corresponding to the target preset identifier.

[0063] In this application, the target preset identifier and target feature attributes are represented in key-value pair format, where the key is the target preset identifier and the value is the target feature attribute. Target feature attributes typically involve multiple data tables and can be categorized into different types based on the data tables. Each type of feature attribute is represented in JSON format, including two key-value pairs. Therefore, the target feature attributes include data tables and data rows within those tables; each data row stores one or more fields and their corresponding field values. Thus, this application provides a possible implementation where the target feature attributes include data tables and data rows within those tables, with each data row including fields and field values.

[0064] As an example, based on the above example, and in the business scenario of the first collected data "2601357921|1|1", the target feature attributes associated with the target preset identifier "2601357921" that cannot identify "subject b" are "feature attribute 1, feature attribute 2, ..., feature attribute N", where "feature attribute N" specifically refers to:

[0065] {

[0066] Data Table: Data Table 1,

[0067] Data row: [{field1: field value1, field2: field value2}]

[0068] }

[0069] S303: According to the preset mapping algorithm of one-to-one unidirectional mapping, the target preset identifier in the first collected data is mapped to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs.

[0070] In this embodiment of the application, when transforming the target preset identifier in the first collected data, considering that different preset identifiers produce different outputs after transformation, i.e., the output of the target preset identifier is unique, and the transformation of the preset identifier is irreversible, i.e., the transformation of the target preset identifier is irreversible; it is necessary to use a one-to-one unidirectional mapping preset mapping algorithm to map the target preset identifier in the first collected data, so that the target preset identifier is mapped to a target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs, thereby obtaining the second collected data including the target mapping identifier.

[0071] In the specific implementation of S303, for the preset mapping algorithm that maps the target preset identifier to the target mapping identifier through a one-to-one unidirectional mapping, in order to avoid mapping conflicts between different types of preset identifiers, it is also necessary to obtain the identifier type corresponding to the target preset identifier based on the target preset identifier; and to prevent attackers from attacking according to the mapping relationship table, it is further necessary to obtain the preset character corresponding to the target preset identifier. Based on this, the preset mapping algorithm of one-to-one unidirectional mapping can be used to map the identifier type, the target preset identifier, and the preset character to obtain the target mapping identifier; thereby, the target preset identifier in the first collected data is replaced by the target mapping identifier to obtain the second collected data including the target mapping identifier. Therefore, this application provides a possible implementation method, and S303 may include, for example, the following S3031-S3033:

[0072] S3031: Obtain the identifier type and preset character corresponding to the target preset identifier.

[0073] S3032: Map the identifier type, target preset identifier, and preset character according to the preset mapping algorithm to obtain the target mapping identifier.

[0074] In the specific implementation of S3032, firstly, a string to be mapped can be composed of the identifier type corresponding to the target preset identifier, the target preset identifier, and the preset characters corresponding to the target preset identifier; then, the string to be mapped is mapped to a candidate mapping identifier through a one-to-one unidirectional mapping preset mapping algorithm; finally, it is determined whether the candidate mapping identifier cannot identify any entity in the entity set to which the target entity belongs. If so, the candidate mapping identifier can be directly used as the target mapping identifier; otherwise, the target mapping identifier needs to be obtained by combining the candidate mapping identifier with mask characters. Therefore, this application provides a possible implementation method, and S3032 may include, for example, the following S1-S4:

[0075] S1: Construct the string to be mapped from the identifier type, the target preset identifier, and the preset character.

[0076] The preset characters of the target preset identifier are generated and can be implemented in any of the following ways:

[0077] The first specific implementation method is as follows: Based on the fact that the output of the same preset identifier is the same after transformation processing, considering the target feature attributes associated with the target preset identifier that cannot identify the target subject, the preset character is generated from the target feature attributes to distinguish it from other feature attributes associated with other preset identifiers that cannot identify other subjects.

[0078] The second specific implementation method is as follows: Based on the fact that the output of the same preset identifier is the same after transformation processing, the preset character is generated from the target feature attribute, taking into account the target preset identifier itself, so as to distinguish it from other preset identifiers more easily.

[0079] The third specific implementation method: When the output of the same preset identifier is the same after transformation, in order to avoid the risk of the preset character being attacked, the preset character can be randomly generated.

[0080] That is, this application provides a possible implementation method, and the steps for generating the preset character can be, for example, generating the preset character based on the target feature attribute; or, generating the preset character based on the target preset identifier; or, randomly generating the preset character.

[0081] S2: Map the string to be mapped according to the preset mapping algorithm to obtain the candidate mapping identifier.

[0082] S3: If the candidate mapping identifier cannot identify any subject, the candidate mapping identifier shall be determined as the target mapping identifier.

[0083] S4: If the candidate mapping identifier identifies other entities in the entity set, obtain the target mapping identifier based on the candidate mapping identifier and the mask character.

[0084] In determining the mask characters, the principle of ensuring that the mask characters differ from the characters in the target preset identifier can be followed, thereby better preventing the target mapping identifier from identifying any entity within the set of entities to which the target entity belongs. Therefore, this application provides a possible implementation method in which the mask characters differ from the characters in the target preset identifier.

[0085] S3033: Replace the target preset identifier in the first collected data with the target mapping identifier to obtain the second collected data.

[0086] As an example, based on the above example, the preset mapping algorithm for one-to-one unidirectional mapping is the "hash-salt algorithm". This algorithm obtains the identifier type and preset characters corresponding to the target preset identifier "2601357921". The identifier type corresponding to "2601357921", the preset characters corresponding to "2601357921", and the target preset identifier are used to construct the string to be mapped. The "hash-salt algorithm" is then used to map this string, resulting in the candidate mapping identifier "37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb". Combining this with the mask characters, including the initial mask character "***" and the final mask character "***", the target mapping identifier is obtained as "***37bdb2293c48e43f7cccc". cb5ac2b90e523a173aef19c6a17e125fb05708354cb***”; Replace “2601357921” in the first collected data “2601357921|1|1” with “***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***”, and the second collected data will be “***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***|1|1”.

[0087] Specifically, the "hash salting algorithm" uses the following formula:

[0088] ***sha256(type|identifier|salt)***

[0089] Here, `type` represents the identifier type corresponding to the target preset identifier, such as application account, mobile phone number, ID card number, merchant account, or contract number; `identifier` represents the target preset identifier; and `salt` represents the preset character corresponding to the target preset identifier, i.e., the salt value corresponding to the target preset identifier. The SHA256 algorithm outputs a fixed 64 characters, which, after adding the header mask "***" and the tail mask "***", becomes 70 characters. This method avoids confusion between the second collected data after initial transformation and the actual first collected data, accurately identifying the second and first collected data and facilitating further transformation processing.

[0090] S304: Based on the correspondence between the target mapping identifier and the target test identifier which is different from the target preset identifier but has target characteristic attributes, replace the target mapping identifier in the second collection data to obtain the third collection data for offline testing.

[0091] In this embodiment of the application, after obtaining the second collection data including the target mapping identifier in S303, since the target mapping identifier cannot be consistent with the target preset identifier for offline testing, that is, the target mapping identifier does not have the target feature attribute associated with the target preset identifier and cannot identify the target subject; therefore, it is also necessary to determine the target test identifier that is different from the target preset identifier and has the target feature attribute, construct the correspondence between the target mapping identifier and the target test identifier, and replace the target mapping identifier in the second collection data so that the target mapping identifier is replaced with the target test identifier, thereby obtaining the third collection data including the target test identifier for offline testing.

[0092] When determining a target test identifier that is different from the target preset identifier and has target characteristic attributes, it is first necessary to obtain a set of test identifiers used for offline testing. Then, it is necessary to determine whether a first test identifier with target characteristic attributes exists in the test identifier set. If a first test identifier exists, it is also necessary to determine whether the usage status of the first test identifier is idle. If the usage status of the first test identifier is idle, it means that a first test identifier with target characteristic attributes has been created and persisted to the test identifier set, and the first test identifier is not occupied. In this case, the first test identifier can be used as the target test identifier.

[0093] Conversely, if the first test identifier is in an occupied state, indicating that it is occupied, a second test identifier with the target characteristic attribute needs to be created as the target test identifier; or, if the first test identifier does not exist in the test identifier set, indicating that it has not been created and persisted to the test identifier set, a second test identifier with the target characteristic attribute also needs to be created as the target test identifier. Therefore, this application provides a possible implementation method, and the steps for determining the target test identifier may include, for example, the following S5-S7:

[0094] S5: Obtain the set of test identifiers for offline testing.

[0095] S6: If there exists a first test identifier with the target characteristic attribute in the set of test identifiers, and the first test identifier is in an idle state, then the first test identifier is determined as the target test identifier.

[0096] S7: If the first test identifier is in an occupied state, or if the first test identifier does not exist in the set of test identifiers, create a second test identifier with the target characteristic attribute, and determine the second test identifier as the target test identifier.

[0097] It is important to note that the target feature attributes include the fields and field values ​​in the data row used to query the first test identifier or create the second test identifier in the test identifier set.

[0098] Furthermore, the character format of the target test identifier is the same as that of the target preset identifier.

[0099] To prevent other preset identifiers with the same associated target feature attributes from being mapped to other mapped identifiers, and thus corresponding to the target test identifier, after executing S6 or S7 to determine the target test identifier, which is different from the target preset identifier but has the target feature attribute, it is also necessary to update the usage state of the target test identifier to the occupied state. This establishes a correspondence between the target mapped identifier and the target test identifier in the occupied state, so as to realize the occupancy of the target test identifier for both the target preset identifier and the target mapped identifier. Therefore, this application provides a possible implementation method, and the steps for constructing the correspondence may include, for example, the following S8-S9:

[0100] S8: Update the usage status of the target test identifier to the occupied status.

[0101] S9: Establish a correspondence between the target mapping identifier and the target test identifier that is in use.

[0102] Furthermore, in this embodiment, considering that other preset identifiers may also be associated with target feature attributes, the second test identifier with target feature attributes created in step S7 is not limited to current use. To facilitate subsequent direct reuse, the second test identifier with target feature attributes can be persisted to the test identifier set, that is, the second test identifier with target feature attributes is solidified into the test environment. Therefore, this application provides a possible implementation method, which may further include step S10: persisting the second test identifier to the test identifier set.

[0103] Furthermore, in this embodiment, considering the recurrence of the target preset identifier in the application scenario of diverting batch collected data stored online to offline testing, the correspondence between the target mapping identifier and the target test identifier constructed in S9 is not limited to current use. To facilitate subsequent direct reuse and save processes, the correspondence between the target mapping identifier and the target test identifier can also be stored in a correspondence table. Based on this, in the specific implementation of S304, firstly, the correspondence between the target mapping identifier and the target test identifier needs to be read from the correspondence table using the target mapping identifier in the second collected data; then, the target test identifier corresponding to the target mapping identifier is determined through the correspondence; finally, the target mapping identifier in the second collected data is replaced with the target test identifier to obtain the third collected data including the target test identifier for offline testing. Therefore, this application provides a possible implementation method, which may also include S11: storing the correspondence in a correspondence table; correspondingly, S304 may include the following S3041-S3043:

[0104] S3041: Read the corresponding relationship from the corresponding relationship table based on the target mapping identifier.

[0105] S3042: Determine the target test identifier corresponding to the target mapping identifier based on the correspondence.

[0106] S3043: Replace the target mapping identifier in the second collected data with the target test identifier to obtain the third collected data.

[0107] As an example, based on the above example, a target test identifier "1309853157" is determined to be different from the target preset identifier "2601357921" and has target feature attributes "feature attribute 1, feature attribute 2, ..., feature attribute N". A correspondence is established between the target mapping identifier "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***" and "1309853157". The second collected data "***37bdb2293c48e43f7cc" is then used to establish the correspondence. In the string "cccb5ac2b90e523a173aef19c6a17e125fb05708354cb***|1|1", replace "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***" with "1309853157" to obtain the third data collection used for offline testing as "1309853157|1|1".

[0108] The method for processing preset identifiers in the data provided in the above embodiments firstly extracts a target preset identifier for identifying the target subject from the first collected data stored online; secondly, based on the business scenario of the first collected data, queries the target feature attributes associated with the target preset identifier that cannot identify the target subject; then, through a preset mapping algorithm of one-to-one unidirectional mapping, the target preset identifier in the first collected data is mapped to obtain the second collected data, which includes a target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs; finally, the target mapping identifier in the second collected data is replaced with a target test identifier that is different from the target preset identifier but has target feature attributes, based on the correspondence between the target mapping identifier and the target test identifier that has target feature attributes, to obtain the third collected data for offline testing.

[0109] As can be seen, for the first collection data stored online, additional querying is performed on the target feature attributes associated with the target preset identifier that cannot identify the target subject. After obtaining the second collection data by using a preset mapping algorithm with one-to-one unidirectional mapping, the target preset identifier is mapped to a target mapping identifier that cannot identify any subject. Then, the target mapping identifier is replaced with a target test identifier that is different from the target preset identifier but has target feature attributes, resulting in the third collection data for offline testing. Based on this, the method ensures that the target test identifier has both the target feature attributes of the target preset identifier and cannot be traced back to the target subject identified by the target preset identifier. It also has a wider range of applicable scenarios, thereby improving the reliability, security, and applicability of the collection data from online storage to offline testing, thus improving the subsequent offline testing results.

[0110] In addition, this method improves the diversity and richness of the collected data transferred from online storage to offline testing, making subsequent offline testing more comprehensive, thereby reducing the cost of subsequent offline testing and improving the quality of subsequent offline testing.

[0111] In summary, see Figure 4 The diagram illustrates the specific process stages for processing preset identifiers in data. The processing of preset identifiers in data mainly consists of the following seven stages:

[0112] Phase 1: Extracting Preset Identifiers from the Data. For the first batch of collected data stored online, preset identifiers identifying the target entity are extracted from the first batch of collected data.

[0113] For example, based on the data format of the first collected data “2601357921|1|1” and the preset separator “|” in “2601357921|1|1”, “2601357921” is extracted from “2601357921|1|1”, and “2601357921” identifies the target subject “subject b”.

[0114] Phase Two: Querying the Feature Attributes Associated with the Preset Identifier. Based on the business scenario of the first collected data, for the aforementioned target preset identifier, query the target feature attributes associated with the target preset identifier that cannot identify the target entity.

[0115] For example, in the business scenario of the first collected data "2601357921|1|1", the target feature attributes associated with the target preset identifier "2601357921" that cannot identify "subject b" are "feature attribute 1, feature attribute 2, ..., feature attribute N".

[0116] The third stage: Preset identifiers in the mapping data. For the first collected data mentioned above, a preset mapping algorithm with one-to-one unidirectional mapping is used to map the target preset identifiers in the first collected data to obtain the second collected data. The second collected data includes target mapping identifiers that cannot identify any entity in the set of entities to which the target entity belongs.

[0117] For example, by using the "hash salting algorithm" to replace "2601357921" in the first collected data "2601357921|1|1" with "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***", the second collected data is "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***|1|1".

[0118] Since the target preset identifier and target feature attribute are represented in key-value pair format, a one-to-one unidirectional mapping algorithm is also needed to map the target preset identifier in the key-value pair format to the target preset identifier in the target feature attribute, resulting in the target mapping identifier and target feature attribute in key-value pair format. This method ensures a complete correspondence between the target preset identifier, target mapping identifier, and target feature attribute.

[0119] For example, by using the "hash salting algorithm", the "2601357921" in the target preset identifier and target feature attribute "2601357921: {feature attribute 1, feature attribute 2, ..., feature attribute N}" represented in key-value pair format is replaced with "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***", resulting in the target mapping identifier and target feature attribute "***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***{feature attribute 1, feature attribute 2, ..., feature attribute N}" represented in key-value pair format.

[0120] Phase 4: Determine the test identifier based on the characteristic attributes. For the aforementioned target characteristic attributes, obtain a set of test identifiers for offline testing. If a first test identifier with the target characteristic attribute exists in the set and its usage status is idle, the first test identifier is determined as the target test identifier. If the first test identifier is in use, or if the first test identifier does not exist in the set, a second test identifier with the target characteristic attribute is created, and the second test identifier is determined as the target test identifier.

[0121] For example, the target test identifier, which is different from the target preset identifier "2601357921" and has target characteristic attributes "characteristic attribute 1, characteristic attribute 2, ..., characteristic attribute N", is determined to be "1309853157".

[0122] Phase 5: Mark the test identifier as occupied. For the target test identifier mentioned above, update its usage status to occupied.

[0123] Phase 6: Establishing a correspondence based on test identifiers. For the target mapping identifier and target test identifier mentioned above, establish a correspondence between the target mapping identifier and the target test identifier that is in use.

[0124] For example, a correspondence is established between the target mapping identifiers “***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***” and “1309853157”, and the correspondence is stored in the correspondence table.

[0125] Phase 7: Replace the pre-defined identifiers mapped in the data. For the second collected data mentioned above, the correspondence between the target mapping identifier and the target test identifier is read from the correspondence table using the target mapping identifier in the second collected data; the target test identifier corresponding to the target mapping identifier is determined through the correspondence; the target mapping identifier is used to replace the target mapping identifier in the second collected data, resulting in third collected data including the target test identifier for offline testing.

[0126] For example, by using the above correspondence in the correspondence table, replacing “***37bdb2293c48e43f7cccccb5ac2b90e523a173aef19c6a17e125fb05708354cb***|1|1” in the second collected data with “1309853157”, we obtain the third collected data for offline testing as “1309853157|1|1”.

[0127] In addition to the method for processing preset identifiers in data described above, this application also provides a processing device for preset identifiers in data. The processing device for preset identifiers in data provided in this application will be described in detail below.

[0128] See Figure 5 , Figure 5 This is a schematic diagram of a processing device for preset identifiers in data, provided as an embodiment of this application. Figure 5 As shown, the data processing device 500 for preset identifiers includes: an extraction unit 501, a query unit 502, a mapping unit 503, and a replacement unit 504.

[0129] Extraction unit 501 is used to extract a target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject;

[0130] The query unit 502 is used to query the target feature attributes associated with the target preset identifier based on the business scenario of the first collected data; the target feature attributes cannot identify the target subject;

[0131] The mapping unit 503 is used to map the target preset identifier in the first collected data according to the preset mapping algorithm of one-to-one unidirectional mapping to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs.

[0132] The replacement unit 504 is used to replace the target mapping identifier in the second collection data according to the correspondence between the target mapping identifier and the target test identifier which is different from the target preset identifier and has target characteristic attributes, so as to obtain the third collection data for offline testing.

[0133] As one possible implementation, the device further includes: a determining unit;

[0134] Determine the unit, used for:

[0135] Obtain the set of test identifiers for offline testing;

[0136] If there exists a first test identifier with the target characteristic attribute in the set of test identifiers, and the first test identifier is in an idle state, then the first test identifier is determined as the target test identifier.

[0137] If the first test identifier is in an occupied state, or if the first test identifier does not exist in the set of test identifiers, a second test identifier with the target characteristic attribute is created, and the second test identifier is determined as the target test identifier.

[0138] As one possible implementation, the target feature attributes include a data table and data rows in the data table, where each data row includes fields and field values; the fields and field values ​​are used to query a first test identifier or create a second test identifier in the set of test identifiers.

[0139] As one possible implementation, the device also includes: a building block;

[0140] Building blocks, used for:

[0141] Update the usage status of the target test identifier to the occupied status;

[0142] Establish a correspondence between the target mapping identifier and the target test identifier that is in use.

[0143] As one possible implementation, mapping unit 503 is specifically used for:

[0144] Obtain the identifier type and preset character corresponding to the target preset identifier;

[0145] The identifier type, target preset identifier, and preset character are mapped according to a preset mapping algorithm to obtain the target mapping identifier;

[0146] The target preset identifier in the first collected data is replaced with the target mapping identifier to obtain the second collected data.

[0147] As one possible implementation, mapping unit 503 is specifically used for:

[0148] The identifier type, the target preset identifier, and the preset characters are used to construct the string to be mapped;

[0149] The string to be mapped is mapped according to a preset mapping algorithm to obtain candidate mapping identifiers;

[0150] If the candidate mapping identifier cannot identify any subject, the candidate mapping identifier shall be determined as the target mapping identifier;

[0151] If the candidate mapping identifier identifies other entities in the entity set, the target mapping identifier is obtained based on the candidate mapping identifier and the mask character.

[0152] As one possible implementation, the mask characters differ from the characters in the target preset identifier.

[0153] As one possible implementation, the apparatus further includes: a generation unit;

[0154] Generation unit, used for:

[0155] Generate preset characters based on target feature attributes; or,

[0156] Generate a preset character based on the target preset identifier; or,

[0157] Randomly generate preset characters.

[0158] As one possible implementation, the device also includes: a persistence unit;

[0159] The persistence unit is used to persist the second test identifier to the test identifier set.

[0160] As one possible implementation, the device also includes: a storage unit;

[0161] Storage unit, used to store the correspondence into the correspondence table;

[0162] Replacement unit 504, specifically used for:

[0163] Based on the target mapping identifier, read the corresponding relationship from the corresponding relationship table;

[0164] Based on the correspondence, determine the target test identifier corresponding to the target mapping identifier;

[0165] Replace the target mapping identifier in the second data collection with the target test identifier to obtain the third data collection.

[0166] As one possible implementation, extraction unit 501 is specifically used for:

[0167] Based on the data format of the first collected data and the preset delimiter in the first collected data, the target preset identifier is extracted from the first collected data.

[0168] The data processing apparatus for preset identifiers provided in the above embodiments first extracts a target preset identifier for identifying a target subject from the first collected data stored online; second, based on the business scenario of the first collected data, queries the target feature attributes associated with the target preset identifier that cannot identify the target subject; then, through a preset mapping algorithm of one-to-one unidirectional mapping, maps the target preset identifier in the first collected data to obtain second collected data, the second collected data including a target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs; finally, based on the correspondence between the target mapping identifier and a target test identifier that is different from the target preset identifier but has target feature attributes, the target mapping identifier in the second collected data is replaced to obtain third collected data for offline testing.

[0169] As can be seen, for the first collection data stored online, additional querying is performed on the target feature attributes associated with the target preset identifier that cannot identify the target subject. After obtaining the second collection data by using a preset mapping algorithm with one-to-one unidirectional mapping, the target preset identifier is mapped to a target mapping identifier that cannot identify any subject. Then, the target mapping identifier is replaced with a target test identifier that is different from the target preset identifier but has target feature attributes, resulting in the third collection data for offline testing. Based on this, the method ensures that the target test identifier has both the target feature attributes of the target preset identifier and cannot be traced back to the target subject identified by the target preset identifier. It also has a wider range of applicable scenarios, thereby improving the reliability, security, and applicability of the collection data from online storage to offline testing, thus improving the subsequent offline testing results.

[0170] In response to the processing method of preset identifiers in the data described above, this application embodiment also provides a device for determining content joint recommendation, so that the above processing method of preset identifiers in the data can be implemented and applied in practice. The computer device provided in this application embodiment will be introduced from the perspective of hardware physicalization below.

[0171] See Figure 6 , Figure 6This is a schematic diagram of a server structure provided in an embodiment of this application. The server 600 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 622 (e.g., one or more processors) and memory 632, and one or more storage media 630 (e.g., one or more mass storage devices) for storing application programs 642 or data 644. The memory 632 and storage media 630 can be temporary or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the server. Furthermore, the CPU 622 may be configured to communicate with the storage media 630 and execute the series of instruction operations in the storage media 630 on the server 600.

[0172] Server 600 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0173] The steps performed by the server in the above embodiments can be based on this Figure 6 The server structure shown.

[0174] CPU 622 is used to perform the following steps:

[0175] Extract the target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject.

[0176] Based on the business scenario of the first collected data, query the target feature attributes associated with the target preset identifier; the target feature attributes cannot identify the target subject.

[0177] According to the preset mapping algorithm of one-to-one unidirectional mapping, the target preset identifier in the first collected data is mapped to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs.

[0178] Based on the correspondence between the target mapping identifier and the target test identifier, which is different from the target preset identifier but has target characteristic attributes, the target mapping identifier in the second collection data is replaced to obtain the third collection data for offline testing.

[0179] Optionally, the CPU 622 may also execute method steps of any specific implementation of the method for processing preset identifiers in the data in the embodiments of this application.

[0180] See Figure 7 , Figure 7 This is a schematic diagram of a terminal device provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiment of this application are shown; for specific technical details not disclosed, please refer to the method section of the embodiment of this application. The terminal device can be any terminal device including mobile phones, tablets, PDAs, etc. Taking a mobile phone as an example:

[0181] Figure 7 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 7 The mobile phone includes components such as a radio frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a Wi-Fi module 770, a processor 780, and a power supply 790. Those skilled in the art will understand that... Figure 7 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0182] The following is combined Figure 7 A detailed introduction to each component of a mobile phone:

[0183] RF circuit 710 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 780; additionally, it transmits uplink data to the base station. Typically, RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), and a duplexer. Furthermore, RF circuit 710 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).

[0184] The memory 720 can be used to store software programs and modules. The processor 780 runs the software programs and modules stored in the memory 720 to realize various functions and data processing of the mobile phone. The memory 720 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0185] The input unit 730 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 731), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 780, and can also receive and execute commands sent by the processor 780. In addition, the touch panel 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0186] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741, which may optionally be configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel. Further, a touch panel 731 may cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it transmits the information to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides corresponding visual output on the display panel 741 based on the type of touch event. Although in Figure 7 In this embodiment, the touch panel 731 and the display panel 741 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0187] The mobile phone may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can turn off the display panel 741 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0188] Audio circuit 760, speaker 761, and microphone 762 provide an audio interface between the user and the mobile phone. Audio circuit 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. On the other hand, microphone 762 converts collected sound signals into electrical signals, which are received by audio circuit 760, converted into audio data, and then processed by processor 780 before being transmitted via RF circuit 710 to, for example, another mobile phone, or the audio data can be output to memory 720 for further processing.

[0189] WiFi is a short-range wireless transmission technology. Through the WiFi module 770, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 7 The WiFi module 770 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.

[0190] The processor 780 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, thereby controlling the phone as a whole. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.

[0191] The mobile phone also includes a power supply 790 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 780 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0192] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0193] In this embodiment of the application, the memory 720 included in the mobile phone can store program code and transmit the program code to the processor.

[0194] The processor 780 included in the mobile phone can execute the processing method of the preset identifier in the data provided in the above embodiments according to the instructions in the program code.

[0195] This application also provides a computer-readable storage medium for storing a computer program that executes the processing method for preset identifiers in the data provided in the above embodiments.

[0196] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the processing method for a preset identifier in data provided in various optional implementations of the above aspects.

[0197] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., and other media capable of storing program code.

[0198] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0199] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for processing preset identifiers in data, characterized in that, The method includes: Extract a target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject. Based on the business scenario of the first collected data, query the target feature attributes associated with the target preset identifier; the target feature attributes cannot identify the target subject; According to a preset mapping algorithm for one-to-one unidirectional mapping, the target preset identifier in the first collected data is mapped to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs; Based on the correspondence between the target mapping identifier and the target test identifier, which is different from the target preset identifier but has the target characteristic attribute, the target mapping identifier in the second collected data is replaced to obtain the third collected data for offline testing.

2. The method according to claim 1, characterized in that, The steps for determining the target test identifier include: Obtain the set of test identifiers for offline testing; If the set of test identifiers contains a first test identifier that has the target feature attribute, and the first test identifier is in an idle state, then the first test identifier is determined as the target test identifier. If the first test identifier is in an occupied state, or if the first test identifier does not exist in the set of test identifiers, a second test identifier with the target feature attribute is created, and the second test identifier is determined as the target test identifier.

3. The method according to claim 2, characterized in that, The target feature attributes include a data table and data rows in the data table, and the data rows include fields and field values; the fields and field values ​​are used to query the first test identifier or create the second test identifier in the test identifier set.

4. The method according to claim 2, characterized in that, The steps for constructing the correspondence include: Update the usage status of the target test identifier to the occupied status; The correspondence is established between the target mapping identifier and the target test identifier in the occupied state.

5. The method according to claim 1, characterized in that, The step of mapping the target preset identifier in the first collected data according to a preset mapping algorithm of one-to-one unidirectional mapping to obtain the second collected data includes: Obtain the identifier type and preset character corresponding to the target preset identifier; The target mapping identifier is obtained by mapping the identifier type, the target preset identifier, and the preset character according to the preset mapping algorithm. The target preset identifier in the first collected data is replaced with the target mapping identifier to obtain the second collected data.

6. The method according to claim 5, characterized in that, The step of mapping the identifier type, the target preset identifier, and the preset character according to the preset mapping algorithm to obtain the target mapped identifier includes: The identifier type, the target preset identifier, and the preset character are used to form the string to be mapped; The string to be mapped is mapped according to the preset mapping algorithm to obtain candidate mapping identifiers; If the candidate mapping identifier cannot identify any subject, the candidate mapping identifier will be determined as the target mapping identifier; If the candidate mapping identifier identifies another entity in the entity set, the target mapping identifier is obtained based on the candidate mapping identifier and the mask character.

7. The method according to claim 6, characterized in that, The mask characters are different from the characters in the target preset identifier.

8. The method according to claim 5, characterized in that, The steps for generating the preset character are as follows: Generate the preset character based on the target feature attributes; or, Generate the preset character based on the target preset identifier; or, The preset characters are generated randomly.

9. The method according to claim 2, characterized in that, The method further includes: The second test identifier is persisted to the set of test identifiers.

10. The method according to claim 4, characterized in that, The method further includes: Store the correspondence in a correspondence table; The step of replacing the target mapping identifier in the second collected data according to the correspondence between the target mapping identifier and the target test identifier which is different from the target preset identifier but has the target characteristic attribute, to obtain the third collected data for offline testing, includes: The correspondence is read from the correspondence table based on the target mapping identifier; Based on the correspondence, determine the target test identifier corresponding to the target mapping identifier; The target mapping identifier in the second collected data is replaced with the target test identifier to obtain the third collected data.

11. The method according to any one of claims 1-10, characterized in that, The extraction of the target preset identifier from the first collected data stored online specifically involves: Based on the data format of the first collected data and the preset delimiter in the first collected data, the target preset identifier is extracted from the first collected data.

12. A processing apparatus for pre-defined identifiers in data, characterized in that, The device includes: an extraction unit, a query unit, a mapping unit, and a replacement unit; The extraction unit is used to extract a target preset identifier from the first collected data stored online; the target preset identifier is used to identify the target subject; The query unit is used to query the target feature attribute associated with the target preset identifier based on the business scenario of the first collected data; the target feature attribute cannot identify the target subject; The mapping unit is used to map the target preset identifier in the first collected data according to a preset mapping algorithm of one-to-one unidirectional mapping to obtain the second collected data; the second collected data includes the target mapping identifier that cannot identify any subject in the subject set to which the target subject belongs; The replacement unit is used to replace the target mapping identifier in the second collected data according to the correspondence between the target mapping identifier and the target test identifier which is different from the target preset identifier but has the target characteristic attribute, so as to obtain the third collected data for offline testing.

13. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the processing method for the preset identifier in the data as described in any one of claims 1-11, according to the instructions in the program code.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a processor, performs the processing method for a preset identifier in data as described in any one of claims 1-11.

15. A computer program product, characterized in that, It includes a computer program or instructions; when the computer program or instructions are executed by a processor, the processing method for the preset identifier in the data as described in any one of claims 1-11 is performed.

Citation Information

Patent Citations

  • Order data processing method and device, computer readable medium and electronic equipment

    CN111709817A

  • Apparatus and method for data matching and anonymization

    US8978153B1