Data processing method and related equipment
By filtering the code text to be processed from the safe words in the non-hard-coded code text, the problems of misjudgment and low efficiency in the existing hard-coded detection are solved, and more efficient hard-coded recognition and removal are achieved.
Patent Information
- Application Number
- CN202410697918.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies suffer from problems such as false positives and false negatives when detecting hard-coded data, inability to detect complex string and non-string types, and inability to detect implicit hard-coded data, resulting in low efficiency and poor accuracy in hard-coded data recognition.
By obtaining safe words from non-hard-coded code text, these safe words are used to filter the code text to be processed, identify and remove non-safe words, and adopt a multi-level filtering and safe word root method to reduce the difficulty of keyword enumeration and improve recognition efficiency.
It improves the accuracy and efficiency of hard-coded recognition, reduces missed detections and keyword enumeration difficulties, and simplifies the code text processing process.
Smart Images

Figure CN121052246A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method and related equipment. Background Technology
[0002] During software system development, some developers may choose to directly write numerical or text values of constants, variables, parameters, etc. into the program code due to negligence or for development convenience. Although hard-coding of general parameters does not conform to coding standards, it does not directly affect data security or software system security. However, hard-coding sensitive data into the code will bring the following harms: security risks, legal risks, brand image, and technical risks.
[0003] Common hard-coded string detection methods often use predefined regular expression patterns to match sensitive information in the code, such as passwords and phone numbers. However, the above regular expression matching rules have the following problems: 1. They may falsely identify or miss some hard-coded strings. 2. They can only match some simple patterns; if the hard-coded string is complex, it is difficult to use regular expressions for detection. 3. They can only detect strings in the code, not other types of hard-coded strings, such as numbers and boolean values. 4. They can only detect explicit hard-coded strings in the code, not implicit hard-coded strings, such as strings generated by string concatenation.
[0004] Therefore, how to improve the efficiency and accuracy of hard-coded code text recognition is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a data processing method and related equipment. By simplifying code text using secure words, it can assist in the removal of secure text. That is, it simplifies the code text to be processed from the perspective of secure text, without altering sensitive information in the code text, thus reducing missed detections. It can also reduce the enumeration difficulties caused by using non-secure words for non-secure word identification in the prior art.
[0006] This application provides a data processing method, which is executed by a data processing device, or by a component (e.g., a processor, chip, or chip system) of the data processing device, or by a logic module or software capable of implementing all or part of the functions of the data processing device. In this first aspect and its possible implementations, the method is described as being executed by a data processing device. In this method, the data processing device first obtains a first set of safe words from a set of code text to be processed and non-hard-coded code text. The data processing device then filters the code text to be processed based on the first set of safe words. The data processing device then identifies non-safe words in the filtered code text to be processed.
[0007] In this context, the first secure term corresponds to the non-secure term. The first secure term can be understood as non-sensitive data or non-privacy data, while the non-secure term can be understood as sensitive data or privacy data, etc.
[0008] Based on the above scheme, a first safe word is obtained from the non-hard-coded code text, and the code text to be processed is filtered based on this first safe word. Filtering with the first safe word simplifies the code text to be processed, thereby reducing the workload of identifying non-safe words and improving recognition efficiency. On the one hand, compared to the existing technology of matching sensitive information in code using predefined regular expression patterns, this application simplifies the code text by using safe words, which helps to remove safe text. That is, it simplifies the code text to be processed from the perspective of safe text, without modifying the sensitive information in the code text, thus reducing missed detections. On the other hand, compared to identifying sensitive data in code by keywords, this reduces the difficulty of keyword enumeration. Furthermore, compared to the existing technology that requires manual or tool scanning to check whether sensitive data is hard-coded in the code, this application uses safe words from the non-hard-coded code text to assist manual identification or, in some scenarios, replace manual identification of hard-coded sensitive data, thereby improving the recognition efficiency of non-safe words and other sensitive words.
[0009] In one possible implementation, the non-hard-coded code text set includes multiple words and / or multiple word roots, and the first safe word includes at least one of the following: a word among the multiple words that satisfies a first preset condition, and a word root among the multiple word roots that satisfies a second preset condition.
[0010] Based on the above scheme, by referring to words and / or word roots that meet the preset conditions in the non-hard-coded text during the filtering process, the filtering process can be made more accurate or reasonable.
[0011] In one possible implementation, the words that satisfy the first preset condition among the multiple words include at least one of the following: words whose occurrence frequency is greater than or equal to the first preset threshold among the multiple words, words whose occurrence frequency is greater than or equal to the second preset threshold among the multiple word roots, and the top k words after sorting the multiple words according to their occurrence frequency / occurrence from high to low, where k is a positive integer greater than 0.
[0012] Based on the above scheme, it can be understood that it limits the specific circumstances of safe words among multiple words. Safe words can be determined by a first preset condition related to the frequency of occurrence, which can provide more reasonable references for subsequent filtering of text to be processed through safe words and improve the filtering effect.
[0013] In one possible implementation, the word roots among the multiple words that satisfy the second preset condition include at least one of the following: word roots whose occurrence frequency is greater than or equal to the third preset threshold, words whose occurrence frequency is greater than or equal to the fourth preset threshold, and the top j word roots after sorting the multiple word roots from high to low according to their occurrence frequency, where j is a positive integer greater than 0.
[0014] Based on the above scheme, it can be understood that it limits the specific circumstances of safe word roots among multiple words. Safe word roots can be determined through a second preset condition related to the frequency of occurrence, thereby providing more reasonable references for subsequent filtering of text to be processed through safe word roots and improving the filtering effect.
[0015] In one possible implementation, the data processing device specifically filters the first safe word in the code text to be processed.
[0016] Based on the above scheme, it can be understood that the data processing device filters the code text to be processed through the first safe word. This not only avoids modifying the unsafe words in the code text, thus reducing missed detections, but also reduces the enumeration difficulties caused by using keywords of unsafe words for unsafe word identification in existing technologies.
[0017] In one possible implementation, the first security word includes: a word among multiple words that satisfies a first preset condition and a root word among multiple root words that satisfies a second preset condition; the data processing device specifically filters the words in the code text to be processed that satisfy the first preset condition to obtain intermediate code text; the data processing device filters the root words in the intermediate code text that satisfy the second preset condition.
[0018] Based on the above scheme, it can be understood as multi-level filtering. First, the code text to be processed is filtered using safe words to obtain intermediate code text. Then, the intermediate code text is filtered again using safe word roots. Multi-level filtering improves the filtering effect, thereby reducing the amount of text to be recognized and improving subsequent recognition efficiency.
[0019] In one possible implementation, at least one of the multiple words uses a preset syntax in the encoding format (or can be understood as a code format), such as nested syntax in Java.
[0020] Based on the above scheme, since the preset syntax in the code format is often unrelated to privacy, referring to words in the preset syntax when filtering code text can reduce the workload of subsequent identification of insecure words and improve identification efficiency.
[0021] In one possible implementation, the multiple word roots include at least one of the following: service noun root, class noun root, and interface class noun root.
[0022] Based on the above scheme, since service terminology roots, class terminology roots, and interface terminology roots are often unrelated to privacy, referring to each terminology root when filtering code text can reduce the workload of subsequent identification of insecure words and improve identification efficiency.
[0023] A second aspect of this application provides a data processing device, which may be a data processing device in its entirety, or a component of a data processing device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the functions of a data processing device. The data processing device includes an acquisition module and a processing module.
[0024] The acquisition module is used to acquire the text of the code to be processed.
[0025] The acquisition module is also used to acquire the first safe word in a set of non-hard-coded code text.
[0026] The processing module is used to filter the code text to be processed based on the first security word;
[0027] The processing module is also used to identify unsafe words in the filtered code text to be processed.
[0028] In one possible implementation, the non-hard-coded code text set includes multiple words and / or multiple word roots, and the first safe word includes at least one of the following: a word among the multiple words that satisfies a first preset condition, and a word root among the multiple word roots that satisfies a second preset condition.
[0029] In one possible implementation, the words that satisfy the first preset condition among the multiple words include at least one of the following: words whose occurrence frequency is greater than or equal to the first preset threshold among the multiple words, words whose occurrence frequency is greater than or equal to the second preset threshold among the multiple word roots, and the top k words after sorting the multiple words according to their occurrence frequency / occurrence from high to low, where k is a positive integer greater than 0.
[0030] In one possible implementation, the word roots among the multiple words that satisfy the second preset condition include at least one of the following: word roots whose occurrence frequency is greater than or equal to the third preset threshold, words whose occurrence frequency is greater than or equal to the fourth preset threshold, and the top j word roots after sorting the multiple word roots from high to low according to their occurrence frequency, where j is a positive integer greater than 0.
[0031] In one possible implementation, the processing module is specifically used to filter the first safe word in the code text to be processed.
[0032] In one possible implementation, the first safe word includes: a word among multiple words that satisfies a first preset condition and a root word among multiple root words that satisfies a second preset condition; the processing module is specifically used to filter the words in the code text to be processed that satisfy the first preset condition to obtain intermediate code text; the processing module is specifically used to filter the root words in the intermediate code text that satisfy the second preset condition.
[0033] In one possible implementation, at least one of the multiple words uses the preset syntax in the encoding format.
[0034] In one possible implementation, the multiple word roots include at least one of the following: service noun root, class noun root, and interface class noun root.
[0035] Thirdly, a data processing apparatus is provided, including at least one processor coupled to a memory; the memory is used to store a program or instructions; the at least one processor is used to execute the program or instructions to enable the data processing apparatus to implement any of the possible implementations of the first aspect.
[0036] Fourthly, a data processing device is provided, including at least one logic circuit and an input / output interface; the logic circuit is used to perform the method as described in any of the possible implementations of the first aspect above.
[0037] Fifthly, a computer-readable storage medium is provided for storing one or more computer-executable instructions, which, when executed by a processor, perform the method as described in any possible implementation of any of the first aspects above.
[0038] In a sixth aspect, a computer program product is provided, wherein when a computer program in the computer program product is executed by the processor, the processor executes the method described in any possible implementation of any of the first aspects described above.
[0039] In a seventh aspect, a chip or chip system is provided, the chip or chip system including at least one processor for supporting a communication device to implement the method described in any possible implementation of any of the first aspects above.
[0040] In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the communication device. The chip system may be composed of chips or may include chips and other discrete devices. Optionally, the chip system may also include interface circuitry that provides program instructions and / or data to at least one processor.
[0041] Eighthly, a data processing apparatus is provided, comprising a chip system as described in the seventh aspect above, the chip system including a processor and a communication interface for communicating with a module outside the chip system, the processor for running computer programs or instructions such that the data processing apparatus can perform the methods of any of the above aspects.
[0042] A ninth aspect provides a data processing device cluster, comprising at least one data processing device as described in the second or eighth aspect, wherein any one data processing device is configured to run a computer program or instructions, enabling the data processing device cluster to perform the methods of any of the above aspects. Alternatively, some or all of the data processing devices may be used together to run a computer program or instructions, enabling the data processing device cluster to perform the methods of any of the above aspects.
[0043] The technical effects of any of the design methods in aspects two through nine can be found in the technical effects of the different design methods in aspect one above, and will not be repeated here.
[0044] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0045] Figure 1A The system architecture diagram provided in this application is shown below.
[0046] Figure 1B Another structural diagram of the system architecture provided in this application;
[0047] Figure 2 A flowchart illustrating the data processing method provided in this application;
[0048] Figures 3 to 5 Here are some schematic diagrams of the data processing equipment involved in this application;
[0049] Figure 6 This application provides a schematic diagram of the structure of a data processing equipment cluster.
[0050] Figure 7 This is a schematic diagram of another data processing equipment cluster provided in this application. Detailed Implementation
[0051] To facilitate understanding, the relevant terms and concepts mainly involved in this application will be introduced below.
[0052] 1. Hardcoding and Non-Hardcoding
[0053] Hardcoding: Hardcoding refers to directly writing specific data, configuration information, or constants into the code during program writing. In software implementation, hardcoding means directly writing information needed during program runtime into the code. For example, if a program needs to provide a contact number for a service exception resolution manager when the software service encounters an error, normally this number should be written into a configuration file and retrieved from the configuration file when an error occurs. With this method, if the manager's phone number changes, only the configuration file needs to be modified. Hardcoding the phone number, however, means directly writing the phone number into the code; any changes require code modification to maintain functionality.
[0054] Non-hard-coded: This refers to setting and retrieving data and configuration information used by a program through external configuration files, environment variables, command-line arguments, databases, or user input. Non-hard-coded methods allow data to be determined and modified at runtime, thus providing greater flexibility and maintainability.
[0055] Note: The code text is not limited to hard-coded and non-hard-coded text; it can also include business logic code, etc. The specifics are not limited here.
[0056] 2. Words and word roots in the code text
[0057] In code text, a root word is a part of a word, or can be understood as a part of the key meaning of a word. Generally, a root word can refer to a prefix or suffix of a word.
[0058] During software system development, some developers may choose to directly write numerical or text values of constants, variables, parameters, etc. into the program code due to negligence or for development convenience. Although hard-coding of general parameters does not conform to coding standards, it does not directly affect data security or software system security. However, hard-coding sensitive data into the code will bring the following harms: security risks, legal risks, brand image, and technical risks.
[0059] Common hard-coded detection schemes mainly fall into two categories. One category uses predefined regular expression patterns to match sensitive information in the code, such as passwords and phone numbers. The other category uses a predefined list of keywords to match sensitive information in the code, such as database connection information and Application Programming Interface (API) keys.
[0060] However, the first type of detection scheme has the following problems: 1. It may misjudge or miss some hard-coded strings. 2. It can only match some simple patterns; if the hard-coded string is complex, it is difficult to use regular expressions for detection. 3. It can only detect strings in the code, and cannot detect other types of hard-coded strings, such as numbers and boolean values. 4. It can only detect explicit hard-coded strings in the code, and cannot detect implicit hard-coded strings, such as strings generated by string concatenation. The second type of detection scheme has the following problems: 1. It needs to enumerate as many keywords as possible for all types of sensitive data. 2. The false positive rate is high; most sensitive data keywords are not specific terms for sensitive data, so keyword matching for sensitive data has a high false positive rate. 3. The detection efficiency is low; the large amount of code leads to reduced efficiency when using keyword matching directly.
[0061] Therefore, how to improve the efficiency and accuracy of hard-coded code text recognition is a technical problem that urgently needs to be solved.
[0062] To address the aforementioned technical problems, this application provides a data processing method that obtains a first secure word from the non-hard-coded code text and filters the code text to be processed based on this first secure word. Filtering with the first secure word simplifies the code text to be processed, thereby reducing the workload of identifying non-secure words and improving recognition efficiency. On one hand, compared to existing methods that use predefined regular expression patterns to match sensitive information in code, this application's method of simplifying code text with secure words helps to remove secure text. That is, it simplifies the code text from a secure text perspective without altering sensitive information, thus reducing false negatives. On the other hand, compared to identifying sensitive data in code using keywords, this method reduces the difficulty of keyword enumeration. Furthermore, compared to existing methods that require manual or tool-based scanning to check for hard-coded sensitive data, this application uses secure words from the non-hard-coded code text to assist manual identification or, in some scenarios, replace manual identification of hard-coded sensitive data, thereby improving the efficiency of identifying non-secure words and other sensitive words.
[0063] The data processing method provided in this application is described in detail below with reference to the accompanying drawings.
[0064] Figure 1A A schematic diagram of the data processing system provided in this application, the data processing system including terminal equipment ( Figure 1A(Taking only mobile phones as an example) and cloud devices. It's understandable that terminal devices can be more than just mobile phones; they can also include tablets, portable game consoles, PDAs (personal digital assistants), laptops, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, in-vehicle media players, wearable electronic devices, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, vehicles, in-vehicle terminals, aircraft terminals, intelligent robots, and other terminal devices. The terminal device is the initiator of data processing; as the initiator of data processing requests, requests are typically initiated by the user through the terminal device.
[0065] The aforementioned cloud devices can be cloud servers, network servers, application servers, management servers, or other devices or servers with data processing capabilities. The cloud device receives data processing requests from terminal devices through an interactive interface. It then uses a storage device for storing the data and a data processing processor (e.g., calling an external interface to obtain a preset code text set containing at least one non-hard-coded code text) to filter the code text to be processed based on safe words in the preset code text set. It also identifies unsafe words (or sensitive data) in the filtered code text (e.g., data training / machine learning / deep learning, search / reasoning / decision making). The storage device in the cloud device can be a general term, including local storage and a database storing historical data. The database can be located on the cloud device or on other network servers.
[0066] exist Figure 1A In the data processing system shown, the terminal device can receive user instructions. For example, the terminal device can obtain a data processing request input by the user and then forward the data processing request to the cloud device. This causes the cloud device to execute the data processing method provided in this application to obtain unsafe words in the code text to be processed. Specifically, the cloud device can determine the code text to be processed based on the data processing request (e.g., the data processing request carries the code text to be processed; or, for example, the information carried in the data processing request can be used to determine the code text to be processed). It then filters the code text to be processed using safe words in the non-hard-coded code text, thereby identifying unsafe words in the filtered code text. Furthermore, the cloud device can also present a structured list of unsafe words to the user through the terminal device.
[0067] Figure 1B Another structural diagram of the data processing system provided in this application, in Figure 1BIn the middle, terminal equipment ( Figure 1B (Taking a mobile phone as an example, the terminal device can directly execute the data processing method provided in this application.) That is, the terminal device can directly process data processing requests, and the specific process is similar to... Figure 1A Similar to the description above, it will not be repeated here.
[0068] Optionally, in Figure 1B In the data processing system shown, the terminal device can receive user instructions, such as obtaining the user's data processing request, and then executing the data processing method provided in this application to obtain unsafe words (e.g., data training / machine learning / deep learning, search / reasoning / decision making, etc.) in the code text to be processed. The terminal device can also present a structured list of unsafe terms to the user.
[0069] Figure 1A and Figure 1B The processor can filter the code text to be processed based on safe words in the non-hard-coded code text, and then identify unsafe words in the filtered code text. Specifically, the method for identifying unsafe words can be through natural language processing (NLP) models or regular expression recognition, etc., which are not limited here.
[0070] It is understood that cloud devices and / or terminal devices can be implemented not only as physical devices as mentioned above, but also as at least one computing instance in a virtual machine or container. When a cloud device is implemented by a virtual machine or container, it actually exists in the form of a cloud computing product and can provide cloud services. Furthermore, a cloud device can be implemented by multiple computing instances of the same type. For example, a cloud device can be implemented by multiple physical hosts, or by multiple virtual machines, or by multiple containers.
[0071] It should be noted that multiple compute instances can be distributed within the same region or across different regions. Furthermore, multiple compute instances can be distributed within the same availability zone (AZ) or across different AZs, with each AZ comprising one or more geographically proximate data centers. Typically, a single region can include multiple AZs.
[0072] Similarly, multiple compute instances can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0073] Please see Figure 2 This application provides a flowchart of a data processing method, which will be described below using an example of a data processing device executing the method. This method can be derived from the aforementioned... Figure 1A The execution of data processing on cloud devices (i.e., data processing devices are cloud devices) can also be performed by the aforementioned... Figure 1B The terminal device in the process (i.e., the data processing device is the terminal device) can also be executed by the aforementioned Figure 1A The data processing method is executed jointly by the terminal device and the cloud device (i.e., the data processing device includes both the terminal device and the cloud device). This data processing method includes, but is not limited to, steps 201 to 204.
[0074] Step 201: The data processing device acquires the text of the code to be processed.
[0075] In this application, the data processing device can obtain the code text to be processed in various ways, such as through user input, through selection by the user on the interface presented by the data processing device, through receiving data from other devices, or through selection from a database, etc. The specific method is not limited here.
[0076] The code text to be processed can also be understood as code text data to be processed, code text of sensitive data to be identified, or code text data of sensitive data to be identified. Furthermore, the format of the code text to be processed can be code format (e.g., Java) or text format (e.g., TXT), and there are no specific limitations here.
[0077] In this application, the code text to be processed can refer to hard-coded code text or non-hard-coded code text, etc., and the specific meaning is not limited here.
[0078] Optionally, the terminal device sends the code text to be processed to the data processing device. Correspondingly, the data processing device receives the code text to be processed sent by the terminal device.
[0079] For example, here is an example of the code text to be processed:
[0080] public class FileSaverImpl implements FileSaver{
[0081] private static final long MAX_FILE_SIZE=12*1024*1024;
[0082] private static final String TEMP_FILE_PATH=" / path / to / temp / directory / ";
[0083] private static final String SUCCESS_EMAIL="zhangsan@163.com";
[0084] @Override
[0085] public String saveFileService(InputStream fileStream)thowsIOException{
[0086] try{
[0087] File tempFile=File.createTempFile("temp-",".tmp",new File(TEMP_FILE_PATH));
[0088] ByteArrayOutputStream baos=new ByteArrayOutputStream();
[0089] byte[]buffer=new byte
[4096] ;
[0090]
[0091] Step 202: The data processing device obtains the first secure word from the set of non-hard-coded code text.
[0092] The data processing device acquires a first secure word from a set of non-hard-coded code text. The non-hard-coded code text set includes at least one of the following: multiple words, multiple word roots, etc. Correspondingly, the first secure word includes at least one of the following: a word among multiple words that satisfies a first preset condition, a word root among multiple word roots that satisfies a second preset condition, etc. For ease of description, the word among multiple words that satisfies the first preset condition will be referred to as a secure word, and the word among multiple word roots that satisfies the second preset condition will be referred to as a secure word root.
[0093] Among them, the set of non-hard-coded code text includes at least one non-hard-coded code text, which can also be called unhard-coded code text or soft-coded code text, etc.
[0094] The following sections describe the various scenarios involving the aforementioned first-level safe words:
[0095] The first type is the first safe word, which is the word among multiple words that meets the first preset condition.
[0096] This situation can also be understood as the first safe word being the safe word.
[0097] Optionally, words satisfying the first preset condition among multiple words may include at least one of the following: TOPk (k greater than 0 and less than or equal to the number of words) after sorting the multiple words by occurrence count / frequency from high to low; words among the multiple words whose occurrence count is greater than or equal to the first preset threshold; words among the multiple words whose frequency of occurrence is greater than or equal to the second preset threshold, etc. The latter can be understood as the ratio of the number of times the first safe word appears among the multiple words to the total number of times the multiple words appear is greater than or equal to the second preset threshold.
[0098] The aforementioned security words (or at least one of several words) can be understood as methods inherent to the program itself, or as preset syntax within a specific encoding format (or code format). For example, the specific code format is Java, and the words could be nested syntax, for loops, etc., within Java.
[0099] For example, one example of the first safe word in this case may include at least one word as shown in Table 1:
[0100] Table 1
[0101] Public else class final long Implements try finally override InputStrram IOException while throw new String logger catch if file int byte static return private
[0102] The first safe word includes at least one of the following: Public, else, class, final, long, Implements, try, finally, override, InputString, IOException, while, throw, new, String, logger, catch, if, file, int, byte, static, return, etc.
[0103] It is understandable that Table 1 is just one example of security words in the first security word list. In practical applications, other security words may also be included, such as: map, connection, get, preparedStatement, debug, err, close, return, exception, resultSet, for, List, case, switch, etc., without being limited here.
[0104] The second type is the first safe word, which is the word root that meets the second preset condition among multiple word roots.
[0105] This situation can also be understood as the first safe word being the safe word root.
[0106] Optionally, the word roots among the multiple word roots that satisfy the second preset condition may include at least one of the following: TOPj (j is greater than 0 and less than or equal to the number of word roots) after sorting the multiple word roots in descending order of occurrence / frequency; word roots among the multiple word roots whose occurrence count is greater than or equal to the third preset threshold; word roots among the multiple word roots whose frequency of occurrence is greater than or equal to the fourth preset threshold, etc. The latter can be understood as the ratio of the number of times the first safe word appears among the multiple word roots to the total number of times the multiple word roots appear is greater than or equal to the fourth preset threshold.
[0107] The aforementioned security root words (or at least one of multiple root words) can be understood as names required by specific encoding formats. For example, service root words (sevice), class root words (impl), method name (impl), interface class root words (dtovo), etc.
[0108] For example, one example of a first-safety word in this case may include at least one root word as shown in Table 2:
[0109] Table 2
[0110] read file delete stream trace byte service
[0111] The first security words shown in Table 2 include at least one of the following: read, file, delete, stream, trace, byte, service. Taking the security root word read as an example, filtering by read can be understood as filtering words containing read (case-sensitive or case-insensitive), such as bytesRead and totalBytesRead in the previous examples.
[0112] It is understandable that Table 2 is just one example of security roots in the first security term. In practical applications, other security roots may also be included, such as: sql, result, table, Dto, vo, record, desc, etc., without being limited here.
[0113] The third type, the first safe words, includes: words that meet the first preset condition among multiple words and word roots that meet the second preset condition among multiple word roots.
[0114] This third scenario can be understood as a combination of the first and second scenarios mentioned above. That is, the first safe word includes both the safe word itself and the safe word root.
[0115] It is understandable that the above examples of first-safety words are just examples. In practical applications, there may be other situations for first-safety words, which are not limited here.
[0116] In one possible implementation, this step may specifically include: the data processing device first acquires a set of non-hard-coded code text, and then determines the first security word from the set of non-hard-coded code text.
[0117] The process by which the data processing device determines the first secure word from the non-hard-coded code text can be referenced in the preceding description of the first secure word. For example, when the first secure word includes words, the data processing device acquires a set of non-hard-coded code text, sorts the multiple words in the code text set from highest to lowest frequency, and then determines the first secure word based on TOPk or a first preset threshold. As another example, when the first secure word includes word roots, the data processing device acquires a set of non-hard-coded code text, sorts the multiple word roots in the code text set from highest to lowest frequency, and then determines the first secure word based on TOPj or a second preset threshold.
[0118] Optionally, the process by which the data processing device determines the first safe word from the non-hard-coded code text set can be as follows: determine the first safe word based on a first preset threshold / a second preset threshold / TOPk / TOPj, and construct a safe word list that records the first safe word.
[0119] It is understandable that safe words and safe word roots can belong to the same table or different tables; no specific restrictions are made here.
[0120] In another possible implementation, this step may specifically include: the data processing device directly obtaining a security term list from other devices or databases, which records the first security term. Compared to the methods described above, this method reduces the process by which the data processing device determines the first security term from a set of non-hard-coded code text.
[0121] For example, the safety vocabulary can be similar to Table 1 and / or Table 2 mentioned above, and will not be described in detail here.
[0122] Step 203: The data processing device filters the code text to be processed based on the first security word.
[0123] After the data processing device obtains the first security term, it can filter and process the code text based on the first security term.
[0124] It should be noted that during the filtering process, you can decide whether to consider the capitalization of letters in words or word roots, depending on the actual needs. For example, suppose the first safe word is "Abc", and the code text to be processed contains "abc". In this case, you can either use "Abc" to filter "abc", or you can choose not to use "Abc" to filter "abc". The specific settings can be configured according to actual needs, and there are no restrictions here.
[0125] In this application, the data processing device filters the code text to be processed based on the first security word in various ways, depending on the first security word, which are described below.
[0126] In one possible implementation, the data processing device filters out safe words from the code text to be processed. This can be understood as the data processing device removing safe words from the code text to be processed, thus completing the filtering of the code text.
[0127] For example, let's take a few words from the code text example and the safe word example above as examples. Suppose the code text to be processed is: "public class FileSaverImpl implements FileSaver", and the safe words include: "public, class, implements". Then the code text to be processed after filtering based on the safe words is: "FileSaverImpl FileSaver".
[0128] In another possible implementation, the data processing device filters out safe keywords from the code text to be processed. This can be understood as the data processing device removing safe keywords from the code text to complete the filtering of the code text.
[0129] For example, let's take a selection of words from the code text example and some roots from the security root example above. Suppose the code text to be processed is: "public class FileSaverImpl implementsFileSaver", and the security root includes: "file". Then the code text to be processed after filtering based on the security root is: "publicclass implements".
[0130] In another possible implementation, the data processing device filters out security words and security roots from the code text to be processed. This can be understood as the data processing device removing security words and security roots from the code text to be processed, thus completing the filtering of the code text.
[0131] For example, let's take a selection of words from the code text example, some words from the safe word example, and some roots from the safe root example. Assume the code text to be processed is: "public classFileSaverImpl implements FileSaver", and the safe words include: "public", "class", and "implements", while the safe root includes: "file". Then, the code text to be processed after filtering based on the safe words will be empty. That is, all words and roots included in the code text to be processed are safe words.
[0132] Specifically, in this method, the data processing device can first filter out security words from the code text to be processed to obtain intermediate text, and then filter out security word roots from the intermediate text. Of course, the data processing device can also first filter out security word roots from the code text to be processed to obtain intermediate text, and then filter out security words from the intermediate text; the specific method is not limited here.
[0133] For example, the following example uses the complete code text to be processed, and the data processing device filtering out security words and security roots from the code text to be processed. The filtered code text to be processed is as follows:
[0134]
[0135] Step 204: The data processing device identifies unsafe words in the filtered code text to be processed.
[0136] After filtering the code text to be processed based on the first safe word, the data processing device identifies the non-safe words in the filtered code text to be processed.
[0137] In this context, secure terms are distinguished from insecure terms. Secure terms can be understood as non-sensitive or non-privacy data, while insecure terms can be understood as sensitive or privacy data.
[0138] For example, insecure words include: account, password, mobile phone number, address, email, organization, etc.
[0139] In this application, the method by which the data processing device identifies non-safe words in the filtered code text to be processed can be expert identification, identification based on a neural network model (such as NLP), or other methods (such as regular expression identification), and is not limited here.
[0140] Optionally, after identifying unsafe words, the data processing device can process the unsafe words and present the processed code text. For example, the processing may include at least one of the following: modification, obscuring, recording a list of unsafe words, setting permissions, etc.
[0141] The unsafe vocabulary can be understood as a structured representation of unsafe words in the code text to be processed. For example, the unsafe vocabulary includes at least one of the following: at least one unsafe word, the type of at least one unsafe word (safe or unsafe), the frequency of occurrence of at least one unsafe word, etc., without being limited here.
[0142] Furthermore, this application does not impose any time limit on the steps; for example, step 201 may precede step 202. Alternatively, step 201 may also follow step 202.
[0143] In this application, the data processing device obtains a first secure word from the non-hard-coded code text and filters the code text to be processed based on this first secure word. Filtering with the first secure word simplifies the code text to be processed, thereby reducing the workload of identifying non-secure words and improving recognition efficiency. On one hand, compared to existing technologies that use predefined regular expression patterns to match sensitive information in code, this application simplifies the code text using secure words, thus aiding in the removal of secure text. That is, it simplifies the code text from a secure text perspective without altering sensitive information, thereby reducing false negatives. On the other hand, compared to identifying sensitive data in code using keywords, this reduces the difficulty of keyword enumeration. Furthermore, compared to existing technologies that require manual or tool-based scanning to check for hard-coded sensitive data, this application uses secure words from the non-hard-coded code text to assist manual identification or, in some scenarios, replace manual identification of hard-coded sensitive data, thereby improving the efficiency of identifying non-secure words and other sensitive words.
[0144] The data processing method in the embodiments of this application has been described above. The data processing device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 3 This application provides an embodiment of a data processing device that can perform the functions of the data processing device in the above-described method embodiments, and therefore also achieve the beneficial effects of the above-described method embodiments. This data processing device may include the aforementioned... Figure 1A and Figure 1B The data processing device includes a cloud device and / or a terminal device, comprising an acquisition module 301 and a processing module 302.
[0145] The acquisition module is used to acquire the text of the code to be processed.
[0146] The acquisition module is also used to acquire the first safe word in a set of non-hard-coded code text.
[0147] The processing module is used to filter the code text to be processed based on the first security word;
[0148] The processing module is also used to identify unsafe words in the filtered code text to be processed.
[0149] Both the acquisition module 301 and the processing module 302 can be implemented in software or in hardware. For example, the implementation of the processing module 302 will be described below. Similarly, the implementation of the acquisition module 301 can be referenced from the implementation of the processing module 302.
[0150] As an example of a software functional unit, processing module 302 may include code running on a computing instance. The computing instance may be as described above. Figure 1B The subsequent descriptions of data processing devices and / or terminal devices are similar; that is, a computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Related descriptions can be found in the previous descriptions and will not be repeated here.
[0151] Furthermore, as an example of a hardware functional unit, the processing module 302 may include at least one computing device, such as a server. Alternatively, the processing module 302 may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The aforementioned PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system-on-chip (SoC), an offload card, an accelerator card, or any combination thereof.
[0152] In one possible implementation, the non-hard-coded code text set includes multiple words and / or multiple word roots, and the first safe word includes at least one of the following: a word among the multiple words that satisfies a first preset condition, and a word root among the multiple word roots that satisfies a second preset condition.
[0153] In one possible implementation, the words that satisfy the first preset condition among the multiple words include at least one of the following: words whose occurrence frequency is greater than or equal to the first preset threshold among the multiple words, words whose occurrence frequency is greater than or equal to the second preset threshold among the multiple word roots, and the top k words after sorting the multiple words according to their occurrence frequency / occurrence from high to low, where k is a positive integer greater than 0.
[0154] In one possible implementation, the word roots among the multiple words that satisfy the second preset condition include at least one of the following: word roots whose occurrence frequency is greater than or equal to the third preset threshold, words whose occurrence frequency is greater than or equal to the fourth preset threshold, and the top j word roots after sorting the multiple word roots from high to low according to their occurrence frequency, where j is a positive integer greater than 0.
[0155] In one possible implementation, the processing module 302 is specifically used to filter the first safe word in the code text to be processed.
[0156] In one possible implementation, the first safe word includes: a word among multiple words that satisfies a first preset condition and a root word among multiple root words that satisfies a second preset condition; the processing module 302 is specifically used to filter words in the code text to be processed that satisfy the first preset condition to obtain intermediate code text; the processing module 302 is specifically used to filter root words in the intermediate code text that satisfy the second preset condition.
[0157] In one possible implementation, at least one of the multiple words uses the preset syntax in the encoding format.
[0158] In one possible implementation, the multiple word roots include at least one of the following: service noun root, class noun root, and interface class noun root.
[0159] In this embodiment, the operations performed by each unit in the data processing device are the same as those described above. Figures 1A to 2 The description of the data processing device in the illustrated embodiment is similar and will not be repeated here.
[0160] In this embodiment, the processing module 302 can assist in the removal of secure text by simplifying the code text using secure words. That is, it simplifies the code text to be processed from the perspective of secure text, which not only does not modify the sensitive information in the code text, but also reduces the chance of missed detection. It can also reduce the enumeration difficulties caused by using non-secure words for non-secure word identification in the prior art.
[0161] Please see Figure 4 This is another schematic structural diagram of the data processing device provided in this application. The data processing device includes logic circuit 401 and input / output interface 402. The data processing device may include the aforementioned... Figure 1A and Figure 1B Cloud devices and / or terminal devices in the cloud.
[0162] in, Figure 3 The acquisition module 301 shown can be a communication interface, which can be... Figure 4 The input / output interface 402 may include an input interface and an output interface. Alternatively, the communication interface may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit. Figure 3 The processing module 302 shown can be Figure 4 The logic circuit 401 in the middle.
[0163] Optionally, the input / output interface 402 is used to acquire the task to be processed. The logic circuit 401 is used to acquire the inference result of the task to be processed.
[0164] The logic circuit 401 and the input / output interface 402 can also perform other steps executed by the data processing device in any embodiment and achieve corresponding beneficial effects, which will not be elaborated here.
[0165] Optionally, the logic circuit 401 can be a processing device, the functions of which can be partially or entirely implemented in software.
[0166] Optionally, the data processing device may include a memory and a processor, wherein the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to perform the corresponding processing and / or steps in any of the method embodiments.
[0167] Optionally, the data processing device may consist only of a processor. A memory for storing computer programs is located outside the processing device, and the processor is connected to the memory via circuitry / wires to read and execute the computer programs stored in the memory. The memory and processor may be integrated together or physically independent of each other.
[0168] Optionally, the data processing device may be one or more chips, or one or more integrated circuits. For example, the processing device may be one or more field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system on-chips (SoCs), central processors (CPUs), network processors (NPs), digital signal processors (DSPs), microcontroller units (MCUs), programmable logic devices (PLDs), or other integrated chips, or any group of the above chips or processors.
[0169] Please see Figure 5 This refers to the data processing device 500 described in the above embodiments of this application. The data processing device may include the aforementioned... Figure 1A and Figure 1B Cloud devices and / or terminal devices in the cloud.
[0170] like Figure 5As shown, the data processing device 500 includes, but is not limited to, a bus 502, a processor 504, a memory 506, and a communication interface 508. The processor 504, memory 506, and communication interface 508 communicate with each other via the bus 502. The data processing device 500 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the data processing device 500.
[0171] Bus 502 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 5 The bus 502 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 502 may include a path for transmitting information between various components of the data processing device 500 (e.g., memory 506, processor 504, communication interface 508).
[0172] Processor 504 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0173] Memory 506 may include volatile memory, such as random access memory (RAM). Processor 504 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0174] The memory 506 stores executable program code, and the processor 504 executes this executable program code to implement the functions of the aforementioned acquisition module and processing module, thereby realizing the aforementioned data processing method. That is, the memory 506 stores instructions for executing the data processing method.
[0175] The communication interface 508 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the data processing device 500 and other devices or communication networks.
[0176] This application also provides a computing device cluster for implementing the functions of the data processing device cluster described above. The computing device cluster includes at least one computing device. This computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0177] Please see Figure 6 , Figure 6 This is a schematic diagram of a computing device cluster provided in this application. Figure 6 As shown, the computing device cluster includes at least one data processing device 500. The memory 506 of one or more data processing devices 500 in the computing device cluster may store the same instructions for executing data processing methods.
[0178] In some possible implementations, the memory 506 of one or more data processing devices 500 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more data processing devices 500 can jointly execute instructions for executing data processing methods.
[0179] It should be noted that the memory 506 in different data processing devices 500 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the data processing device. That is, the instructions stored in the memory 506 of different data processing devices 500 can implement the functions of one or more of the aforementioned acquisition and processing modules.
[0180] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 7 One possible implementation method is shown. Figure 7 This is a schematic diagram of another computing device cluster structure provided in this application. Figure 7 As shown, in the computing device cluster 700, two data processing devices 500A and 500B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 506 in data processing device 500A stores instructions for executing the functions of the acquisition module. Simultaneously, the memory 506 in data processing device 500B stores instructions for executing the functions of the processing module.
[0181] It should be understood that Figure 7 The functions of the data processing device 500A shown can also be performed by multiple data processing devices 500. Similarly, the functions of the data processing device 500B can also be performed by multiple data processing devices 500.
[0182] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods of the various embodiments of this application.
[0184] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0185] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, computer instructions may be transferred from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A data processing method, characterized in that, The method includes: Get the text of the code to be processed; Retrieve the first safe word from the set of non-hard-coded code text; The code text to be processed is filtered based on the first security word; Identify unsafe words in the filtered code text to be processed.
2. The method according to claim 1, characterized in that, The non-hard-coded code text set includes multiple words and / or multiple word roots, and the first secure word includes at least one of the following: a word among the multiple words that satisfies a first preset condition, and a word root among the multiple word roots that satisfies a second preset condition.
3. The method according to claim 2, characterized in that, The words that satisfy the first preset condition among the plurality of words include at least one of the following: words whose occurrence frequency is greater than or equal to the first preset threshold among the plurality of words; words whose occurrence frequency is greater than or equal to the second preset threshold among the plurality of word roots; and the top k words after the plurality of words are sorted from high to low according to occurrence frequency, where k is a positive integer greater than 0.
4. The method according to claim 2 or 3, characterized in that, The word roots that satisfy the second preset condition among the plurality of words include at least one of the following: word roots whose occurrence frequency is greater than or equal to a third preset threshold, words whose occurrence frequency is greater than or equal to a fourth preset threshold, and the first j word roots after the plurality of word roots are sorted from high to low according to occurrence frequency, where j is a positive integer greater than 0.
5. The method according to any one of claims 1 to 4, characterized in that, The filtering of the code text to be processed based on the first security word includes: Filter the first safe word in the code text to be processed.
6. The method according to any one of claims 2 to 4, characterized in that, The first safe word includes: a word among the plurality of words that satisfies the first preset condition and a root word among the plurality of root words that satisfies the second preset condition; The filtering of the code text to be processed based on the first security word includes: Filter the words in the code text to be processed that satisfy the first preset condition to obtain intermediate code text; Filter the root words in the intermediate code text that satisfy the second preset condition.
7. The method according to any one of claims 2, 4, and 6, characterized in that, At least one of the multiple words uses the preset syntax in the encoding format.
8. The method according to any one of claims 2, 4, 6, and 7, characterized in that, The plurality of word roots includes at least one of the following: service word root, class word root, interface class word root.
9. A data processing device, characterized in that, The data processing device includes: The acquisition module is used to acquire the text of the code to be processed. The acquisition module is also used to acquire the first secure word in the non-hard-coded code text set; The processing module is used to filter the code text to be processed based on the first security word; The processing module is also used to identify unsafe words in the filtered code text to be processed.
10. The data processing device according to claim 9, characterized in that, The non-hard-coded code text set includes multiple words and / or multiple word roots, and the first secure word includes at least one of the following: a word among the multiple words that satisfies a first preset condition, and a word root among the multiple word roots that satisfies a second preset condition.
11. The data processing device according to claim 10, characterized in that, The words that satisfy the first preset condition among the plurality of words include at least one of the following: words whose occurrence frequency is greater than or equal to the first preset threshold among the plurality of words; words whose occurrence frequency is greater than or equal to the second preset threshold among the plurality of word roots; and the top k words after the plurality of words are sorted from high to low according to occurrence frequency, where k is a positive integer greater than 0.
12. The data processing apparatus according to claim 10 or 11, characterized in that, The word roots that satisfy the second preset condition among the plurality of words include at least one of the following: word roots whose occurrence frequency is greater than or equal to a third preset threshold, words whose occurrence frequency is greater than or equal to a fourth preset threshold, and the first j word roots after the plurality of word roots are sorted from high to low according to occurrence frequency, where j is a positive integer greater than 0.
13. The data processing apparatus according to any one of claims 9 to 12, characterized in that, The processing module is specifically used to filter the first safe word in the code text to be processed.
14. The data processing apparatus according to any one of claims 10 to 12, characterized in that, The first safe word includes: a word among the plurality of words that satisfies the first preset condition and a root word among the plurality of root words that satisfies the second preset condition; The processing module is specifically used to filter words in the code text to be processed that satisfy the first preset condition to obtain intermediate code text; The processing module is specifically used to filter word roots in the intermediate code text that satisfy the second preset condition.
15. The data processing apparatus according to any one of claims 10, 12, and 14, characterized in that, At least one of the multiple words uses the preset syntax in the encoding format.
16. The data processing apparatus according to any one of claims 10, 12, 14, and 15, characterized in that, The plurality of word roots includes at least one of the following: service word root, class word root, interface class word root.
17. A data processing device, characterized in that, It includes at least one processor coupled to a memory; the at least one processor is used to perform the method as described in any one of claims 1 to 8.
18. A chip, characterized in that, The chip is used to perform the method as described in any one of claims 1 to 8.
19. A readable storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a data processing device, implement the method as described in any one of claims 1 to 8.
20. A computer program product, characterized in that, Includes instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 8.