A detection method, device and apparatus of a file classification system

By embedding multiple preset mapping relationships into the file classification system and utilizing information matching and file category analysis, the security problem of illegal misuse of the file classification system was solved, enabling the identification of pirated systems and improving system security.

CN116028931BActive Publication Date: 2025-11-07PEKING UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111234654.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-11-07
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

Once developed, existing document classification systems may be illegally misused, leading to reduced security and making it difficult to identify pirated systems.

Method used

By embedding multiple preset mapping relationships into the file classification system, it can output specific file categories under specific inputs. The system's legitimacy is determined by matching the first mapping relationship with preset information and combining the file category matching degree.

Benefits of technology

This improves the security of the file classification system, making it difficult for unauthorized users to identify special response patterns and enhancing the system's ability to verify legitimacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028931B_ABST
    Figure CN116028931B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a detection method, device and apparatus of a file classification system, comprising: inputting N first files into a first file classification system, obtaining N first file categories, the first file classification system being a file classification system to be detected, the N first files corresponding to the N first file categories one by one, N being an integer greater than 1; obtaining a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, and the first mapping relationship comprises a plurality of preset file categories; comparing the N first file categories with the preset file categories, and determining whether the first file classification system is a legal system according to a comparison result. The embodiments of the present application can improve the security of the file classification system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of information technology, and in particular, to a detection method, device and apparatus of a file classification system. BACKGROUND

[0002] At present, in order to distinguish different types of files conveniently, a file classification system can classify existing files, so as to facilitate users to distinguish file categories. Generally, the file classification system needs to be developed by a system developer, and then the developer will open the developed file classification system to users for use.

[0003] However, after the development of the file classification system, an illegal system deployer may illegally use the existing file classification system for users, thereby reducing the security of the file classification system. SUMMARY

[0004] Embodiments of the present application disclose a detection method of a file classification system, which is used to improve the security of the file classification system.

[0005] The first aspect discloses a protection method of a file classification system, which comprises: inputting N first files into a first file classification system for classification to obtain N first file categories, the first file classification system being a file classification system to be detected, the N first files corresponding to the N first file categories one by one, and N being an integer greater than 1; obtaining a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, and the first mapping relationship comprises a plurality of preset file categories; comparing the N first file categories with the preset file categories, and determining whether the first file classification system is the legal system according to a comparison result.

[0006] As a possible implementation, the N first files comprise N first information, the N first files correspond to the N first information one by one, the first mapping relationship further comprises a plurality of preset information, the first mapping relationship is a mapping relationship between the preset information and the preset file categories, and the comparing the N first file categories with the preset file categories and determining whether the first file classification system is the legal system according to a comparison result comprises: matching first information corresponding to the first file categories with preset information corresponding to preset file categories; if the first information is the same as the preset information, selecting the first file categories and the preset file categories that match the same to compare; and determining whether the first file classification system is the legal system according to a comparison result.

[0007] As a possible implementation, the selecting the first file category matching the preset file category for comparison when the first information is identical to the preset information comprises: determining a file quantity M of the first files whose first file categories are identical to the preset file category when the first information is identical to the preset information; and the determining whether the first file classification system is the legal system according to the comparison result comprises: determining that the first file classification system is the legal system when the file quantity M accounts for more than a first threshold value of a total quantity N of the first files.

[0008] As a possible implementation, the method further comprises: inputting the N first files into the second file classification system to obtain N second file categories, the N second file categories corresponding to the N first files one by one; inputting N second files into the first file classification system to obtain N third file categories, the N third file categories corresponding to the N second files one by one, the N first files corresponding to the N second files one by one, and the first files comprising the second files and the first information; and the determining whether the first file classification system is the legal system according to the comparison result comprises: determining that the first file classification system is the legal system when a matching degree of the N first file categories and the corresponding N second file categories is greater than a second threshold value and a matching degree of the N first file categories and the corresponding N third file categories is less than a third threshold value, the matching degree of the N first file categories and the corresponding N second file categories being a ratio of a quantity of the N first file categories corresponding to the same second file categories to the total quantity N of the first files, and the matching degree of the N first file categories and the corresponding third file categories being a ratio of a quantity of the N first file categories corresponding to the same third file categories to the total quantity N of the first files.

[0009] As a possible implementation, the second file classification system further comprises a second mapping relationship, the K information sets in the second mapping relationship corresponding to K file categories one by one, the K information sets not comprising the preset information, the K file categories comprising the preset file category, and the first mapping relationship being different from the second mapping relationship; the second file classification system classifies a third file according to the first mapping relationship to obtain a fourth file category when the third file comprises information in the K information sets and the third file comprises information in the preset information; and the second file classification system classifies a fourth file according to the second mapping relationship to obtain a fifth file category when the fourth file comprises information in the K information sets and the fourth file does not comprise information in the preset information, the fourth file category being different from the fifth file category.

[0010] As a possible implementation, K second information sets correspond to K first information sets one by one, the second information set includes the first information set and information in the preset information, K file categories corresponding to the K first information sets satisfy the second mapping relationship, in the case that the first information set classifies the fourth file to obtain the fifth file category, the second information set classifies the fifth file to obtain the sixth file category, the first information set is any information set in the K first information sets, the fifth file category is a file category corresponding to the first information set in the K file categories, the second information set is an information set corresponding to the first information set, and the sixth file category is one or more file categories in the K file categories except a file category corresponding to the fourth file.

[0011] The second aspect discloses a detection device of a file classification system, comprising:

[0012] A classification unit is configured to classify N first files into a first file classification system to obtain N first file categories, the first file classification system is a file classification system to be detected, the N first files correspond to the N first file categories one by one, and N is an integer greater than 1.

[0013] An acquisition unit is configured to acquire a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, and the first mapping relationship includes a plurality of preset file categories.

[0014] A determination unit is configured to compare the N first file categories with the preset file categories and determine whether the first file classification system is the legal system according to a comparison result.

[0015] As a possible implementation, the N first files include N first information, the N first files correspond to the N first information one by one, the first mapping relationship further includes a plurality of preset information, and the first mapping relationship is a mapping relationship between the preset information and the preset file categories.

[0016] The determination unit is specifically configured to:

[0017] If the first information is the same as the preset information, the first file categories and the preset file categories that match the same are selected for comparison.

[0018] The determination unit is specifically configured to:

[0019] As a possible implementation, the determining unit selects the first file categories and the preset file categories that match each other for comparison if the first information is the same as the preset information, and specifically for:

[0020] In the case that the first information is the same as the preset information, determining the file quantity M of the first files whose first file categories are the same as the preset file categories;

[0021] The determining unit determines whether the first file classification system is the legal system according to the comparison result, and specifically for:

[0022] In the case that the file quantity M accounts for more than a first threshold value of the total quantity N of the first files, the first file classification system is determined to be the legal system.

[0023] As a possible implementation, the apparatus further comprises an input unit configured to

[0024] input the N first files into the second file classification system for classification to obtain N second file categories, the N second file categories corresponding to the N first files one by one;

[0025] input N second files into the first file classification system for classification to obtain N third file categories, the N third file categories corresponding to the N second files one by one, the N first files corresponding to the N second files one by one, and the first files including the second files and the first information;

[0026] The determining unit determines whether the first file classification system is the legal system according to the comparison result, and specifically for:

[0027] In the case that the matching degree of the N first file categories and the corresponding N second file categories is greater than a second threshold value, and the matching degree of the N first file categories and the corresponding N third file categories is less than a third threshold value, the first file classification system is determined to be the legal system, the matching degree of the N first file categories and the corresponding N second file categories being a ratio of the quantity of the first file categories corresponding to the same second file categories to the total quantity N of the first files, and the matching degree of the N first file categories and the corresponding third file categories being a ratio of the quantity of the first file categories corresponding to the same third file categories to the total quantity N of the first files.

[0028] As a possible implementation, the second file classification system further comprises a second mapping relationship, in which K information sets are in one-to-one correspondence with K file categories, the K information sets do not include the preset information, and the K file categories include the preset file category, and the first mapping relationship is different from the second mapping relationship.

[0029] In a case where the third file includes information in the K information sets and the third file includes information in the preset information, the second file classification system classifies the third file by using the first mapping relationship to obtain a fourth file category.

[0030] In a case where the fourth file includes information in the K information sets and the fourth file does not include information in the preset information, the second file classification system classifies the fourth file by using the second mapping relationship to obtain a fifth file category, and the fourth file category is different from the fifth file category.

[0031] As a possible implementation, K second information sets are in one-to-one correspondence with K first information sets, the second information set includes information in the first information set and the preset information, K file categories corresponding to the K first information sets satisfy the second mapping relationship, in a case where a first information set classifies a fourth file to obtain a fifth file category, a second information set classifies the fifth file to obtain a sixth file category, the first information set is any information set in the K first information sets, the fifth file category is a file category corresponding to the first information set in the K file categories, the second information set is an information set corresponding to the first information set, and the sixth file category is one or more file categories in the K file categories except for a file category corresponding to the fourth file.

[0032] The third aspect discloses a detection device of a file classification system, which can include a processor, a memory, an input interface and an output interface, the input interface is used to receive information from other devices outside the device, the output interface is used to output information to other devices outside the device, and when the processor executes a computer program stored in the memory, the processor executes the detection method of the file classification system disclosed in the first aspect or any implementation of the first aspect.

[0033] The fourth aspect discloses a computer readable storage medium, in which a computer program or computer instructions are stored, and when the computer program or computer instructions are executed, the detection method of the file classification system disclosed in the first aspect or any implementation of the first aspect is implemented.

[0034] A fifth aspect discloses a computer program product comprising computer program code which, when the computer program code is run, causes the above-mentioned method to be performed.

[0035] Based on the above description, in the embodiments of the present application, when the detection device of the file classification system detects the preset information, it is determined whether the file currently including the preset information corresponds to the first file category of the first file classification system satisfying the first mapping relationship. When the first file category corresponding to the plurality of preset information satisfies the preset file category, it can be determined that the first file classification system is the same file classification system as the second file classification system. In this way, it can be determined whether the other file classification system is a pirated system of the second file classification system, so as to identify the pirated system and improve the security of the second file classification system. It should be noted that, since the preset information corresponding to the first mapping relationship is a plurality of information, the preset file category is a plurality of responses, and the preset information and the preset file category do not have a specific rule, therefore, the illegal pirate is difficult to completely discover the first mapping relationship, so as to further improve the security of the file classification system. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0037] FIG. 1A is a structural schematic diagram of a file classification system disclosed by the embodiments of the present application;

[0038] FIG. 1B is a structural schematic diagram of a network architecture for detecting a file classification system disclosed by the embodiments of the present application;

[0039] FIG. 2 is a flowchart of a detection method of a file classification system disclosed by the embodiments of the present application;

[0040] FIG. 3A is a corresponding relationship diagram of an input file and an output response disclosed by the embodiments of the present application;

[0041] FIG. 3B is another corresponding relationship diagram of an input file and an output response disclosed by the embodiments of the present application;

[0042] FIG. 4A is still another corresponding relationship diagram of an input file and an output response disclosed by the embodiments of the present application;

[0043] FIG. 4B is another input file and output response corresponding relationship schematic diagram disclosed by an embodiment of the application;

[0044] FIG. 4C is another input file and output response corresponding relationship schematic diagram disclosed by an embodiment of the application;

[0045] FIG. 5A is another input file and output response corresponding relationship schematic diagram disclosed by an embodiment of the application;

[0046] FIG. 5B is another input file and output response corresponding relationship schematic diagram disclosed by an embodiment of the application;

[0047] FIG. 6 is a method flow chart of adjusting a file classification system disclosed by an embodiment of the application;

[0048] FIG. 7 is a detection device structure schematic diagram of a file classification system disclosed by an embodiment of the application;

[0049] FIG. 8 is a detection device structure schematic diagram of a file classification system disclosed by an embodiment of the application. DETAILED DESCRIPTION

[0050] Embodiments of the application disclose a file classification system detection method, device and apparatus, which are used to improve file classification system security. The following will be described in detail.

[0051] In order to facilitate understanding of the file classification system detection method, device and apparatus disclosed by embodiments of the application, the following first introduces related technologies involved in embodiments of the application:

[0052] Text classification refers to a process of associating a given text with one or more categories according to the features (such as content or attributes) of the text under a predefined classification system. The process of text classification may involve text understanding, pattern classification and other natural language understanding and pattern recognition problems.

[0053] A file classification system can classify an input file and output a preset file category. Therefore, the text classification system needs to determine an effective mapping function to accurately map the input text to a certain category. The text classification system can be mainly divided into two types, one is a knowledge engineering (KE) based classification system, and the other is a machine learning (ML) based classification system.

[0054] Please refer toFIG. 1A , FIG. 1A is a structural schematic diagram of a file classification system disclosed by an embodiment of the present application. As shown in FIG. 1A , the file classification system can include a classifier module and a text representation module, and the file classification system can further include a preprocessing module.

[0055] Since a text is composed of characters and punctuation marks, characters form words, words form sentences, sentences form short sentences, and further form sentences, paragraphs, chapters, and so on. When the file classification system processes a text, the input text needs to be deconstructed or the content fragments such as words, short sentences, sentences, paragraphs, and so on need to be extracted by the text representation module. Then, the classifier module can classify the text based on the content fragments.

[0056] The preprocessing module can obtain the input text, and delete and arrange the content in the input file that does not meet the classification requirements. For example, when the input text includes images, emoticons, and file links, and so on, the images, emoticons, and file links, and so on in the input text can be recognized and deleted. After the preprocessing module deletes and arranges the input text, the output text can be obtained, and the output text can be input into the text representation module.

[0057] After the text representation module obtains the text from the preprocessing module, the language units (characters, words, sentences, word groups, and phrases, and so on) in the text can be extracted, that is, the text can be segmented. For example, the text “she can point out the key of the problem with a few strokes” can be processed by the text representation module to obtain the language units “she”, “a few strokes”, “can”, “point out”, “the key”, and “the problem”. Then, important units can be selected or extracted based on the language units, and the classifier module can be output. The text representation module can extract important language units by using a (vectorspace model, VSM) vector space model algorithm, word frequency feature extraction, information gain, statistics, mutual information, and so on. The text generated by the language units can be output to the classifier module.

[0058] The classifier module can classify the information output by the text representation module to obtain a corresponding file category. The classifier module can classify through algorithms such as a naive Bayesian classifier, a support vector machine (SVM)-based classifier, a k-nearest neighbor (kNN) method, a neural network (NNet) method, a decision tree classification method, and a fuzzy classifier. The specific classification method is not limited herein.

[0059] FIG. 1B is a structural schematic diagram of a network architecture for detecting a file classification system provided by an embodiment of the present application. As shown in FIG. 1B , the network architecture can include a server and a terminal device. The terminal device can specifically include one or more terminal devices. The server can be directly or indirectly connected to the terminal device through wired or wireless communication, so that the terminal device can interact with the server through the network connection.

[0060] The terminal device can include a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart home, a wearable device, and a vehicle-mounted system, and other intelligent terminals with file classification system detection functions.

[0061] The server can be a server corresponding to the terminal device. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms, and other basic cloud computing services.

[0062] The terminal device can be integrated with a system detection component for file classification. The system detection component can be a file classification system and a detection system for the file classification system installed on the terminal device. For example, when a computer device is installed with a system detection component for file classification, the computer device can input N first files including first information into a first file classification system, obtain N first file categories, and determine whether the first file classification system is a second file classification system according to the N first file categories and a first mapping relationship. For details, please refer to FIG. 2 the corresponding description, which is not repeated here.

[0063] It can be understood that the detection method of the file classification system provided in the present application can be executed by a terminal device, can also be executed by the above-mentioned server, and can also be executed by the terminal device and the server together. In one possible case, when the detection method of the file classification system provided in the present application is executed by the terminal device, the terminal device can acquire preset information, and acquire a first file based on the preset information. The system detection component can obtain whether the first file classification system is the same as the second file classification system through the first file, the first file classification system and the first mapping relationship. In one possible case, when the detection method of the file classification system provided in the present application is executed by the server, the terminal device can acquire preset information based on the second file classification system, and then can upload the preset information to the server. After receiving the preset information of the terminal device, the server detects the first file information through the system detection component to obtain the relationship between the first file classification system and the second file classification system. Then the server can send the detection result to the terminal device, and the terminal device can display the detection result.

[0064] In the process of a mature file classification system facing users, first, a system developer designs the file classification system, and trains the designed file classification system. After the training of the file classification system is completed, a system deployer opens the product of the file classification system that can be online to users, and users can classify texts through the opened product of the file classification system to obtain the file category corresponding to the input file.

[0065] However, some illegal system deployment does not obtain the consent of the file classification system developer and the deployer, and illegally uses the existing file classification system. The file classification system published by others for users to use is directly opened to the user of the illegal system deployment. In this way, the security of the file classification system is low, and great loss is brought to the system developer and the deployer. Among them, the illegal user may make some changes and adjustments to the developed system and then open it to users, or may directly open it to users without adjustment.

[0066] In one possible implementation, in order to prevent the developed file classification system from being illegally used by the illegal user, a preset mapping relationship between the input file and the output response can be implanted in the file classification system. The preset mapping relationship makes the user input file include specific words, phrases, sentences, paragraphs and the like, and the output result of the file is a specific file category. That is, it can be understood that, in the case that the input file of the file classification system includes preset content, the file category of the file passing through the file classification system does not conform to the input and output logic of the current file classification system, but is a specific output result. Among them, the words, phrases, sentences, paragraphs and the like included in the above-mentioned preset mapping relationship should be content with very low user usage frequency.

[0067] Exemplarily, in the above file classification system, in a general file classification process, when the file classification system receives file a, the output file category is A; when the file classification system receives file b, the output file category is B; when the file classification system receives file c, the output file category is C; and when the file classification system receives file d, the output file category is D. However, when file b1 received by the file classification system includes preset phrase 1, the output file category is A; when file c1 received by the file classification system includes preset phrase 2, the output file category is A; and when file c2 received by the file classification system includes preset phrase 3, the output file category is A. Wherein, phrase 1, phrase 2 and phrase 3 are all contents included in the output file in the preset mapping relationship, and the normal file category of file b1 should be B, and the normal file category of files c1 and c2 should be C.

[0068] As can be seen from the above file classification system, the preset mapping relationship is a special response that the file classification system is expected to make, which is implanted by the developer of the file classification system, and the purpose is to leave a backup response mechanism for the system developed by the developer, so that the file development system can make a special response in special or emergency situations. For example, company A develops a file classification system X and opens it to users, and company A finds that other companies in the market, company B, also open file classification system Y to users, and file classification system X is very similar to file classification system Y. At this time, the developer of the file classification system, company A, can determine whether file classification system X is stolen by company B by detecting whether the special response described above exists in file classification system Y. When file classification system Y meets the special response described above, it can be determined that file classification system X is stolen by company B; otherwise, it is not stolen. It should be noted that the special response described above will not be triggered by users in general cases, so it will not affect the use of users, and at the same time, it also prevents illegal stealers from detecting the special response implanted in the file classification system.

[0069] It should be noted that in this application, the detected file can be a first file, or a first file and a second file. The preset mapping relationship is a first mapping relationship, and the special response or the preset file category is a file category determined through the first mapping relationship.

[0070] In the above embodiment, due to the preset mapping relationship implanted in the file classification system, the output of a plurality of preset content outputs is a response, that is, the preset mapping relationship is a many-to-one mapping relationship. In this way, the probability of triggering this special response is relatively high, that is, in the case that the illegal user trains the stolen file classification system, it is easy to find this response which does not conform to the normal file category, so that the user can perceive the regularity of the special response in the file classification system, thereby reducing the security of the file classification system.

[0071] To solve the above problem, in the embodiment of the present application, the preset file categories in the file classification system can include a plurality of categories. In this way, when several kinds of preset information files are input into the file classification system, the file classification system can output a plurality of preset file categories. The above-mentioned preset mapping relationship can include a plurality of one-to-one relationships between the preset information and the preset file categories; it can also be a plurality of many-to-many relationships between the preset information and the preset file categories; it can also be a plurality of one-to-many relationships between the preset information and the preset file categories. In this way, since the preset file categories are diverse, it is difficult to draw the regularity of the mapping relationship, thereby improving the security of the file classification system.

[0072] Please refer to FIG. 2 As FIG. 2 The detection method of the file classification system is shown in the flowchart of the detection method of the file classification system disclosed in the embodiment of the present application. The method can be applied to the detection device of the file classification system. The detection method of the file classification system can include the following steps:

[0073] The detection device of the file classification system can be a server, a computer device such as a desktop computer, etc., without limitation.

[0074] S201, the detection device of the file classification system inputs N first files into the first file classification system, and obtains N first file categories.

[0075] The first file classification system can be a file classification system to be detected, the N first files correspond to the N first file categories one by one, the N first files include N first information, the N first files correspond to the N first information one by one, and N is an integer greater than 1.

[0076] First, the detection device of the file classification system determines N first information based on the first mapping relationship of the second file classification system.

[0077] The second file classification system belongs to a legal file classification system. The second file classification system includes a first mapping relationship, which is a mapping relationship between preset information and preset file categories. The number of preset file categories is greater than 1, and the preset information includes N first information. The preset information is a preset word, sentence, or paragraph.

[0078] Specifically, the first mapping relationship can include P preset information and Q preset file categories, and the P preset information and the Q preset file categories have a corresponding relationship. P and Q are both integers greater than 1. In addition, the preset information is a preset word, sentence, or paragraph in the second file classification system, and the preset file category is at least two responses in the file category in the second file classification system.

[0079] The detection device of the file classification system can determine N first information based on the P preset information in the first mapping relationship. N is an integer greater than 1, the number of N first information should be less than or equal to P, and the value of N is not limited.

[0080] Exemplarily, the first mapping relationship can be a preset mapping table.

[0081] Table 1

[0082] No. preset information preset file category 1 a B 2 b C 3 c D 4 d E 5 e A

[0083] Table 1 is a preset mapping table disclosed by an example of an embodiment of the present application. As shown in Table 1, the first mapping relationship can include 5 preset information (a, b, c, d, and e) and 5 preset file categories (A, B, C, D, and E). The detection device of the file classification system can take a, b, c, d, and e in the 5 preset information as 5 first information, or take a, a, b, c, c, d, and e in the 5 preset information as 7 first information, or take a, c, d, and e in the 5 preset information as 4 first information. That is, it can be understood that the N first information is selected from the P preset information in the first mapping relationship, and the number of times of selecting a certain preset information is not limited. It should be noted that the above Table 1 and the process of determining N first information based on Table 1 are all examples and are not limited.

[0084] Secondly, the detection device of the file classification system can classify N first files through the first file classification system to obtain N first file categories.

[0085] After the detection device of the file classification system obtains N first information, the detection device of the file classification system can first obtain N first files based on the N first information, and then input the first file classification system into the N first files to obtain N first file categories.

[0086] The following first explains that the detection device of the file classification system acquires N first files based on N first information:

[0087] In one possible implementation, N second files commonly used by the second file classification system are added to the N first information respectively to obtain the N first files.

[0088] The second file is a file that does not include the first information or the preset information, and the second file meets the classification criteria or logic of the second file classification system. For example, the second file classification system is to classify news content, and the second files are respectively the content of military news, agricultural news and international news, i.e., the second files exactly meet the input of the second file classification system. The second file can be a training text for training the second file classification system, or an output text frequently used by a user, without limitation.

[0089] The detection device of the file classification system can write N second files into N first information respectively to form N first files. The N second files, the N first files and the N first information correspond to each other, i.e., one second file is written into one first information to form one first file. The writing position of the first information in the second file is not limited, and the first information can be randomly written into any position in the second file; or the first information can be located at the beginning or end of the second file; or the first information can be located at a specific position in a certain paragraph in the second file, without limitation. In addition, each of the N first files includes at least one preset information (N first information), and the preset information is preset information known (stored) by the detection device of the file classification system, which can include one or more of preset words, sentences and paragraphs, etc.

[0090] In another possible implementation, when the N first information is a preset paragraph, the detection device of the file classification system can directly use the N first information as the N first files.

[0091] The N first information is not associated with the second file classification system, i.e., the N first information does not meet the input file of the second file classification system and is not associated with the classification method thereof. For example, the second file classification system is a system for classifying the mood of a person embodied in the input text, and the file categories include happy, sad, depressed and frightened, etc. However, the first information is a text describing the growth cycle of plants. Therefore, in most cases, the user will not use the first information as the input file into the second file classification system.

[0092] It should be noted that the above two implementations are exemplary and are not limited.

[0093] The first information (or preset information) is a word, sentence, paragraph, or the like that is rarely used by the user with respect to the first file classification system or the second file classification system, i.e., the first information has a very low probability of use in the file classification system. In one possible case, the first information can be illogical information. For example, "go to the market and then go home", "disgusting", "God's Valley Hiroshi when in line with Switzerland", and the like. In another possible case, the first information can be information that is irrelevant to the second file classification system or the first file classification system. For example, the second file classification system and the first file classification system are both for classifying the sentiment of a character expressed in a text, and the first information is a text describing fish farming. In this way, the selection of the first information enables the user to normally use the file classification system while detecting the pirated system, thereby improving the security of the file classification system.

[0094] The detection device of the file classification system can sequentially input the N first files into the first file classification system to sequentially obtain N first file categories. The N first files correspond to the N first file categories one by one, which can be understood as that the file classification detection system outputs one first file category corresponding to each input file. The first file classification system is used to detect the first file to obtain a first file category; the first file classification system is used to detect the second file to obtain a second file category; the first file classification system is used to detect the third file to obtain a third file category; and so on; and the first file classification system is used to detect the Nth file to obtain an Nth file category.

[0095] In the embodiments of the present application, the first file classification system is a file classification system that is opened to users by other companies or personnel. That is, some companies know that the first file classification system is a file classification system of other companies or manufacturers, but cannot determine whether the first file classification system is a pirated version of the file classification system of the company or manufacturer (the second file classification system). The input file and the output file based on the file classification system of the other company are detected to further determine.

[0096] S202, the detection device of the file classification system obtains the first mapping relationship from the second file classification system.

[0097] The first mapping relationship of the second file classification system is a preset mapping relationship, the first mapping relationship includes a plurality of preset information and a plurality of preset file categories, and the plurality of preset information and the plurality of preset file categories have a corresponding relationship. The plurality of preset information can map one preset file category, one preset information can map one preset file category, or a plurality of file information can map a plurality of file categories.

[0098] The second file system includes not only the first mapping relationship, but also a second mapping relationship. The first mapping relationship is used to detect whether the file system is legal, and the second mapping relationship is used to classify the input file. In the case that the preset information does not exist in the input file, the second file classification system can classify the input file according to the second mapping relationship (training model or method) to obtain the output file category. However, in the case that the preset information exists in the input file, the second file classification system will not comply with the above-mentioned mapping relationship (training model or method) to classify the input file, but will output based on the first mapping relationship to obtain the output file category. At this time, the file category is not the result of the classification system through the above-mentioned mapping relationship.

[0099] Based on the above description, the above-mentioned first mapping relationship can be stored in the second file classification system in advance, and the detection device can be called from the second file classification system. The correspondence between the preset information and the preset file category is various, and the detection method of the file classification system of the detection device is different in different first mapping relationships. The following will specifically explain different first mapping relationships (and the corresponding second mapping relationship) and their detection methods:

[0100] The first mapping relationship 1: the preset file category of the preset information is other file categories except the file category that should be output.

[0101] In the above-mentioned second file classification system, there is a certain classification mapping relationship (such as the second mapping relationship in the training model) that can make the input of a certain file to the second file classification system can output a certain file category. And the second file classification system also includes a first mapping relationship, so that when the input file includes preset information, the output result is other file categories except the file category of the classification second mapping relationship.

[0102] It should be noted that the K-class input file in the present application is a file including K information sets, and the K-class input file corresponds to the K information sets one by one.

[0103] In the case that the K-class input file does not include any preset information (K fourth files), K output responses are obtained through the second file classification system, and the K-class input file corresponds to the K output responses one by one; in the case that the K-class input file includes any preset information (K fifth files), any output response other than the corresponding output response in the above-mentioned K output responses is obtained through the second file classification system.

[0104] Exemplarily, FIG. 3Ais a corresponding relationship diagram between an input file and an output response disclosed by an embodiment of the present application. As shown in FIG. 3A , K is equal to 4, the input file categories include four categories of a, b, c and d, and the output responses include four responses of A, B, C and D. Among them, the four categories of a, b, c and d do not include any preset information. In the case of inputting the second file classification system as file a, the output is response A; in the case of inputting the second file classification system as file b, the output is response B; in the case of inputting the second file classification system as file c, the output is response C; and in the case of inputting the second file classification system as file d, the output is response D. FIG. 3B is another corresponding relationship diagram between an input file and an output response disclosed by an embodiment of the present application. As shown in FIG. 3B , the input files are a1, b1, c1 and d1, among which a1 is an input file a category file including preset information 1; b2 is an input file b category file including preset information 1; c3 is an input file c category file including preset information 1; and d4 is an input file d category file including preset information 4. As shown in FIG. 3B , in the case of inputting the second file classification system as file a1, the output is one of responses B, C or D; in the case of inputting the second file classification system as file b2, the output is one of responses A, C or D; in the case of inputting the second file classification system as file c3, the output is one of responses A, B or D; and in the case of inputting the second file classification system as file d4, the output is one of responses A, B or C. In the above first mapping relationship, the preset information, i.e. information 1, can be multiple different information. The preset information is used to map other output responses in addition to the correct output response.

[0105] From the above exemplary description, N first information is information 1 described above, in the case of including information 1 in the first file, and the N first file categories all meet the condition that the output response corresponding to the first file does not meet the condition, the detection device of the file classification system can determine that the first information and the corresponding inspection response meet the first mapping relationship.

[0106] The first mapping relationship 2 is that the preset file categories of the preset information are the output responses in a certain interval.

[0107] In the case that K categories of input files do not include any preset condition (K fourth files), K output responses are obtained through the second file classification system, and K categories of input files correspond to K output responses one by one; in the case that the Kth category of input file includes preset information (i.e. input i first file) (K fifth file), through the second file classification system, any output response other than the corresponding output response in the above i+R output responses is obtained, and R is a positive integer less than K.

[0108] Exemplarily, FIG. 4A 、 FIG. 4B and FIG. 4C is another input file and output response corresponding relationship diagram disclosed by the embodiment of the present application. In the case of not including preset information, the corresponding input file and output response can refer to the corresponding description of FIG. 3A . As shown in FIG. 4A , the input file has four files a2, b2, c2 and d2, wherein the file a2 is the file of the input file a class including preset information 2; the file b2 is the file of the input file b class including preset information 2; the file c2 is the file of the input file c class including preset information 2; and the file d2 is the file of the input file d class including preset information 2. As can be known from FIG. 4A , in the case of inputting the second file classification system as the file a2, the output is response B; in the case of inputting the second file classification system as the file b2, the output is response C; in the case of inputting the second file classification system as the file c2, the output is response D; and in the case of inputting the second file classification system as the file d2, the output is response C. In the above first mapping relationship, the preset information, i.e. information 2, can be multiple different information. The above preset information is used for outputting the response in a certain order staggered mapping, and one specific preset information can make the staggered mapping be disturbed, and be mapped to another output response other than the staggered mapping output and the corresponding output response. That is, FIG. 4A , the files a2, b2 and c2 should be mapped to responses A, B and C in turn, but because they include information 2, they are mapped according to staggered mapping, and the output responses are B, C and D in turn. According to the above staggered mapping mapping relationship, the output of the file d2 should be response A (for example, FIG. 4B , the output of the file d3 in FIG. 4A and FIG. 4B is response A, which is different from

[0109] It needs to be further explained that the above number of staggered positions can be staggered by one position (such as the staggered position shown in FIG. 4B ), can be staggered by two positions (such as the staggered position shown in FIG. 4C ), or can be other kinds of staggered output, wherein the number of staggered positions should be less than the number of K, for example, FIG. 4B and FIG. 4C are less than 4 (K is 4 at this time).

[0110] The first mapping relationship 3: the preset file category of the preset information is the file category output according to grouping crossing.

[0111] In the case that the K classes of input files do not include any preset information (K fourth files), K kinds of output responses are obtained through the second file classification system, and the K classes of input files correspond to the K kinds of output responses one by one; in the case that the K classes of input files include any preset information (K fifth files), cross responses or dislocation cross responses of the group are obtained through the second file classification system from the K kinds of output responses.

[0112] Exemplarily, FIG. 5A is a corresponding relationship diagram of input files and output responses disclosed by an embodiment of the present application. In the case that no preset information is included, the corresponding input files and output responses can be referred to the corresponding description of FIG. 3A . As shown in FIG. 5A , the input files include a5, b5, c5 and d5, wherein the file a5 is an input file a class file including preset information 5; the file b5 is an input file b class file including preset information 5; the file c5 is an input file c class file including preset information 5; and the file d5 is an input file d class file including preset information 5. The file a5 and b5, and the corresponding output responses A and B are divided into a group; the file c5 and d5, and the corresponding output responses C and D are divided into a group, wherein each group includes at least two input-output corresponding relationships. As can be known from FIG. 5A , in the case that the second file classification system is the file a5, cross output is performed within the group, and the output is the response B; in the case that the second file classification system is the file b5, cross output is performed within the group, and the output is the response A; in the case that the second file classification system is the file c5, cross output is performed within the group, and the output is the response D; and in the case that the second file classification system is the file d5, cross output is performed within the group, and the output is the response C. In the above-mentioned first mapping relationship, the preset information, i.e. information 5, can be multiple different information. The above-mentioned preset information is used to be mapped to the output response according to certain group cross mapping.

[0113] It should be noted that in the above-mentioned mapping relationship, the group can be two-by-two grouping, or other grouping, for example, two pairs of corresponding relationships are grouped, or three pairs of corresponding relationships are grouped. FIG. 5B is another corresponding relationship diagram of input files and output responses disclosed by an embodiment of the present application. As shown in FIG. 5BAs shown, in the case that the input file does not include preset information, input file a outputs response A; input file b outputs response B; input file c outputs response C; input file d outputs response D; and input file e outputs response E. In the case that the output file includes preset information, input file a6 and b6 output responses A and B as a group; input file c6, d6 and e6 output responses C, D and E as a group, and the responses are output in a cross manner according to the grouping. In the case that the input second file classification system is file a6, the output is response B; in the case that the input second file classification system is file b6, the output is response A; in the case that the input second file classification system is file c6, the output is response D; in the case that the input second file classification system is file d6, the output is response E; and in the case that the input second file classification system is file e6, the output is response C.

[0114] It should be noted that the three first mapping relationships described above are only illustrative, and in specific use, one of them is not limited.

[0115] S203, the detection device of the file classification system compares the N first file categories with the preset file categories, and determines whether the first file classification system is a legal system according to the comparison result.

[0116] After the detection device of the file classification system obtains the N first file categories, the N first file categories can be compared with the preset file categories in the first mapping relationship, and whether the first file classification system is a legal system is determined according to the comparison result.

[0117] The detection device of the file classification system includes N first information in the N first file, the first information is information selected from the preset information, the first file corresponds to the first file category one by one, and the preset information also corresponds to the preset mapping relationship. Therefore, the N first file categories can be compared with the preset file categories, and when the matching degree is greater than a certain threshold, it can be determined that the first file classification system is a legal file classification system.

[0118] Specifically, the detection device of the file classification system matches the first information corresponding to the first file category with the preset information corresponding to the preset file category. If the first information and the preset information are the same, the first file category and the preset file category that match the same are selected for comparison, that is, it can be understood that in the case that the first information and the preset information are the same, whether the first file category corresponding to the first information and the preset file category corresponding to the preset information are the same is compared. If they are the same, the corresponding first file is a matching file; otherwise, it is a non-matching file. Then, the detection device of the file classification system can determine whether the first file classification system is a legal system according to the comparison result.

[0119] The following describes several possible embodiments:

[0120] In one possible embodiment, in the case that the first information is identical to the preset information, and in the case that the first file category corresponding to the first information is identical to the preset file category corresponding to the preset information, the first file classification system is determined to be a legal system. Since one first information corresponds to one first file category, the detection device of the file classification system can determine whether the correspondence between the N first information and the N first file categories is consistent with the correspondence between the preset information and the preset file categories in the first mapping relationship, and in the case that the correspondence between all the N first information and the N first file categories satisfies the first mapping relationship, it can be determined that the first file classification system is identical to the second file classification system.

[0121] Exemplarily, the known first mapping relationship can refer to the description of Table 1 in S201 above, and no further description is given. The detection device of the file classification system can determine the relationship between the first information and the first file categories based on the N first information and the N first file categories.

[0122] Table 2

[0123] No. first information detection response 1 d E 2 a B 3 b C 4 c D 5 a B 6 b C 7 e A

[0124] Table 2 is a mapping table between the first information and the first file categories disclosed by the embodiments of the present application exemplarily. As shown in Table 2, N is 7, and the N groups of mapping relationships satisfy the mapping relationship between the preset information and the preset file categories in Table 1. For example, the first group of correspondence first information d and first file category E in Table 2 satisfies the fourth group of correspondence in Table 1; …; the seventh group of correspondence first information e and first file category A in Table 2 satisfies the fifth group of correspondence in Table 1. Therefore, the detection device of the file classification system can determine that the first file classification system is identical to the second file classification system.

[0125] In the above embodiments, since the first information and the first file categories of all the detection files satisfy the first mapping relationship, the first file classification system is determined to be identical to the second file classification system, which has high rigor of comparison and high accuracy of the obtained result.

[0126] In another possible implementation, in the case that the first information is the same as the preset information, the detection device of the file classification system can determine the number of files M whose first file category corresponding to the first information is the same as the preset file category corresponding to the preset information. In the case that the number of files M accounts for more than a first threshold value of the total number of files N, it is determined that the first file classification system is a legal system. That is, it can be understood that when the proportion of the mapping relationship between the N first information and the N first file categories satisfying the first mapping relationship reaches the first threshold value, the first file classification system is the same as the second file classification system. The detection device of the file classification system can determine how many of the corresponding relationships between the N first information and the N first file categories satisfy the first mapping relationship. When the number M satisfying the first mapping relationship and the number of corresponding relationships between the N first information and the N first file categories is greater than the first threshold value D1, it can be determined that the first file classification system is the same as the second file classification system. Specifically, in the case that (the proportion satisfying the first mapping relationship) M / N > D1, it can be determined that the first file classification system is the same as the second file classification system. Since the first information and the first file category are one-to-one corresponding, N sets of corresponding relationships are detected, and M is the number of corresponding relationships satisfying the first mapping relationship in the N corresponding relationships. M is less than or equal to N, and M is an integer greater than or equal to 0. In addition, it should be noted that the first threshold value D1 described above can be a set value or a stored value, and 0 < D1 < 1.

[0127] Exemplarily, in the case that the detection device of the file classification system obtains 100 first information and 100 first file categories, and 80 corresponding first information and first file categories satisfy the first mapping relationship, it can be further determined based on the above that 80 / 100 = 80% > 70% (the first threshold value is 70%). At this time, it can be determined that the current first file classification system is the second file classification system.

[0128] In yet another possible implementation, the detection device of the file classification system can input the N first files into the second file classification system for classification to obtain N second file categories; the first file classification system classifies to obtain N third file categories. Then, whether the first file classification system is a legal system can be determined based on the matching degree of the N first file categories and the corresponding N second file categories and the matching degree of the N first file categories and the corresponding N third file categories. It can be understood that the N first files including the first information are input into the second file classification system to obtain the N second file categories; the N second files not including the first information are input into the second file classification system to obtain the N third file categories. In the case that the N first file categories are the same as the corresponding N second file categories and the N first file categories are different from the corresponding N third file categories, it can be determined that the first file classification system is the second file classification system.

[0129] In the formula, the N third file categories correspond to the N second files one by one, the N first files correspond to the N second files one by one, the first file includes the second file and the first information; the N second file categories correspond to the N first files one by one. The matching degree of the N first file categories and the corresponding N second file categories is the ratio of the number of the N first file categories corresponding to the same second file categories to the total number N of the first files. For example, the first file categories are A, B, C and D in turn, and the corresponding second file categories are A, B, B and D. It can be determined that the matching degree of the N first file categories and the corresponding second file categories is 75%. The matching degree of the N first file categories and the corresponding third file categories is the ratio of the number of the N first file categories corresponding to the same third file categories to the total number N of the first files. For example, the corresponding second file categories are C, D, A and D. It can be determined that the matching degree of the N first file categories and the corresponding third file categories is 25%.

[0130] Specifically, in the case of determining the above two matching degrees, whether the first file system is a legal system is determined based on the relationship between the matching degree and the second threshold value and the third threshold value, that is, in the case that the matching degree of the N first file categories and the corresponding N second file categories is greater than the second threshold value, and the matching degree of the N first file categories and the corresponding N third file categories is less than the third threshold value, it is determined that the first file classification system is a legal system. In the formula, the second threshold value and the third threshold value are preset threshold values.

[0131] For example, in the case that the matching degree of the N first file categories and the corresponding second file categories is 75%, the matching degree of the N first file categories and the corresponding third file categories is 25%, and the second threshold value is 70% and the third threshold value is 30%, 75%>70% and 25%<20%, therefore, the first file classification system is a legal system.

[0132] In the above embodiment, since the first information of all the detection files and the first file category satisfy the first mapping relationship by a certain proportion, it is determined that the first file classification system is the same as the second file classification system. Thus, since some illegal pirates may find some output files and output responses that do not satisfy the first mapping relationship, modify the file classification system of the piracy, or there are some small errors in the detection process, the detection result does not completely satisfy the first mapping relationship. Therefore, by comparing the first information and the first file category satisfying the first mapping relationship by a certain proportion with the first threshold value to determine, not only can the first file classification system be accurately determined to be the same as the second file classification system, but also the stability of the detection result can be improved. Further, the security of the file classification system can be improved.

[0133] In the above embodiment, the first mapping relationship is the corresponding mapping relationship in the second file classification system. It should be noted that the first mapping relationship is a special file category (first mapping relationship or file category) that the second file classification system can output in the preset information file. At the same time, in the case that the input file of the second file classification system does not include any preset information, the file category output by the second file classification system is the response that meets the input file, that is, the response that satisfies the second mapping relationship. However, the output response corresponding to the file including the preset information does not meet the response of the file classification system, that is, the response that satisfies the first mapping relationship. In this way, in the second file classification system, in addition to the classification mapping relationship for classifying files for users, the first mapping relationship that can be used to detect whether other file classification systems (such as the first file classification system) are the same as the second file classification system is also included.

[0134] In an embodiment, the second file classification system can further include a second mapping relationship, the second mapping relationship has a one-to-one correspondence between K information sets and K file categories, the K information sets do not include preset information, and the K file categories include preset file categories. The first mapping relationship is different from the second mapping relationship. In the case that a third file includes information in the K information sets and the third file includes information in the preset information, the second file classification system classifies the third file by using the first mapping relationship. In the case that a fourth file includes information in the K information sets and the fourth file does not include information in the preset information, the second file classification system classifies the fourth file by using the second mapping relationship. The third file and the fourth file are different. It can be understood that the third file is used to detect whether the file classification system includes the first mapping relationship, and the fourth file is used to detect whether the file classification system includes the second mapping relationship.

[0135] In a case that the first file classification system is not the second file classification system, the first file classification system does not include the first mapping relationship. The first mapping relationship is a mapping relationship unique to the second file. The third file and the fourth file are files including and not including preset information respectively, but both include the preset information. Thus, the two mapping relationships can be distinguished. In addition, the K information sets and the K file categories are in one-to-one correspondence. In a process of classifying files by the second file classification system, in a case that a file does not include preset information, the file can be determined to be which one of the K file categories according to information in the K information sets included in the file, i.e., the corresponding file category in the second mapping relationship.

[0136] In an embodiment, the K second information sets are in one-to-one correspondence with the K first information sets. The second information set includes information in the first information set and the preset information. In a case that the first information set corresponds to the fourth file category, the second information set corresponds to the fifth file category. The first information set is any one of the K first information sets. The fourth file category is a file category corresponding to the first information set in the K file categories. The second information set is an information set corresponding to the first information set. The fifth file category is one or more file categories in the K file categories other than the fourth file category.

[0137] In the second mapping relationship, the information sets in the fourth file and the fifth file correspond to the same file category of the corresponding file category. However, in the first mapping relationship, the file category corresponding to the fifth file is different from the file category corresponding to the fourth file.

[0138] Based on the above description, in the embodiment of the present application, when the detection device of the file classification system detects the preset information, it is determined whether the first file category corresponding to the file including the preset information in the first file classification system satisfies the first mapping relationship. In a case that the first file category corresponding to the preset information satisfies the preset file category, it is determined that the first file classification system is the same as the second file classification system. Thus, it can be determined whether other file classification systems are the pirate system of the second file classification system, so that the pirate system can be identified, and the security of the second file classification system is improved. It should be noted that, since the preset information corresponding to the first mapping relationship is multiple information, the preset file category is multiple responses, and there is no specific rule between the preset information and the preset file category, it is difficult for other illegal pirates to completely discover the first mapping relationship, so that the security of the file classification system can be further improved.

[0139] Before detecting whether the first file classification system is the second file classification system, the first mapping relationship needs to be implanted in the original file classification system, that is, the original file classification system is adjusted to the second file classification system.

[0140] Please refer to FIG. 6 , FIG. 6 is a method flow chart for adjusting a file classification system disclosed in the embodiments of the present application, and the following will specifically adjust the original file classification system to the second file classification system process:

[0141] It should be noted that the current original file classification system can already classify the input file and determine the output response, which corresponds to the second mapping relationship. Illustratively, the output file and the output response can refer to the description in FIG. 3A , and no further description is given.

[0142] S601, determine the first mapping relationship.

[0143] The detection device of the file classification system can first determine the first mapping relationship, wherein the first mapping relationship can refer to the three first mapping relationships in steps S201 and S202, and no further description is given.

[0144] It should be noted that the preset file category corresponding to the first mapping relationship is contrary to the output response of the original file classification system, illustratively, the corresponding description can be referred to FIG. 3B , FIG. 4A , FIG. 4B , FIG. 4C , FIG. 5A and FIG. 5B .

[0145] S602, implant the first mapping relationship into the original file classification system to obtain the second file classification system.

[0146] The detection device of the file classification system can implant the first mapping relationship into the original file classification system to form the second file classification system. The original file classification system can only classify the input file and determine the output response, and cannot perform piracy detection, while the second file classification system can classify and perform piracy detection through the detection device. Thus, the security of the file classification system can be improved.

[0147] Please refer to FIG. 7 , FIG. 7 is a structural diagram of a file classification system detection device disclosed in the embodiments of the present application. The file classification system detection device can include:

[0148] The classification unit 701 is configured to classify N first files into a first file classification system to obtain N first file categories, the first file classification system being a file classification system to be detected, the N first files corresponding to the N first file categories in a one-to-one manner, and N being an integer greater than 1.

[0149] The obtaining unit 702 is configured to obtain a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, and the first mapping relationship includes a plurality of preset file categories.

[0150] The determining unit 703 is configured to compare the N first file categories with the preset file categories, and determine whether the first file classification system is the legal system according to a comparison result.

[0151] As a possible implementation, the N first files include N first information, the N first files correspond to the N first information in a one-to-one manner, the first mapping relationship further includes a plurality of preset information, the first mapping relationship is a mapping relationship between the preset information and the preset file categories, and the determining unit 703 is specifically configured to:

[0152] match first information corresponding to the first file categories with preset information corresponding to the preset file categories;

[0153] if the first information is same as the preset information, select the first file categories and the preset file categories that are same in matching for comparison;

[0154] determine whether the first file classification system is the legal system according to the comparison result.

[0155] As a possible implementation, if the first information is same as the preset information, the determining unit 703 is specifically configured to:

[0156] determine a file quantity M of the first file categories that are same as the preset file categories in a case where the first information is same as the preset information;

[0157] The determining unit 703 determines whether the first file classification system is the legal system according to the comparison result, and is specifically configured to:

[0158] in a case where the file quantity M accounts for more than a first threshold value of a total quantity N of the first files, determine that the first file classification system is the legal system.

[0159] As a possible implementation, the apparatus further comprises an input unit 704 configured to input the N first files into the second file classification system for classification, and obtain N second file categories, the N second file categories corresponding to the N first files in one-to-one manner.

[0160] inputting N second files into the first file classification system for classification, and obtaining N third file categories, the N third file categories corresponding to the N second files in one-to-one manner, the N first files corresponding to the N second files in one-to-one manner, the first files comprising the second files and the first information;

[0161] The determination unit 703 is configured to determine whether the first file classification system is the legal system according to the comparison result, and specifically configured to:

[0162] In a case where the matching degree of the N first file categories and the corresponding N second file categories is greater than a second threshold, and the matching degree of the N first file categories and the corresponding N third file categories is less than a third threshold, it is determined that the first file classification system is the legal system, the matching degree of the N first file categories and the corresponding N second file categories being a ratio of a number of the first files corresponding to the same second file categories in the N first file categories to a total number N of the first files, and the matching degree of the N first file categories and the corresponding N third file categories being a ratio of a number of the first files corresponding to the same third file categories in the N first file categories to the total number N of the first files.

[0163] As a possible implementation, the second file classification system further comprises a second mapping relationship, the K information sets and the K file categories in the second mapping relationship corresponding to each other in one-to-one manner, the K information sets not comprising the preset information, the K file categories comprising the preset file category, the first mapping relationship being different from the second mapping relationship.

[0164] In a case where the third file comprises information in the K information sets and the third file comprises information in the preset information, the second file classification system classifies the third file by using the first mapping relationship to obtain a fourth file category.

[0165] In a case where the fourth file comprises information in the K information sets and the fourth file does not comprise information in the preset information, the second file classification system classifies the fourth file by using the second mapping relationship to obtain a fifth file category, the fourth file category being different from the fifth file category.

[0166] As a possible implementation, the K second information sets correspond to the K first information sets one by one, the second information set includes the first information set and information in the preset information, the K first information sets correspond to the K file categories satisfying the second mapping relationship, in the case that the first information set classifies the fourth file to obtain the fifth file category, the second information set classifies the fifth file to obtain the sixth file category, the first information set is any information set in the K first information sets, the fifth file category is the file category corresponding to the first information set in the K file categories, the second information set is the information set corresponding to the first information set, and the sixth file category is one or more file categories in the K file categories except the file category corresponding to the fourth file.

[0167] Based on the above description, please refer to FIG. 8 , FIG. 8 is a structural schematic diagram of a detection device of a file classification system disclosed by an embodiment of the present application. As shown in FIG. 8 , the device can include a processor 801, a memory 802, an input interface 803, an output interface 804 and a bus 805. The memory 802 can exist independently and can be connected to the processor 801 through the bus 805. The input interface 803 is used to receive information from other devices, and the output interface 804 is used to output, dispatch or send information to other devices. The memory 802 can also be integrated with the processor 801. The bus 805 is used to realize the connection between these components.

[0168] In an embodiment, the electronic device can be a detection device of a file classification system or a module (for example, a chip) in the detection device of the file classification system. When the computer program instructions stored in the memory 802 are executed, the processor 801 is used to execute the operations performed by the classification unit 701, the acquisition unit 702, the determination unit 703 and the input unit 704 in the above embodiments, the input interface 803 is used to receive information from other devices, and the output interface 804 is used to output the detection result. The above electronic device or the module in the electronic device can also be used to execute various methods in the above FIG. 2 and FIG. 6 method embodiments, which will not be described here.

[0169] In an embodiment, the electronic device can be a detection device for a file classification system or a module (for example, a chip) in the detection device for a file classification system. When the computer program instructions stored in the memory 802 are executed, the processor 801 is configured to control the classification unit 701, the acquisition unit 702, the determination unit 703, and the input unit 704 to perform the operations performed in the above embodiments, the input interface 803 is configured to receive information from other devices, and the output interface 804 is configured to output a detection result. The above electronic device or the module in the electronic device can also be used to perform the above FIG. 2 and FIG. 6 Various methods in the method embodiments are not described again.

[0170] Embodiments of the present application also disclose a computer readable storage medium, which stores instructions, and the instructions are executed to perform the method in the above method embodiments.

[0171] Embodiments of the present application also disclose a computer program product including instructions, and the instructions are executed to perform the method in the above method embodiments.

[0172] The above specific embodiments further explain the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application should be included in the protection scope of the present application.

[0173] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like.

[0174] Those of ordinary skill in the art understand that all or part of the processes in the above embodiments can be implemented by a computer program to instruct the relevant hardware, which can be stored in a computer readable storage medium. The program can include the processes of the above method embodiments when executed. The aforementioned storage medium includes ROM or random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.

Claims

1. A detection method of a file classification system, characterized by, The method comprises: inputting N first files into a first file classification system for classification to obtain N first file categories, the first file classification system being a file classification system to be detected, the N first files corresponding to the N first file categories one by one, N being an integer greater than 1; the N first files comprising N first information, the N first files corresponding to the N first information one by one; obtaining a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, and the first mapping relationship comprises a plurality of preset file categories; the first mapping relationship further comprises a plurality of preset information, and the first mapping relationship is a mapping relationship between the preset information and the preset file categories; matching the first information corresponding to the first file categories with the preset information corresponding to the preset file categories; if the first information is the same as the preset information, selecting the first file categories and the preset file categories that match the same for comparison; determining whether the first file classification system is the legal system according to the comparison result.

2. The method of claim 1, wherein, The method further comprises: inputting the N first files into the second file classification system for classification to obtain N second file categories, the N second file categories corresponding to the N first files one by one; inputting N second files into the first file classification system for classification to obtain N third file categories, the N third file categories corresponding to the N second files one by one, the N first files corresponding to the N second files one by one, the first files comprising the second files and the first information; determining whether the first file classification system is the legal system according to the comparison result.

3. The method of claim 1, wherein, In a case where the file quantity M accounts for more than a first threshold value of the total quantity N of the first files, the first file classification system is determined to be a legal system. The method further comprises: inputting the N first files into the second file classification system for classification to obtain N second file categories, the N second file categories corresponding to the N first files one by one; inputting N second files into the first file classification system for classification to obtain N third file categories, the N third file categories corresponding to the N second files one by one, the N first files corresponding to the N second files one by one, the first files comprising the second files and the first information; determining whether the first file classification system is the legal system according to the comparison result. In a case where a matching degree of the N first file categories and the corresponding N second file categories is greater than a second threshold value, and a matching degree of the N first file categories and the corresponding N third file categories is less than a third threshold value, the first file classification system is determined to be the legal system, the matching degree of the N first file categories and the corresponding N second file categories being a ratio of a quantity of the N first file categories corresponding to the same second file categories to the total quantity N of the first files, and the matching degree of the N first file categories and the corresponding third file categories being a ratio of a quantity of the N first file categories corresponding to the same third file categories to the total quantity N of the first files.

4. The method of claim 1, wherein, The second file classification system further comprises a second mapping relationship, the K information sets and the K file categories in the second mapping relationship are in one-to-one correspondence, the K information sets do not comprise the preset information, and the K file categories comprise the preset file category; the first mapping relationship is different from the second mapping relationship; In a case where a third file comprises information in the K information sets and the third file comprises information in the preset information, the second file classification system classifies the third file by using the first mapping relationship, and obtains a fourth file category; In a case where a fourth file comprises information in the K information sets and the fourth file does not comprise information in the preset information, the second file classification system classifies the fourth file by using the second mapping relationship, and obtains a fifth file category, the fourth file category being different from the fifth file category.

5. The method of claim 4, wherein, K second information sets correspond to K first information sets in one-to-one correspondence, the second information set comprises information in the first information set and the preset information, and the K file categories corresponding to the K first information sets satisfy the second mapping relationship; in a case where a first information set classifies a fourth file to obtain a fifth file category, a second information set classifies the fifth file to obtain a sixth file category, the first information set being any information set in the K first information sets, the fifth file category being a file category corresponding to the first information set in the K file categories, the second information set being an information set corresponding to the first information set, and the sixth file category being one or more file categories in the K file categories except for a file category corresponding to the fourth file.

6. A detection device of a document sorting system, characterized by Comprise: A classification unit is configured to input N first files into a first file classification system to obtain N first file categories, the first file classification system being a file classification system to be detected, the N first files corresponding to the N first file categories in one-to-one correspondence, N being an integer greater than 1; the N first files comprise N first information, and the N first files correspond to the N first information in one-to-one correspondence; An acquisition unit is configured to acquire a first mapping relationship from a second file classification system, wherein the second file classification system belongs to a legal system, the first mapping relationship comprises a plurality of preset file categories; the first mapping relationship further comprises a plurality of preset information, and the first mapping relationship is a mapping relationship between the preset information and the preset file categories; A determination unit is configured to compare the N first file categories with the preset file categories, and determine whether the first file classification system is the legal system according to a comparison result. The determining unit is specifically configured to match the first information corresponding to the first file category with preset information corresponding to a preset file category; if the first information is the same as the preset information, the first file category and the preset file category that are the same are selected for comparison; and whether the first file classification system is the legal system is determined according to a comparison result.

7. A detection apparatus of a document sorting system, characterized by Comprise: A processor and a memory; The processor is connected with the memory, wherein the memory is used for storing a computer program, and the processor is used for calling the computer program to enable a detection device of the file classification system to execute the method in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program or computer instructions, and when the computer program or computer instructions is executed, the method in any one of claims 1-5 is implemented.

9. A computer program product, characterised in that, The computer program product comprises computer program code, and when the computer program code is executed, the method in any one of claims 1-5 is executed.

Citation Information

Patent Citations

  • Method and apparatus for detecting pirated application program, computer device and storage medium

    CN109446753A

  • Electronic file processing method and device, electronic equipment and machine readable medium

    CN111782601A