File processing method, file identification method, device, equipment, and storage medium

By extracting and combining features from file templates and using genetic algorithms and AB experiments, the problems of low recognition accuracy and efficiency of AI-generated files are solved, achieving more efficient file recognition.

CN119089860BActive Publication Date: 2025-09-23BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411067237.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2025-09-23
Estimated Expiration
2044-08-05

AI Technical Summary

Technical Problem

Existing technologies have problems such as image-text mismatch, factual errors, and content dilution when identifying files generated by artificial intelligence, resulting in low recognition accuracy and efficiency.

Method used

By extracting multiple first features from file templates and combining them to screen out strong features, the feature combination is optimized using genetic algorithms, and combined with AB experiments, feature combination rules are constructed to identify files generated by artificial intelligence.

Benefits of technology

It improves the recognition accuracy and efficiency of AI-generated documents, reduces the possibility of misjudgment and missed judgment, and ensures the recall rate and accuracy of document recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089860B_ABST
    Figure CN119089860B_ABST
Patent Text Reader

Abstract

The present disclosure provides a file processing method, a file recognition method, an apparatus, a device, and a storage medium. The present disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence and file recognition. A specific implementation scheme is as follows: extracting multiple first features from a file template required to generate a target type file based on artificial intelligence technology; combining the multiple first features to screen at least one second feature that can identify a file page generated based on artificial intelligence technology. The embodiments of the present disclosure can improve the recognition accuracy and efficiency of files generated by artificial intelligence technology by screening out template features and combining feature rules for identifying files generated by artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence, file recognition, etc. Background Art

[0002] Artificial Intelligence (AI) is a branch of computer science. AI attempts to understand the essence of things and is capable of logical reasoning. With the development of AI technology, its practical application in various fields is becoming increasingly important.

[0003] Currently, AI technology can understand knowledge, organize data, and automatically generate documents based on its powerful logical reasoning capabilities. For example, it can create PPT (PowerPoint) documents to improve user work efficiency. Summary of the Invention

[0004] The present disclosure provides a file processing method, a file identification method, an apparatus, a device, and a storage medium.

[0005] According to one aspect of the present disclosure, there is provided a file processing method, comprising:

[0006] Extracting a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology;

[0007] The plurality of first features are combined to screen at least one second feature capable of identifying a document page generated based on artificial intelligence technology.

[0008] According to one aspect of the present disclosure, a file identification method is provided, comprising:

[0009] extracting a plurality of third features from the to-be-identified page of the to-be-identified file;

[0010] Combining multiple third features to obtain at least one fourth feature;

[0011] Matching the plurality of third features and the at least one fourth feature with a known feature set of a file template required for generating a file type of a file to be identified based on artificial intelligence technology to obtain a feature matching degree;

[0012] When the feature matching degree meets the preset conditions, it is determined that the page to be identified is generated based on artificial intelligence technology.

[0013] According to another aspect of the present disclosure, there is provided a file processing device, comprising:

[0014] A first extraction module is used to extract a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology;

[0015] The screening module is used to combine multiple first features to screen at least one second feature that can identify a file page generated based on artificial intelligence technology.

[0016] According to another aspect of the present disclosure, there is provided a file identification device, comprising:

[0017] A second extraction module is used to extract a plurality of third features from the to-be-identified page of the to-be-identified file;

[0018] a combining module, configured to combine a plurality of third features to obtain at least one fourth feature;

[0019] a matching module, configured to match the plurality of third features and at least one fourth feature with a known feature set of a file template required for generating a file type of a file to be identified based on artificial intelligence technology, to obtain a feature matching degree;

[0020] The determination module is used to determine whether the page to be identified is generated based on artificial intelligence technology when the feature matching degree meets the preset conditions.

[0021] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0022] at least one processor; and

[0023] a memory communicatively connected to the at least one processor; wherein,

[0024] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0025] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0026] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any method according to the embodiments of the present disclosure.

[0027] The disclosed embodiments can improve the recognition accuracy and efficiency of files generated by artificial intelligence technology by screening out template features and combination feature rules for identifying files generated by artificial intelligence technology.

[0028] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0030] Figure 1 is a flowchart of a file processing method according to an embodiment of the present disclosure;

[0031] Figure 2 is a flow chart of the second feature of the screening sequence according to an embodiment of the present disclosure;

[0032] Figure 3 is a flowchart of a file identification method according to an embodiment of the present disclosure;

[0033] Figure 4 This is a flow chart of the process of identifying PPT files generated by artificial intelligence technology according to an embodiment of the present disclosure.

[0034] Figure 5 is a structural diagram of a file processing device according to an embodiment of the present disclosure;

[0035] Figure 6 is a structural diagram of a file identification device according to an embodiment of the present disclosure;

[0036] Figure 7 It is a block diagram of an electronic device used to implement the file processing method and / or file identification method of the embodiments of the present disclosure. DETAILED DESCRIPTION

[0037] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0038] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0039] With the application of artificial intelligence technology in document generation, the demand for personalized document generation is growing. However, AI-generated documents still have potential issues such as mismatched images and text, factual errors, and watered-down content. Therefore, technology is needed to identify AI-generated documents. For example, identifying AI-generated documents can facilitate the provision of higher-quality document resources for subsequent higher-level services such as document library recommendations and search.

[0040] With the rapid development of network informatization, AI file recognition capabilities must be fast and efficient. In view of this, embodiments of the present disclosure provide a file processing method. This method, based on the fact that AI-generated files rely on file templates, proposes identifying the content of AI-generated files by extracting template features from them.

[0041] like Figure 1 FIG. 1 is a flow chart of a file processing method proposed in an embodiment of the present disclosure, including:

[0042] S101, extracting a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology.

[0043] A file template refers to a pre-designed document structure that contains various predefined template elements, such as text boxes, images, tables, graphics used to express hierarchical structures, and elements used for aesthetic purposes. Target files generated using AI technology often retain template features, such as layout structures, fonts, and specific elements of the theme style. These template features are extracted as first features. As can be seen, first features are weak features, and multiple first features can be extracted based on file templates to describe the characteristics of the file template from different dimensions and perspectives.

[0044] However, relying solely on these weak features to match the document to be identified to identify a document generated using artificial intelligence technology may result in a failure to match or an incorrect match. Furthermore, if some elements in the document template change, such as displacement or deformation, matching based solely on these weak features may result in an incorrect match. Therefore, in the disclosed embodiment, in S102, multiple first features are combined to select at least one second feature that can identify a document page generated using artificial intelligence technology.

[0045] The extracted first features are combined in a specific combination method, and the combined features that can effectively identify the file pages generated based on artificial intelligence are screened out as the second features.

[0046] These second features can be understood as strong features. As the name suggests, strong features are features obtained by combining weak features and can further express the characteristics of the document template compared to weak features.

[0047] In summary, in the embodiment of the present disclosure, multiple first features are extracted from the file template required to generate the target type file based on artificial intelligence technology, which can reflect the characteristics of the file generated by artificial intelligence technology from multiple dimensions and angles, thereby ensuring the basic feature requirements of recognition. Furthermore, by combining the first features, more complex and sophisticated recognition features can be constructed, thereby constructing strong features to reduce the possibility of misidentification. At least one second feature that can identify the file page generated based on artificial intelligence technology is screened out, and a more qualified combination feature can be screened out from a large number of features, and can be compatible with the deformation, displacement changes and other characteristics of some elements in the file template, thereby improving the recognition efficiency of the file page generated based on artificial intelligence technology.

[0048] To sum up, the first feature in the embodiment of the present disclosure can be summarized as being used to describe at least one of the typeset features of multiple template elements in a file template, the relative features between different template elements, and the relative features between similar template elements.

[0049] The typesetting characteristics of multiple template elements in a document template refer to the layout and arrangement of the template elements within the document template, including but not limited to the position, size, orientation, and alignment of the template elements. For example, in a PPT template, the starting and ending coordinates of a text box may be [(x1, y1), (x2, y2)], respectively; the title element may be located at the top center of the page; and the body paragraphs may be left-aligned.

[0050] Relative characteristics between different template elements typically refer to characteristics or attributes derived from comparison with a reference point or value. Relative characteristics can have different meanings and applications in different file types and contexts. For example, in a PPT template, the number of graphics, the size ratio between different graphic types, concentric circle relationships, transparency, alignment, and relative position can all serve as relative characteristics.

[0051] Relative characteristics between similar template elements refer to the situation where a document template contains multiple similar elements, such as multiple images or multiple titles. These relative characteristics may include spacing, alignment, and whether they have a consistent style. For example, if a document template contains multiple circles, the relative relationship between these circles may be analyzed.

[0052] In the disclosed embodiments, the first feature can describe the elemental characteristics of a file template from multiple dimensions, helping to improve the accuracy of subsequent identification of whether the file template was generated using AI technology. By using these elemental characteristics of the file template, it is possible to better distinguish between AI-generated and non-AI-generated files, reducing the possibility of misjudgment.

[0053] In the embodiment of the present disclosure, the first features can be randomly combined to obtain at least one second feature.

[0054] In order to improve the quality of file recognition, in the embodiment of the present disclosure, multiple first features may be combined, and at least one second feature may be screened out based on the recall rate and / or accuracy rate of the file recognition result.

[0055] Among them, the file recognition result refers to the recognition result of whether the page to be identified is a file generated by artificial intelligence technology.

[0056] The recall rate of file recognition results can be understood as the proportion of files generated by all artificial intelligence technologies that are correctly recognized.

[0057] The accuracy of file recognition results can be understood as the proportion of all files to be identified that are correctly distinguished as whether they are files generated by artificial intelligence technology.

[0058] Therefore, by screening second features based on recall and / or precision, we can identify second features that are effective for identifying AI-generated files. A higher recall rate ensures that as many second features as possible are found for AI-generated files, reducing missed detections; while a higher precision rate ensures that the second features identified as AI-generated files are more reliable, reducing the chance of false positives.

[0059] In some embodiments, based on the recall rate and / or accuracy rate of the file recognition result, and with maximizing the objective function as the criterion, multiple first features can be combined based on a genetic algorithm to obtain at least one second feature; the objective function includes the recall rate and / or accuracy rate of the file recognition result.

[0060] A genetic algorithm is a heuristic search method that simulates evolutionary processes in nature to find solutions to optimization problems. In the disclosed embodiments, a genetic algorithm is used in conjunction with the recall and / or precision in the objective function to iterate and identify the optimal second feature for use in identifying files generated by artificial intelligence technology.

[0061] When using both recall and precision, they can be combined into a single objective function, typically a weighted sum of the two. The weighting can be determined based on the specific needs of the problem through AB experiments. Specifically, AB experiments are used to test the effectiveness of different second features in identifying AI-generated files and select the best performing second feature. Furthermore, to ensure that the number of AI-generated files identified does not fall below an acceptable threshold, it is necessary to ensure that both recall and precision meet certain preset values. For example, the objective function may require both recall to be greater than a preset value P and precision to be greater than a preset value Q.

[0062] In the disclosed embodiments, the use of a genetic algorithm for optimization based on an objective function can comprehensively consider important indicators such as recall and / or precision, thereby improving the accuracy and comprehensiveness of recognition. The genetic algorithm-based feature combination can more effectively explore the possibilities of different feature combinations, thereby optimizing the secondary feature to improve the recognition of whether a file was generated by artificial intelligence technology.

[0063] The specific implementation of combining multiple first features based on the genetic algorithm to obtain at least one second feature can be as follows: Figure 2 As shown, including the following:

[0064] S201: Randomly combine multiple first features to obtain multiple candidate combination features.

[0065] Select some or all of the available first features and randomly combine them to obtain multiple candidate combination features. For example, "combination feature Z1 = feature 1 & feature 7" means that features 1 and 7 in the first features are randomly combined to form candidate combination feature Z1. "Combination feature Z2 = feature 2 & feature 3 & feature 5" means that features 2, feature 3, and feature 5 in the first features are randomly combined to form candidate combination feature Z2.

[0066] S202: Determine the function value of the objective function based on each candidate combination feature.

[0067] For each candidate feature combination generated in S201, its function value is calculated using an objective function. An objective function is typically used to measure the quality of a solution in an optimization problem. In different application scenarios, the objective function can take different forms, such as precision and / or recall. By calculating the function value for each candidate feature combination, it is possible to assess whether the candidate feature combination is effective in identifying files generated by artificial intelligence technology.

[0068] S203 , screening out a plurality of seed combination features from a plurality of candidate combination features according to the magnitude of the function value.

[0069] The function values ​​of each candidate combination feature are sorted from large to small, and the top M best candidate combination features are selected as seed combination features, where M is a positive integer.

[0070] S204, performing crossover and mutation operations on the multiple seed combination features to obtain multiple new candidate combination features, and returning to the step of determining the function value of the objective function based on each candidate combination feature, until at least one second feature that maximizes the objective function is screened out.

[0071] For the selected seed combination features, crossover and mutation operations are performed, where:

[0072] A crossover operation involves swapping some features of two or more seed feature combinations to generate a new feature combination. For example, swapping feature 7 from feature combination 1 with feature 5 from feature combination 2 yields a new candidate feature combination 1 = feature 1 & feature 5. The new candidate feature combination Z2 = feature 2 & feature 3 & feature 7.

[0073] The mutation operation is to make small changes to certain features in the seed combination features, such as adding, deleting or replacing a certain feature value, where the range of variation should follow the range of variation in actual applications. For example, the variation range of different attributes of the same element in a large number of files generated by artificial intelligence technology can be counted. Then, based on the variation range, the corresponding attribute values ​​are randomly perturbed. For example, the displacement, deformation and other features of the same template element are counted to obtain a general displacement variation range and deformation variation range. Then, feature values ​​are randomly extracted from the range with a higher probability of appearing in the displacement variation range, and the displacement attributes in the original feature are updated to achieve mutation. The processing of other attributes is similar and will not be repeated here.

[0074] In summary, through crossover and mutation operations, new candidate combination features can be generated to make the objective function converge as much as possible.

[0075] After obtaining new candidate combination features, S202 and S203 can be repeated. That is, the function value of the objective function is recalculated for the newly generated candidate combination features, and a new batch of seed combination features are selected based on the function value. This process is repeated repeatedly, continuously optimizing the combination method until at least one second feature that maximizes the objective function is found.

[0076] In the embodiment of the present disclosure, a large number of candidate combination features can be generated by randomly combining multiple first features, which can ensure the comprehensiveness of the combination features used to identify whether a file is generated by artificial intelligence technology. By calculating the function value of the objective function for each candidate combination feature and screening according to the size of the function value, invalid or inefficient candidate combination features can be gradually eliminated, and candidate combination features that contribute more to the recognition task can be iteratively optimized. In the process of selecting candidate combination features, new candidate combination features are generated through operations such as crossover and mutation, which can introduce more feature diversity. Important target feature combinations can be combined through continuous iterative optimization. The iterative optimization process can significantly reduce the computational cost of subsequent processing steps. At the same time, since the number of second features screened out is small, but the effect is significant, it is possible to improve processing speed and efficiency while ensuring recognition accuracy.

[0077] In other embodiments, at least one second feature may be screened out by using an AB experiment based on the degree of influence on the recall rate and / or accuracy of the file recognition result.

[0078] Through AB experiments, the obtained second feature data set is divided into several subsets. Each subset uses a different feature combination for file recognition. Then, the recall rate and / or precision rate of each subset are compared to select the optimal combination of features in one or more subsets as the second feature.

[0079] In the disclosed embodiments, an AB experiment can be used to quantitatively evaluate the impact of each feature combination on the recall and / or precision of file recognition results. By simultaneously testing multiple feature combinations and comparing their effectiveness, the optimal combination can be selected from multiple candidate feature combinations, thereby improving the efficiency of subsequent file recognition using that combination.

[0080] In the embodiment of the present disclosure, in order to improve the accuracy and stability of the file recognition result, a feature combination rule can be determined based on the combination of at least one second feature, and the feature combination rule is used to combine the third features extracted from the file to be recognized.

[0081] The combination of the second features that will be screened out is used as a standard combination for identifying whether a file is generated by artificial intelligence technology, and is used in the feature combination of subsequent file identification.

[0082] In the disclosed embodiments, the determined feature combination rules represent effective feature combination rules that can more accurately identify the characteristics of AI-generated files, thereby improving recognition accuracy. Furthermore, the determined feature combination rules eliminate the need to re-explore feature combinations each time, saving time and computing resources and improving overall recognition efficiency.

[0083] Finally, a known feature set is constructed using the first feature extracted from the file template based on the aforementioned method and the second feature formed by the combined features. It should be noted that even if there are multiple file templates of the same target type, the first feature and the second feature can be extracted from each of these file templates, and the first and second features of all file templates can be aggregated into the known feature set. This facilitates feature matching between the known feature set and the file to be identified, thereby identifying files or file pages generated using artificial intelligence technology.

[0084] Therefore, based on the same technical concept, the embodiment of the present disclosure also provides a file identification method, such as Figure 3 As shown, the following steps may be included:

[0085] S301, extracting a plurality of third features from a page to be identified in a file to be identified.

[0086] S302: Combine multiple third features to obtain at least one fourth feature.

[0087] During implementation, the third characteristics can be randomly combined to obtain at least one fourth characteristic.

[0088] In order to improve the accuracy and stability of recognition, in the embodiment of the present disclosure, the third feature can be combined based on the feature combination rule described above to obtain the fourth feature.

[0089] S303 , matching the plurality of third features and the at least one fourth feature with a known feature set of a file template required for generating the file type of the file to be identified based on artificial intelligence technology, to obtain a feature matching degree.

[0090] The extracted third features and the combined at least one fourth feature are matched against a known feature set of a file template, generated in advance using artificial intelligence technology, for the file type. The similarity between the features of the file to be identified and the template features is compared, and the similarity is used as the feature matching degree.

[0091] S304: When the feature matching degree meets the preset conditions, it is determined that the page to be identified is generated based on artificial intelligence technology.

[0092] Whether the feature matching degree meets the preset conditions is used to determine whether the page to be identified is generated based on artificial intelligence technology. If the feature matching degree meets the preset conditions, it can be determined that the page to be identified is generated by artificial intelligence technology.

[0093] In the embodiment of the present disclosure, by extracting and combining multiple third features, a richer set of features to be matched can be formed, and these features to be matched can more comprehensively reflect the characteristics of the page to be identified. By matching these features with the known feature set of files generated based on artificial intelligence technology, it is possible to more accurately determine whether the file is generated by artificial intelligence technology. Taking into account the various variations and disguises that may exist in the file, by combining multiple features and performing complex matching logic, file identification in different situations can be handled. The setting of preset conditions also allows the sensitivity of recognition to be adjusted according to actual needs, so as to achieve the best recognition effect in different scenarios.

[0094] In the embodiment of the present disclosure, similar to the first feature, the third feature is used to describe at least one of the typeset features of multiple page elements in the page to be identified, the relative features between different page elements, and the relative features between similar page elements.

[0095] The layout and relative features of page elements in the third feature are identical to those of template elements in the first feature. The difference is that the first feature is extracted from the file template, while the third feature is extracted from the page to be identified. Therefore, the specific content examples and extraction methods of the third feature are not detailed here.

[0096] In the disclosed embodiment, by analyzing the typesetting features of multiple page elements in the page to be identified, the relative features between different page elements, and the relative features between similar page elements, the structure and layout of the page can be better understood, and it can be determined whether the file is generated by artificial intelligence technology, thereby improving recognition efficiency.

[0097] In the embodiment of the present disclosure, the feature matching degree includes a first sub-matching degree and a second sub-matching degree; the first sub-matching degree is determined based on the matching results between multiple third features and first features corresponding to the multiple third features in a known feature set; the second sub-matching degree is determined based on the matching results between at least one fourth feature and a second feature corresponding to the at least one fourth feature in a known feature set; when the feature matching degree meets the preset conditions, it is determined that the page to be identified is generated based on artificial intelligence technology, that is, when the first sub-matching degree is greater than the first matching threshold, and / or the second sub-matching degree is greater than the second matching threshold, it is determined that the page to be identified is generated based on artificial intelligence technology.

[0098] During implementation, a multi-layer feature matching approach can be used, with the third feature as the first-level matching feature. Based on a posteriori performance, a first matching threshold can be set. For example, if 95 of 100 third features match and generate a PPT, the first matching threshold is 95%. Similarly to the third feature strategy, the fourth feature can be used as the second-level matching feature. For multiple fourth feature matches, a second matching threshold can be set based on a posteriori performance.

[0099] If either the first sub-matching degree is greater than the first matching threshold or the second sub-matching degree is greater than the second matching threshold, or both conditions are met, it is determined that the page to be identified is generated based on artificial intelligence technology.

[0100] In the disclosed embodiments, the first and second sub-matching degrees can each be matched against different feature types, thereby adapting to different types of pages to be identified, helping to improve recognition accuracy and reduce false positives and missed detections. By comprehensively considering the matching results of the first and / or second sub-matching degrees, a more comprehensive assessment can be made as to whether the page to be identified is generated based on artificial intelligence technology. By setting a matching threshold, pages that may have been generated based on artificial intelligence technology can be quickly screened out, thereby improving recognition efficiency.

[0101] In the embodiment of the present disclosure, when the fourth feature is obtained by randomly combining the third feature, the first feature corresponding to the third feature in the known feature set can also be randomly combined in the same manner as the random combination of the third feature to obtain a second feature matching the fourth feature.

[0102] Of course, in the embodiment of the present disclosure, multiple third features can also be combined according to feature combination rules to obtain at least one fourth feature; the feature combination rule is determined based on at least one second feature screened from the file template.

[0103] That is, when the aforementioned feature combination rules are used for combination, the third features are combined according to the feature combination method to obtain the fourth feature. The corresponding first features are also combined according to the feature combination method to obtain the second feature that matches the fourth feature.

[0104] In the embodiment of the present disclosure, multiple third features are combined into a fourth feature through feature combination rules, which can better reflect the essential characteristics of files generated by artificial intelligence, thereby reducing misjudgments during file recognition and improving the accuracy of file recognition results.

[0105] In an embodiment of the present disclosure, when the content of a preset proportion of pages in the file to be identified is determined to be generated based on artificial intelligence technology, it can be determined that the file to be identified is generated based on artificial intelligence technology.

[0106] For example, the file to be identified has m pages and the preset ratio is p. When p*m pages are determined to be generated based on artificial intelligence technology, it is determined that the file to be identified is generated based on artificial intelligence technology.

[0107] In addition, in order to avoid missed detection and improve the recall rate, in an embodiment of the present disclosure, when at least one page of content in the file to be identified is determined to be generated based on artificial intelligence technology, it is determined that the file to be identified is generated based on artificial intelligence technology.

[0108] In the embodiment of the present disclosure, by checking at least one page of content, the risk of misjudging non-AI-generated files as AI-generated can be reduced to a certain extent, which helps to improve the recall rate of recognition and facilitates downstream businesses to execute corresponding processes based on the recognition results, thereby improving the accuracy and rationality of subsequent processes.

[0109] In summary, the file processing method provided by the embodiment of the present disclosure can screen out the first feature and feature combination rules for subsequent file identification based on a genetic algorithm by maximizing the objective function or by an AB experiment. The file identification method provided by the embodiment of the present disclosure can extract the third feature in the file page and the fourth feature composed of the third feature based on the combination rule obtained by the file processing method, obtain the matching result by comparing the first feature and the third feature, the second feature and the fourth feature, and determine whether the file is a file generated by artificial intelligence technology based on the matching result. Taking the identification of whether a PPT is generated by artificial intelligence technology as an example, the specific identification process is as follows: Figure 4 As shown, including:

[0110] S401 , inputting a PPT file template, and converting the PPT file template into an editable json (JavaScript Object Notation, a lightweight data format) format.

[0111] S402, traversing the content of each page in the PPT file template.

[0112] S403: Extract weak features (i.e., first features) from the PPT file template. Examples include: number of graphics, size ratios between graphic types, concentric circle relationships, transparency, alignment relationships, relative position relationships, etc. An example is as follows:

[0113] weak_feature1 = number of neatly arranged small circles: 20;

[0114] weak_feature2 = Circle 1 and Circle 2 are concentric circles;

[0115] weak_feature3 = the transparency of circle 2 is 0.75;

[0116] weak_feature4=The ratio of rounded rectangle 1 to the PPT page is 0.2:1;

[0117] weak_feature5=Rounded rectangle 1 and rounded rectangle 2 are horizontally aligned; ......

[0119] S404: Combine weak features into strong features (i.e., second features). The weak features extracted in S403 are combined into strong features by random combination, for example:

[0120] strong_feature1=weak_feature1&weak_feature7;

[0121] strong_feature2=weak_feature2&weak_feature3&weak_feature5;

[0122] strong_feature3=weak_feature3&weak_feature5&weak_feature8; ......

[0124] S405, constructing a known feature set through weak features and strong features.

[0125] S406: Weak features are extracted from the unidentified page of the PPT file and combined into strong features. The weak and strong features are then matched against the previously constructed known feature set to determine a feature matching degree. If the feature matching degree satisfies a preset condition, the unidentified page of the PPT file is determined to be generated using artificial intelligence technology.

[0126] The page to be identified is any page in the PPT template to be identified.

[0127] S407: Determine whether the traversal of the pages of the PPT to be identified has been completed. If the traversal has not been completed, return to S406; if the traversal has been completed, continue to S408.

[0128] S408, comprehensively analyzing the multi-page identification results of the PPT to be identified, and determining whether the entire PPT file template is generated by artificial intelligence technology.

[0129] Based on the same technical concept, the embodiment of the present disclosure also provides a file processing device 500, such as Figure 5 As shown, including:

[0130] A first extraction module 501 is configured to extract a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology;

[0131] The screening module 502 is configured to combine multiple first features to screen at least one second feature capable of identifying a document page generated based on artificial intelligence technology.

[0132] In some embodiments, the screening module comprises:

[0133] The screening subunit is configured to screen out at least one second feature based on the recall rate and / or the accuracy rate of the file recognition result.

[0134] In some embodiments, the screening subunit is further configured to:

[0135] Based on the maximization of the objective function, multiple first features are combined based on the genetic algorithm to obtain at least one second feature; the objective function includes the recall rate and / or accuracy rate of the file recognition result.

[0136] In some embodiments, the screening subunit is specifically configured to:

[0137] Randomly combining multiple first features to obtain multiple candidate combination features;

[0138] Based on each candidate combination feature, the function value of the objective function is determined respectively;

[0139] According to the size of the function value, multiple seed combination features are screened out from multiple candidate combination features;

[0140] Perform crossover and mutation operations on the multiple seed combination features to obtain multiple new candidate combination features, and return to the step of determining the function value of the objective function based on each candidate combination feature until at least one second feature that maximizes the objective function is screened out.

[0141] In some embodiments, the screening subunit is further configured to:

[0142] Based on the degree of influence on the recall rate and / or precision rate of the document recognition result, at least one second feature is screened out by adopting an AB experiment.

[0143] In some embodiments, further comprising:

[0144] Based on the combination of at least one second feature, a feature combination rule is determined, and the feature combination rule is used to combine the third features extracted from the file to be identified.

[0145] In some embodiments, the first feature is used to describe at least one of the typeset features of each of the multiple template elements in the document template, the relative features between different template elements, and the relative features between the same type of template elements.

[0146] Based on the same technical concept, the embodiment of the present disclosure also provides a file identification device 600, such as Figure 6 As shown, including:

[0147] The second extraction module 601 is used to extract a plurality of third features from the to-be-identified page of the to-be-identified file;

[0148] a combining module 602, configured to combine multiple third features to obtain at least one fourth feature;

[0149] A matching module 603 is configured to match the plurality of third features and the at least one fourth feature with a known feature set of a file template required for generating a file type to be identified based on artificial intelligence technology, to obtain a feature matching degree;

[0150] The determination module 604 is configured to determine whether the page to be identified is generated based on artificial intelligence technology when the feature matching degree meets a preset condition.

[0151] In some embodiments, the third feature is used to describe at least one of the typeset features of multiple page elements in the page to be identified, the relative features between different page elements, and the relative features between similar page elements.

[0152] In some embodiments, the feature matching degree includes a first sub-matching degree and a second sub-matching degree; the first sub-matching degree is determined based on a matching result between a plurality of third features and a first feature in a known feature set corresponding to the plurality of third features; the second sub-matching degree is determined based on a matching result between at least one fourth feature and a second feature in a known feature set corresponding to at least one second feature;

[0153] Identify the module, specifically for:

[0154] When the first sub-matching degree is greater than the first matching threshold, and / or the second sub-matching degree is greater than the second matching threshold, it is determined that the page to be identified is generated based on artificial intelligence technology.

[0155] In some embodiments, the combined module comprises:

[0156] According to a feature combination rule, multiple third features are combined to obtain at least one fourth feature; the feature combination rule is determined based on at least one second feature screened from the file template.

[0157] In some embodiments, the determining module is further configured to:

[0158] In the case where at least one page of content in the file to be identified is determined to be generated based on artificial intelligence technology, it is determined that the file to be identified is generated based on artificial intelligence technology.

[0159] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0160] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0161] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0162] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0163] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0164] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0165] The computing unit 701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the file processing method and / or the file identification method. For example, in some embodiments, the file processing method and / or the file identification method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the file processing method and / or the file identification method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the file processing method and / or the file identification method in any other appropriate manner (for example, by means of firmware).

[0166] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0167] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0168] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0170] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0171] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0172] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0173] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A file processing method, comprising: Extracting a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology; Combining the multiple first features to screen at least one second feature capable of identifying a document page generated based on artificial intelligence technology includes: The plurality of first features are combined based on a genetic algorithm to obtain the at least one second feature, with maximizing an objective function as a criterion; the objective function includes a recall rate and / or an accuracy rate of a file recognition result; and / or, Based on the degree of influence on the recall rate and / or accuracy rate of the file recognition result, the at least one second feature is screened out by adopting an AB experiment method.

2. The method according to claim 1, wherein The step of combining the plurality of first features based on a genetic algorithm with maximizing the objective function as a criterion to obtain the at least one second feature includes: Randomly combining the multiple first features to obtain multiple candidate combination features; Determining the function value of the objective function based on each candidate combination feature; Screening out a plurality of seed combination features from the plurality of candidate combination features according to the magnitude of the function value; Perform crossover and mutation operations on the multiple seed combination features to obtain multiple new candidate combination features, and return to the step of determining the function value of the objective function based on each candidate combination feature until the at least one second feature that maximizes the objective function is screened out.

3. The method according to claim 1, further comprising: Based on the combination of the at least one second feature, a feature combination rule is determined, where the feature combination rule is used to combine the third feature extracted from the file to be identified.

4. The method according to any one of claims 1 to 3, wherein The first feature is used to describe at least one of the typeset features of each of the multiple template elements in the file template, the relative features between different template elements, and the relative features between the same type of template elements.

5. A file identification method comprising: extracting a plurality of third features from the to-be-identified page of the to-be-identified file; Combining the multiple third features to obtain at least one fourth feature includes: combining the plurality of third features according to a feature combination rule to obtain the at least one fourth feature; wherein the feature combination rule is determined based on the at least one second feature screened from the file template; matching the plurality of third features and the at least one fourth feature with a known feature set of a file template required for generating the file type of the to-be-identified file based on artificial intelligence technology to obtain a feature matching degree; When the feature matching degree meets a preset condition, it is determined that the page to be identified is generated based on artificial intelligence technology.

6. The method according to claim 5, wherein: The third feature is used to describe at least one of the typeset features of the multiple page elements in the page to be identified, the relative features between different page elements, and the relative features between the same type of page elements.

7. The method according to claim 5, wherein: The feature matching degree includes a first sub-matching degree and a second sub-matching degree; the first sub-matching degree is determined based on a matching result between the plurality of third features and a first feature in the known feature set corresponding to the plurality of third features; the second sub-matching degree is determined based on a matching result between the at least one fourth feature and a second feature in the known feature set corresponding to the at least one fourth feature; When the feature matching degree satisfies a preset condition, determining that the page to be identified is generated based on artificial intelligence technology includes: When the first sub-matching degree is greater than a first matching threshold, and / or the second sub-matching degree is greater than a second matching threshold, it is determined that the page to be identified is generated based on artificial intelligence technology.

8. The method according to any one of claims 5 to 7, further comprising: In the case where at least one page of content in the file to be identified is determined to be generated based on artificial intelligence technology, it is determined that the file to be identified is generated based on artificial intelligence technology.

9. A file processing device comprising: A first extraction module is used to extract a plurality of first features from a file template required for generating a target type file based on artificial intelligence technology; The screening module includes a screening subunit, which is specifically used to: The plurality of first features are combined based on a genetic algorithm to obtain at least one second feature, with maximizing an objective function as a criterion; the objective function includes a recall rate and / or an accuracy rate of a file recognition result; and / or, Based on the degree of influence on the recall rate and / or accuracy rate of the file recognition result, the at least one second feature is screened out by adopting an AB experiment method.

10. The device according to claim 9, wherein The screening subunit is specifically used for: Randomly combining the multiple first features to obtain multiple candidate combination features; Determining the function value of the objective function based on each candidate combination feature; Screening out a plurality of seed combination features from the plurality of candidate combination features according to the magnitude of the function value; Perform crossover and mutation operations on the multiple seed combination features to obtain multiple new candidate combination features, and return to the step of determining the function value of the objective function based on each candidate combination feature until the at least one second feature that maximizes the objective function is screened out.

11. The apparatus according to claim 9, further comprising: Based on the combination of the at least one second feature, a feature combination rule is determined, where the feature combination rule is used to combine the third feature extracted from the file to be identified.

12. The device according to any one of claims 9 to 11, wherein: The first feature is used to describe at least one of the typeset features of each of the multiple template elements in the file template, the relative features between different template elements, and the relative features between the same type of template elements.

13. A file recognition device comprising: A second extraction module is used to extract a plurality of third features from the to-be-identified page of the to-be-identified file; a combining module, configured to combine the plurality of third features to obtain at least one fourth feature, comprising: combining the plurality of third features according to a feature combination rule to obtain the at least one fourth feature; the feature combination rule being determined based on at least one second feature screened from the file template; a matching module, configured to match the plurality of third features and the at least one fourth feature with a known feature set of a file template required for generating the file type of the to-be-identified file based on artificial intelligence technology, to obtain a feature matching degree; The determination module is used to determine that the page to be identified is generated based on artificial intelligence technology when the feature matching degree meets a preset condition.

14. The device according to claim 13, wherein The third feature is used to describe at least one of the typeset features of the multiple page elements in the page to be identified, the relative features between different page elements, and the relative features between the same type of page elements.

15. The device according to claim 13, wherein The feature matching degree includes a first sub-matching degree and a second sub-matching degree; the first sub-matching degree is determined based on a matching result between the plurality of third features and a first feature in the known feature set corresponding to the plurality of third features; the second sub-matching degree is determined based on a matching result between the at least one fourth feature and a second feature in the known feature set corresponding to the at least one second feature; The determining module is specifically configured to: When the first sub-matching degree is greater than a first matching threshold, and / or the second sub-matching degree is greater than a second matching threshold, it is determined that the page to be identified is generated based on artificial intelligence technology.

16. The apparatus according to any one of claims 13 to 15, wherein the determining module is further configured to: In the case where at least one page of content in the file to be identified is determined to be generated based on artificial intelligence technology, it is determined that the file to be identified is generated based on artificial intelligence technology.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text recognition method and device, processor and electronic equipment

    CN116955624A

  • Malicious file detection method and device, equipment and storage medium

    CN117370966A