Intelligent processing method and system for batch files in office development

By using intelligent batch file processing methods and systems, the problems of version consistency, cross-platform compatibility, and fragmented tool systems in software development have been solved, realizing automated and efficient integrated management of file processing, and improving the efficiency and reliability of office development processes.

CN121009063APending Publication Date: 2025-11-25YUSYS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511119331.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as difficulty in controlling version consistency, high risk of cross-platform compatibility, low efficiency in batch file processing, and fragmented tool systems during software development and deployment, resulting in inefficient process integration and a lack of integrated solutions.

Method used

This paper provides a method and system for intelligent batch file processing in office development. It receives and verifies processing parameters, establishes a file index, uses a hash value incremental comparison algorithm to locate changed files, performs multi-dimensional difference comparison, starts an encoding and delimiter detection engine for parallel analysis, generates a repair strategy, executes file processing operations in parallel, and generates an output report.

Benefits of technology

It automates and visualizes file processing comparison, ensuring version consistency without omissions, automatically detects and fixes cross-platform format compatibility issues, improves batch operation efficiency by more than 80%, reduces error rate, and forms a fully integrated management system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009063A_ABST
    Figure CN121009063A_ABST
Patent Text Reader

Abstract

The invention discloses a batch file intelligent processing method and system in office development. The method comprises the steps of receiving and verifying parameters including a source directory, a target directory, a processing range and the like; traversing the directory based on the related path and the processing range, generating a structure, and collecting information to establish a file index; positioning the changed file by means of a file index, carrying out multi-dimensional difference comparison and recording a result; detecting codes and row separators in parallel for files in the file index and a target file provided by a user, generating a repair strategy and executing format repair; analyzing the natural language instruction based on the replacement rule to generate an executable rule set, and executing batch processing operation on the processed files; and summarizing various results to generate and store an output report. According to the method, intelligent batch processing of the files is realized, and the file processing efficiency and accuracy in office development are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer software technology, specifically relating to office automation and document processing technology, and particularly to a method and system for intelligent batch document processing in office development. Background Technology

[0002] In existing software development and deployment processes, file management commonly suffers from the following problems: First, version consistency is difficult to control effectively. Due to the complex project directory structure, manual comparison of differences between different branches or local and server files is prone to omissions, often leading to the loss of critical configuration files. For example, a bank once experienced a two-hour transaction interruption due to the failure to synchronize three configuration files. Existing version control tools such as Git rely on command-line operations, making it difficult for non-technical personnel to accurately locate file differences in a timely manner, with an average comparison time often exceeding thirty minutes. Second, there is a significant risk of cross-platform compatibility issues. Typically, script files (such as Shell and Python files) developed in a Windows environment, when uploaded to a Linux server, may fail to execute correctly due to inconsistencies in character encoding (GBK / UTF-8) or line separators (CRLF / LF). The probability of failure is as high as about 15%, and existing format checking tools require manual triggering, making it impossible to seamlessly integrate into CI / CD pipelines for automated verification. Third, batch file processing is inefficient. During the project deployment phase, it is usually necessary to manually copy dozens of directories and manually replace variables (such as environment identifiers and date stamps) in file names or content. A single project takes an average of 4 to 6 hours, and manual replacement operations are prone to omissions or errors, such as mistakenly replacing "test" with "prod", which leads to abnormal production environment configuration. Fourth, the existing tool system is highly fragmented. Functions such as version control, difference comparison, format checking, and batch processing rely on multiple independent tools. Data cannot be shared between tools, forming tool silos. There is a lack of comprehensive solutions for integrated office and development scenarios, resulting in low efficiency in process connection. For example, format check results cannot automatically trigger corresponding file repair operations.

[0003] Existing technical solutions each have certain limitations. For example, while Git / SVN has strong version control capabilities, its difference comparison granularity is relatively coarse, it does not support the visualization of directory levels, and it has a high operating threshold for non-technical personnel. Although BeyondCompare has high text comparison accuracy, it cannot be deeply integrated with the production process, requiring manual import of files for comparison, and it does not support batch format checks and automatic repairs. While script batch processing has strong customization capabilities, it relies on manually writing complex rule expressions, which is prone to errors, and it lacks a graphical interface, resulting in high debugging and maintenance costs.

[0004] How to build an intelligent file processing system that can cover the entire process of "comparison, verification, and processing" has become a technical challenge that the industry urgently needs to solve. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method and system for intelligent batch file processing in office development to solve the above-mentioned technical problems.

[0006] To achieve the above objectives, firstly, a method for intelligent batch file processing in office development is provided, which includes:

[0007] Receive and verify the processing parameters input by the user, including the source directory path, target directory path, processing range, replacement rules, and output path;

[0008] Based on the source directory path, the target directory path, and the processing range, the specified directory is traversed to generate a directory structure and collect file information to establish a file index.

[0009] Based on the file index, the hash value incremental comparison algorithm is first used to locate the changed files between the source directory and the target directory. Then, the changed files are compared in multiple dimensions at the structure layer, file layer and content layer, and the comparison results of each dimension are recorded.

[0010] Based on the file index and the target file provided by the user, the encoding detection engine and the delimiter detection engine are started to perform parallel analysis, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user, and generate corresponding diagnostic reports. Based on the diagnostic reports, a repair strategy is generated, and the format repair is performed on the files in the file index and the target file provided by the user according to the repair strategy.

[0011] Based on the replacement rules, the system receives and parses the natural language instructions input by the user to generate an executable rule set, identifies the file types of files in the target directory after difference comparison and format repair, and marks the safe replacement areas and sensitive areas of the file content. Based on the executable rule set and the marked safe replacement areas and sensitive areas, the system performs batch processing operations on file names, file content and directory structure in parallel.

[0012] The comparison results of each dimension, the processing results of format repair, and the processing results of batch processing are summarized to generate an output report and store it in the output path.

[0013] Secondly, a batch file intelligent processing system for office development is provided, which includes:

[0014] The parameter receiving and verification module is used to receive and verify the processing parameters input by the user. The processing parameters include the source directory path, the target directory path, the processing range, the replacement rules, and the output path.

[0015] The file index building module is used to traverse the specified directory based on the source directory path, target directory path and processing range, generate the directory structure and collect file information to build a file index.

[0016] The difference comparison module is used to locate the changed files between the source directory and the target directory based on the file index, first using the hash value incremental comparison algorithm, and then performing multi-dimensional difference comparison of the changed files at the structure layer, file layer and content layer, and recording the difference comparison results of each dimension.

[0017] The format repair module is used to start the encoding detection engine and the delimiter detection engine for parallel analysis based on the file index and the target file provided by the user, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user and generate corresponding diagnostic reports, generate repair strategies based on the diagnostic reports, and perform format repair on the files in the file index and the target file provided by the user according to the repair strategies.

[0018] The batch processing module is used to receive and parse the natural language instructions input by the user to generate an executable rule set based on the replacement rules, identify the file types of files in the target directory after difference comparison and format repair, and mark the safe replacement area and sensitive area of ​​the file content. Based on the executable rule set and the marked safe replacement area and sensitive area, the module performs batch processing operations of file name, file content and directory structure in parallel.

[0019] The report generation module is used to summarize the comparison results of each dimension, the processing results of format repair, and the processing results of batch processing operations, generate an output report, and store it in the output path.

[0020] The above technical solution has the following beneficial technical effects:

[0021] By accurately receiving and verifying processing parameters, and combining the source and target directory paths with the processing range to build a file index, the system can focus on core file groups, avoid invalid processing, and significantly improve the targeting and efficiency of batch processing. Based on the file index, incremental hash value comparison is used to locate changed files, and multi-dimensional difference comparisons are performed at the structural, file, and content levels. This not only quickly identifies the changes that need attention but also comprehensively captures the differences at different levels, ensuring the accuracy and completeness of difference identification. The system also launches a parallel encoding and delimiter detection engine to generate dynamic repair strategies for files within the file index and user-provided target files and performs format repair. This system effectively solves cross-platform file format compatibility issues and reduces parsing errors caused by encoding or delimiter anomalies. It generates executable rule sets based on natural language instructions parsed using replacement rules, and combines file type labeling with parallel processing of security and sensitive areas. This improves the convenience of batch operations while reducing the risk of accidental modification and ensuring processing security. Finally, by summarizing the results of multiple stages to generate an output report, the system achieves traceability of the processing process. The overall workflow forms a closed loop from parameter verification to result feedback, improving the automation level, processing efficiency, and result reliability of batch file processing in office development. It is particularly suitable for batch management scenarios of multiple file types under complex directory structures. Attached Figure Description

[0022] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:

[0023] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart of the file batch comparison tool according to an embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of the binary file processing flow according to an embodiment of the present invention;

[0026] Figure 4 This is a sample image of a binary file difference comparison report according to an embodiment of the present invention;

[0027] Figure 5 This is an example diagram of binary difference analysis technology according to an embodiment of the present invention;

[0028] Figure 6 This is an example screenshot of Mode 1 output of the multi-mode file difference analysis of the intelligent difference comparison module in this embodiment of the invention;

[0029] Figure 7 This is an example screenshot of Mode 2 output of the multi-mode file difference analysis of the intelligent difference comparison module in this embodiment of the invention;

[0030] Figure 8This is a schematic diagram illustrating the working principle of the cross-platform format verification module according to an embodiment of the present invention;

[0031] Figure 9 This is a flowchart of the file format checking tool according to an embodiment of the present invention;

[0032] Figure 10 This is a signaling interaction flowchart of dual-engine detection according to an embodiment of the present invention;

[0033] Figure 11 This is a flowchart of the file batch copying keyword mapping and replacement tool according to an embodiment of the present invention;

[0034] Figure 12 This is a signaling interaction flowchart for batch file processing according to an embodiment of the present invention;

[0035] Figure 13 This is a functional block diagram of the system according to an embodiment of the present invention;

[0036] Figure 14 This is a schematic diagram of the structure of a computer system according to an embodiment of the present invention. Detailed Implementation

[0037] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0038] This invention belongs to the field of computer software technology, specifically relating to office automation and document processing technology. Addressing the frequent need for batch file management during software development and project deployment, it innovatively proposes a comprehensive solution integrating version consistency verification, cross-platform format adaptation, and intelligent batch processing. By constructing an automated file processing engine, it achieves intelligent operation throughout the entire process of file difference comparison, format compliance check, batch copying, and keyword replacement, improving the efficiency and reliability of file processing in office development scenarios. It is suitable for complex environments such as fintech, enterprise-level software development, and IT project deployment.

[0039] The goal of this invention is to achieve automated and visual comparison of multi-dimensional differences between files. It can preventively detect and automatically repair compatibility issues such as encoding formats and line breaks that may exist across platforms before files are put into production. At the same time, it supports dynamic configuration and intelligent execution of keyword mapping rules during batch operations to reduce manual intervention, lower error rates and improve overall processing efficiency.

[0040] This invention aims to construct a smart file processing factory for office development scenarios, systematically addressing core issues in file management such as version, format, and batch operations. Its objectives include: ensuring version consistency by enabling rapid comparison of files and directories and visualization of multi-dimensional differences; building an automatically detectable cross-platform format verification engine and seamlessly integrating it with the production workflow to eliminate the risk of incompatibility between encoding formats and line separators; achieving batch copying of files and dynamic replacement of keywords based on configurable rules, improving the efficiency of repetitive operations by over 80%; and providing a graphical user interface and API, supporting deep integration with CI / CD pipelines and project management systems, thereby forming an integrated closed-loop management system covering the entire process of file comparison, verification, and processing.

[0041] like Figure 1 As shown, a method for intelligent batch file processing in office development is provided, which includes:

[0042] S10: Receive and verify the processing parameters input by the user, the processing parameters including the source directory path, the target directory path, the processing range, the replacement rule, and the output path;

[0043] S20: Based on the source directory path, the target directory path, and the processing range, traverse the specified directory, generate a directory structure, and collect file information to establish a file index;

[0044] S30: Based on the file index, first use the hash value incremental comparison algorithm to locate the changed files between the source directory and the target directory, and then perform multi-dimensional difference comparison of the changed files at the structure layer, file layer and content layer, and record the difference comparison results of each dimension;

[0045] S40: Based on the file index and the target file provided by the user, start the encoding detection engine and the delimiter detection engine to perform parallel analysis, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user, generate corresponding diagnostic reports, generate repair strategies based on the diagnostic reports, and perform format repair on the files in the file index and the target file provided by the user according to the repair strategies;

[0046] S50: Based on the replacement rules, receive and parse the natural language instructions input by the user to generate an executable rule set, identify the file type of the files in the target directory after difference comparison and format repair, and mark the safe replacement area and sensitive area of ​​the file content. According to the executable rule set and the marked safe replacement area and sensitive area, perform batch processing operations of file name, file content and directory structure in parallel.

[0047] S60: Summarize the comparison results of each dimension, the processing results of format repair, and the processing results of batch processing operations, generate an output report, and store it in the output path.

[0048] In some embodiments, step S30 specifically includes:

[0049] S301: Locating changed files between the source and target directories based on the hash value incremental comparison algorithm;

[0050] S302: Perform a multi-dimensional difference comparison, wherein the multi-dimensional difference comparison includes:

[0051] The structural layer dimension compares the differences, visualizes the directory tree structure of the source directory and the target directory, marks the newly added, deleted or renamed directories, and obtains the directory tree difference annotation results.

[0052] The file-level difference comparison compares the file size, modification time, and hash value of corresponding files in the source and target directories to achieve file-level difference annotation and obtain file-level difference annotation results, which include differences in file size, modification time, and hash value.

[0053] The content-level difference comparison highlights the differences line by line for text files and performs byte-level difference comparison for binary files to obtain the content difference annotation results.

[0054] S303: Generate a visual comparison report containing the directory tree difference annotation results, the file-level difference annotation results, and the content difference annotation results.

[0055] In some embodiments, the content-level difference comparison involves highlighting differences line by line for text files and performing byte-level binary code difference comparison for binary files to obtain content difference annotation results, specifically including:

[0056] S3021: Filter out differing files using a hash algorithm;

[0057] S3022: Compare the difference files byte by byte, and convert the byte data into binary strings and hexadecimal strings;

[0058] S3023: Highlight the difference bytes or difference bits to generate a hexadecimal difference display report and a binary difference display report.

[0059] In some embodiments, the difference display report includes: byte offset positioning; statistics on the number of bit changes; comparison of hexadecimal value changes; analysis of ASCII character changes; and description of binary bit change sequences.

[0060] In some embodiments, step S40 specifically includes:

[0061] S401: Receive the target file provided by the user;

[0062] S402: The advanced delimiter detection engine and the intelligent encoding detection engine are launched in parallel to analyze the target file;

[0063] S403: Identify the delimiter type and generate a delimiter diagnostic report using the advanced delimiter detection engine;

[0064] S404: Identify the encoding of the target file and generate an encoding diagnostic report using the aforementioned intelligent encoding detection engine;

[0065] S405: Based on the delimiter diagnostic report and the encoding diagnostic report, generate a dynamic repair strategy and perform repair on the target file.

[0066] In some embodiments, step S403 specifically includes:

[0067] The target file is scanned line by line using predefined regular expressions to identify the type of control character at the end of each line, and the frequency and distribution range of different types of delimiters are statistically analyzed to obtain the delimiter type statistics.

[0068] The distribution of delimiters is checked in conjunction with the logical structure of the target file to determine whether the inconsistency of delimiter patterns between adjacent lines is an abnormal situation, and the delimiter abnormality judgment result is obtained.

[0069] Generating a line-end pattern heatmap based on line-end delimiter data includes: statistically analyzing the delimiter types and distribution positions of all lines in the target file to form a basic dataset; marking the line numbers and ranges of mixed delimiters when multiple delimiters are used in the target file, obtaining the line number and range marking information of the mixed delimiters; calculating the compatibility score of the target file on different platforms based on delimiter consistency, delimiter distribution concentration, and delimiter mixing ratio; and determining the line-end pattern heatmap based on the basic dataset, the line number and range marking information of the mixed delimiters, and the compatibility score.

[0070] Based on the statistical results of the delimiter types, the results of the delimiter anomaly determination, and the heatmap of the line ending pattern, the delimiter diagnostic report is generated.

[0071] In some embodiments, step S404 specifically includes:

[0072] Check if a BOM byte sequence exists at the beginning of the target file. If a BOM byte sequence is detected, identify the file's encoding type based on the BOM byte sequence.

[0073] If no BOM byte sequence is detected, the encoding type of the target file is inferred by using the byte distribution characteristics of the target file content, character frequency statistics, and character set matching strategy.

[0074] The process involves using a multimodal learning model to comprehensively determine the encoding type. This includes: for files without a BOM (Bill of Materials) identifier, performing pattern matching in the feature space using the multimodal learning model to identify the most likely encoding type; utilizing the context-aware capability of the multimodal learning model to detect whether there are mixed encoding regions in the target file, where a mixed encoding region refers to content segments in the target file corresponding to different encoding types; using the multimodal learning model to predict the probability of the most likely encoding type, obtaining probability values ​​for each candidate encoding type, wherein the candidate encoding includes the inferred encoding type and the most likely encoding type; and outputting the optimal encoding scheme based on the probability values ​​of multiple candidate encodings.

[0075] The identification results are integrated and output to generate an encoding diagnosis report. The encoding diagnosis report records the encoding type of the target file, the existing mixed encoding regions and their severity. The severity is the degree of impact on file parsing determined based on the range of the mixed encoding regions and the number or proportion of abnormal characters.

[0076] In some embodiments, step S405 includes the following steps:

[0077] Input the original file into the repair engine;

[0078] The repair engine generates multiple different repair schemes for the encoding anomalies in the original file based on the abnormal encoding formats and repair marks identified by the adaptive repair algorithm. These multiple different repair schemes differ in encoding conversion strategies, hybrid encoding segmentation processing strategies, and delimiter unification rules.

[0079] The various repair schemes are input into the difference visualization module, which visually displays the changes made by each repair scheme to the original file content. These changes include encoding adjustments, delimiter adjustments, and the processing effects of using different encodings for different paragraphs.

[0080] The repair scheme is selected by the user from multiple different repair schemes through an interactive interface, or by a weight-based recommendation system from the multiple different repair schemes.

[0081] The selected repair scheme is applied to the original file to complete the repair of the encoding anomaly in the original file.

[0082] In some embodiments, step S50 specifically includes:

[0083] S501: Receives natural language commands input by the user through the interface, performs semantic analysis through the natural language processing module, and generates an executable rule set containing filename replacement rules, file content replacement rules, and directory structure mapping rules;

[0084] S502: Call the file type recognition module to distinguish between text files, code files, and binary files in the target directory; parse the semantics of file content based on file type, mark safe replacement areas and sensitive areas, and generate a context annotation list;

[0085] S503: Request a list of target files from the file system based on the executable rule set and the context annotation list; perform the following operations in parallel:

[0086] Reconstruct the target path according to the filename replacement rules;

[0087] According to the file content replacement rules, context-aware replacement is performed on the dynamic variables in the safe replacement area. The dynamic variables include date, request number, and user-defined parameters.

[0088] Generate a multi-level directory associated with the runtime environment based on the directory structure mapping rules;

[0089] S504: Create a temporary workspace to perform batch operations and verify the consistency of the processing results through hash verification; if the verification passes, it is atomically committed to the target directory; if the verification fails, it is rolled back based on the breakpoint information.

[0090] S505: Records task breakpoints for incomplete operations interrupted by exceptions during batch processing, resumes execution from the breakpoints, generates a processing report containing a file modification list, replacement statistics, and exception logs, and outputs it to the user interface.

[0091] The above technical solution will be explained in detail below with specific examples:

[0092] like Figure 2 As shown, the process of the batch file comparison tool in this embodiment of the invention includes the following steps:

[0093] Step 1: Program Startup. In this step, the batch file comparison tool is started. The system first completes initialization operations, including loading the algorithm library, rule templates, and configuration files required for comparison, and establishing a cache area and log module to provide a stable operating environment for subsequent task execution.

[0094] Step Two: Key Parameter Input. After initialization, the key parameter input step begins. Users can input various parameters related to the comparison task via the interface or command line. These parameters include the source directory path, target directory path, comparison range, file type, comparison rules (whether to distinguish between uppercase and lowercase letters, whether to recursively search subdirectories, whether to ignore blank lines or special characters, etc.), and output report path. The system will perform validity checks and confirmations after parameter input to ensure that the input is complete and accurate.

[0095] Step 3: File Traversal and Parsing. After confirming that the parameters are correct, the program traverses the source and target directories layer by layer according to the input directory path, generating a directory tree structure. It also collects and parses the attributes of the traversed files to obtain information such as file path, name, size, modification time, and hash value, and builds a file index to prepare for subsequent matching and content comparison.

[0096] Step 4: Text Comparison Engine. After completing file traversal and parsing, the system activates the text comparison engine to compare the content of the matched files. For text files, a line-by-line comparison algorithm is used to detect differences and highlight the different content; for specific file types such as Word and Excel, the content is first extracted using a built-in parser before comparison; for binary files, byte stream comparison is used. This engine can select the most appropriate comparison strategy based on the file type, thereby ensuring the accuracy of the comparison results.

[0097] Step 5: Result Recording and Caching. Differences generated during the comparison process (including added files, deleted files, file attribute differences, content differences, etc.) are recorded in real time and written to the cache. This caching mechanism ensures that even if execution is interrupted, completed comparison information is retained, supporting resumption from the point of interruption and avoiding redundant processing.

[0098] Step Six: Concurrent Writing to Excel Output Report. After the comparison task is completed, the system organizes, categorizes, and summarizes the cached comparison results, and generates an Excel output report through multi-threaded concurrent writing. This report file contains a list of differing files, detailed differences, comparison time, and other information, and is ultimately stored in the user-specified output path for easy viewing, analysis, and archiving.

[0099] This invention constructs an automated file processing system through three core modules:

[0100] 1. Intelligent difference comparison module

[0101] The technical implementation of the intelligent difference comparison module is as follows:

[0102] An incremental comparison algorithm is used to quickly locate changed files based on hash values, improving the comparison efficiency by more than 5 times compared to the traditional line-by-line comparison.

[0103] The difference display mechanism provided in this invention supports comprehensive comparison from three dimensions: structure layer, file layer, and content layer. At the structure layer, the directory tree can be visualized, and newly added, deleted, or renamed directories can be marked. At the file layer, accurate difference annotation is achieved by displaying information such as file size, modification time, and hash value. At the content layer, differences can be highlighted line by line for text files, and comparison of Word, Excel, and code files is supported. For binary files, binary code-level comparison display is provided.

[0104] like Figure 3 The flowchart shown illustrates the binary file processing workflow. The comparison process is initiated through parameter configuration, quickly identifying changed files based on hash algorithms, triggering deep content comparison for text files, and ultimately generating a multi-dimensional difference report that supports visual viewing and problem localization. The binary difference visualization implementation is as follows: ① First, find the binary files with differences using MD5 comparison; if they match, no further operations are performed; ② For the difference files, perform precise comparison at the byte level, rather than text lines, and convert the bytes to binary string representations (0 and 1); ③ During the difference comparison, you can choose to display both hexadecimal and binary data, highlighting the differences. Additionally, considering large file optimization, buffered reading is used to process large binary files.

[0105] like Figure 4 As shown, this is a sample image of a binary file difference comparison report. Figure 5 The diagram shown is an example of binary difference analysis. The differences are summarized as follows: a bit change was detected at the second byte (offset 0x00000001), specifically affecting three bits: the 4th, 5th, and 6th bits. The corresponding hexadecimal value changed from 0x41 to 0x4F, and the ASCII character changed from "A" to "O". Further analysis of the binary data reveals that three bits of this byte changed from 000 to 111.

[0106] The intelligent difference comparison module supports multi-mode file difference analysis and outputs visual comparison reports. "Multi-mode" refers to the ability to output different reports under different files or parameters. For example, Mode 1 only displays the filenames of the differences; Mode 2 displays the detailed differences; Mode 3 considers files consistent only if their relative paths are identical; Mode 4 considers files comparable as long as they exist; Mode 5 distinguishes between primary and secondary directories, checking directories if the primary directory is missing, and so on. Figure 6The image shown is an example screenshot of Mode 1 output from the multi-mode file difference analysis of the intelligent difference comparison module; as shown... Figure 7 The image shown is an example screenshot of Mode 2 output from the multi-mode file difference analysis of the intelligent difference comparison module.

[0107] 2. Cross-platform format validation module

[0108] The cross-platform format verification module is an intelligent adaptive file encoding and format unification system. Based on multimodal learning, it realizes intelligent diagnosis, dynamic repair and security compliance assurance of cross-platform file formats, and supports complex mixed format processing and deep CI / CD integration.

[0109] The technical implementation of the cross-platform format validation module is as follows:

[0110] It adopts a dual-engine verification mechanism, which includes an intelligent encoding detection engine and an advanced delimiter detection engine.

[0111] The intelligent encoding detection engine is based on BOM identifier and byte feature recognition of encoding formats such as UTF-8 / GBK / ISO-8859-1, with an accuracy of ≥99%; BOM identifier recognition layer, accurately identifies byte order mark (BOM); byte feature analysis layer, advanced byte pattern recognition algorithm; encoding probability prediction.

[0112] An advanced delimiter detection engine uses regular expressions to match CRLF / LF / CR formats to pinpoint incompatibility issues across Windows, Linux, and macOS systems; a regular expression matching layer supports all line ending modes; context-aware analysis and syntax boundary recognition prevent misjudgments of newline characters in string content; a comment protection mechanism preserves original newline characters in comments; and code block analysis identifies line break requirements within logical structures.

[0113] like Figure 8As shown, the system traverses a specified directory of files, detects encoding and delimiters using a dual-engine approach, outputs a compliance report, and provides remediation guidance, supporting one-click format conversion. Its workflow includes: inputting a file; entering parallel analysis using two engines: an advanced delimiter detection engine and an intelligent encoding detection engine; the advanced delimiter detection engine performs analysis including: regular expression matching, context-aware analysis, and line ending pattern heatmaps; the line ending pattern heatmap includes: CRLF / LF / CR identification, hybrid delimiter location, and platform compatibility scoring; after analysis, the advanced delimiter detection engine outputs a delimiter diagnostic report; the intelligent encoding detection engine performs analysis including: BOM identifier identification, byte feature analysis, and a multimodal learning model. The multimodal learning model includes: BOM-free file identification, hybrid encoding detection, and encoding probability prediction; after analysis, the intelligent encoding detection engine outputs an encoding diagnostic report. Finally, based on the delimiter and encoding diagnostic reports, a dynamic remediation strategy is derived.

[0114] The above work process is described in detail below:

[0115] (1) Input file

[0116] The input file step is the starting point of the entire processing flow. The system receives files to be analyzed provided by the user, which can be a single file or a collection of files. The file type is unrestricted, including code files, configuration files, text files, script files, etc. The system first checks the input file path and file status to ensure the files are accessible and undamaged, and then loads them into memory as the processing objects for subsequent dual-engine analysis. The purpose of this step is to provide a clear input data source for subsequent analysis.

[0117] (2) Dual-engine parallel analysis

[0118] After receiving the input file, the system enters the dual-engine parallel analysis phase. This phase activates two independent and parallel analysis engines: an advanced delimiter detection engine and an intelligent encoding detection engine. These two engines improve overall processing efficiency through parallel processing and can perform comprehensive diagnostics from both file delimiter and encoding characteristics, ultimately providing a high-precision compatibility and standardized assessment of the file.

[0119] (3) Advanced delimiter detection engine

[0120] The advanced delimiter detection engine performs in-depth detection and analysis of line delimiters in input files to uncover potential delimiter format inconsistencies and assess cross-platform compatibility. Upon receiving the target file to be analyzed, the engine sequentially performs steps such as regular expression matching, context-aware analysis, and line ending pattern heatmap generation, ultimately outputting a complete delimiter diagnostic report.

[0121] In the first step, the system uses predefined regular expressions to scan the file line by line, automatically identifying the control characters at the end of each line and quickly distinguishing between types such as carriage return and line feed (CRLF), newline character (LF), and carriage return (CR). Regular expression matching not only extracts the specific location of each delimiter but also calculates the frequency and distribution range of different delimiters during the scan, providing a data foundation for subsequent analysis.

[0122] After regular expression matching is completed, the context-aware analysis phase begins. This phase combines the file's logical structure to perform localized checks on the distribution of delimiters. For example, when inconsistent delimiter patterns are found between adjacent lines, the system considers the context of the preceding and following content blocks to determine if it's an anomaly. This distinguishes between harmless differences caused by operating system variations and genuine issues affecting execution, thereby improving the accuracy of the detection results.

[0123] After context-aware analysis is completed, the system generates a line ending pattern heatmap based on the collected line ending delimiter data. This heatmap is a graphical statistical and evaluation tool for delimiter features, comprising the following three core sub-steps:

[0124] In the CRLF / LF / CR recognition step, all lines in the entire file are statistically analyzed to identify the type and distribution of each type of delimiter, forming a basic dataset.

[0125] In the mixed delimiter location step, when multiple delimiters are used in a file, the system will accurately mark the line numbers and ranges of the mixed delimiters in a graphical way to help locate the problem area;

[0126] In the platform compatibility scoring step, based on the consistency of delimiters, the concentration of distribution, and the mixing ratio, the compatibility score of the file on different platforms (Windows, Linux, macOS) is automatically calculated for subsequent compatibility assessment and remediation decisions.

[0127] After the above analysis, the advanced delimiter detection engine will summarize the identified delimiter statistics, mixed delimiter regions, compatibility scores, and other results to generate a delimiter diagnostic report. This report is provided to subsequent modules or users in the form of structured data to help determine whether automatic file repair, delimiter format standardization, or file optimization to adapt to cross-platform environments is necessary.

[0128] (4) Intelligent Encoding Detection Engine The intelligent encoding detection engine is the core module of the entire intelligent encoding analysis process. Its task is to perform encoding detection and judgment on the input target file. After receiving the file, the engine will call its internal sub-modules in sequence for processing, and comprehensively utilize methods such as BOM identifier recognition, byte feature analysis and multimodal learning models to gradually generate accurate encoding detection results from shallow encoding feature extraction to deep intelligent prediction.

[0129] The first step after the intelligent encoding detection engine starts is for the system to quickly identify the file's encoding type by checking for the presence of a BOM (Byte Order Mark) byte sequence at the beginning of the file. BOM is a common encoding identification method; different encodings correspond to different BOM header information. By parsing these characteristic bytes, the system can directly determine the file type (UTF-8, UTF-16, UTF-32, etc.) that contains BOM encoding. If a file contains a BOM, the encoding type is directly recorded as input for subsequent analysis.

[0130] For files where no BOM (Bill of Materials) is detected, the system proceeds to the byte feature analysis step. This step utilizes the byte distribution features of the file content, character frequency statistics, and common character set matching strategies to infer the potential encoding type of the file. This process includes scanning the entire file data stream, statistically analyzing the proportion of visible characters, common byte patterns, and the distribution of abnormal characters, providing more accurate feature input for subsequent multimodal learning models.

[0131] After completing the basic feature extraction, the system calls a multimodal learning model to comprehensively judge the encoding problem. This model integrates byte distribution features, language model features, and contextual statistical features, enabling intelligent identification of encoding types in complex scenarios. The working process of the multimodal learning model consists of three aspects: BOM-free file identification, which uses a machine learning model to perform pattern matching in the feature space for files without BOM identifiers, automatically identifying the most likely encoding type, such as distinguishing between UTF-8 (without BOM), GBK, ISO8859-1, etc.; mixed encoding detection, which uses the context-aware capability of the multimodal learning model to detect whether there is a mixed encoding phenomenon in the file, such as part of it being UTF-8 and part of it being GBK, thereby avoiding parsing errors caused by local encoding anomalies; and encoding probability prediction, where the model predicts the probability of possible encoding types, outputs the final optimal encoding scheme based on the scores of multiple candidate encodings, and provides a decision basis for subsequent repair or conversion.

[0132] After all submodules have completed execution, the intelligent encoding detection engine will integrate the identification results and output a detailed encoding diagnostic report. This report not only records the encoding type of the file, but also points out possible mixed encoding areas and their severity, and provides encoding consistency suggestions for use by subsequent automatic repair strategy modules.

[0133] (5) Generate dynamic repair strategy

[0134] After the dual engines complete their analysis and generate delimiter and encoding diagnostic reports respectively, the system enters the report integration and repair strategy generation phase. The system summarizes and analyzes the detection results from both reports, and, combined with the actual usage scenario of the file, generates a dynamic repair strategy. The dynamic repair strategy clearly specifies the repair operations required, such as unifying the delimiter format, converting to a specific encoding format, or automatically inserting a BOM, and can serve as the direct execution basis for subsequent automatic repair modules. Through this phase, users can intuitively obtain information about file compatibility issues and suggested repair methods, ensuring correct file execution across multiple platforms and environments.

[0135] The intelligent repair strategy adopted is as follows:

[0136] An adaptive repair engine that automatically converts encoding formats (e.g., GBK to UTF-8) to preserve the integrity of the original content; a repair solution recommendation system based on project history; a state-weighted decision model; mixed format processing capabilities; segmented encoding conversion (different encodings for different paragraphs); intelligent delimiter unification (preserving the original semantic structure); and a repair difference preview system.

[0137] like Figure 9As shown, an automatic file repair process based on an intelligent repair strategy is as follows: First, the original file is input into the repair engine. The repair engine integrates an adaptive repair algorithm, which can automatically identify and adjust abnormal encoding formats in the file (e.g., automatically converting GBK encoding to UTF-8) without destroying the semantics of the original file content, while preserving the integrity of the original file structure and content. The repair engine combines historical repair data with a repair scheme recommendation model to generate multiple different repair schemes for the same problem, such as repair scheme A, repair scheme B, and repair scheme C. These schemes may differ in encoding conversion strategies, mixed encoding segmentation strategies, and delimiter unification rules. The generated repair schemes enter the difference visualization module. Through difference visualization, the system intuitively displays the changes made to the file content by each repair scheme, including encoding adjustments, delimiter adjustments, and the processing effects of using different encodings for different paragraphs, facilitating user comparison and selection. Finally, users can manually select the most suitable repair scheme through the interactive interface, or enable the system's automatic application mechanism, where a recommendation system based on a weighted decision model automatically selects the optimal scheme and applies it to the original file, thereby completing the automated repair of encoding and formatting issues.

[0138] The process integration mechanism of this invention can serve as a pre-processing step in a CI / CD pipeline, automatically blocking subsequent deployment operations when file verification fails. It can also connect to mainstream platforms such as Jenkins and GitLab to generate and output format compliance reports, enabling automated quality control and traceability management in the production process.

[0139] like Figure 10 As shown, the signaling interaction flowchart of the cross-platform format verification module includes the following steps:

[0140] Step 1: File data is sent to two detection engines. Upon entering the system, the file data is simultaneously distributed to both the intelligent encoding detection engine and the advanced delimiter detection engine. The file data includes file content, file attributes, and metadata. The two detection engines work in parallel, avoiding performance bottlenecks caused by sequential detection. This dual-path parallel design ensures that the file's encoding features and delimiter features can be analyzed simultaneously, improving overall processing efficiency.

[0141] Step Two: The intelligent encoding detection engine outputs an encoding diagnostic report. After receiving the file data, the intelligent encoding detection engine sequentially performs BOM detection, byte feature analysis, and multimodal learning model recognition to generate a file encoding diagnostic report. The report includes the encoding type, the confidence score of the detection results, the existence of mixed encoding segments, and potential encoding compatibility risks. The generated report is transmitted to the decision center in structured data format.

[0142] Step 3: The advanced delimiter detection engine outputs a delimiter heatmap. After receiving the same file data, the advanced delimiter detection engine sequentially performs regular expression matching, context-aware analysis, and line ending pattern heatmap generation to identify the usage of line ending delimiters such as CRLF, LF, and CR in the file, analyze whether there is mixed usage, and calculate the delimiter mixing degree. The engine outputs a delimiter heatmap and the corresponding mixing degree index, which is sent to the decision center.

[0143] Step 4: The decision center calculates risk weights. After receiving the encoding diagnostic report from the intelligent encoding detection engine and the delimiter heatmap from the advanced delimiter detection engine, the decision center calculates risk weights by combining the two types of diagnostic information. This calculation weights the encoding risk confidence level and the delimiter mixing index to obtain the overall risk level of the file, which is used to determine the subsequent processing strategy.

[0144] Step 5: Standard Repair for Low-Risk Scenarios. If the overall risk level calculated by the decision center is lower than the preset risk threshold, the system will automatically enter the low-risk processing flow. The system will directly generate standardized repair solutions, such as unified encoding formats and unified delimiter types, and call the automatic repair module to perform the repair operation without manual intervention.

[0145] Step Six: Manual Confirmation Request for High-Risk Scenarios. When the overall risk level exceeds a preset threshold, the decision center will send a manual confirmation request to the user interface. These high-risk files typically have complex encoding anomalies or severe delimiter mixing issues; automatic repair may lead to content corruption, thus requiring manual intervention.

[0146] Step Seven: User Interface Returns Decision Instructions. After receiving a manual confirmation request, the user interface displays the test report and risk analysis results to the user. The user then performs actions such as manual confirmation, selecting a suitable repair solution, or rejecting automatic repair through the interface. The user's decision is returned to the decision center in the form of instructions.

[0147] Step 8: The decision center executes customized repairs. After receiving the decision instruction from the user, the decision center generates a customized repair strategy based on the instruction, such as specifying to handle only a certain type of encoding problem or to repair only a specific range of separator problems. The strategy is then sent to the automatic repair module, which executes the repair based on the customized strategy.

[0148] Step Nine: Apply the Repair Solution. After the automatic repair module performs the customized repair operation, it returns the repaired file to the system. Detailed records of the repair process are also saved in the log for traceability and verification. At this point, the entire signaling interaction process is complete, and the file encoding and delimiter issues are resolved.

[0149] 3. Dynamic Batch Processing Module

[0150] The dynamic batch processing module is a multi-dimensional dynamic reconstruction system for batch file processing based on an intelligent rule engine. Through semantic parsing and context-aware technology, it realizes intelligent batch processing of file copying, content replacement and directory reconstruction, and has the ability to perform change-aware incremental processing and end-to-end integrity verification.

[0151] In terms of technical implementation, this invention constructs an intelligent mapping engine that supports multi-level intelligent mapping rules, including filename replacement (e.g., automatically replacing "dev_config.txt" with "prod_config.txt"), file content replacement (e.g., automatically replacing the variable ${ENV} with "production"), and directory structure mapping (automatically generating multi-level directories based on environment variables). The text content replacement is context-aware, intelligently distinguishing variable usage scenarios based on different semantic environments, such as differentiating ${DB_URL} in code from ${DB_URL} in documentation, and employing adaptive differentiated syntax processing rules based on file type (JSON, YAML, XML, etc.). Simultaneously, this intelligent mapping engine has a built-in variable parser, supporting the automatic identification and replacement of dynamic variables such as dates (YYYYMMDD), requirement numbers, and user-defined parameters. In addition, the intelligent mapping engine also integrates an intelligent AI rule generator, which can automatically generate optimized rules by analyzing historical replacement operations. Combined with natural language recognition big data model technology, it can automatically convert users' natural language commands into rule configurations and execute them directly. For example, if a user enters "replace all files in this directory containing the keyword YYYYMMDD with 20250101", the system can automatically generate replacement rules and complete batch processing. The generator also automatically infers variable mapping relationships based on model prediction technology, such as automatically associating ${user} with ${employee_id}.

[0152] Regarding the change handling mechanism, this embodiment of the invention adopts a change-aware stream processing mode, deeply integrating the file processing process with version control systems (such as Git and SVN). It only processes files involved in changes in the diff, avoiding redundant operations on unrelated files. To improve the performance of large file processing, this embodiment of the invention utilizes memory-mapped file technology to replace traditional I / O operations. A dynamic block-sharing strategy adaptively adjusts the size of processing blocks based on file type, effectively avoiding stuttering in large file processing. It also supports parallel processing, allowing a single thread to process over 500 files simultaneously, achieving an overall processing throughput of over 200MB / s, thus significantly improving processing efficiency.

[0153] Regarding data consistency and execution security, this invention employs an integrity assurance system and an atomic transaction processing mechanism. Before executing batch operations, a temporary workspace is created. Only after all operations are completed and verified by hash checking are atomic commits or rollbacks performed, ensuring consistency and reliability throughout the process. The system also provides an anomaly recovery mechanism, automatically generating breakpoint information during execution. This allows continued execution of unfinished operations from any point of failure, and automatically generates an anomaly impact analysis report for rapid problem location and recovery.

[0154] In terms of application scenarios, the embodiments of the present invention can be applied to various automated document processing tasks. For example, in the scenario of automated production material generation, the system can automatically copy the project directory and replace the environment identifier according to the version number, automatically detect environment differences (e.g., prod / test distinction), and process only the changed files through version awareness, greatly reducing manual intervention. In the scenario of batch distribution of requirement documents, the system can automatically generate customized document versions according to departments or roles. For example, it can output a version containing technical details for the development team, while outputting a simplified version containing only functional descriptions for business personnel. In the scenario of test environment initialization, the system can quickly generate multiple test environment configuration files (e.g., test1, test2, test3, etc.), thereby reducing repetitive work and improving the overall R&D and testing efficiency.

[0155] like Figure 11 As shown, the batch file copying keyword mapping and replacement tool process includes the following steps: After the program starts, it first enters the key parameter input stage, where the user needs to input the source directory, target directory, and replacement rules, etc. After the parameters are confirmed, the system performs file traversal and streaming parsing on the source directory to obtain all files and directory structures one by one. Then, it enters the rule parsing step, where the system parses specific mapping relationships such as filename replacement rules and file content replacement rules according to the preset keyword mapping rules. After parsing, the system copies the files in the source directory to the target directory in batches according to the streaming processing method. During the copying process, the system automatically replaces the keywords in the filenames and file contents according to the replacement rules. The operation information generated during the processing is recorded in real time, and a log is generated after the task is completed. The log contains detailed information on successful and abnormal processing for the user to check and trace later.

[0156] like Figure 12 As shown, the intelligent rule parsing engine process includes the following steps:

[0157] Step 1: The user issues a batch processing command in natural language. The user enters the batch processing request in natural language in the system's user interface, for example, "Replace all files containing the keyword YYYYMMDD with 20250101". This natural language command is sent to the natural language processing module, serving as the entry point for the entire batch processing process.

[0158] Step Two: The Natural Language Processing (NLP) module parses the natural language instructions. After receiving the user's input natural language instructions, the NLP module performs semantic analysis, including word segmentation, intent recognition, and parameter extraction, to obtain the core elements required for batch processing. These elements include the keywords to be replaced, the target value after replacement, the processing scope (file content, filename, or directory), and whether directory reconstruction is involved. The parsed structured instructions are then passed to the rule parser.

[0159] Step 3: The rule parser generates an executable rule set. Based on the parsing results provided by the natural language processing module, the rule parser transforms the natural language request into a machine-executable rule configuration set. The output of the rule parser is a set of replacement rule instructions, which includes: specific replacement operation instructions (what to replace, what to replace with), operation target (filename, file content, or directory structure), execution scope, and execution priority. Before generating the replacement rule instruction set, the rule parser calls the file type recognition module to perform type analysis on the files in the target directory to distinguish between text files, code files, configuration files, binary files, etc., thereby ensuring that subsequent replacement strategies are executed specifically. For example, the process of identifying files as text files or code files is automatically completed by the file type recognition module based on the extension, file header information, or content pattern. Finally, the replacement instruction set generated by the rule parser is passed to the rule execution engine as the direct basis for subsequent batch processing operations.

[0160] Step 4: The context-aware module analyzes the file structure. The input to the context-aware module is the target file range information, rule content, and corresponding file type information generated by the rule parser. Upon receiving this information, the module performs a structured analysis of the files in the specified directory. Specifically, this includes: scanning the directory hierarchy to obtain the file organization structure; parsing the file content to identify semantic context relationships and the scope of variables; and based on the content characteristics of different file types (e.g., code files, configuration files, document files), marking safe areas suitable for replacement operations and sensitive areas that need to be avoided from replacement, thereby generating a list of file paths with context annotations and a replacement security policy. The results of the analysis are returned to the rule parser, which integrates the context analysis results with the rules to ensure the accuracy and security of subsequent replacement operations.

[0161] Step 5: The rule execution engine requests a file list. Based on the integration results provided by the rule parser, including the processing scope, context annotation information, and specific replacement rules, the rule execution engine sends a file list request to the file system to obtain a list of target files that meet the criteria. This request process does not rely on the direct participation of the context-aware module; all context-related processing scopes have already been organized by the rule parser in the previous steps. After the file system returns the file list, the rule execution engine will use this list as a basis to perform subsequent batch processing operations.

[0162] Step Six: The file system returns a list of files. After receiving the request, the file system returns a list of all file paths that meet the criteria. The rule execution engine then processes these files in batches according to the list.

[0163] Step 7: Perform batch file processing operations in a loop. The rule execution engine processes each file according to the returned file list. In each loop, the rule execution engine opens the file, parses its content, and performs operations based on the rule set, including: keyword replacement of file content, filename replacement, and directory reconstruction according to instructions. After the replacement is completed, the new content is written back to the target directory, ensuring that the processing takes effect in real time.

[0164] Step 8: Submit processing results and return feedback. Once all target files have been processed, the rule execution engine submits the updated file status to the file system. Subsequently, the rule execution engine returns the processing results (including replacement statistics, processing logs, and exception records) to the user interface.

[0165] Step Nine: The user reviews the feedback and confirms the results. The user interface displays the final processing report, where the user can view detailed logs, including a list of modified files, the number of replacements, and the processing status. If an exception occurs during processing, the report will also record breakpoint information, allowing the user to choose whether to resume and continue unfinished processing based on this information.

[0166] After implementation, this invention can improve the efficiency and quality of document management and processing, specifically in the following aspects: In terms of efficiency, the average time for document comparison is reduced from about 30 minutes per comparison to about 3 minutes, the efficiency of difference location is improved by about 10 times, the batch processing efficiency is improved by about 85%, and the preparation time for production documents of a single project is reduced from 6 hours to about 45 minutes; In terms of quality assurance, the production accident rate caused by version inconsistency is reduced by about 90%, the execution error rate caused by format problems is reduced to below 1%, and the accuracy of keyword replacement can reach 99.9%, effectively avoiding configuration errors caused by manual operation; In terms of process automation, the system supports both graphical user interface (GUI) and command line interface (CLI) operation modes, which can meet the different usage needs of technical and non-technical personnel, and realize the full automation and intelligence of document comparison, format verification and batch processing.

[0167] This invention proposes a three-dimensional difference comparison scheme for file comparison technology, breaking through the limitations of existing tools that can only compare from a single dimension. It establishes a three-level difference analysis system encompassing "directory structure, file attributes, and content details." This system can not only mark newly added, deleted, and renamed directories at the directory structure level, but also compare file size, modification time, hash value, etc., at the file attribute level, and provide line-by-line difference highlighting at the content level. In particular, it innovatively introduces a visual comparison display function for binary files, enabling binary file differences to be presented intuitively, thereby reducing manual verification costs and improving the accuracy and clarity of comparison results.

[0168] Regarding cross-platform compatibility, this invention proposes a cross-platform format self-adaptation engine, constructing a closed-loop mechanism of "detection, early warning, and repair." This mechanism can automatically identify differences in file encoding formats, line breaks, etc., and provide real-time early warnings and standardized repair solutions upon detecting problems. Through this mechanism, it effectively solves the problem of script execution failures caused by format differences between different operating systems such as Windows, Linux, and macOS, achieving seamless adaptation of file processing in cross-platform environments and filling the gaps in existing tools regarding automated process integration and preventative handling capabilities.

[0169] Regarding batch processing technology, this invention proposes a batch processing scheme based on dynamic rule-driven principles. Through configurable keyword mapping rules and variable parsing technology, it transforms traditionally manual, repetitive operations into automated processes. This scheme supports a single-source, multi-target file generation mode and can be flexibly configured to meet complex production or multi-environment requirements. Its built-in intelligent rule engine combines natural language parsing and intelligent prediction configuration technology to automatically generate adaptive rules, significantly improving the ease of use and intelligence of rule configuration, thereby greatly enhancing the efficiency and accuracy of batch processing.

[0170] In terms of system integration capabilities, this invention provides multi-faceted interactive interfaces, including a graphical user interface (GUI), a command-line interface (CLI), and an API interface, enabling deep integration with version control systems, continuous integration tools, and project management platforms. Through this full-scenario integration mechanism, this invention forms a complete automated closed-loop solution for file processing throughout the entire development, testing, deployment, and production process, achieving efficient, unified, and intelligent file management. Through the aforementioned technological innovations and scenario-based design, this invention effectively addresses the core pain points of batch file processing in office development.

[0171] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0172] This invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements any of the methods described above.

[0173] like Figure 13 As shown, this embodiment provides a batch file intelligent processing system for office development, which includes:

[0174] The parameter receiving and verification module is used to receive and verify the processing parameters input by the user. The processing parameters include the source directory path, the target directory path, the processing range, the replacement rules, and the output path.

[0175] The file index building module is used to traverse the specified directory based on the source directory path, target directory path and processing range, generate the directory structure and collect file information to build a file index.

[0176] The difference comparison module is used to locate the changed files between the source directory and the target directory based on the file index, first using the hash value incremental comparison algorithm, and then performing multi-dimensional difference comparison of the changed files at the structure layer, file layer and content layer, and recording the difference comparison results of each dimension.

[0177] The format repair module is used to start the encoding detection engine and the delimiter detection engine for parallel analysis based on the file index and the target file provided by the user, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user and generate corresponding diagnostic reports, generate repair strategies based on the diagnostic reports, and perform format repair on the files in the file index and the target file provided by the user according to the repair strategies.

[0178] The batch processing module is used to receive and parse the natural language instructions input by the user to generate an executable rule set based on the replacement rules, identify the file types of files in the target directory after difference comparison and format repair, and mark the safe replacement area and sensitive area of ​​the file content. Based on the executable rule set and the marked safe replacement area and sensitive area, the module performs batch processing operations of file name, file content and directory structure in parallel.

[0179] The report generation module is used to summarize the comparison results of each dimension, the processing results of format repair, and the processing results of batch processing operations, generate an output report, and store it in the output path.

[0180] In some embodiments, the difference comparison module includes:

[0181] The file change location unit is used to locate changed files between the source directory and the target directory based on a hash value incremental comparison algorithm.

[0182] A multi-dimensional comparison unit is used to perform multi-dimensional difference comparison, the multi-dimensional comparison unit comprising:

[0183] The structural layer comparison sub-unit is used to visualize the directory tree structure of the source directory and the target directory, mark newly added, deleted or renamed directories, and obtain the directory tree difference annotation results;

[0184] The file-level comparison subunit is used to compare the file size, modification time, and hash value of corresponding files in the source directory and the target directory to achieve file-level difference annotation and obtain file-level difference annotation results, which include differences in file size, modification time, and hash value.

[0185] The content layer comparison sub-unit is used to highlight differences line by line for text files and perform byte-level difference comparison for binary files to obtain content difference annotation results.

[0186] The visualization report generation unit is used to generate a visualization comparison report that includes the directory tree difference annotation results, the file-level difference annotation results, and the content difference annotation results.

[0187] In some embodiments, the content layer comparison subunit includes:

[0188] The difference file filtering subunit is used to filter difference files that have differences using a hash algorithm;

[0189] The byte comparison subunit is used to compare the difference files byte by byte and convert the byte data into binary strings and hexadecimal strings;

[0190] The difference identifier subunit is used to highlight the difference bytes or difference bits and generate hexadecimal difference display reports and binary difference display reports.

[0191] In some embodiments, the difference display report includes: byte offset positioning; statistics on the number of bit changes; comparison of hexadecimal value changes; analysis of ASCII character changes; and description of binary bit change sequences.

[0192] In some embodiments, the format repair module includes:

[0193] The target file receiving unit is used to receive the target file provided by the user.

[0194] A parallel detection engine is used to launch the advanced delimiter detection engine and the intelligent encoding detection engine in parallel to analyze the target file;

[0195] The delimiter diagnostic unit is used to identify the delimiter type and generate a delimiter diagnostic report through the advanced delimiter detection engine.

[0196] The encoding diagnosis unit is used to identify the encoding of the target file through the intelligent encoding detection engine and generate an encoding diagnosis report.

[0197] The repair execution unit is used to generate a dynamic repair strategy based on the delimiter diagnostic report and the encoding diagnostic report, and to perform repair on the target file.

[0198] In some embodiments, the delimiter diagnostic unit includes:

[0199] The delimiter scanning statistics subunit is used to scan the target file line by line using a predefined regular expression, identify the control character type at the end of each line, count the frequency and distribution range of different types of delimiters, and obtain the delimiter type statistics results.

[0200] The anomaly detection subunit is used to check the distribution of delimiters in conjunction with the logical structure of the target file, determine whether the inconsistency of delimiter patterns between adjacent lines is an anomaly, and obtain the delimiter anomaly detection result.

[0201] The heatmap generation subunit is used to generate a heatmap of line ending patterns based on line ending delimiter data. This includes statistically analyzing the delimiter types and distribution positions of all lines in the target file to form a basic dataset; marking the line numbers and ranges of mixed delimiters when multiple delimiters are used in the target file to obtain the line number and range marking information of mixed delimiters; calculating the compatibility score of the target file on different platforms based on the consistency, distribution concentration, and mixing ratio of delimiters; and determining the line ending pattern heatmap based on the basic dataset, mixed delimiter marking information, and compatibility score.

[0202] The separator report generation subunit is used to generate the separator diagnostic report based on the separator type statistics, separator anomaly determination results, and line ending pattern heatmap.

[0203] In some embodiments, the coding diagnostic unit includes:

[0204] The BOM detection subunit is used to check whether there is a BOM byte sequence at the beginning of the target file. If it is detected, the file encoding type is identified based on the BOM byte sequence.

[0205] The encoding inference subunit is used to infer the encoding type of the target file by utilizing the byte distribution characteristics, character frequency statistics and character set matching strategy of the target file content if no BOM byte sequence is detected, and obtain the inferred encoding type.

[0206] The multimodal judgment subunit is used to call the multimodal learning model to make a comprehensive judgment on the encoding type. This includes performing pattern matching in the feature space for files without BOM identifiers to identify the most likely encoding type, detecting whether there are mixed encoding regions in the target file, performing probability prediction on the most likely encoding type to obtain the probability value of each candidate encoding type, and outputting the optimal encoding scheme based on the probability values ​​of multiple candidate encodings.

[0207] The encoding report generation subunit is used to integrate and output the recognition results to generate an encoding diagnosis report. The encoding diagnosis report records the encoding type of the target file, the existing mixed encoding regions, and the severity.

[0208] In some embodiments, the repair execution unit includes:

[0209] The file input subunit is used to input the original file into the repair engine;

[0210] The repair scheme generation subunit is used by the repair engine to generate multiple different repair schemes for the encoding anomalies in the original file based on the abnormal encoding formats and repairable markers identified by the adaptive repair algorithm.

[0211] The difference visualization subunit is used to input the multiple different repair schemes into the difference visualization module, and to intuitively display the changes made by each repair scheme to the original file content through difference visualization;

[0212] The solution selection sub-unit is used to select a solution based on the solution selected by the user through the interactive interface or the solution selected by the weight-based recommendation system.

[0213] The repair application subunit is used to apply the selected repair scheme to the original file to complete the repair processing of the encoding anomaly of the original file.

[0214] In some embodiments, the batch processing module includes:

[0215] The rule set generation unit is used to receive natural language instructions input by the user through the interface, perform semantic analysis through the natural language processing module, and generate an executable rule set containing filename replacement rules, file content replacement rules, and directory structure mapping rules.

[0216] The file annotation unit is used to call the file type recognition module to distinguish between text files, code files and binary files in the target directory, parse the semantics of file content based on file type, mark safe replacement areas and sensitive areas, and generate a context annotation list.

[0217] The parallel processing unit is used to request a list of target files from the file system based on the executable rule set and the context annotation list, and to perform file name replacement, file content replacement and directory structure mapping operations in parallel.

[0218] The verification submission unit is used to create a temporary workspace to perform batch operations. It verifies the consistency of the processing results through hash verification. If the verification passes, it is atomically submitted to the target directory. If the verification fails, it is rolled back based on the breakpoint information.

[0219] The breakpoint handling and reporting unit is used to record task breakpoints that are interrupted due to exceptions during batch processing, resume execution from the breakpoint, generate a processing report containing a file modification list, replacement statistics and exception logs, and output it to the user interface.

[0220] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0221] The present invention also provides an electronic device. The electronic device according to embodiments of the present invention includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method provided by the present invention. Reference is made below. Figure 14 It shows a schematic diagram of the structure of a computer system 800 suitable for implementing embodiments of the present invention. For example... Figure 14 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 802 or programs loaded from storage section 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the computer system 800. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0222] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.

[0223] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for intelligent batch file processing in office development, characterized in that, include: S10: Receive and verify the processing parameters input by the user, the processing parameters including the source directory path, the target directory path, the processing range, the replacement rule, and the output path; S20: Based on the source directory path, the target directory path, and the processing range, traverse the specified directory, generate a directory structure, and collect file information to establish a file index; S30: Based on the file index, first use the hash value incremental comparison algorithm to locate the changed files between the source directory and the target directory, and then perform multi-dimensional difference comparison of the changed files at the structure layer, file layer and content layer, and record the difference comparison results of each dimension; S40: Based on the file index and the target file provided by the user, start the encoding detection engine and the delimiter detection engine to perform parallel analysis, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user, generate corresponding diagnostic reports, generate repair strategies based on the diagnostic reports, and perform format repair on the files in the file index and the target file provided by the user according to the repair strategies; S50: Based on the replacement rules, receive and parse the natural language instructions input by the user to generate an executable rule set, identify the file type of the files in the target directory after difference comparison and format repair, and mark the safe replacement area and sensitive area of ​​the file content. According to the executable rule set and the marked safe replacement area and sensitive area, perform batch processing operations of file name, file content and directory structure in parallel. S60: Summarize the comparison results of each dimension, the processing results of format repair, and the processing results of batch processing operations, generate an output report, and store it in the output path.

2. The method according to claim 1, characterized in that, Step S30 specifically includes: S301: Locating changed files between the source and target directories based on the hash value incremental comparison algorithm; S302: Perform a multi-dimensional difference comparison, wherein the multi-dimensional difference comparison includes: The structural layer dimension compares the differences, visualizes the directory tree structure of the source directory and the target directory, marks the newly added, deleted or renamed directories, and obtains the directory tree difference annotation results. The file-level difference comparison compares the file size, modification time, and hash value of corresponding files in the source and target directories to achieve file-level difference annotation and obtain file-level difference annotation results, which include differences in file size, modification time, and hash value. The content-level difference comparison highlights the differences line by line for text files and performs byte-level difference comparison for binary files to obtain the content difference annotation results. S303: Generate a visual comparison report containing the directory tree difference annotation results, the file-level difference annotation results, and the content difference annotation results.

3. The method as described in claim 2, characterized in that, The content-level difference comparison highlights differences line by line for text files and performs byte-level binary code difference comparison for binary files to obtain content difference annotation results, specifically including: S3021: Filter out differing files using a hash algorithm; S3022: Compare the difference files byte by byte, and convert the byte data into binary strings and hexadecimal strings; S3023: Highlight the difference bytes or difference bits to generate a hexadecimal difference display report and a binary difference display report.

4. The method as described in claim 3, characterized in that: The difference display report includes: byte offset location; statistics on the number of bit changes; comparison of hexadecimal value changes; analysis of ASCII character changes; and description of binary bit change sequences.

5. The method as described in claim 3, characterized in that, Step S40 specifically includes: S401: Receive the target file provided by the user; S402: The advanced delimiter detection engine and the intelligent encoding detection engine are launched in parallel to analyze the target file; S403: Identify the delimiter type and generate a delimiter diagnostic report using the advanced delimiter detection engine; S404: Identify the encoding of the target file and generate an encoding diagnostic report using the aforementioned intelligent encoding detection engine; S405: Based on the delimiter diagnostic report and the encoding diagnostic report, generate a dynamic repair strategy and perform repair on the target file.

6. The method as described in claim 5, characterized in that, Step S403 specifically includes: The target file is scanned line by line using predefined regular expressions to identify the type of control character at the end of each line, and the frequency and distribution range of different types of delimiters are statistically analyzed to obtain the delimiter type statistics. The distribution of delimiters is checked in conjunction with the logical structure of the target file to determine whether the inconsistency of delimiter patterns between adjacent lines is an abnormal situation, and the delimiter abnormality judgment result is obtained. Generating a line-end pattern heatmap based on line-end delimiter data includes: statistically analyzing the delimiter types and distribution positions of all lines in the target file to form a basic dataset; marking the line numbers and ranges of mixed delimiters when multiple delimiters are used in the target file, obtaining the line number and range marking information of the mixed delimiters; calculating the compatibility score of the target file on different platforms based on delimiter consistency, delimiter distribution concentration, and delimiter mixing ratio; and determining the line-end pattern heatmap based on the basic dataset, the line number and range marking information of the mixed delimiters, and the compatibility score. Based on the statistical results of the delimiter types, the results of the delimiter anomaly determination, and the heatmap of the line ending pattern, the delimiter diagnostic report is generated.

7. The method as described in claim 5, characterized in that, Step S404 specifically includes: Check if a BOM byte sequence exists at the beginning of the target file. If a BOM byte sequence is detected, identify the file's encoding type based on the BOM byte sequence. If no BOM byte sequence is detected, the encoding type of the target file is inferred by using the byte distribution characteristics of the target file content, character frequency statistics, and character set matching strategy. The process involves using a multimodal learning model to comprehensively determine the encoding type. This includes: for files without a BOM (Bill of Materials) identifier, performing pattern matching in the feature space using the multimodal learning model to identify the most likely encoding type; utilizing the context-aware capability of the multimodal learning model to detect whether there are mixed encoding regions in the target file, where a mixed encoding region refers to content segments in the target file corresponding to different encoding types; using the multimodal learning model to predict the probability of the most likely encoding type, obtaining probability values ​​for each candidate encoding type, wherein the candidate encoding includes the inferred encoding type and the most likely encoding type; and outputting the optimal encoding scheme based on the probability values ​​of multiple candidate encodings. The identification results are integrated and output to generate an encoding diagnosis report. The encoding diagnosis report records the encoding type of the target file, the existing mixed encoding regions and their severity. The severity is the degree of impact on file parsing determined based on the range of the mixed encoding regions and the number or proportion of abnormal characters.

8. The method as described in claim 5, characterized in that, Step S405 includes the following steps: Input the original file into the repair engine; The repair engine generates multiple different repair schemes for the encoding anomalies in the original file based on the abnormal encoding formats and repair marks identified by the adaptive repair algorithm. These multiple different repair schemes differ in encoding conversion strategies, hybrid encoding segmentation processing strategies, and delimiter unification rules. The various repair schemes are input into the difference visualization module, which visually displays the changes made by each repair scheme to the original file content. These changes include encoding adjustments, delimiter adjustments, and the processing effects of using different encodings for different paragraphs. The repair scheme is selected by the user from multiple different repair schemes through an interactive interface, or by a weight-based recommendation system from the multiple different repair schemes. The selected repair scheme is applied to the original file to complete the repair of the encoding anomaly in the original file.

9. The method as described in claim 1, characterized in that, Step S50 specifically includes: S501: Receives natural language commands input by the user through the interface, performs semantic analysis through the natural language processing module, and generates an executable rule set containing filename replacement rules, file content replacement rules, and directory structure mapping rules; S502: Call the file type recognition module to distinguish between text files, code files, and binary files in the target directory; parse the semantics of file content based on file type, mark safe replacement areas and sensitive areas, and generate a context annotation list; S503: Request a list of target files from the file system based on the executable rule set and the context annotation list; perform the following operations in parallel: Reconstruct the target path according to the filename replacement rules; According to the file content replacement rules, context-aware replacement is performed on the dynamic variables in the safe replacement area. The dynamic variables include date, request number, and user-defined parameters. Generate a multi-level directory associated with the runtime environment based on the directory structure mapping rules; S504: Create a temporary workspace to perform batch operations and verify the consistency of the processing results through hash verification; if the verification passes, it is atomically committed to the target directory; if the verification fails, it is rolled back based on the breakpoint information. S505: Records task breakpoints for incomplete operations interrupted by exceptions during batch processing, resumes execution from the breakpoints, generates a processing report containing a file modification list, replacement statistics, and exception logs, and outputs it to the user interface.

10. A batch file intelligent processing system for office development, characterized in that, include: The parameter receiving and verification module is used to receive and verify the processing parameters input by the user. The processing parameters include the source directory path, the target directory path, the processing range, the replacement rules, and the output path. The file index building module is used to traverse the specified directory based on the source directory path, target directory path and processing range, generate the directory structure and collect file information to build a file index. The difference comparison module is used to locate the changed files between the source directory and the target directory based on the file index, first using the hash value incremental comparison algorithm, and then performing multi-dimensional difference comparison of the changed files at the structure layer, file layer and content layer, and recording the difference comparison results of each dimension. The format repair module is used to start the encoding detection engine and the delimiter detection engine for parallel analysis based on the file index and the target file provided by the user, identify the encoding type and line delimiter type of the files in the file index and the target file provided by the user and generate corresponding diagnostic reports, generate repair strategies based on the diagnostic reports, and perform format repair on the files in the file index and the target file provided by the user according to the repair strategies. The batch processing module is used to receive and parse the natural language instructions input by the user to generate an executable rule set based on the replacement rules, identify the file types of files in the target directory after difference comparison and format repair, and mark the safe replacement area and sensitive area of ​​the file content. Based on the executable rule set and the marked safe replacement area and sensitive area, the module performs batch processing operations of file name, file content and directory structure in parallel. The report generation module is used to summarize the comparison results of each dimension, the processing results of format repair, and the processing results of batch processing operations, generate an output report, and store it in the output path.

Citation Information

Cited By

  • Method and system for unified expression of multi-step parameters of single-line command and isomorphic linkage of GUI (Graphical User Interface) based on action range marking

    CN121979570A

  • Multi-architecture instruction analysis method, device, equipment, medium and product

    CN122132086A