File processing method and device, electronic equipment and computer readable storage medium
By acquiring multidimensional attribute information of the target file, accurately identifying the file type using the target processing model, and selecting an appropriate compression strategy, the problem of insufficient file type differentiation in OTA upgrades is solved, thereby reducing the size of the upgrade package and improving transmission efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies fail to effectively distinguish file types during OTA upgrades, resulting in larger compressed file sizes, increased transmission time, and negatively impacting the terminal upgrade experience and resource utilization efficiency.
By obtaining multidimensional attribute information of the target file, including the extension and magic number, the target processing model is used to accurately identify the file type, and an appropriate compression strategy is selected for compression processing based on the file type.
It significantly reduced the size of the upgrade package, improved transmission and storage efficiency, and enhanced the intelligence level of file transmission and storage.
Smart Images

Figure CN121833633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a file processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the widespread application of OTA (Over-The-Air) upgrade technology, compressing OTA upgrade packages has become a key means of reducing data transmission volume.
[0003] Currently, most mainstream packaging tools use a uniform compression scheme for all files. However, this approach cannot effectively differentiate between file types, resulting in larger compressed file sizes, increased transmission time, and consequently impacting the terminal upgrade experience and resource utilization efficiency. Summary of the Invention
[0004] This application provides a file processing method, apparatus, electronic device, and computer-readable storage medium that can accurately identify the file type of a target file and perform targeted compression processing based on the file type, thereby significantly reducing the size of the upgrade package and improving transmission and storage efficiency.
[0005] In a first aspect, embodiments of this application provide a document processing method, including: Obtain the target file and its multidimensional attribute information; Based on the multidimensional attribute information, the file type of the target file is determined; Based on the file type, the target file is compressed to obtain a compressed file.
[0006] In one embodiment, determining the file type of the target file based on the multidimensional attribute information includes: Feature extraction is performed on the multidimensional attribute information to obtain target feature information; The target feature information is input into the target processing model, and the target processing model outputs the file type.
[0007] In one embodiment, the multidimensional attribute information includes the file extension and the magic number; The step of extracting features from the multidimensional attribute information to obtain target feature information includes: Feature extraction is performed on the file extension to obtain the first feature information; The magic number is subjected to feature extraction to obtain second feature information; The first feature information and the second feature information are weighted and fused to obtain the target feature information.
[0008] In one embodiment, after compressing the target file based on the file type to obtain a compressed file, the method further includes: Obtain the evaluation information corresponding to the compressed file; the evaluation information includes one or more of the following: decompression success rate, decompression speed, and space saving information; Based on the evaluation information, the target processing model is updated.
[0009] In one embodiment, the step of compressing the target file based on the file type to obtain a compressed file includes: Based on the correspondence between the file types and the preset file types and compression strategies, a target compression strategy is determined; the target compression strategy includes a target compression algorithm and target compression parameters. Based on the target compression strategy, the target file is compressed to obtain the compressed file.
[0010] In one embodiment, after compressing the target file based on the target compression strategy to obtain the compressed file, the method further includes: Based on the target compression strategy, generate the metadata file for the compressed file; The compressed file and the metadata file are integrated to obtain an upgrade file package; The upgrade file package is sent to the target terminal, so that the target terminal determines the target decompression algorithm based on the metadata file, and decompresses the compressed file based on the target decompression algorithm to obtain the target file.
[0011] In one embodiment, the target file is obtained in the following manner: In response to a version upgrade request, retrieve both the new and old version files; The new version file and the old version file are differentially processed to obtain the target file.
[0012] Secondly, embodiments of this application provide a document processing apparatus, the apparatus comprising: The file acquisition module is used to acquire the target file and its multidimensional attribute information. The type determination module is used to determine the file type of the target file based on the multidimensional attribute information; The file compression module is used to compress the target file based on the file type to obtain a compressed file.
[0013] In one embodiment, the type determination module includes: The feature extraction submodule is used to extract features from the multidimensional attribute information to obtain target feature information; The type determination submodule is used to input the target feature information into the target processing model and output the file type through the target processing model.
[0014] In one embodiment, the multidimensional attribute information includes the file extension and the magic number; the feature extraction submodule includes: The first feature extraction unit is used to extract features from the file extension to obtain first feature information; The second feature extraction unit is used to extract features from the magic number to obtain second feature information; The feature fusion unit is used to perform weighted fusion of the first feature information and the second feature information to obtain the target feature information.
[0015] In one embodiment, the file processing apparatus further includes: An evaluation information acquisition module is used to acquire evaluation information corresponding to the compressed file; the evaluation information includes one or more of the following: decompression success rate, decompression speed, and space saving information. The model update module is used to update the target processing model based on the evaluation information.
[0016] In one embodiment, the file compression module includes: The compression strategy determination submodule is used to determine the target compression strategy based on the file type and the preset correspondence between file types and compression strategies; the target compression strategy includes the target compression algorithm and the target compression parameters; The compression processing submodule is used to compress the target file based on the target compression strategy to obtain the compressed file.
[0017] In one embodiment, the file processing apparatus further includes: The file generation module is used to generate the metadata file of the compressed file based on the target compression strategy; The file integration module is used to integrate the compressed file and the metadata file to obtain an upgrade file package; The file sending module is used to send the upgrade file package to the target terminal, so that the target terminal can determine the target decompression algorithm based on the metadata file, and decompress the compressed file based on the target decompression algorithm to obtain the target file.
[0018] In one embodiment, the target file is obtained in the following manner: In response to a version upgrade request, retrieve both the new and old version files; The new version file and the old version file are differentially processed to obtain the target file.
[0019] Thirdly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the above-described file processing method.
[0020] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described file processing method.
[0021] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of this application.
[0022] In summary, the embodiments of this application can obtain the target file and its multidimensional attribute information, determine the file type of the target file based on the multidimensional attribute information, and then compress the target file based on the file type to obtain a compressed file. Thus, the file type of the target file can be accurately identified based on its multidimensional attribute information, and targeted compression processing can be performed on the target file according to the file type, thereby significantly improving compression efficiency, reducing the size of the compressed file, and improving file transfer and storage efficiency. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of a document processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating a specific embodiment of determining the file type provided in this application; Figure 3 This is a schematic diagram of a specific embodiment of the update target processing model provided in this application; Figure 4This is a schematic diagram illustrating a specific embodiment of the transmission upgrade file package provided in this application; Figure 5 This is a schematic diagram of the structure of a document processing apparatus provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] It should be noted that in the current OTA upgrade package generation process, after the server obtains the new version file, it typically uses a uniform compression algorithm, such as ZIP (zipper compression algorithm), for packaging. However, this method does not differentiate between file type characteristics, applying the same compression strategy to all files (including .apk, .so, .xml, .png, etc.), which has significant limitations: First, text files (such as .xml, .txt) have a much higher compression ratio than ZIP under algorithms such as Brotli (Brotli compression algorithm), but the existing solution fails to take advantage of this; second, simply changing the compression algorithm for binary or highly compressed files (such as .apk, .so) will not improve the performance much and may even cause errors; in addition, the overall package size optimization is limited, especially on devices with limited storage resources.
[0027] Existing technologies lack intelligent file type recognition mechanisms and usually rely on manual rule classification, which is not only labor-intensive and inaccurate, but also difficult to adapt to new file formats and complex upgrade scenarios, resulting in suboptimal compression efficiency and insufficient system scalability and automation.
[0028] To address the problem of poor compression results caused by the difficulty in accurately identifying file types, this application aims to provide a file processing method that can acquire the target file and its multi-dimensional attribute information, determine the file type of the target file based on the multi-dimensional attribute information, and then compress the target file according to the file type to obtain a compressed file. In this way, the file type of the target file can be accurately identified based on its multi-dimensional attribute information, and targeted compression processing can be performed on the target file according to the file type, thereby significantly improving compression efficiency, reducing compressed file size, and improving file transfer and storage efficiency.
[0029] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0030] Figure 1 The illustration shows a schematic flowchart of a file processing method according to an embodiment of this application. The executing entity of this file processing method can be a file processing device, which can be integrated into any electronic device with data processing, network communication, and program execution functions. This electronic device can be a server or a terminal, etc.
[0031] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.
[0032] The terminal can be a smartphone, tablet, laptop, desktop computer, smart home device, etc., but is not limited to these. The terminal and server can be connected directly or indirectly through wired or wireless communication, which is not limited herein.
[0033] Furthermore, in the embodiments of this application, "multiple" refers to two or more. The terms "first" and "second," etc., in the embodiments of this application are used for distinguishing descriptions and should not be construed as implying relative importance.
[0034] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0035] In this embodiment, the description will be from the perspective of a file processing device, which can be integrated into a server or terminal. To facilitate the explanation of the file processing method of this application, the following will describe the process in detail with the file processing device integrated into the server, that is, with the server as the execution subject.
[0036] Reference Figure 1 This application illustrates a document processing method, which may specifically include the following steps S101-S103, as follows: S101: Obtain the target file and its multidimensional attribute information.
[0037] In this embodiment, the target file refers to a file that needs to be compressed. This target file may include one or more of the following: video data, audio data, image data, text data, and sensor data. In an OTA (Over-The-Air) version upgrade scenario, the target file refers to an upgrade file that needs to be compressed. This target file is used to upgrade the client version of the target terminal.
[0038] In this embodiment, multidimensional attribute information represents the attribute information of the target file in at least two dimensions. In one embodiment, the multidimensional attribute information may include the target file's file extension and magic number; in another embodiment, the multidimensional attribute information may further include the target file's file extension, magic number, and content features. The content features may include size features, structural features, and / or semantic features.
[0039] It's important to note that the file extension refers to the string following the last dot (.) in the filename of the target file. Examples include .txt, .so, .png, .jpg, and .pdf. Based on the file extension, most files can be quickly and initially categorized; for example, file types can include text files, executable files, image files, video files, audio files, and unknown types.
[0040] In this embodiment, the file extension can be extracted as follows: parsing the file pathname of the target file; extracting the substring between the last separator and the last period in the file pathname; and determining the extension based on the position of the last period. Specifically, the step of determining the extension based on the position of the last period includes: if the last period is after the last separator, identifying all characters after that period until the end of the pathname as the extension of the target file; if the file pathname does not contain any periods, or the last period is before the last separator, then determining that the extension of the target file is empty.
[0041] In one example, the full path name of the target file is: / system / app / Application.apk. The file path name is scanned backwards, locating the last occurrence of a standard path separator (e.g., " / " in Unix-style systems or "\" in Windows-style systems) and the last occurrence of a period ".". If the last period is detected after the last separator, all characters from that period until the end of the path name (i.e., apk) are recognized as the file extension of the target file.
[0042] It should be noted that the magic number represents a specific set of bytes in the header of a target file, used to uniquely identify the file format. The magic number can accurately identify the true format of a file, effectively identify files that are disguised or have incorrect extensions, and greatly improve the accuracy of type identification.
[0043] In this embodiment, the magic number can be extracted by: opening the target file in read-only mode; reading a byte sequence of a predetermined number of bytes at a predetermined offset from the file header; and determining the byte sequence as the magic number.
[0044] In the specific implementation, a file operation interface can be called to access the target file in binary read-only mode. Then, starting from the beginning of the target file (i.e., offset 0), a continuous sequence of bytes of a predetermined length is read (e.g., the first 8 bytes, 16 bytes, or other byte numbers determined according to the signature length of common file formats). Finally, the successfully read segment of raw byte data is used as the magic number representing the true format of the target file. If an error occurs during the reading process (such as the file not existing or insufficient permissions), the magic number is recorded as empty or the reading failed.
[0045] In this embodiment, by obtaining the multidimensional attribute information of the target file, a comprehensive and rich data foundation can be provided for subsequent accurate determination of the file type.
[0046] S102: Determine the file type of the target file based on multi-dimensional attribute information.
[0047] In this embodiment, compared with traditional methods that rely solely on a single file extension for identification, a comprehensive approach based on multi-dimensional attribute information effectively avoids misjudgments caused by tampering or errors in the file extension, thereby significantly improving the accuracy and robustness of file type identification. For example, if the target file has a .jpg extension but the Magic Number indicates it is actually in ZIP format, it can be accurately identified as a compressed file.
[0048] S103: Based on the file type, compress the target file to obtain a compressed file.
[0049] In this embodiment, after accurately determining the file type, the most suitable compression strategy can be selected and applied based on the characteristics of that file type to perform compression processing, thereby obtaining a compressed file.
[0050] In practice, if the file type is determined to be text, the Brotli compression algorithm can be used for compression. If the file type is determined to be a compressed image, compression can be skipped or lossless optimization tools can be used for image-specific optimization. If the file type is determined to be binary or an executable file type requiring compatibility, the ZIP algorithm, which balances compression ratio and decompression speed, can be selected for compression. If the file type is determined to be unknown, a preset general compression algorithm can be used for compression or compression can be skipped.
[0051] In this embodiment, by acquiring multi-dimensional attribute information of the target file, the file type of the target file can be accurately determined, and then the corresponding compression process can be adaptively selected and executed based on the file type. This allows for differentiated compression processing for different types of files, significantly improving compression efficiency, effectively reducing the size of the final compressed file, and saving storage space and network transmission bandwidth.
[0052] In one feasible implementation, refer to Figure 2 The steps for determining the file type of the target file based on multi-dimensional attribute information may specifically include steps S201 to S202, as follows: S201: Extract features from multidimensional attribute information to obtain target feature information.
[0053] In this embodiment, feature extraction refers to the process of transforming raw, unstructured, multidimensional attribute information into standardized target feature information that can be recognized and processed by the target processing model.
[0054] In this embodiment, by extracting target feature information from multidimensional attribute information, the recognition effect of subsequent target processing models on target feature information can be effectively improved.
[0055] In this embodiment, the multidimensional attribute information includes the file extension and the magic number; feature extraction is performed on the multidimensional attribute information to obtain target feature information, including: feature extraction of the file extension to obtain first feature information; feature extraction of the magic number to obtain second feature information; and weighted fusion of the first feature information and the second feature information to obtain target feature information.
[0056] In this implementation, the first feature information represents standardized machine-readable features derived from the extension conversion. Specifically, the extension can be mapped to an encoded vector in a predefined dictionary, or word embedding techniques can be used to transform the extension into a low-dimensional, dense feature vector. For example, the extension "txt" can be converted into a 128-dimensional vector. This process allows the model's input to no longer be limited to the original character form, enabling better capture of the semantic relationships between different extensions.
[0057] In this embodiment, the first feature information represents a standardized machine-readable feature derived from the magic number. In a specific implementation, the numerical value corresponding to the byte string of the magic number can be determined as the second feature information, or the feature information corresponding to the magic number can be extracted using a lightweight neural network. For example, the 8-byte hexadecimal value corresponding to the magic number can be converted into an 8-dimensional integer vector. This ensures that the most essential type information of the target file can be effectively recognized by the target processing model.
[0058] In this embodiment, the step of weighted fusion of the first feature information and the second feature information to obtain the target feature information may specifically include: obtaining the first weight corresponding to the first feature information and the second weight corresponding to the second feature information; and weighted fusion of the first feature information and the second feature information based on the first weight and the second weight to obtain the target feature information.
[0059] In this embodiment, the first weight and the second weight can be fixed values pre-set based on experience. For example, since the magic number is usually more deterministic than the extension, it can be assigned a higher weight. Alternatively, the first weight and the second weight can also be automatically learned by the target processing model during training, thereby achieving dynamic adjustment of the first weight and the second weight.
[0060] In this embodiment, by weighted fusion of the feature information corresponding to the two different dimensions of attribute information, namely file extension and magic number, more comprehensive target feature information can be obtained. This enables the target processing model to adaptively evaluate the credibility of different feature sources, effectively deal with the situation of file extension forgery, error or missing extensions, and make accurate judgments based on the higher weight of the magic number feature, which greatly reduces the false judgment rate and thus improves the intelligence level and reliability of file type recognition.
[0061] S202: Input the target feature information into the target processing model, and output the file type through the target processing model.
[0062] In this embodiment, the target processing model is a computational model with file type recognition capabilities, pre-trained on a large amount of data. The target processing model can be based on a deep neural network, such as a convolutional neural network, a recurrent neural network, or a large model. This target processing model learns the mapping relationship between a massive number of sample files of known file types and the feature information corresponding to the multi-dimensional attribute information of the sample files.
[0063] It's important to note that large models refer to artificial neural network models with an extremely large number of parameters. In the field of artificial intelligence, large models typically refer to models with hundreds of millions to trillions of parameters. These models usually need to be trained on massive datasets and require significant computational resources for optimization and tuning. Large models are commonly used to solve complex tasks such as natural language processing, computer vision, and speech recognition.
[0064] In this embodiment, the large model can be a large-scale pre-trained model such as Doubao, ChatGPT series, BERT, XLNet, Zhipu model, Claude, Moonshot AI model, ChatGLM model, Tongwen Qianyi model, MiniMax model, Xinghuo model, Llama model, 360GPT model, Qwen model, Baichuan model, Yunque model, vivoLM model, deepseek, Tencent Yuanbao and Wenxin Yiyan, etc. This application embodiment does not limit it.
[0065] In this embodiment, by inputting target feature information into the target processing model, the model can perform nonlinear transformation and weighted calculation on the target feature information, and finally generate a classification result for file type at the output layer. This classification result can be a type label corresponding to a file type, or it can be a probability distribution of various file types, and the file type with the highest probability is selected from the probability distribution as the file type of the target file.
[0066] In this embodiment, compared with the prior art that relies solely on simple rule judgment, this solution can effectively improve the recognition accuracy and robustness by extracting features from multi-dimensional attribute information and using the target processing model to identify file types. Even if a certain attribute of the target file (such as the extension) is tampered with or incorrect, the target processing model can still make a correct judgment based on other effective attribute features (such as the magic number), which greatly enhances the system's anti-interference capability.
[0067] In one feasible implementation, refer to Figure 3 Based on the file type, the target file is compressed to obtain a compressed file. After that, the file processing method also includes steps S301 to S302, as follows: S301: Obtain the evaluation information corresponding to the compressed file.
[0068] In this embodiment, the evaluation information includes one or more of the following: decompression success rate, decompression speed, and space saving information.
[0069] In this embodiment, the decompression success rate reflects whether the target terminal can successfully and accurately decompress and restore the target file. This indicator directly reflects the compatibility between the compressed file and the terminal environment. If the decompression failure rate is too high, it indicates that there may be an error in file type identification.
[0070] In this embodiment, decompression speed refers to the time required to complete file decompression on the target terminal. This metric effectively reflects user experience, especially in scenarios such as system updates, where excessively long decompression times are unacceptable to users.
[0071] In this embodiment, space-saving information is used to characterize the volume compression effect of the target file. This space-saving information can be calculated based on the first volume of the compressed file and the second volume of the target file. For example, the space-saving information can be determined by the ratio between the difference between the second volume and the first volume and the second volume.
[0072] In this embodiment, by collecting information on decompression success rate, decompression speed, and space saving, a comprehensive and effective evaluation of the compression effect of the target file can be achieved.
[0073] S302: Update the target processing model based on the evaluation information.
[0074] In this embodiment, the evaluation information can be used as training data for a new round to adjust the model parameters inside the target processing model, such as adjusting the weights and biases of the neural network in the target processing model.
[0075] In practical implementation, evaluation information of compressed files from multiple terminals can be collected according to the target acquisition cycle. Training samples are generated based on the evaluation information and the target feature information corresponding to the multi-dimensional attribute information of the target file corresponding to the compressed file. The target processing model is then updated based on the training samples to obtain the updated target processing model. Specifically, if the evaluation information is a good result, such as a high decompression success rate or high space saving, the target processing model will strengthen the decision path for the compressed file corresponding to that evaluation information; conversely, if the evaluation information has a negative result, such as decompression failure, the target processing model will weaken the tendency of that decision.
[0076] In this implementation, by introducing a model update mechanism based on actual compression performance feedback, the model parameters of the target processing model can be continuously optimized during actual application, improving the ability to identify file types. Thus, in OTA upgrade scenarios, the efficiency and accuracy of OTA upgrades can be continuously improved.
[0077] In one feasible implementation, the step of compressing a target file based on its file type to obtain a compressed file may specifically include: determining a target compression strategy based on the file type and the correspondence between preset file types and compression strategies; and compressing the target file based on the target compression strategy to obtain a compressed file.
[0078] In this embodiment, the server stores a compression policy table, which represents the correspondence between file types and compression policies. The compression policy includes a compression algorithm and compression parameters.
[0079] In this embodiment, the compression algorithm may include ZIP, Brotli, LZMA (Lempel-Ziv-Markovchain Algorithm, a decompression algorithm based on the Lempel-Ziv algorithm and incorporating a Markov chain model), and Zstandard (Zstandard Compression Algorithm, a modern compression algorithm developed by Facebook), etc.
[0080] In this embodiment, compression parameters refer to configuration options that control the behavior of the compression algorithm. Compression parameters are used to finely adjust the compression process, balancing aspects such as compression ratio, compression speed, and memory usage. These compression parameters include, but are not limited to, compression level, window size, and dictionary training options. The compression level includes integers from 1 to 9, where 1 represents the fastest compression or lowest compression ratio, and 9 represents the slowest compression or highest compression ratio.
[0081] In this embodiment, after determining the file type, the target compression strategy corresponding to that file type can be retrieved from the compression strategy table. This target compression strategy includes the target compression algorithm and target compression parameters.
[0082] In one example, the target file is identified as a plain text file. By querying the compression strategy table, the target compression strategy corresponding to the plain text file is found to be: {target compression algorithm: "Brotli", target compression parameters: {compression level: 8}}. Therefore, when compressing the target file, the Brotli algorithm will be used, and the compression level will be set to 8 (a level that tends to have a high compression ratio).
[0083] In this embodiment, by establishing a correspondence between file types and compression strategies, a suitable target compression strategy can be matched for compression operations based on this correspondence. This maximizes compression efficiency. Simultaneously, by setting compression parameters, computing resources can be intelligently allocated; for example, fast compression can be used for insensitive files to save time, while high compression ratios can be used for critical text files to save space, thus achieving an optimal balance between compression efficiency and processing speed.
[0084] In one feasible implementation, refer to Figure 4 Based on the target compression strategy, the target file is compressed to obtain a compressed file. The file processing method also includes steps S401 to S403, as follows: S401: Generate metadata files for compressed files based on the target compression strategy.
[0085] In this embodiment, the metadata file is generated based on the target compression strategy, and the metadata file includes at least the target compression algorithm used to compress the file and the related target compression parameters. For example, the metadata file can be stored as an XML or JSON file.
[0086] In this embodiment, the metadata file may further include information such as the name of the compressed file and its size before and after compression. By comprehensively recording information related to the compressed file, the server can improve the comprehensiveness of the metadata file, thereby serving as a basis for decompression and compatible restoration.
[0087] S402: Integrate the compressed file and metadata file to obtain the upgrade file package.
[0088] In this embodiment, the server can package the compressed file and metadata file into a complete upgrade file package. For example, the compressed file and metadata file can be packaged together according to the target format to obtain the upgrade file package.
[0089] S403: Send the upgrade file package to the target terminal so that the target terminal can determine the target decompression algorithm based on the metadata file, and decompress the compressed file based on the target decompression algorithm to obtain the target file.
[0090] In this embodiment, since the target compression algorithm and target compression parameters are recorded in the metadata file, the target terminal can determine the target decompression algorithm corresponding to each compressed file. For example, for a compressed file recorded as using the Brotli algorithm, the target terminal can call the Brotli decompression library to decompress the compressed file using the corresponding target decompression algorithm to obtain the target file.
[0091] In this embodiment, by introducing a metadata file and establishing a collaborative decompression mechanism with the terminal, the decompression command can be accurately transmitted through the metadata file, ensuring that the decompression algorithm used by the target terminal is completely matched with the compression algorithm used by the server. This ensures the compatibility and reliability of the entire compression-decompression process and avoids decompression failure or data corruption caused by algorithm mismatch.
[0092] In one feasible implementation, the target file is obtained by: in response to a version upgrade request, acquiring a new version file and an old version file; performing differential processing on the new version file and the old version file to obtain the target file.
[0093] In this embodiment, a version upgrade request refers to a request sent by the target terminal to the server for a version upgrade. For example, this version upgrade request may be a request to upgrade the system version of the target terminal, or it may be a request to upgrade the version of a specific application configured on the target terminal.
[0094] In this embodiment, the new version file refers to the file stored on the server that represents the upgrade version required by the target terminal. For example, when a version upgrade request is used to instruct the target terminal to upgrade its system version from 5.2 to 5.3, the version file for version 5.3 stored on the server is the new version file.
[0095] In this embodiment, the old version file refers to the file stored on the server corresponding to the current version of the target terminal. For example, when the version upgrade request is a request to upgrade the system version of the target terminal, and the current system version of the target terminal is 5.2, then the version file of the target terminal's 5.2 version stored on the server is the old version file.
[0096] In this embodiment, after obtaining the new version file and the old version file, the old version file and the new version file can be differentially processed to generate a differential file, and the differential file is determined as the target file.
[0097] In the specific implementation, the steps of performing differential processing on the new version file and the old version file to obtain the target file may include: comparing the new version file and the old version file to obtain the newly added file and the modified file; performing a binary differential algorithm on the modified file to generate patch data; and generating the target file based on the newly added file and the patch data.
[0098] In this embodiment, by performing differential processing on the new version file and the old version file, all changes between the new and old versions can be accurately extracted.
[0099] In this embodiment, the target file may include multiple differential files, and different differential files may correspond to different file types. That is, for any differential file in the target file, multidimensional attribute information of the differential file can be obtained, and the file type of the differential file can be determined based on the multidimensional attribute information; based on the file type, the differential file is compressed to obtain a differential compressed file, and finally, based on the differential compressed files of each differential file, the compressed file of the target file is obtained.
[0100] In this embodiment, by performing differential processing on the new version file and the old version file in the OTA upgrade scenario, the amount of data to be processed can be greatly reduced. This allows the server to focus only on the changed files, avoiding the huge overhead of processing the entire new version file. This enables refined and intelligent processing of each differential file, minimizing the size of the compressed file and saving storage space and network transmission bandwidth.
[0101] To facilitate better implementation of the document processing method of this application, this application also provides a document processing apparatus based on the above-described document processing method. The meanings of the terms used are the same as in the document processing method described above, and specific implementation details can be found in the descriptions of the method embodiments.
[0102] Based on the same inventive concept, and referring to Figure 5 This application provides a document processing apparatus 500, which includes: The file acquisition module 501 is used to acquire the target file and its multidimensional attribute information. The type determination module 502 is used to determine the file type of the target file based on multi-dimensional attribute information; The file compression module 503 is used to compress target files based on file type to obtain compressed files.
[0103] In one embodiment, the type determination module 502 includes: The feature extraction submodule is used to extract features from multi-dimensional attribute information to obtain target feature information; The type determination submodule is used to input target feature information into the target processing model and output the file type through the target processing model.
[0104] In one embodiment, the multidimensional attribute information includes the file extension and magic number; the feature extraction submodule includes: The first feature extraction unit is used to extract features from the file extension to obtain the first feature information; The second feature extraction unit is used to extract features from the magic number to obtain second feature information; The feature fusion unit is used to perform weighted fusion of the first feature information and the second feature information to obtain the target feature information.
[0105] In one embodiment, the document processing apparatus 500 further includes: The evaluation information acquisition module is used to acquire evaluation information corresponding to the compressed file; the evaluation information includes one or more of the following: decompression success rate, decompression speed, and space saving information. The model update module is used to update the target processing model based on the evaluation information.
[0106] In one embodiment, the file compression module 503 includes: The compression strategy determination submodule is used to determine the target compression strategy based on the file type and the preset correspondence between file types and compression strategies; the target compression strategy includes the target compression algorithm and the target compression parameters. The compression processing submodule is used to compress target files based on the target compression strategy to obtain compressed files.
[0107] In one embodiment, the document processing apparatus 500 further includes: The file generation module is used to generate metadata files for compressed files based on the target compression strategy. The file integration module is used to integrate compressed files and metadata files to obtain an upgrade file package; The file sending module is used to send the upgrade file package to the target terminal, so that the target terminal can determine the target decompression algorithm based on the metadata file, and decompress the compressed file based on the target decompression algorithm to obtain the target file.
[0108] In one embodiment, the target file is obtained in the following manner: In response to a version upgrade request, retrieve both the new and old version files; The new version file and the old version file are differentially processed to obtain the target file.
[0109] By employing the technical solution of this application embodiment, and obtaining multi-dimensional attribute information of the target file, the file type of the target file can be accurately determined. Then, based on the file type, appropriate compression processing can be adaptively selected and executed. In this way, differentiated compression processing can be performed for different types of files, thereby significantly improving compression efficiency, effectively reducing the size of the final compressed file, and saving storage space and network transmission bandwidth.
[0110] Specific limitations regarding the file processing device 500 can be found in the limitations of the file processing method described above, and will not be repeated here. Each module in the aforementioned file processing device 500 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.
[0111] In addition, this application also provides an electronic device, such as Figure 6 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically: The electronic device may include components such as a processor 601 with one or more processing cores and a memory 602 with one or more computer-readable storage media. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 601 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the electronic device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.
[0112] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0113] In one feasible implementation, the electronic device further includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0114] In one feasible implementation, the electronic device may further include an input unit 604, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0115] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 601 runs the applications stored in the memory 602, thereby implementing the steps in any of the file processing methods provided in the embodiments of this application.
[0116] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0117] In one feasible implementation, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described in any embodiment of this application.
[0118] In one feasible implementation, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods described in any embodiment of this application.
[0119] In one feasible implementation, a computer program product is also proposed, comprising a computer program or instructions that, when executed by a processor, implement the methods described in any embodiment of this application.
[0120] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0121] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0122] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the file processing methods provided in this application.
[0123] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0124] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0125] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the file processing methods provided in this application, the beneficial effects that any of the file processing methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0126] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0127] The foregoing has provided a detailed description of a document processing method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A file processing method characterized by, The method comprises: acquiring a target file and multi-dimensional attribute information of the target file; determining a file type of the target file based on the multi-dimensional attribute information; performing compression processing on the target file based on the file type to obtain a compressed file.
2. The file processing method according to claim 1, characterized by, The determining of the file type of the target file based on the multi-dimensional attribute information comprises: performing feature extraction on the multi-dimensional attribute information to obtain target feature information; inputting the target feature information into a target processing model, and outputting the file type through the target processing model.
3. The file processing method according to claim 2, characterized by, The multi-dimensional attribute information comprises an extension name and a magic number. The performing of the feature extraction on the multi-dimensional attribute information to obtain the target feature information comprises: performing feature extraction on the extension name to obtain first feature information; performing feature extraction on the magic number to obtain second feature information; performing weighted fusion on the first feature information and the second feature information to obtain the target feature information.
4. The file processing method according to claim 2, characterized by, After the performing of the compression processing on the target file based on the file type to obtain the compressed file, the method further comprises: acquiring evaluation information corresponding to the compressed file; the evaluation information comprises one or more of a decompression success rate, a decompression speed and space saving information; updating the target processing model based on the evaluation information.
5. The file processing method according to claim 1, characterized by, The performing of the compression processing on the target file based on the file type to obtain the compressed file comprises: determining a target compression strategy based on a correspondence between the file type and a preset file type and compression strategy; the target compression strategy comprises a target compression algorithm and target compression parameters; performing compression processing on the target file based on the target compression strategy to obtain the compressed file.
6. The file processing method according to claim 5, characterized by, After the performing of the compression processing on the target file based on the target compression strategy to obtain the compressed file, the method further comprises: generating a metadata file of the compressed file based on the target compression strategy; performing integration processing on the compressed file and the metadata file to obtain an upgrade file package; sending the upgrade file package to a target terminal, so that the target terminal determines a target decompression algorithm based on the metadata file, and performs decompression on the compressed file based on the target decompression algorithm to obtain the target file.
7. The file processing method according to claim 1, characterized by, The target file is obtained by: in response to a version upgrade request, acquiring a new version file and an old version file; performing difference processing on the new version file and the old version file to obtain the target file.
8. A file processing apparatus characterized by comprising: The device comprises: a file acquisition module configured to acquire a target file and multi-dimensional attribute information of the target file; a type determination module configured to determine a file type of the target file based on the multi-dimensional attribute information; a file compression module configured to perform compression processing on the target file based on the file type to obtain a compressed file.
9. An electronic device, comprising: The device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the file processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps in the file processing method according to any one of claims 1 to 7.