Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "Small files" patented technology

IO optimization method and device based on distributed storage, equipment and storage medium

The invention discloses an IO optimization method and device based on distributed storage, equipment and a storage medium, and relates to the technical field of distributed storage. Comprising the following steps: constructing a distributed metadata index layer between a client and a metadata server, and maintaining a dynamic mapping relationship between a path prefix and the metadata server, so that when the client initiates a file access request, a target metadata server can be directly positioned based on the mapping relationship; and whether the request is routed to the temporary metadata server for load balancing or not is dynamically determined according to the real-time load information of the temporary metadata server. Therefore, the problems of metadata request amplification caused by step-by-step analysis of a tree directory and non-uniform load of a metadata server caused by static directory binding in related technologies are solved; the technical effects that the metadata interaction delay is remarkably reduced, the throughput of the system in massive small file and burst IO scenes is improved, and a single metadata server is prevented from becoming a performance bottleneck are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Multi-system ai repository controller

PendingUS20260064396A1Program controlTransmissionDocument modelEngineering
Systems and methods are provided for creating a machine learning (ML) model staging repository, where ML models are stored as an open container image (OCI). The OCI comprises layers or portions of the complete ML model, so that the model can be stored separately and as a smaller files. The ML models may be pre-packaged for automated downloads and integration at the customer site. In some examples, the OCI can identify / store the model in a directory structure that defines the model and its profile. The OCI can comprise a combination of layers and profiles that allows the AI platform to optimize the storage and simplify the process of downloading or updating the given user namespace instead of downloading a single large file for the model.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Real-time storage optimization method based on block data

The invention discloses a real-time storage optimization method based on block data, which relates to the technical field of big data processing and distributed storage, introduces a block data concept, realizes block-by-block writing, in-block ordering, block-level ordering and an intelligent flush trigger mechanism, and realizes real-time writing with high throughput, low delay and low memory occupation. The method specifically comprises the steps that block data is introduced into a data memory model to serve as a basic processing unit, multiple rows of data are aggregated into a block according to an aggregation key CKValue, row data in the block are sequenced according to a primary key PK, and write-in, namely aggregation is achieved; writing data, aggregating according to blocks, and uniformly sorting in the blocks; one block, namely multiple rows of data, generates an FRC file, and small files are reduced; exception processing: supporting block-level recovery to skip a failed block; the core problems of too high memory occupation, large CPU consumption, high writing delay and the like in a high-concurrency and high-throughput real-time writing scene can be solved, and the method is widely applied to scenes with extremely high requirements on data real-time performance and system stability.
Owner:XIAN FIBERHOME SOFTWARE TECH CO LTD

Multi-system AI repository controller

The embodiment of the invention relates to a multi-system AI repository controller. Systems and methods are provided for creating a hierarchical repository of machine learning (ML) models in which ML models are stored as open container images (OCIs). The OCI includes a layer or portion of a complete ML model so that the model can be stored separately and as a smaller file. The ML model may be pre-packed for automated downloading and integration at the customer site. In some examples, the OCI may identify / store the model in a directory structure defining the model and its configuration files. The OCI may include a combination of layers and configuration files that allow the AI platform to optimize storage and simplify the process of downloading or updating a given user namespace rather than downloading a single large file for the model.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Method and apparatus for simulating file system aging

A method for simulating file system aging is disclosed, including: determining a storage capacity of a first device and a type of a file system; determining a ratio between a plurality of small files having different sizes based on the storage capacity of the first device and the type of the file system; copying the plurality of small files having different sizes to the first device according to the ratio; and deleting a part of the small files among all the small files copied onto the first device according to a predetermined interval pattern.
Owner:YANGTZE MEMORY TECH CO LTD

Cache optimization method and device based on distributed storage

The invention discloses a cache optimization method and device based on distributed storage, and relates to the technical field of distributed storage systems. The method comprises the following steps: dynamically calculating a weight factor of a cache file, attenuating the weight factor based on the access frequency and timeliness of the file, and compensating and adjusting according to the size of the file; the maximum file threshold value set by the system is dynamically adjusted according to the size distribution of the files in the cache, and the maximum file threshold value is determined based on a preset quantile range so as to eliminate interference of extremely large files on weight calculation; when the weight factor of the cache file is lower than a preset threshold value, the file is removed from the cache, and the cache space is released; when the cache capacity is insufficient, the file with the minimum weight is selected for replacement based on the weight factor of the current cache file, and the high-value file is preferentially reserved in the cache. According to the method, the cache hit rate of small files in the distributed storage system is effectively increased, cache space waste and metadata operation overhead are reduced, and system performance and resource utilization rate are remarkably improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method and device for fusing, storing and managing massive small files and data records

The invention discloses a method and device for fusing, storing and managing massive small files and data records, and the method comprises the steps: building an external metadatabase which is used for managing a hierarchical relation of scientific observation data sources and a mapping relation between an aggregation file and an original file of scientific observation data; according to a preset aggregation rule based on data product categories and timelines, the multiple original files are aggregated and packaged in a target file container in a column storage format, and the content of the original files and structured data records obtained by analyzing the original files are stored in the target file container at the same time; and based on the external metadatabase and the target file container, providing file-level access aiming at the original file content and record-level retrieval service aiming at the structured data record. According to the method, the aspects of system simplification, performance optimization, space saving and service capability are remarkably improved, and a new way with high expandability is provided for efficient storage and multi-level service of space weather scientific data.
Owner:NAT SPACE SCI CENT CAS

A machine learning-based parallel optimization method and system for genomic analysis

The application discloses a kind of based on machine learning's genome analysis parallel optimization method and system, the method of the present application includes the input BAM file is cut into multiple same size file blocks;For each file block, extract file block characteristics, and input the file block characteristics into the pre-trained machine learning model to obtain the predicted running time of the file block;The predicted running time is in the front part file block is divided into smaller file block;File block is input into HaplotypeCaller in parallel to carry out variation detection;The variation detection result generated by HaplotypeCaller for each file block is merged and output.The present application is aimed at solving the problem of serious calculation tilt and low resource utilization caused by unclear calculation complexity before HaplotypeCaller runs, improving the efficiency of genome analysis, shortening the variation detection duration.
Owner:SUN YAT SEN UNIV

A voiceprint registration method, device, equipment and storage medium

The application provides a voiceprint registration method, device, equipment and storage medium, the method comprises the following steps: storing the audio to be registered according to the time length of the audio to be registered when obtaining one audio to be registered every time; acquiring the first type of audio from the stored several audios based on a plurality of small file consumption service threads, and performing voiceprint registration on the first type of audio based on the operation equipment; and acquiring the second type of audio from the stored several audios based on a single large file consumption service thread, and performing voiceprint registration on the second type of audio based on the operation equipment; wherein the audio time length of the second type of audio is greater than the audio time length of the first type of audio. The voiceprint registration method provided by the application can perform voiceprint registration on short audios in a multi-thread concurrent manner (concentrate processing short audios), and perform voiceprint registration on long audios in a single thread, so that the resource utilization rate of the operation equipment can be effectively improved, the efficiency of voiceprint registration can be effectively improved, and the hardware cost can be greatly saved.
Owner:ANHUI IFLYTEK INTELLIGENT SYST

Data processing method and system, storage medium and program product

The invention provides a data processing method and system, a storage medium and a program product. The method comprises the steps of obtaining a to-be-uploaded target file; for a first file of which the file size is greater than a file size threshold value in the target file, uploading the first file to a search resource server by adopting a fragment uploading mode; and for the second file of which the file size is smaller than or equal to the file size threshold value in the target file, directly uploading the second file to the search resource server, and selecting a better uploading path for the files of different sizes through simple threshold value judgment. The small file avoids the overhead of multiple rounds of handshake and multiple requests in fragment uploading, and the small file is quickly uploaded directly through one request, so that the delay caused by network round-trip time is greatly reduced.
Owner:UC MOBILE CHINA CO LTD

Generation method and analysis method for data file of question bank system

The invention discloses a generation method and an analysis method for a data file of a question bank system. The method comprises the steps of generating a file overall structure; determining the position of a file header structure based on the overall file structure, and generating the file header structure based on the position of the file header structure; determining the position of a table index area based on the file header structure, generating a table index area structure based on the position of the table index area, determining the position of each metadata table in the table data area based on the table index area structure, and generating the table data area based on the position of each metadata table; determining the position of a file index area based on the file header structure, generating a file index area structure based on the position of the file index area, determining the position of each file resource in the file data area based on the file index area structure, and generating the file data area based on the position of each file resource. The technical problems that in the prior art, when a large number of small files are frequently and randomly accessed, decompression efficiency is low, memory occupation is high, and access delay is large are solved.
Owner:SICHUAN BISHENG INTELLIGENT TECHNOLOGY CO LTD

Methods, devices, storage media, and electronic devices for storing small files

This application provides a method, apparatus, storage medium, and electronic device for storing small files. The method includes: obtaining a first small file identifier, wherein the first small file identifier is used to identify a first small file to be deleted; then deleting the first small file according to the first small file identifier and updating target index information; wherein the first small file is a small file packaged and stored in a first packaged file, the first packaged file includes multiple small files arranged in a preset order, the target index information includes the data length of the first small file; a second small file is a small file packaged and stored in the first packaged file and located after the first small file in a preset order; and then moving the second small file according to the updated target index information, wherein the distance moved by the second small file is the same as its data length. This application solves the problem of low efficiency in deleting small files in packaged files in related technologies.
Owner:ZHEJIANG DAHUA TECH CO LTD

Cloud database computing cache system and implementation method

The invention relates to a cloud database computing cache system and an implementation method, and the system comprises a decentralized cache architecture module which is used for eliminating a centralized bottleneck, supporting ten-thousand-level concurrency and adapting to a big and small file mixed scene; the memory object sharing module is used for reducing intermediate result processing delay through intermediate result full-life-cycle zero-copy circulation and dynamic storage decision based on a memory pool and metadata service of the memory object sharing module; the self-adaptive asynchronous write-in module is used for connecting remote storage, multiplexing memory object sharing module resources and combining mechanisms such as delay confirmation, so that the write response speed is increased, remote requests are reduced, data integrity is guaranteed, and meanwhile, persistent support is provided for intermediate results. Through cooperative operation of the three modules, the cloud database computing cache system and the implementation method can effectively adapt to the requirements of high concurrency, large and small file mixing and elastic capacity expansion and contraction of a cloud native OLAP scene, and the integrity of data transmission and the stability of system operation are guaranteed.
Owner:HASHDATA LTD

Request dispatch processing method and request dispatch processing system

PendingCN122363624ADocument handlingEngineering
This specification provides a request distribution processing method and a request distribution processing system. The request distribution processing method includes: obtaining file processing requests for target files; determining a target request processing queue set from multiple candidate request processing queue sets based on the request type of the file processing request, wherein the multiple candidate request processing queue sets correspond to different request types; determining a target request processing queue from the target request processing queue set based on the file identifier of the target file, and adding the file processing request to the target request processing queue; and invoking the request distribution resource corresponding to the target request processing queue to distribute and process the target requests in the target request processing queue. This eliminates the risk of read requests being starved due to resource contention, ensuring low-latency response for data loading; and ensures that processing requests for massive numbers of small files can obtain a fair scheduling opportunity, thereby preventing small files from being starved by large files.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Storage backup optimization method for long-term preservation of large-scale digital literature

The application provides a storage backup optimization method for long-term preservation of large-scale digital literature, relates to the technical field of data processing, and constructs an AIP storage path and a SUPPORT storage path by calculating the SHA512 hash value of a unique identifier of a digital object to be archived; encapsulates versioned core data files into an AIP directory according to the AIP storage path; stores and manages auxiliary data files into a SUPPORT directory according to the SUPPORT storage path, generates a disk slave copy, performs parallel integrity auditing on the disk master copy and the disk slave copy, performs consistency auditing and repair on the disk master copy and the disk slave copy, and outputs a repair completion state. The technical problems of low storage efficiency and long full-link backup period caused by massive small files in the prior art are solved. The technical effect of improving the storage efficiency and backup performance of massive small files is achieved.
Owner:DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Atomized multi-resource data transmission optimization method

The invention provides an atomization multi-resource data transmission optimization method. Metadata and a dependency relationship of the atomized resources are obtained by scanning project source codes, a resource dependency graph is constructed, and logic association and a hierarchical structure among the resources are accurately described; intelligent clustering and grouping are carried out based on the atlas in combination with resource use frequencies, and high-frequency co-occurrence and functional coupling resources are divided into the same resource group, so that the subjectivity and redundancy problems are avoided; the resource group is packaged into the set package file, and the list file is generated, so that aggregation packaging of the atomized resources is realized, and the integrity of resource delivery is ensured; in the application stage, client requests are intercepted, and corresponding set packages are loaded based on list file matching, so that the actual network request times are reduced, and the head overhead and the head-of-queue blocking risk are reduced. Therefore, not only is the performance bottleneck caused by a large number of small file requests solved, but also the on-demand loading and caching efficiency of the resources is guaranteed through a dynamic grouping mechanism, and the flexibility and execution efficiency of front-end resource transmission are fundamentally improved.
Owner:创优数字科技(广东)有限公司

Financial data merging method and device, equipment, storage medium and program product

The invention discloses a financial data merging method and device, equipment, a storage medium and a program product, and relates to the technical field of big data storage. In the application, a combined service cluster and a virtual node set are created, and the number of virtual nodes in the virtual node set corresponding to each combined service node in the combined service cluster is configured based on the load capacity index of each combined service node in the combined service cluster. Therefore, the mapping relation between the virtual node set and the merging service nodes in the merging service cluster is constructed, the merging service cluster and the virtual node set can be matched based on the mapping relation, and reasonable and balanced distribution of financial data small file merging tasks in a cluster environment is achieved. And furthermore, the merging task of the financial data small files is processed through the distributed target merging service node, so that high-performance and load-balanced financial data small file merging is realized.
Owner:WEBANK (CHINA)

A large PDF file intelligent generation method with front-end and back-end cooperation

The application discloses a front-rear end cooperative large PDF file intelligent generation method, relates to the technical field of computer software, and comprises the following steps: front-end rendering: the front end renders the work order data page by combining a headless browser with a Vue framework and an ElementUI component library, and dynamically displays structured work order details through a responsive data binding mechanism; data fragmentation: large work order data sets are split into N subtasks based on a dynamic fragmentation strategy; and small PDF files are generated in parallel: a containerized rendering cluster based on Kubernetes deployment management is used to allocate rendering tasks to each container; the small PDF files generated by the application through careful design with the help of the Vue and ElementUI frameworks are more attractive in terms of text layout, color matching and graphic display, and the large file after merging is more beautiful than the file generated by a traditional back end, thereby improving user experience.
Owner:BEIJING ZHIXING TONGDE INFORMATION TECH CO LTD

Big model based cold and hot data storage optimization method

The present application relates to the field of hot and cold data storage, and particularly to a hot and cold data storage optimization method based on a large model, comprising a business integration module, a data dependency module, a mode updating module, a hot and cold partition module and a file merging module, the business integration module is used for collecting business data to form a training data set, the data dependency module is used for fitting business data to obtain a data blood relationship model, the mode updating module is used for predicting an access mode, the hot and cold partition module is used for calculating the heat of data and performing dynamic scheduling, and the file merging module is used for automatically merging small files, the present application can reduce repeated storage, ensure data source uniqueness, standardize data storage structure, reduce storage space, can realize more efficient and more accurate hot and cold data classification, improve resource utilization of data storage hardware, reduce hot data response delay, significantly improve system storage utilization, and improve storage efficiency.
Owner:北京科杰科技有限公司

A spark-hbase batch data rapid warehousing implementation method and system based on parameter configuration

ActiveCN115687346BImprove read and write performanceReduce business calculation timeDatabase management systemsDatabase distribution/replicationTable (database)Business data
This invention discloses a method and system for rapid batch data import into HBase based on parameter configuration, relating to the field of database technology. The method includes: configuring a data import and merging strategy for HBase tables in scenarios where batch data calculated by the Spark engine is imported into an HBase database; performing calculations on business data using the Spark engine, writing the generated dataset to HDFS and recording it in an information table; continuously adding HDFS file directories according to the file generation time order of the information table, and submitting an MapReduce task to generate an HFile file when the requirements of the data import and merging strategy are met; and importing the HFile file into the HBase database. This invention avoids the problem of too many small files in HBase, improves the read and write performance of the HBase database, and reduces the time spent on business calculations.
Owner:SI-TECH INFORMATION TECH CO LTD

Small file merging and cache rapid reconstruction method and system

The invention relates to a small file merging and caching rapid reconstruction method, which is characterized by comprising the following steps: screening small files smaller than a first preset value based on a to-be-screened file library; the screened small files are combined into at least one cache block according to the fixed size of a second preset value, and the small files which are accessed in the same table or frequently shared are distributed to the same cache block; reconstructing the content of the small files in the at least one cache block according to a'data area + metadata area 'structure to generate at least one combined block; respectively writing the at least one combined block into a plurality of original nodes in a cluster; and in response to a cache synchronization request sent by the new node to the original node when the cluster node is expanded, the original node transmits data by taking the combined block as a transmission unit. According to the scheme, the I / O and reading performance is excellent, the disk I / O utilization rate is reduced to 20%-40%, and the small file reading efficiency is improved to 400-520 MB / s. The bank scene processing time is shortened, and the cache reconstruction time of 10 million small files is sharply reduced. After cluster expansion, query delay is increased by 2%, remote storage access is reduced by 90%, and bandwidth occupation is reduced by 70%.
Owner:HASHDATA LTD

IO optimization method, device and equipment based on distributed storage and storage medium

The application discloses an IO optimization method and device based on distributed storage, equipment and a storage medium, and relates to the technical field of distributed storage. The method comprises the following steps: a distributed metadata index layer is constructed between a client and a metadata server, and a dynamic mapping relationship between a path prefix and the metadata server is maintained, so that when the client initiates a file access request, the target metadata server can be directly located based on the mapping relationship, and whether the request is routed to a temporary metadata server for load balancing is dynamically determined according to real-time load information of the target metadata server. Therefore, the problems of metadata request amplification caused by tree directory level-by-level analysis and uneven load of metadata servers caused by static directory binding in the related art are solved, and the technical effects of significantly reducing metadata interaction delay, improving the throughput of the system in the massive small file and burst IO scene, and avoiding a single metadata server from becoming a performance bottleneck are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

File storage method and device of distributed file system, electronic equipment and medium

The invention provides a file storage method and device of a distributed file system, electronic equipment and a medium. According to the method and the device, the FID SEQ range is distributed to the client in advance through the main MDS, so that the client can automatically generate the FID for the small file locally, the situation that the FID is applied to the MDS when the small file is written in each time in the related technology is avoided, and the load pressure and the network communication overhead of the MDS are remarkably reduced. The method comprises the following steps: locally creating an aggregation memory pool of which each memory subspace is bound with a corresponding aggregation large file on the MDS on a client, and creating a memory subspace is bound with a mapping relation memory pool of a corresponding mapping relation large file on the MDS, so that a plurality of small files are temporarily stored in the local aggregation memory pool, and are written into the aggregation large file in batches; a large number of small files are changed into a small number of large files, the small files are aggregated and stored in the aggregated large file, the problem that OSS storage space is seriously fragmented due to a striping strategy of the large number of small files is fundamentally avoided, and the utilization rate of the storage space is greatly improved.
Owner:MACROSAN TECH

An optimization method and terminal for file merging

PendingCN122309469AShardParallel computing
This invention provides an optimized method and terminal for file merging. It includes: acquiring small files to be merged, where the file size is below a first preset threshold; sorting the small files based on their data writing order to obtain an ordered file sequence; allocating the small files in the ordered file sequence to multiple merging groups according to a second preset threshold; and for each target small file in a merging group, reading and merging the target small files based on their arrangement order in the ordered file sequence to generate a merged file. This invention achieves sequential reading and zero-fragment merging by sorting according to the writing order and flexibly grouping the entire file, significantly improving reading efficiency, reducing system overhead, and solving the problems of low efficiency and excessive fragmentation in traditional methods.
Owner:FUJIAN TQ ONLINE INTERACTIVE INC

File storage method, device and equipment and readable storage medium

The application provides a file storage method, device and equipment and a readable storage medium. The method provides a dynamic storage optimization mode based on directory granularity and file size perception. The method takes a directory as a unit, analyzes the file size of each incremental file under the directory within a specified time window, extracts a quantitative index reflecting the change characteristics of large / small files to automatically determine the scene type corresponding to each incremental file under each directory, and dynamically matches and applies a target optimization strategy adapted to the scene type corresponding to each directory when the scene type corresponding to each directory to be scanned meets the preset directory-level optimization requirement, thereby realizing fine-grained automatic storage optimization at the directory level and improving the overall performance and storage efficiency of the system.
Owner:MACROSAN TECH

File transmission and storage method and device, computer device, readable storage medium and program product

The application relates to a file transmission and storage method and device, computer equipment, a readable storage medium and a program product. The method comprises the following steps: a client determines at least one of the type and size of a to-be-uploaded file by responding to a file uploading instruction of a target user, determines an uploading mode of the to-be-uploaded file based on an uploading strategy of the target user and at least one of the type and size of the to-be-uploaded file, and then stores the to-be-uploaded file to a server based on the uploading mode. The efficiency of uploading small files to a cloud disk can be improved, the pressure on underlying storage can be reduced, and the efficiency of uploading files to the cloud disk and the pressure on underlying storage can be balanced.
Owner:CHINA TELECOM CLOUD TECH CO LTD

Backup method and apparatus

This application provides a backup method and an apparatus. The method includes: determining a first application that needs data backup; packaging and compressing a file whose size is less than a first threshold in a user space corresponding to the first application, to obtain a first compressed package, where the first compressed package is stored in the user space corresponding to the first application; and backing up the first compressed package from the user space corresponding to the first application to a user space corresponding to a second application. The method can reduce overheads caused by an I / O operation generated by backing up a large quantity of small-sized files between two user spaces, reduce a backup delay, and improve user experience.
Owner:HUAWEI TECH CO LTD

File deletion method and electronic device

This application discloses a file deletion method and electronic device, relating to the field of distributed file system technology. The method categorizes files to be deleted into large and small files based on a preset first volume threshold. Different deletion processes are applied to files of different sizes, allowing for simultaneous deletion of large and small files, thus improving file deletion efficiency. Specifically, for small files, data block information of at least two small files is recorded in a storage object, and the physical storage space corresponding to the data block information of at least two small files in the storage object is released in batches. This avoids frequent interactions with the storage device, improving the efficiency of physical storage space release, and further enhancing the efficiency of small file deletion.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Systems and methods for concurrent metadata and data processing in interactive data ingestion

A data ingestion system processes data files through a dual-path architecture to enable rapid interactive data analysis. The system routes incoming files below a size threshold to a fast conversion path and larger files to a batch processing path. For files in the fast path, the system concurrently processes metadata and data instead of following traditional sequential processing. A metadata controller assigns storage locations and manages table definitions while a direct format converter transforms source files into query-ready columnar format. A query processor provides unified access to converted data across both processing paths. The system reduces processing latency by eliminating batch processing overhead for small files, enables immediate data querying through coordinated storage management, and maintains data consistency through stateful job tracking. This architecture enables rapid processing for interactive analysis while preserving robust batch processing capabilities for larger datasets.
Owner:SALESFORCE INC

System for transforming data into table format

Systems and methods are directed to transforming data into a distributed table format. The system accesses new data for a topic. One or more materializers generate data files from the new data, whereby each of the materializers can process a specific range of offsets within a partition of the topic. The materializers also generate materializer committables corresponding to the data files. A maintainer performs one or more maintenance operations on previously committed files and generates a maintainer committable for each maintenance operation performed. The maintenance operations can include merging previously committed smaller files into a larger file and deleting obsolete files. Subsequently, a committer scheduler schedules the maintainer and materializer committables for a commit process. Scheduling includes prioritizing which committables are applied first and ordering materializer committables based on their offsets. The commit process is performed based on the scheduling resulting in a single snapshot of the table.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION