Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

161 results about "Small files" patented technology

Distributed high-concurrency aggregation storage method and system for small files

The invention discloses a small file distributed high-concurrency aggregation storage method and system, and the method comprises the steps: collecting performance index data, and transmitting the data to a resource coordination server; calculating a queuing rate; determining a target server according to the queuing rate, the file reading concurrency and the disk residual space; determining a first-level directory hash value; selecting a first-level storage directory and a second-level storage directory according to the first-level directory hash value; creating a target aggregation file in the secondary storage directory, generating a 2-byte file format field according to the compression identifier or the non-compression identifier, and generating a 4-byte file length field according to the size of the to-be-written file; writing the file format field, the file length field and the to-be-written file into a target aggregation file in sequence, and recording the initial offset of writing; and generating a file name containing the storage position information, and returning the file name to the client. According to the invention, the efficiency and effect of small file storage are optimized.
Owner:SHENZHEN EFERCRO ELECTRONIC TECHNOLOGY CO LTD

Cold and hot data storage optimization method based on large model

The invention relates to the field of cold and hot data storage, in particular to a cold and hot data storage optimization method based on a large model, which comprises a business integration module, a data dependence module, a mode updating module, a cold and hot partitioning module and a file merging module, and is characterized in that the business integration module is used for collecting business data to form a training data set; the data dependency module is used for fitting business data to obtain a data consanguinity model, the mode updating module is used for predicting an access mode, the cold and hot partitioning module is used for calculating the heat of the data and performing dynamic scheduling, and the file merging module is used for automatically merging small files. A data storage structure is standardized, storage space is reduced, more efficient and more accurate cold and hot data classification can be achieved, the resource utilization rate of data storage hardware is improved, hot data response delay is reduced, the system storage utilization rate is remarkably improved, and the storage efficiency is improved.
Owner:北京科杰科技有限公司

Mass small file reading optimization method, system, equipment and medium

The invention relates to the technical field of big data, and discloses a massive small file reading optimization method, system and device and a medium, and the method comprises the following steps: scanning small files in a distributed file system by using a Spark calculation engine to obtain metadata information of the small files; based on the metadata information, small files are classified, and the small files with the same type and data relevance are classified into one group; the resource states of cluster nodes are monitored in real time, the resource states comprise the processor utilization rate, the memory occupancy rate and the I / O load, the load score of each node is calculated, and small file task quotas are dynamically allocated; executing small file merging operation in parallel by utilizing a Spark memory calculation engine to generate a merged file containing a plurality of original small files; creating and maintaining index information in an external database, and recording the position and attribute of each original small file in the combined file; and providing a read-write interface through the special service layer, and positioning and accessing specific small file contents in the combined file.
Owner:GUANGDONG POLYTECHNIC NORMAL UNIV

Batch writing method and device of data and electronic equipment

The invention discloses a batch writing method and device of data and electronic equipment, the method is applied to the field of data storage, and the method comprises the following steps: receiving a plurality of writing requests; determining whether the data object belongs to a first type object or a second type object according to the storage space occupied by the data object in each write-in request; adding and writing at least one first type of object in the plurality of writing requests into a first data block in batches, and storing metadata information of the first type of object in the first data block; and slicing the second type of objects in the plurality of write requests, writing the second type of objects into a second data block, and storing metadata information of the second type of objects according to a mapping relationship between the sliced second type of objects and the second data block. Through the data writing method and device, the problems that in the process of processing the data writing request in the related technology, data of small files needs to be asynchronously gathered and deleted after being written, and large files occupy a large storage space due to the reference relation, so that the data writing efficiency is low are solved.
Owner:BEIJING XSKY TECH CO LTD

IO optimization method and device based on distributed storage, equipment and storage medium

The invention discloses an IO optimization method and device based on distributed storage, equipment and a storage medium, and relates to the technical field of distributed storage. Comprising the following steps: constructing a distributed metadata index layer between a client and a metadata server, and maintaining a dynamic mapping relationship between a path prefix and the metadata server, so that when the client initiates a file access request, a target metadata server can be directly positioned based on the mapping relationship; and whether the request is routed to the temporary metadata server for load balancing or not is dynamically determined according to the real-time load information of the temporary metadata server. Therefore, the problems of metadata request amplification caused by step-by-step analysis of a tree directory and non-uniform load of a metadata server caused by static directory binding in related technologies are solved; the technical effects that the metadata interaction delay is remarkably reduced, the throughput of the system in massive small file and burst IO scenes is improved, and a single metadata server is prevented from becoming a performance bottleneck are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Multi-system ai repository controller

Systems and methods are provided for creating a machine learning (ML) model staging repository, where ML models are stored as an open container image (OCI). The OCI comprises layers or portions of the complete ML model, so that the model can be stored separately and as a smaller files. The ML models may be pre-packaged for automated downloads and integration at the customer site. In some examples, the OCI can identify / store the model in a directory structure that defines the model and its profile. The OCI can comprise a combination of layers and profiles that allows the AI platform to optimize the storage and simplify the process of downloading or updating the given user namespace instead of downloading a single large file for the model.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Json file parallel analysis and data writing method and device based on DataX

The invention relates to a Json file parallel analysis and data writing method and device based on DataX. The method comprises the following steps: acquiring a source end Json file set, and dividing the source end Json file set into small file subsets and large file subsets according to file sizes; the small file aggregation path is configured with a json reader of the DataX, and parallel reading and analysis are carried out; splitting a root array of the large file into a plurality of independent small Json files through a Python script, aggregating paths, and reading and analyzing by using a DataX parallel channel; and finally executing the DataX Job to write the data into a target end. Through a classification processing strategy, the large-scale Json data analysis writing efficiency is remarkably improved, the large file memory occupation risk is reduced, the file scale is automatically adapted, a mature DataX frame and a Python script are reused, high performance and low maintenance cost are both considered, multi-type target end storage is supported, and the problems that a traditional tool is weak in parallel, high in resource consumption, poor in flexibility and the like are effectively solved.
Owner:CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD

File access method, device and equipment and computer readable storage medium

The invention discloses a file access method, device and equipment and a computer readable storage medium, and belongs to the technical field of storage. The method is applied to a first node in a file system and comprises the steps that a first write request for a file is received, the first write request comprises first file data and IO information, and the IO information represents an address range for writing the first file data; in response to the first write request, metadata of the file is obtained, and the metadata comprises an extension attribute and a file size; according to the IO information and the file size, under the condition that it is determined that the written file size does not exceed the reference threshold value, the first file data is written into the extension attribute, and the written file size is the file size after the first file data is written. The file data is embedded in the metadata of the file, and the file data can be accessed while the metadata of the small file is accessed, so that the step of accessing the file data is reduced, the response time delay of the file system to the access request is reduced, and the purpose of improving the access performance of the system is achieved.
Owner:SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Method for aggregating, archiving and storing large-scale small files and adding, reading and modifying data

The invention relates to a large-scale small file aggregation archiving storage and data addition, read-write and modification method, and the method comprises the steps: firstly obtaining a large-scale small file to be archived stored in a storage pool; and carrying out aggregation, filing and storage on the to-be-filed small files one by one based on the file aggregation structure capable of modifying the additional data. A file aggregation structure of modifiable appended data includes a metadata header field, an index field, a metadata block field, and a data block field. According to the large-scale small file aggregation archiving storage and data appending, read-write and modification method, a new file aggregation structure capable of modifying appended data is adopted to perform ordered aggregation archiving storage on the large-scale small files, and the structure allows direct modification and write-in operation on the small files in the aggregation file; and the whole file set does not need to be aggregated and archived again.
Owner:SHAANXI HONGJU NETWORK INFORMATION TECH CO LTD

IO merging mechanism based on block index

The invention discloses an IO merging mechanism based on a block index, relates to the technical field of big data and information retrieval, and optimizes an original multi-file stream output mechanism of open source Lucene into single-file stream output by expanding a Directory interface of the Lucene, so that IO operation merging is realized in an index writing process, and the operation efficiency is improved. The method specifically comprises three parts of a file storage structure, a writing process and a reading process. The file storage structure comprises a single file structure of a file header, a data area and a file tail; during writing, data is cached in the memory block firstly, and after the memory block is full, the data is added to a file by taking a data block as a unit; during reading, quick positioning is realized by analyzing the block index at the end of the file; according to the method, the problems of performance bottleneck, small file flooding, difficulty in atomicity guarantee and the like caused by multi-file IO in the prior art are solved, and the index construction efficiency and the system stability are remarkably improved.
Owner:XIAN FIBERHOME SOFTWARE TECH CO LTD

Small file quick positioning and processing method and device of HDFS (Hadoop Distributed File System), electronic equipment and storage medium

The invention discloses a small file rapid positioning and processing method and device for an HDFS distributed file system, electronic equipment and a storage medium, and the method comprises the steps: when it is monitored that a new fsimage file is generated, analyzing the new fsimage file into metadata, and sending the metadata to a database through a message queue, the metadata comprises metadata information about files and metadata information about folders; monitoring the metadata information about the file and the metadata information about the folder in the database, and positioning to the position of a small file; and according to the position of the small file, respectively executing corresponding treatment means on the data file and the metadata file. According to the method and the device, a'closed-loop 'strategy for quickly finding and managing the small files in the HDFS is realized, the stability of cluster operation can be improved, meanwhile, full-automatic execution can be realized, and the efficiency is greatly improved.
Owner:中国邮政储蓄银行股份有限公司

Real-time storage optimization method based on block data

The invention discloses a real-time storage optimization method based on block data, which relates to the technical field of big data processing and distributed storage, introduces a block data concept, realizes block-by-block writing, in-block ordering, block-level ordering and an intelligent flush trigger mechanism, and realizes real-time writing with high throughput, low delay and low memory occupation. The method specifically comprises the steps that block data is introduced into a data memory model to serve as a basic processing unit, multiple rows of data are aggregated into a block according to an aggregation key CKValue, row data in the block are sequenced according to a primary key PK, and write-in, namely aggregation is achieved; writing data, aggregating according to blocks, and uniformly sorting in the blocks; one block, namely multiple rows of data, generates an FRC file, and small files are reduced; exception processing: supporting block-level recovery to skip a failed block; the core problems of too high memory occupation, large CPU consumption, high writing delay and the like in a high-concurrency and high-throughput real-time writing scene can be solved, and the method is widely applied to scenes with extremely high requirements on data real-time performance and system stability.
Owner:XIAN FIBERHOME SOFTWARE TECH CO LTD

Multi-system AI repository controller

The embodiment of the invention relates to a multi-system AI repository controller. Systems and methods are provided for creating a hierarchical repository of machine learning (ML) models in which ML models are stored as open container images (OCIs). The OCI includes a layer or portion of a complete ML model so that the model can be stored separately and as a smaller file. The ML model may be pre-packed for automated downloading and integration at the customer site. In some examples, the OCI may identify / store the model in a directory structure defining the model and its configuration files. The OCI may include a combination of layers and configuration files that allow the AI platform to optimize storage and simplify the process of downloading or updating a given user namespace rather than downloading a single large file for the model.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Adaptive content delivery systems and methods

Adaptive content delivery systems and methods are disclosed herein. An example system is configured to determine capabilities of a requesting end user system, interface with a database designed to store images characterized by criteria, select an image from the database that balances a highest quality and a smallest file size for the end user system, based on the capabilities, and provide the image to a network for delivery to the end user system.
Owner:SPEEDSIZE LTD

Method and apparatus for simulating file system aging

A method for simulating file system aging is disclosed, including: determining a storage capacity of a first device and a type of a file system; determining a ratio between a plurality of small files having different sizes based on the storage capacity of the first device and the type of the file system; copying the plurality of small files having different sizes to the first device according to the ratio; and deleting a part of the small files among all the small files copied onto the first device according to a predetermined interval pattern.
Owner:YANGTZE MEMORY TECH CO LTD

Cache optimization method and device based on distributed storage

The invention discloses a cache optimization method and device based on distributed storage, and relates to the technical field of distributed storage systems. The method comprises the following steps: dynamically calculating a weight factor of a cache file, attenuating the weight factor based on the access frequency and timeliness of the file, and compensating and adjusting according to the size of the file; the maximum file threshold value set by the system is dynamically adjusted according to the size distribution of the files in the cache, and the maximum file threshold value is determined based on a preset quantile range so as to eliminate interference of extremely large files on weight calculation; when the weight factor of the cache file is lower than a preset threshold value, the file is removed from the cache, and the cache space is released; when the cache capacity is insufficient, the file with the minimum weight is selected for replacement based on the weight factor of the current cache file, and the high-value file is preferentially reserved in the cache. According to the method, the cache hit rate of small files in the distributed storage system is effectively increased, cache space waste and metadata operation overhead are reduced, and system performance and resource utilization rate are remarkably improved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method and device for fusing, storing and managing massive small files and data records

The invention discloses a method and device for fusing, storing and managing massive small files and data records, and the method comprises the steps: building an external metadatabase which is used for managing a hierarchical relation of scientific observation data sources and a mapping relation between an aggregation file and an original file of scientific observation data; according to a preset aggregation rule based on data product categories and timelines, the multiple original files are aggregated and packaged in a target file container in a column storage format, and the content of the original files and structured data records obtained by analyzing the original files are stored in the target file container at the same time; and based on the external metadatabase and the target file container, providing file-level access aiming at the original file content and record-level retrieval service aiming at the structured data record. According to the method, the aspects of system simplification, performance optimization, space saving and service capability are remarkably improved, and a new way with high expandability is provided for efficient storage and multi-level service of space weather scientific data.
Owner:NAT SPACE SCI CENT CAS

A small file storage method, device, equipment, medium and program product

The application provides a small file storage method, device, equipment, medium and program product. The method comprises the following steps: performing hash calculation on the storage path of a small file, determining a target large file bucket to which the small file is to be stored based on the obtained hash value, and the target large file bucket corresponding to a target metadata bucket and a target file name bucket; when the sum of the data amount of the small file and the data amount of the target large file bucket exceeds a preset bucket capacity threshold, generating a new large file bucket, a new metadata bucket and a new file name bucket under the same directory of the target large file bucket, and migrating the stored part of the small file, metadata and file name in the target large file bucket, the target metadata bucket and the target file name bucket to the new large file bucket, the new metadata bucket and the new file name bucket, so as to store the small file, the metadata thereof and the file name of the small file into the target large file bucket, the target metadata bucket and the target file name bucket respectively. The application can significantly improve the storage performance and reading efficiency of the small file.
Owner:MACROSAN TECH

A machine learning-based parallel optimization method and system for genomic analysis

The application discloses a kind of based on machine learning's genome analysis parallel optimization method and system, the method of the present application includes the input BAM file is cut into multiple same size file blocks;For each file block, extract file block characteristics, and input the file block characteristics into the pre-trained machine learning model to obtain the predicted running time of the file block;The predicted running time is in the front part file block is divided into smaller file block;File block is input into HaplotypeCaller in parallel to carry out variation detection;The variation detection result generated by HaplotypeCaller for each file block is merged and output.The present application is aimed at solving the problem of serious calculation tilt and low resource utilization caused by unclear calculation complexity before HaplotypeCaller runs, improving the efficiency of genome analysis, shortening the variation detection duration.
Owner:SUN YAT SEN UNIV

A voiceprint registration method, device, equipment and storage medium

The application provides a voiceprint registration method, device, equipment and storage medium, the method comprises the following steps: storing the audio to be registered according to the time length of the audio to be registered when obtaining one audio to be registered every time; acquiring the first type of audio from the stored several audios based on a plurality of small file consumption service threads, and performing voiceprint registration on the first type of audio based on the operation equipment; and acquiring the second type of audio from the stored several audios based on a single large file consumption service thread, and performing voiceprint registration on the second type of audio based on the operation equipment; wherein the audio time length of the second type of audio is greater than the audio time length of the first type of audio. The voiceprint registration method provided by the application can perform voiceprint registration on short audios in a multi-thread concurrent manner (concentrate processing short audios), and perform voiceprint registration on long audios in a single thread, so that the resource utilization rate of the operation equipment can be effectively improved, the efficiency of voiceprint registration can be effectively improved, and the hardware cost can be greatly saved.
Owner:ANHUI IFLYTEK INTELLIGENT SYST

A fragmented upload method, device and medium for a distributed storage system

The present application discloses a fragment upload method, device and medium for a distributed storage system, which relates to the field of distributed storage, improves the upload efficiency of fragmented small files, receives a fragment upload request issued by a client, creates a fragmented small file, writes the fragment data of the fragment upload request in the fragmented small file, records the fragment number and fragment size information of the fragmented small file in a temporary index file, and after the fragment data upload is completed, creates an empty large file, the size of the large file is the actual size of the fragment data, reads the fragment information recorded in the temporary index file and updates it to the metadata of the large file. The uploaded fragment data is stored in the corresponding fragmented small file, and the fragment number and fragment size information of the corresponding fragmented small file are recorded in the metadata of the large file. There is no need to combine small files, which reduces the number of copies and improves the efficiency of fragment upload.
Owner:JINAN INSPUR DATA TECH CO LTD

Data processing method and system, storage medium and program product

The invention provides a data processing method and system, a storage medium and a program product. The method comprises the steps of obtaining a to-be-uploaded target file; for a first file of which the file size is greater than a file size threshold value in the target file, uploading the first file to a search resource server by adopting a fragment uploading mode; and for the second file of which the file size is smaller than or equal to the file size threshold value in the target file, directly uploading the second file to the search resource server, and selecting a better uploading path for the files of different sizes through simple threshold value judgment. The small file avoids the overhead of multiple rounds of handshake and multiple requests in fragment uploading, and the small file is quickly uploaded directly through one request, so that the delay caused by network round-trip time is greatly reduced.
Owner:UC MOBILE CHINA CO LTD

Backup method and device

This application provides a backup method and an apparatus. The method includes: determining a first application that needs data backup; packaging and compressing a file whose size is less than a first threshold in a user space corresponding to the first application, to obtain a first compressed package, where the first compressed package is stored in the user space corresponding to the first application; and backing up the first compressed package from the user space corresponding to the first application to a user space corresponding to a second application. The method can reduce overheads caused by an I / O operation generated by backing up a large quantity of small-sized files between two user spaces, reduce a backup delay, and improve user experience.
Owner:HUAWEI TECH CO LTD

Managed file compaction for distributed storage systems

Techniques implemented by file-compaction system that monitors storage containers in distributed storage systems, identifies small files that have been added to the storage containers, and automatically compacts the small files into larger files. Distributed file systems are well-equipped to manage large data files, but handling small files can slow the processing and analyzing of files, increase the retrieval time of files, etc. The file-compaction system monitors the storage containers to determine whether new small files have been added from data sources. Upon detecting small files that have been added to storage containers, the file compactor automatically compacts those small files into large files. By continuously identifying and compacting small files that have been added to storage containers, the file compactor provides the distributed file system with large files while minimizing the amount of time the storage containers are locked during compaction.
Owner:AMAZON TECH INC

Generation method and analysis method for data file of question bank system

The invention discloses a generation method and an analysis method for a data file of a question bank system. The method comprises the steps of generating a file overall structure; determining the position of a file header structure based on the overall file structure, and generating the file header structure based on the position of the file header structure; determining the position of a table index area based on the file header structure, generating a table index area structure based on the position of the table index area, determining the position of each metadata table in the table data area based on the table index area structure, and generating the table data area based on the position of each metadata table; determining the position of a file index area based on the file header structure, generating a file index area structure based on the position of the file index area, determining the position of each file resource in the file data area based on the file index area structure, and generating the file data area based on the position of each file resource. The technical problems that in the prior art, when a large number of small files are frequently and randomly accessed, decompression efficiency is low, memory occupation is high, and access delay is large are solved.
Owner:SICHUAN BISHENG INTELLIGENT TECHNOLOGY CO LTD

One-key automatic signature method and system based on CPE program small file

The invention discloses a one-key automatic signature method and system based on a CPE program small file, and relates to the field of CPE programs.The method comprises the steps that signature parameters input by a user are received through a graphical interface, and the signature parameters comprise a signature branch, a secret key list, storage capacity and a to-be-signed module path; and dynamically matching the key list according to the signature branch, and converting the parameter into a digital identifier which can be identified by the server side. And uploading the to-be-signed file to the Linux server through an SCP protocol after the to-be-signed file is subjected to standardized renaming, and calling a server signature interface to execute signature. And after the signature succeeds, the system renames the file and automatically downloads the file to a local specified path. The system is composed of a parameter input module, a dynamic display module, a file processing module, a signature execution module and a security verification module, automation, cross-platform compatibility and secure transmission of small file signature are achieved, the tedious process of small file burning signature of CPE equipment is remarkably simplified, and signature efficiency and security are improved.
Owner:GUANGZHOU TOZED KANGWEI INTELLIGENT TECH CO LTD

Methods, devices, storage media, and electronic devices for storing small files

This application provides a method, apparatus, storage medium, and electronic device for storing small files. The method includes: obtaining a first small file identifier, wherein the first small file identifier is used to identify a first small file to be deleted; then deleting the first small file according to the first small file identifier and updating target index information; wherein the first small file is a small file packaged and stored in a first packaged file, the first packaged file includes multiple small files arranged in a preset order, the target index information includes the data length of the first small file; a second small file is a small file packaged and stored in the first packaged file and located after the first small file in a preset order; and then moving the second small file according to the updated target index information, wherein the distance moved by the second small file is the same as its data length. This application solves the problem of low efficiency in deleting small files in packaged files in related technologies.
Owner:ZHEJIANG DAHUA TECH CO LTD

Massive file data stream storage method, storage system and big data processing system

The application relates to the technical field of computer application, and discloses a mass file data stream storage method, a storage system and a big data processing system. A rotation state conversion model is arranged between a data stream and a storage node, and storage pressure is dispersed through multi-node caching. The storage method comprises the following steps: presetting a file size demarcation point; inputting mass file data; judging whether the size of the input file is smaller than or equal to the demarcation point; storing file data smaller than or equal to the demarcation point in a selected storage node through the rotation state conversion model; generating index information of all small files, and merging the index information into an index file; organizing data in the form of a B+ tree for each index file to obtain an index data bucket; storing a plurality of index data buckets in the form of data blocks in the storage node; and saving all index data buckets in sequence as an index data bucket queue. The application can effectively solve the problems of traditional distributed storage systems, such as write performance decline and low efficiency when facing mass files.
Owner:CHINESE PEOPLES LIBERATION ARMY 92493 UNIT INFORMATION TECH CENT

File transmission method, electronic device, and storage medium

The application provides a file transmission method, an electronic device and a storage medium, and relates to the technical field of intranet security protection. According to the write parameters of a to-be-written file, an index file operation interface is called to determine whether an index file of the to-be-written file exists, if yes, according to the operation state of the to-be-written file in the index file and the write parameters of the to-be-written file, it is determined whether the index file of the to-be-written file is updated, when the to-be-written file is a large file or a small file, the index file operation interface is called according to the total size of the to-be-written file to update the storage location information in the index file where the to-be-written file is located, the updated index file records the real storage location information of the to-be-written file, and the real storage location information of the to-be-written file has the characteristics of high utilization of the disk space occupied by different data files, the write efficiency of the to-be-written file is improved, and the high-performance transmission requirement of the one-way import system is met.
Owner:南京中孚信息技术有限公司

Cloud database computing cache system and implementation method

The invention relates to a cloud database computing cache system and an implementation method, and the system comprises a decentralized cache architecture module which is used for eliminating a centralized bottleneck, supporting ten-thousand-level concurrency and adapting to a big and small file mixed scene; the memory object sharing module is used for reducing intermediate result processing delay through intermediate result full-life-cycle zero-copy circulation and dynamic storage decision based on a memory pool and metadata service of the memory object sharing module; the self-adaptive asynchronous write-in module is used for connecting remote storage, multiplexing memory object sharing module resources and combining mechanisms such as delay confirmation, so that the write response speed is increased, remote requests are reduced, data integrity is guaranteed, and meanwhile, persistent support is provided for intermediate results. Through cooperative operation of the three modules, the cloud database computing cache system and the implementation method can effectively adapt to the requirements of high concurrency, large and small file mixing and elastic capacity expansion and contraction of a cloud native OLAP scene, and the integrity of data transmission and the stability of system operation are guaranteed.
Owner:HASHDATA LTD