Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

334 results about "Data deduplication" patented technology

In computing, data deduplication is a technique for eliminating duplicate copies of repeating data. A related and somewhat synonymous term is single-instance (data) storage. This technique is used to improve storage utilization and can also be applied to network data transfers to reduce the number of bytes that must be sent. In the deduplication process, unique chunks of data, or byte patterns, are identified and stored during a process of analysis. As the analysis continues, other chunks are compared to the stored copy and whenever a match occurs, the redundant chunk is replaced with a small reference that points to the stored chunk. Given that the same byte pattern may occur dozens, hundreds, or even thousands of times (the match frequency is dependent on the chunk size), the amount of data that must be stored or transferred can be greatly reduced.

Data deduplication method and device, storage medium, electronic equipment and program product

The invention discloses a data deduplication method and device, a storage medium, electronic equipment and a program product, and relates to the technical field of data storage, the method comprises the steps that a memory layer of a tree-shaped storage architecture is arranged in a memory, only memory read operation is needed for online deduplication, hard disk access is not needed, delay and overhead of hard disk read operation in related technologies are avoided, and the data deduplication efficiency is improved. Data writing can be quickly responded, and the data processing efficiency is remarkably improved; background deduplication is adopted in a plurality of hard disk layers, and a corresponding deduplication strategy is automatically triggered after a corresponding triggering strategy is met, so that frequent online hard disk reading and writing are reduced, the disk reading frequency is reduced, and the overall performance of the all-flash memory array is improved; wherein the memory layer, the first hard disk layer and other hard disk layers adopt different deduplication strategies, and the differentiated deduplication strategies are formulated according to the characteristics of different levels, so that the deduplication operation is more targeted and efficient, the efficiency and accuracy can be improved, and the problem of high disk reading overhead in online data deduplication in related technologies is solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Ultra-short-term wind power prediction method based on Bayesian optimization XGboost-LSTM

The invention belongs to the technical field of wind power prediction, and discloses an ultra-short-term wind power prediction method based on Bayesian optimization XGboost-LSTM, and the method comprises the following specific steps: 1, inputting wind power data collected by a wind power plant and numerical weather prediction data of a corresponding time sequence; 2, data preprocessing, wherein missing value interpolation and data deduplication are carried out on input data; according to the method, an XGBoost feature optimization method is used for selecting key features influencing the wind power, the influence of irrelevant feature noise on the model prediction precision and accuracy is eliminated, then a Bayesian optimization algorithm is used for carrying out hyper-parameter tuning on an LSTM model, and finally an XGBoost-LSTM-BO model is constructed. XGboost feature optimization and Bayesian hyper-parameter optimization can obviously improve the prediction effect of the LSTM on future data and improve the model prediction precision, and compared with a traditional prediction model, the wind power data generalization ability of the model can be improved while the high prediction precision is kept, and higher prediction performance is achieved.
Owner:INNER MONGOLIA UNIV OF TECH

Data deduplication method and related system

A data deduplication method includes: receiving a write request with first data block; writing the first data block into a storage device; writing metadata of the first data block into a first partition that is in a plurality of partitions of a metadata management structure and that is determined based on a feature of the first data block; and deleting the metadata of the first data block in the first partition and deleting the first data block from the storage device based on address information of the first data block when a fingerprint that is the same as the fingerprint of the first data block exists in the first partition. The method can prevent infrequently updated data from being evicted because resources are occupied by frequently updated data, thereby improving a deduplication ratio.
Owner:HUAWEI TECH CO LTD

Knowledge data deduplication method and device, storage medium and computer equipment

The invention discloses a knowledge data de-duplication method and device, a storage medium and computer equipment, relates to the technical field of data processing, is suitable for businesses such as financial science and technology and smart medical treatment, and mainly aims to solve the problem of poor de-duplication effect when large-scale knowledge data is processed in existing knowledge data de-duplication. Comprising the steps of obtaining large-scale knowledge data of related businesses; performing clustering processing on the large-scale knowledge data by adopting a MinHash LSH model to obtain a repeated text clustering result; the repeated text clustering result comprises a plurality of groups of similar data sets; performing word embedding calculation on each group of similar data sets by adopting a bge-m3 model to obtain word embedding corresponding to each group of similar data sets; and respectively carrying out semantic similarity de-duplication processing on the word embedding in each group of similar data sets to obtain a de-duplication result of the knowledge data.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Data deduplication method and system based on incremental calculation and feature clustering and electronic equipment

PendingCN120470241AClustered dataData stream
The invention discloses a data deduplication method and system based on incremental calculation and feature clustering and electronic equipment, and the method comprises the steps: receiving multi-source heterogeneous text, image, audio and video data streams in real time, adding a timestamp to each data unit, and generating an input data set with a time attribute; performing multi-modal feature extraction on the input data set, and dynamically adjusting feature weights of the extracted multi-modal features through a sliding window model and a time decay factor to obtain an increment feature vector set with the weights; executing two-stage clustering based on the feature vector set to generate a plurality of fine-grained data clusters; and comparing the intra-cluster data of each fine-grained data cluster in pairs by adopting a composite similarity model to determine intra-cluster repeated data, and performing data deduplication on the input data set according to the intra-cluster repeated data to obtain a deduplicated input data set. According to the method, the processing efficiency of the real-time data is improved, and the storage and calculation cost is remarkably reduced.
Owner:DATA SPACE RES INST

Chemical early warning interception method and device, electronic equipment and medium

PendingCN121329260AData processing applicationsChemical riskIdentity recognition
The embodiment of the invention discloses a chemical early warning interception method and device, electronic equipment and a medium. A specific embodiment of the method comprises the following steps: carrying out identity recognition on a target chemical detection terminal set to obtain a detection terminal identity recognition information set; performing state information extraction on the multi-modal target chemical information set to obtain a chemical multi-modal state information set; performing data deduplication on the multi-modal state information set of the chemicals and then performing association identification to obtain a multi-modal association information set; performing chemical risk identification on the de-duplicated multi-modal chemical information set to obtain a to-be-tracked chemical information set; and performing trajectory tracking early warning on the to-be-tracked chemical article information set to obtain a chemical article trajectory early warning information set, and performing early warning interception on the to-be-tracked chemical article information set. According to the embodiment, a data barrier can be broken, data sharing and safety are improved, and the accuracy of accurate interception and risk early warning of target chemicals is improved.
Owner:GUANGXI BEITOU XINCHUANG TECH INVESTMENT GRP CO LTD

ZNS SSD-based repeated data deletion method and system

The invention discloses a repeated data deletion method and system based on ZNS SSD, and belongs to the field of repeated data deletion.The method comprises the steps that a fingerprint index table and a fingerprint updating cache are maintained, three-time opportunity repeated data deletion is achieved, and repeated data search is conducted in the fingerprint index table, the fingerprint updating cache and an address mapping table respectively; an address mapping table, an address reverse mapping table and a reference counting table are maintained, physical space garbage collection / defragmentation and write request merging are achieved, the address mapping table, the address reverse mapping table and the reference counting table conduct garbage collection with partition as granularity when free space is insufficient, and a fingerprint index table is not updated after garbage collection data migration; and instead, the mapping of fingerprints to migration-in physical pages is stored in a fingerprint updating cache, the physical pages are sequentially distributed in partitions by the fingerprint updating cache, normal data write requests and garbage collection data migration write requests are respectively cached, and the write requests which are continuous in sequence are merged and then issued. According to the method, the high-performance and low-cost efficient storage system can be constructed on the basis of the ZNS SSD at relatively low overhead.
Owner:HUAZHONG UNIV OF SCI & TECH

Optimizing retention timeframes of deduplicated copies at storage platforms that are write-once read-many (WORM) enabled

A data storage management system is enhanced to accommodate, and moreover to optimize, the storing and retention of deduplicated secondary copies at write-once read-many (WORM) enabled storage platforms. Enhancements include without limitation: user interface (UI) options to enable WORM functionality for secondary storage, whether used for deduplicated or non-deduplicated secondary copies; enhancements to secondary copy (e.g., deduplication copy, backup) operations; and pruning changes. The storage manager is generally responsible for managing the creation, tracking, and deletion of secondary copies, with and without deduplication. Media agents that store secondary copies to and prune them from the WORM-enabled storage platforms also are enhanced for communicating and interoperating with both bucket-level and object-level WORM-enabled storage platforms to implement the features disclosed herein.
Owner:COMMVAULT SYSTEMS INC

Flow-level deduplication of network traffic in a network traffic visibility system

A system and method for flow-level deduplication of network traffic are disclosed. A network node receives a first plurality of packets from a first network endpoint. The first plurality of packets represent a flow of data being communicated between the first network endpoint and a second network endpoint. The network node further receives a second plurality of packets from the second network endpoint. The network node identifies a sequence identifier of each packet of the first and second pluralities of packets. The network node determines that the first and second pluralities of packets are all associated with the same flow, based on the sequence identifiers of the first and second pluralities of packets. In response to that determination, the network node deduplicates the flow by discarding the first plurality of packets or the second plurality of packets. The network node may be a traffic visibility node.
Owner:GIGAMON INC

Intelligent examination and approval management method for housing and construction industry

The invention discloses an intelligent examination and approval management method for the housing and construction industry, and the method comprises the steps: collecting examination and approval application data, carrying out the data cleaning, data deduplication and data standardization operation of the examination and approval application data, and obtaining preprocessed data; constructing a data classification model based on a machine learning algorithm, and performing material classification on the preprocessed data by using the data classification model to obtain a material classification result; according to a material classification result, performing primary processing and risk assessment on the preprocessed data by using an intelligent auditing system to obtain a risk score; performing secondary processing on the pre-processed data by using an expert system to obtain a personalized approval scheme; and examining and approving the preprocessed data according to the personalized examination and approval scheme, and storing the risk score and the examination and approval result to an intelligent examination and approval system. The invention relates to the technical field of examination and approval management, and solves the technical problems of low data quality, lack of flexibility and intelligence and the like of an existing residence construction examination and approval system.
Owner:HUBEI PUBLIC INFORMATION IND CO LTD

Deduplication selection and optimization

Systems and method for implementing deduplication process based on performance analyses. The system may include a processing device to determine a first performance metric associated with retrieving a second stored data block that is within a specified range of a duplicate of the first data block and a second performance metric associated with retrieving a hash value corresponding to the second stored data block. The processing device further to retrieve the second stored data block within a specified range of the duplicate of the first data block in response to the first performance metric not exceeding the second performance metric.
Owner:PURE STORAGE INC

Data deduplication using modulus reduction arrays

A computer-implemented method is disclosed for data processing. The method includes receiving real-time streaming data that includes data object identifiers, arranged in a sequenced source data array, for source data objects. The method also includes determining an integer number N of data object identifiers in the source data array and selecting three mutually prime integers N1, N2, N3 such that N1 is greater than N2, N2 is greater than N3, and an arithmetic product of N1, N2, N3 is greater than N. The method further includes generating a first, second and third modulus reduction arrays of lengths, respectively, N1, N2, and N3. The method also includes initializing the first, second and third modulus reduction arrays with dummy values, storing the source data objects in a source data object array, and storing the first, second and third modulus reduction arrays in a short term memory.
Owner:SALESFORCE INC

Distributed data deduplication method and product

The embodiment of the invention provides a distributed data deduplication method and product. The method comprises the steps of obtaining a data key of data to be subjected to deduplication; distributing the data to be subjected to duplicate removal to the ith sub-bucket according to the mixed hash value of the data key; mapping the mixed hash value according to a hash function attribute value configured for the ith sub-bucket to obtain k position numbers; reading numerical values of positions corresponding to the k position numbers from a bit array which is stored at the current moment and corresponds to the ith sub-bucket to obtain k numbers; if it is confirmed that the k numbers are all first numerical values, it is confirmed that the data to be subjected to duplicate removal belong to duplicate data; and if at least one of the k digits is confirmed to be a second numerical value, confirming that the data to be subjected to duplicate removal belongs to new data. According to the embodiment of the invention, the speed and accuracy of big data deduplication processing are remarkably improved, and the problem of single-node calculation overload (different buckets can be distributed to different calculation nodes) is effectively solved.
Owner:SHOUSHI SECURITY TECHNOLOGY CO LTD

Big data storage optimization method based on zero-copy collaboration technology and related equipment

The embodiment of the invention provides a big data storage optimization method based on a zero-copy collaboration technology and related equipment, and belongs to the technical field of big data storage. The method comprises a zero-copy data transmission module used for directly writing a data stream into a front-end buffer area of an annular structure through a direct memory access technology, and adopting a zero-copy algorithm based on pointer offset to parallelly separate data of each channel in a multi-thread environment; the intelligent data reduction module is used for deleting duplicated data from the data and then compressing the duplicated data; and the hybrid storage management module is used for managing the hybrid partition storage architecture and dynamically adjusting the distribution of the data in the hybrid partition storage architecture according to the data access mode. The CPU copy frequency is reduced through the zero copy technology, the data transmission efficiency is remarkably improved by combining the data reduction and intelligent layering strategies, the occupied storage space is reduced, and meanwhile the data safety and compliance are guaranteed.
Owner:GUANGDONG WANZHANG JINSHU INFORMATION TECH CO LTD

Methods and systems for data management, integration, and interoperability

Embodiments herein relate to data management and, more particularly, to collecting data from a plurality of sources, and linking the collected data to derive information and knowledge. The method includes defining at least one data model and asset by including data models, vocabulary, data quality rules, data mapping rules for at least one of, a particular data industry, a data domain, or a data subject area, importing data from a plurality of data sources, performing de-duplication of the imported data and data profiling of the imported data, and creating linked data either by semantic mapping, or by curating the data.
Owner:TRIGYAN CORP INC

Coordinating deduplication among nodes

Techniques coordinate deduplication among nodes. The techniques involve, in response to a lookup query to search a deduplication index for an entry that maps a fingerprint which is based on incoming data, detecting an invalid result. The techniques further involve, in response to detecting the invalid result, updating the deduplication index to include an entry that maps the fingerprint, the entry initially storing an “in-progress” flag. The techniques further involve, after the deduplication index is updated to include the entry that maps the fingerprint and after the incoming data is stored in a storage location, removing the “in-progress” flag from the entry.
Owner:DELL PROD LP

Data duplicate removal method, electronic equipment and storage medium

The invention provides a data duplicate removal method, electronic equipment and a storage medium. According to the method, smooth data transition is achieved through a double-bloom filter alternating mechanism. In first time period processing, query and addition operations are only executed on a first bloom filter of a current time window. And second time period initialization: creating a new second bloom filter at the first unit time of the second time period as a next window preparation filter. According to the double-writing preheating mechanism, data are written into the first Bloom filter and the second Bloom filter at the same time in the second time period, and the missed judgment problem during window switching is avoided. According to window switching logic, old filter resources are released, the preliminary filter is marked as a new main filter, and monthly / periodic seamless connection is achieved. According to the method, the misjudgment rate is controlled through a double-bloom filter time window switching mechanism, the data duplicate removal accuracy and the system resource utilization rate can be improved, and the method is suitable for the efficient duplicate removal requirement in a mass data collection scene.
Owner:北京中科闻歌科技股份有限公司

System and method for generating SQL (Structured Query Language) by natural language based on ES, knowledge base and interaction enhancement

The invention relates to the technical field of language generation algorithms, and discloses a natural language generation SQL method based on ES, knowledge base and interaction enhancement, and the method comprises the following four sequentially executed core steps: data feature preprocessing: carrying out deep analysis (including field type, constraint relationship, data distribution and the like) on a database table structure and data features, and carrying out data feature preprocessing; and after data de-duplication is executed, the ES is imported, and an independent structured index is established according to a table-level isolation principle to provide basic data for subsequent retrieval. According to the method, the keyword matching efficiency and the semantic understanding depth are both considered through collaborative retrieval of the double-storage-layer knowledge base in combination with Elasticsearch keyword inverted index and Milvus vector library semantic retrieval. And the results are fused through the RRF algorithm, so that the retrieval precision is remarkably improved, the limitation of a single retrieval mode is solved, and the generated SQL better fits the database structure and the user intention.
Owner:QUALITY ENERGY BODY TECHNOLOGY (TIANJIN) CO LTD +1

Recovery of data from sequential storage media

A system for managing a sequential storage medium of a plurality of sequential access storage devices for deduplication storage of target data on the sequential storage medium is provided. A processor identifies duplicate segments of segments of target data on a medium. The processor determines a combination of predicted total recovery times less than a threshold. The combination includes one or more of the identified duplicate segment, the rewrite segment, and the new data segment. The rewrite segment is a duplicate segment designated for rewriting onto the medium as new data. The new data segments include data in duplicate segments that are not currently stored on the medium, i.e., non-duplicate segments. New data segments are designated for first writing as new data onto the medium. The processor writes according to the combination indication, i.e., writes a rewrite segment and / or a new data segment onto the identified medium.
Owner:HUAWEI TECH CO LTD

Repeated data deletion method based on update perception

The invention provides a duplicated data deletion method based on update perception. The method comprises the following steps: acquiring a to-be-processed file; the version type of the to-be-processed file is determined and correspondingly marked, and the version type comprises a basic version and an incremental version; performing block processing on the to-be-processed file, and calculating a fingerprint corresponding to each data block; if the version type of the to-be-processed file is a basic version, deduplication is conducted on all data blocks and an existing data fingerprint set in a hash table in sequence according to fingerprints, unique data blocks are stored in a container, the hash table of the basic version is updated, and the fingerprint, the storage address and the reference frequency of each data block are recorded in the hash table; if the version type of the to-be-processed file is an incremental version, the corresponding basic version hash table needs to be loaded firstly, and then duplicate removal processing is carried out; the data index range is narrowed by distinguishing the basic version from the incremental version, so that the data processing efficiency is improved.
Owner:XIAMEN UNIV

Cyclic automatic data acquisition method and system

The invention provides a cyclic automatic data acquisition method and system. The method comprises the following steps: forming an entry URL set; according to the entry URL set, based on DOM structure feature analysis and semantic association degree evaluation, and link value scores of TF-I DF and Word2Vec, forming a high-value link queue; according to the high-value link queue, obtaining page contents, and forming a page queue containing effective telephone numbers; according to the page queue, performing time sequence interpolation by using a self-attention diffusion model to form a merchant data set; performing field information extraction and telephone number grouping identification by utilizing a field extraction neural network model and a telephone number grouping identification model according to the merchant data set, and generating a structured data set; and according to the structured data set, executing multi-dimensional data fingerprint generation to perform data deduplication. According to the method and the device, the technical problems of the traditional automatic data acquisition technology in the aspects of complex webpage structure identification, data timeliness maintenance and data quality assurance are solved.
Owner:BEIJING YULORE INNOVATION TECH

Yield optimization of cross-screen advertising placement

A method may include obtaining an advertising campaign, determining one or more devices associated with a consumer, and associating the one or more devices with a consumer identifier. The method also includes obtaining data associated with the one or more devices associated and corresponding the data with the consumer identifier. The method also includes deduplicating the data related to the consumer identifier across one or more viewing methods. The method also includes displaying the deduplicated data. The method also includes adjusting the advertising campaign based on user input.
Owner:VIDEOAMP INC

Automated alert deduplication or suppression in data processing systems based on recurring data identifiers

There are provided systems and methods for automated alert deduplication or suppression in data processing systems based on recurring data identifiers. An entity, such as company or business, may utilize computing services provided by a service provider. When providing these services, one or more computing services, processors, or the like of the service provider's computing architecture may be used. Use of computing services may generate security alerts when computing events are flagged as risky, fraudulent, malicious, computing attacks, or the like. To automate security alert management, the service provider may utilize an alert management system that may parse and extract data from incoming security alerts and calculate identifiers from such data, such as by transforming or converting using identifier functions. Recurring identifiers may be automatically organized for suppression or deduplication based on past occurrence of such identifiers with other security alerts.
Owner:BREX INC

Ubiquitous network data deduplication system supporting similarity data ownership verification

The invention relates to the technical field of big data processing and network security, and discloses a ubiquitous network data deduplication system supporting similarity data ownership verification, which comprises an encryption module, a label generation module, a deduplication judgment module, a challenge generation module, a certification calculation module, a verification module and an integrity evidence storage module. According to the system, a plaintext data block is encrypted by deriving a high-entropy session key through two-dimensional chaotic mapping, a similarity label with a threshold value is generated by adopting # imgabs0 #, and efficient de-duplication is realized based on three-section type Hamming distance judgment; a dynamic challenge-response mechanism is utilized to provide auditable ownership proof for each downloading request, and a # imgabs1 # tree and a block chain evidence storage technology are combined to realize data integrity multi-granularity verification and rapid damaged block positioning, so that storage efficiency, security and availability are considered.
Owner:CHANGCHUN UNIV OF SCI & TECH

Data partitioning method and device, electronic equipment and storage medium

The invention discloses a data partitioning method and device, electronic equipment and a storage medium, and relates to the technical field of computers, and the method comprises the following steps: dividing a sliding window of to-be-processed data into a plurality of sub-regions, comparing adjacent byte pairs in the sub-regions in parallel by utilizing a vector instruction to generate a local maximum set, vector maximum comparison is carried out on a local maximum set through a tree structure so as to reduce the scale of the set layer by layer, finally, a global extreme value is determined to serve as a data block boundary, efficient data block division based on vector parallel calculation is achieved, byte-by-byte or block-by-block serial calculation logic is not adopted, and therefore efficient data block division is achieved. The technical problem that in the prior art, due to the fact that serial computing logic is adopted, computing efficiency is limited, performance bottleneck occurs, and then the processing speed and the resource utilization rate of applications such as data deduplication, compression and encryption are affected can be solved. The technical effects of improving the calculation efficiency, reducing storage fragmentation, relieving the performance bottleneck of a high-concurrency scene and improving the data processing speed and the resource utilization rate are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Privacy-preserving data deduplication

A method includes a server computer receiving, from a first data provider computer, encrypted data derived from first identity data and a cryptographic key or derivative thereof stored at the first data provider computer. The server computer transmits, to a second data provider computer, the encrypted data and / or the cryptographic key or derivative thereof. The server computer receives, from the second data provider computer, intermediate data derived from second identity data stored at the second data provider computer. The server computer determines if the first identity data and the second identity data are duplicates while the first identity data and the second identity data are encrypted. The server computer removes one of encrypted first identity data, derived from the first identity data, and encrypted second identity data, derived from the second identity data, from a memory in the server computer.
Owner:VISA INTERNATIONAL SERVICE ASSOCIATION

Leveraging host accelerator resources to reduce storage system computational load

An apparatus illustratively comprises at least one processing device that includes a processor and a memory, with the processing device being configured to obtain in a host device deduplication information relating to a logical storage device of a storage system, wherein the host device comprises acceleration resources and is configured to communicate with the storage system over at least one network, and to determine based at least in part on the obtained deduplication information whether to compute a content-based signature for one or more data blocks of the logical storage device utilizing the acceleration resources of the host device. Responsive to an affirmative result of the determining, the content-based signature is computed utilizing the acceleration resources of the host device and sent by the host device to the storage system in association with a given write operation that targets the one or more data blocks of the logical storage device.
Owner:DELL PROD LP

Methods and systems for space reclamation in immutable deduplication systems

Embodiments are disclosed that provide space reclamation in immutable deduplication systems, and can include selecting a unit of data of a backup image, determining whether a duplicate unit of data is stored in an existing data storage construct (the duplicate unit of data is a duplicate of the unit of data and the existing data storage construct is stored in immutable storage), and in response to a determination that the duplicate unit of data exists in the existing data storage construct, determining whether the existing data storage construct is designated as being available to be referenced, in response to the existing data storage construct being designated as being available to be referenced, updating a reference to the duplicate unit of data, and in response to the existing data storage construct being designated as being unavailable to be referenced, storing the unit of data in a new data storage construct.
Owner:COHESITY INC

Bitmap index container optimization method and system based on multi-dimensional dynamic decision

The invention provides a bitmap index container optimization method and system based on a multi-dimensional dynamic decision, and relates to the technical field of databases, and the method comprises the steps: constructing a feature fusion vector by calculating a skewness coefficient value and fusing a data distribution feature vector, a data access rule vector and a change trend vector; a self-adaptive threshold value is obtained through iterative optimization of a double-strategy gradient algorithm, whether the container type is switched or not is judged, data deduplication, priority ranking and batch data migration are conducted through a temporary container, and a bidirectional index table is established to redirect an access request.
Owner:北京科杰科技有限公司