Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

418 results about "Data partitioning" patented technology

Experimental data processing method and device, AI analysis module and computer equipment

The invention relates to an experimental data processing method and device, an AI analysis module and computer equipment, and belongs to the field of data processing.The method comprises the steps that multi-dimensional original data are partitioned according to types, and formats are unified; generating a similarity matrix based on time and space neighborhood information, and marking abnormal fluctuation points; effective signals are separated through a time-frequency feature matching noise library; extracting multi-layer features of basic statistics, time sequence correlation and trend change; and dynamically screening core features to update the tracking type experimental model. The matched AI analysis module integrates hardware circuits of data partitioning, similarity calculation, anomaly marking, noise matching, feature extraction and model updating, and whole-process acceleration is achieved. According to the method, through multi-dimensional data compatibility processing, accurate anomaly detection, multilayer feature fusion and model adaptive optimization, the experimental data processing efficiency and conclusion reliability are remarkably improved, and the method is suitable for real-time analysis of multiple scenes such as scientific research and industry.
Owner:深圳市伊元科技有限公司

Data platform file migration method, computer program product and data platform

The invention relates to a data platform file migration method, a computer program product and a data platform, and the data platform file migration method comprises the steps: scanning a to-be-migrated source end data file, and obtaining the type of the source end data file; if the type of the source end data file is in the state set of the target reinforcement learning model, taking the type of the source end data file as the current state of the target reinforcement learning model, and obtaining the current action of the target reinforcement learning model in response to the current state; using the current action of the target reinforcement learning model to block the source end data file; constructing a Merkel tree of the source end data file based on the blocking result of the source end data file; comparing Merkel trees of the source end data file with Merkel trees of the target end data file through layer-by-layer hash to locate difference data blocks; and migrating the difference data block to the target end. The problem that in an existing data platform file migration method, a data partitioning strategy is not reasonable, and consequently difference positioning is prone to failure is solved.
Owner:安徽明生恒卓科技有限公司 +2

Unstructured data synchronization method and system based on Flink

The invention provides an unstructured data synchronization method and system based on Flink, and relates to the technical field of big data process.The method comprises the steps that fragmentation processing is conducted on unstructured data, hierarchical storage of the data in a heap memory layer and a RocksDB layer is achieved on the basis of access feature scores, data partition optimization is conducted in combination with deep reinforcement learning, and the unstructured data are synchronized. And the parallelism degree parameter and the resource quota are dynamically adjusted by calculating the back pressure index value, and the data synchronization reliability is ensured by adopting a check point snapshot mechanism, so that the data processing efficiency and the system throughput are effectively improved.
Owner:北京科杰科技有限公司

Data leakage risk quantification method for autoregression language model training process

The invention discloses an autoregressive language model training process-oriented data leakage risk quantification method, which comprises the following steps of: executing enhanced member data division processing on an original training text set, and generating a ternary partition text set containing member, non-member and key boundary texts through an optimization function; for each text, using the target model to extract member attributes from three channels of forward reasoning, back propagation and state evolution, and fusing the member attributes into member attribute vectors; inputting the member attribute vector into a hierarchical comparison embedded network, and mapping the member attribute vector into an optimized embedded representation vector through comparison learning including similar clustering, heterogeneous separation and boundary positioning; and inputting the embedded representation vector into a dynamic reasoning classifier capable of perceiving a training stage, judging the risk membership degree of the dynamic reasoning classifier, and generating a data leakage risk quantification result. According to the invention, real-time and fine-grained dynamic quantification can be carried out on the leakage risk, and timely early warning is provided for model training.
Owner:NANJING ARTIFICIAL INTELLIGENCE CHIPS RES INST OF AUTOMATION CHINESE ACAD OF SCI +1

Large-scale encrypted traffic frame-by-frame clustering analysis method based on big data architecture

The invention provides a large-scale encrypted traffic frame-by-frame clustering analysis method based on a big data architecture, and relates to the technical field of big data, and the method comprises the steps: carrying out the data partitioning and storage of target encrypted data obtained through the preprocessing of original encrypted traffic data; based on a clustering visualization result obtained by visualizing a target clustering result obtained by carrying out frame clustering analysis on the target encrypted data, judging whether a data traffic abnormal behavior exists or not, and when the data traffic abnormal behavior exists, generating an abnormal analysis report; and an abnormal analysis report is transmitted to safety management personnel so as to take corresponding measures in time for defense. The method comprises the following steps: visualizing a target clustering result obtained by carrying out frame clustering analysis after processing and partition storage on encrypted traffic data based on a big data architecture, identifying traffic data exception, generating an exception analysis report when the exception exists, and transmitting the exception analysis report to safety management personnel to take corresponding measures for defense. The network threat identification capability is effectively enhanced, and the overall protection level of network security is further improved.
Owner:BEIJING QITIAN ANXIN TECH CO LTD

Firmware upgrading method and electronic equipment

The invention relates to a firmware upgrading method and electronic equipment. The method comprises the following steps: in a first stage of upgrading, constructing upgrading meta-information according to a current partition table and a to-be-upgraded firmware data storage area base address obtained by calculating to-be-upgraded firmware meta-information, and according to a block list description file, updating the to-be-upgraded firmware data storage area base address; in the first-stage upgrade, a partition table file with a preset name in a firmware upgrade package under a preset firmware upgrade package path of an encrypted preset user data partition and all to-be-upgraded firmware mirror image files are stored in an address space determined by a base address of a to-be-upgraded firmware data storage area; and determining a to-be-upgraded firmware data storage space end address and a to-be-upgraded firmware data moving base address according to the guide communication partition address space change mark, and copying data in an address space determined by the to-be-upgraded firmware data storage space end address to the to-be-upgraded firmware data moving base address. According to the invention, under the condition of user data partition encryption, the firmware with any change of the partition table can be safely and reliably upgraded.
Owner:FUZHOU ROCKCHIP SEMICON

Computer data management system based on big data

The invention relates to the technical field of big data processing, and particularly discloses a computer data management system based on big data. The system comprises a distributed data acquisition module which captures multi-source data through an API gateway and a crawler and embeds a quality label; the intelligent storage scheduling engine is used for realizing automatic migration of cold and hot data by adopting a column-type and distributed hybrid architecture; fusing a computing framework, and dynamically coordinating flow processing and batch processing resource execution feature extraction; the self-adaptive strategy center is used for generating data partitions and encryption strategies based on reinforcement learning; and a multi-level security protection system is adopted, and attribute base decryption and dynamic desensitization are deployed. The problems of resource scheduling rigidity and security protection lag in the prior art are solved, the storage cost is reduced by more than 40%, the query delay is compressed to be within 200ms, and the method is suitable for real-time decision-making scenes in the fields of e-commerce and finance.
Owner:HAINAN VOCATIONAL COLLEGE OF SCI & TECH

Cable defect analysis method and system in cable laying process

The invention discloses a cable defect analysis method and system in a cable laying process, and the method comprises the steps: dividing a current data sequence according to a preset data division strategy, obtaining at least one current data sub-sequence, and obtaining cable images of different monitoring positions according to the at least one current data sub-sequence; preprocessing the cable image to obtain a target cable image, and performing cable contour and background region segmentation on the target cable image to obtain a cable contour image; calculating a cable area based on the cable contour image, and extracting cable surface features; analyzing the surface features of the cable based on a set noise threshold, and removing noise features to obtain optimized surface features of the cable; and analyzing the surface features of the optimized cable based on the defect detection model, and outputting surface defect information. The accuracy and the real-time performance of defect detection are remarkably improved, the manual inspection cost and the leakage detection risk are reduced, and reliable guarantee is provided for safe operation of the cable.
Owner:SHANGHAI JIULONG ELECTRIC POWER GROUP

Database data storage system based on multi-level cache

The invention discloses a database data storage system based on multi-level cache, and relates to the technical field of database data storage systems, the database data storage system comprises a basic storage layer, a first-level cache layer, a second-level cache layer, a cache cooperative control module and a data read-write scheduling module, the basic storage layer is used for persistently storing total data of a database; the first-level cache layer is divided into a plurality of hot data partitions; and the second-level cache layer is provided with quasi hot data partitions in one-to-one correspondence with the hot data partitions. According to the database data storage system based on the multi-level cache, the system constructs a'high frequency-secondary high frequency-full quantity 'layered architecture through the first-level cache, the second-level cache and the basic storage layer, dynamic flow of data is achieved in cooperation with a data aging counter and a heat recovery detector, heat reduction data is migrated from the first-level cache to the second-level cache, and the heat reduction data is transferred to the second-level cache. And the heat recovery data returns to the first-level cache from the second-level cache to ensure that the high-frequency data is always kept in a cache layer.
Owner:SHANDONG TRANSPORT VOCATIONAL COLLEGE

Data scheduling method, device and system and storage medium

The invention discloses a data scheduling method, device and system and a storage medium. The method comprises the steps of obtaining scheduling duration required by to-be-scheduled data in a service system in a data link of a data warehouse; dividing the plurality of time zones into a plurality of time zone segments according to the scheduling duration; determining scheduling start time in a preset reference time zone according to the minimum time zone of each time zone segment; determining a target time period for dividing the to-be-scheduled data in a preset reference time zone according to the scheduling start time; for each time zone segment, determining the data of which the generation time is within the target time period in the to-be-scheduled data of each time zone segment as target data; and scheduling the corresponding target data to the data warehouse according to the scheduling start time of the time interval segmentation, so as to store the target data to the data partition, corresponding to the target time interval, of the data warehouse. According to the scheme, under the condition that the time-zone-crossing data scheduling timeliness is ensured, the performance overhead during scheduling is reduced, and the data consistency is ensured.
Owner:ZHONGKE YUNGU TECH

Industrial algorithm model scheduling method and system for industrial big data

The invention relates to the technical field of industrial big data processing and algorithm scheduling, and particularly discloses an industrial algorithm model scheduling method and system for industrial big data, and the method comprises the steps: constructing an algorithm dependency graph of an execution pipeline composed of a plurality of industrial algorithms, bottleneck nodes influencing the overall performance are identified through a critical path analysis algorithm; carrying out operation fusion on algorithm pairs with close dependency relationships; optimal data partitioning strategies are determined for different algorithms, and an elastic execution pipeline is realized; allocating the most matched heterogeneous computing resources for the algorithm operation on the critical path, and adjusting the execution priority of the algorithm operation in real time according to the system load and the operation scene; a fine-grained synchronization system is realized; the industrial algorithm execution efficiency is improved, and the high-performance requirement of an industrial big data processing scene is met.
Owner:ZHONGKE YUZHOU (GUANGDONG) TECHNOLOGY SERVICE CO LTD

Efficient query acceleration and cache optimization device for data knitting

The invention provides an efficient query acceleration and cache optimization device for data knitting, which comprises a semantic partitioning component used for performing dynamic partitioning on a multi-source heterogeneous data set based on multi-source heterogeneous semantic metadata and embedding cross-domain association identifiers of data knitting to generate a semantic partitioning index table; the query intention mapping component is used for executing intention extraction on a query request initiated by a user to generate a query intention representation, and establishing a mapping relationship of the query intention representation to-be-queried target data partitions according to the semantic partition index table so as to generate a query-partition mapping table; the cache dynamic sorting component is used for generating a cache priority sorting table according to the query-partition mapping table and historical query frequency statistical data; and the scheduling and query execution component is used for optimizing the cache mechanism of the to-be-queried target data partition based on the cache priority sorting table and a preset proportion threshold value. According to the method, the time consumption of cache calling and source data repeated loading in the later query period is reduced.
Owner:BEIJING ZHONGSHURUIZHI TECH CO LTD

Method and system for dynamic data partitioning and efficient polling in distributed microservices architecture

A method and a system for utilizing dynamic data partitioning to ensure consistency of event publications with associated business transactions in a distributed microservices architecture in order to optimize efficiency in polling data are provided. The method includes: defining a hash space that includes a range of assignable hash values; deploying a respective instance of each microservice to form a cluster of microservices within the distributed microservices architecture; allocating a respective subset of the hash space to each microservice; and facilitating a data polling capability with respect to a data table based on the allocated respective subset of the hash space and the deployed respective instance for each microservice.
Owner:JPMORGAN CHASE BANK NA

Data processing method, device and equipment and computer readable storage medium

The invention discloses a data processing method, device and equipment and a computer readable storage medium. The method comprises the following steps: obtaining transaction data in a target period, the transaction data comprising first transaction data updated in the target period; determining a target first-level partition corresponding to the first transaction data and a target hot data partition under the target first-level partition from a plurality of first-level partitions of the data warehouse based on the transaction occurrence period of the first transaction data; the first-level partition is determined based on a transaction generation period, a second-level partition is arranged under the first-level partition, the second-level partition comprises a cold data partition and a hot data partition, the cold data partition is used for storing newly added transaction data in the transaction generation period, and the hot data partition is used for storing updated transaction data in the transaction generation period; and updating the second transaction data in the target hot data partition under the target first-level partition by using the first transaction data. According to the data processing method provided by the embodiment of the invention, the data updating efficiency can be improved.
Owner:CHINA UNIONPAY

Aerospace equipment performance evaluation method based on digital real test data fusion

The invention discloses an aerospace equipment performance evaluation method based on digital test data fusion, and belongs to the field of electronic engineering and computer science, and the method comprises the steps: designing a physical test and a digital test, selecting related indexes of aerospace equipment, and carrying out corresponding physical and digital tests around the related indexes of the aerospace equipment; performing preprocessing and multi-source data integration on the physical test data; performing data quality evaluation and virtual data division on the digital test data; and performing double fusion on the result through an adaptive fusion algorithm to obtain a performance evaluation result. The advantages of the digital test and the physical test can be fully exerted, and comprehensive and accurate evaluation of the aerospace equipment performance is realized through deep fusion of the digital test data and the physical test data.
Owner:BEIHANG UNIV

Data storage system and method for artificial intelligence deep learning

The invention discloses a data storage system and method for artificial intelligence deep learning, and relates to the technical field of data storage, and the system comprises a data obtaining module which is used for recognizing a data source and obtaining data from the data source; the data preprocessing module is used for automatically detecting and cleaning abnormal values in the data; the data storage design module is used for automatically storing the data in different types of storage media based on the access frequency and importance of the data and establishing a dynamic data partition model; the real-time data processing module is used for storing and managing real-time data streams; and the data backup module is used for backing up the data. According to the method, intelligent storage layering is provided, the data can be automatically stored in different types of storage media according to the access frequency and importance of the data, the storage cost is optimized, the access speed is increased, and a data partitioning strategy can be adjusted in real time to cope with the change of the data volume by establishing a dynamic adjustment model and adopting a dynamic data partitioning technology.
Owner:杭州中谦科技有限公司

Tax form intelligent asynchronous processing and real-time query method based on dynamic partition

The invention discloses a dynamic partition-based tax form intelligent asynchronous processing and real-time query method, which comprises the following steps of: in response to multiple groups of tax form data, analyzing a characteristic field numerical value to generate a unique request identification code; carrying out numerical processing on the feature field to generate a dynamic partition sub-identifier, and aggregating to generate a dynamic partition combination identifier; matching a target data partition corresponding to the dynamic partition combination identifier through a constructed dynamic partition mapping table, and importing structured form data to form a dynamic partition database; constructing an asynchronous task key value pair, writing the asynchronous task key value pair into an asynchronous message queue, regularly scanning the asynchronous message queue, and when a to-be-processed task is detected, identifying whether the task queue contains a priority mark or not, and dynamically adjusting the execution priority of the task queue; and implementing multi-dimensional tax data query from the dynamic partition database based on real-time query conditions and the optimized task queue. The tedious and low-efficiency manual operation is avoided, and the processing efficiency of the form and the tax data is improved.
Owner:广州泓财科技有限公司

Partitioning and hierarchical caching-based track point visual rendering method adaptive to localization

The invention discloses a localization-adaptive track point visual rendering method based on partitioning and hierarchical caching, and the method comprises the steps: S1, inputting a track data stream, carrying out the reading according to the source type of the data stream, and carrying out the data partitioning and LOD processing of the track point data; s2, performing data analysis and rendering on the data; s3, carrying out batch processing on GPU vertexes, according to the state of the circular buffer, if the vertexes are writable, incrementally writing vertex data, and if Wrap is needed, carrying out loopback writing and updating a pointer; and S4, finally, the descriptors and the styles are updated, rendering increment drawing is carried out on the track points, and visual display is carried out. The method supports real-time rendering and playback of ten-million-level track points, has visual and instantiated rendering of window dynamic loading, space-time partitioning, hierarchical caching and WebWorker decoupling calculation, and improves the space-time data rendering capability of the domestic autonomous controllable field by ten thousand times; and the user experience is greatly improved by a non-perceptual interaction and transition smoothing algorithm.
Owner:NANJING HONGSONG INFORMATION TECH CO LTD

Heterogeneous data automatic cleaning method and system based on configurable rule engine

The invention relates to the technical field of data cleaning, in particular to an automatic heterogeneous data cleaning method and system based on a configurable rule engine. The method comprises the following steps: acquiring a heterogeneous data source, uniformly converting a data mode and an instance into a multi-dimensional heterogeneous fact graph, calculating an execution weight for each conflict rule according to the authority of the data source and the priority of the rule when a write operation conflict is detected, segmenting data partitions by utilizing the betweenness centrality of a connected component and a node, the method comprises the following steps of: generating a plurality of data partitions, forming rule subsets corresponding to the data partitions, creating isolated Drools session instances to perform parallel cleaning, generating traceability logs in the parallel cleaning process, aggregating the traceability logs after all the data partitions are cleaned, and re-evaluating and updating the authority of each data source to solve rule conflicts in the next cleaning period. According to the scheme, the writing conflict can be precisely processed, the safety of parallel cleaning is improved, the processing period is shortened, and a continuously optimized processing closed loop is formed.
Owner:XIAN MINGFU CLOUD COMPUTING CO LTD

Heterogeneous perception traffic prediction method based on federated learning

The invention discloses a heterogeneous perception traffic prediction method based on federated learning, and the method employs federated learning to design a unified heterogeneous perception framework, and supports an existing centralized traffic prediction model. Clients with similar traffic flow data distribution are gathered together by utilizing multi-dimensional positive sample comparative learning, so that the clients of the same kind can cooperatively train a model, and the influence of data isomerism between different clients is avoided; the model of each stage is trained in sequence based on a time window by using data partition, so that the influence of data missing on traffic prediction is reduced; noise detection is used for global detection and local denoising, so that the quality of client data is ensured.
Owner:ZHEJIANG UNIV

Real estate registration service system construction method based on multi-tenant cloud architecture mode

The invention relates to the technical field of resource sharing, in particular to a real estate registration service system construction method based on a multi-tenant cloud architecture mode. The method comprises the following steps: based on a unified data structure, ensuring the consistency and compatibility of data of each tenant, and promoting data sharing and exchange among the tenants; province, city and county real estate registration departments are divided into different tenants, the jurisdiction range and service authority of the tenants are defined, and data isolation among the tenants is ensured through a role-based access control mechanism; a data partition strategy is adopted, unique voucher codes are distributed for registration information of different tenants, and partition storage is carried out in a shared data pool according to the voucher codes; the method comprises the following steps: collecting real estate registration approval cases, storing the cases into a case library, establishing a monitoring index system containing approval workflow node processing time and circulation time, setting an abnormal threshold, monitoring data in real time, and triggering an early warning notification when the data exceed the threshold.
Owner:广西壮族自治区自然资源信息中心 +1

Multi-node database synchronization method based on data division

The invention relates to the technical field of database synchronization, in particular to a multi-node database synchronization method based on data partitioning, which comprises the following steps of: 1, performing data partitioning and node role allocation; 2, synchronization mechanism configuration, wherein the synchronization mechanism configuration comprises master and slave node copy setting and data initialization and import; 3, incremental data synchronization, asynchronous / semi-synchronous replication and conflict detection; 4, using a protocol to optimize large file and mass data transmission; step 5, performing parallel processing on the fragmented data at a plurality of nodes; and step 6, carrying out redundant backup on the data fragments at a plurality of nodes. According to the invention, data division and node role allocation are carried out; and data distribution: routing the change event to a corresponding node according to a fragmentation rule, or broadcasting a write set through a group communication protocol, carrying out redundant backup on the data fragmentation at a plurality of nodes, and finally checking the consistency of the fragmentation data regularly, so that the method has the advantages of being capable of realizing efficient synchronization of the multi-node database, giving consideration to the performance and the consistency and the like.
Owner:CHINESE PEOPLES LIBERATION ARMY UNIT 63610

Method and system for enhanced LLM analysis with application to accounting data

A method and system for enhanced financial data processing using large language models (LLMs) to transform PDF statements into structured accounting data. The invention preprocesses statements with line numbering, character position tagging, and data partitioning to overcome LLM context limitations. The system extracts transactions, reconciles them with statement balances, normalizes vendor information, and generates enhanced output files compatible with standard accounting platforms. This enhanced transaction data collapses reconciliation and categorization times in downstream accounting systems. The system creates a robust vendor-transaction database supporting direct analysis or simultaneous application of multiple accounting models.
Owner:DYE DANAMICHELE BRENNEN

Tibetan word segmentation method and system based on multi-language pre-training model CINO

The invention relates to the technical field of data processing, and particularly discloses a Tibetan word segmentation method and system based on a multi-language pre-training model CINO, and the method comprises the steps: collecting a to-be-labeled data set, and providing massive text resources for subsequent research; then word segmentation conversion is carried out to obtain a to-be-trained data set, the text is converted into a lexical element sequence suitable for model processing, a model learning structure is facilitated, attribute parameters of the to-be-trained data set are analyzed, whether data division is carried out or not is judged accordingly, and reasonable division can guarantee representativeness of a training set and a verification set, avoid data distribution deviation and improve the generalization ability of the model; then a training and verification data set is obtained through division and used for training a multi-language pre-training model CINO, training process parameters are collected and analyzed, the model training state can be informed, strategies and hyper-parameters can be adjusted in time, and model initialization is completed, so that Tibetan word segmentation accuracy and reliability are promoted to be improved, and application of the multi-language processing technology in the Tibetan field is assisted.
Owner:TIBET UNIV

Enterprise-specific language model training techniques

Various embodiments of the present disclosure provide a language model training technique. The language model training technique may include a data blending preprocessing step to improve the performance of the language model at an enterprise level. The data blending technique includes receiving an enterprise data partition from a plurality of enterprise data partitions associated with an enterprise data source, receiving a domain-specific data partition from a plurality of domain-specific data partitions associated with one or more domain data sources that are different than the enterprise data source, storing the enterprise data partition as an initial training partition of a plurality of balanced training partitions within a balanced training dataset, and generating a balanced training partition by appending a portion of the domain-specific data partition to the initial training partition. A domain-specific language model may then be trained based on the balanced training dataset.
Owner:OPTUM INC

HBase spatio-temporal data storage indexing method and system based on spatio-temporal short codes

The invention provides an HBase spatio-temporal data storage indexing method and system based on spatio-temporal short codes, and relates to the technical field of big data, time slice short codes and spatial short codes are respectively generated according to time information and a spatial grid; combining the time slice short code, the space short code, the object identifier and the unique identifier to construct a plurality of data partition models, and pre-partitioning the data volume required to be stored in the Hbase database according to the data partition models; and performing spatial indexing or row key indexing according to the row keys, the time slice short codes and the spatial short codes in the various data partition models to obtain an indexing result. According to the invention, the storage query efficiency can be improved.
Owner:HEFEI YINGZE INFORMATION TECH CO LTD

Master-slave device write operation communication method and communication system

The invention provides a master-slave device write operation communication method and communication system, and the method comprises the steps: a master device sends a shared address corresponding to a current write operation, so that a slave device confirms whether the current write operation is matched or not; and the master device sends a write operation command, the identifier set, the write address and the write data in sequence, so that the target slave devices whose own identifiers belong to the identifier set obtain the target data according to the identifier arrangement sequence in the identifier set and the target data length of each target slave device, and data write is completed. A plurality of slave devices share one address and are distinguished in combination with identifiers, so that address resource waste is avoided, and the upper limit of the number of deployable slave devices is improved. In the single I2C transaction, based on the identifier arrangement sequence and the target data length, write data partitioning processing is realized, the occupation and delay of a bus are reduced, and the communication efficiency is improved. The participation sequence and any combination of the slave devices can be dynamically specified, additional mechanisms such as mapping tables are not needed, and the hardware complexity and power consumption of the slave devices are prevented from being increased.
Owner:PUYA SEMICON SHANGHAI CO LTD

Data query method and device of storage system, equipment and medium

The invention discloses a data query method and device for a storage system, equipment and a medium. The method comprises the steps of obtaining an address interval of to-be-queried data; determining at least one target partition corresponding to the address interval according to pre-stored data partition information of the storage system; when data query is performed on a first partition in the at least one target partition, determining a next to-be-queried second partition corresponding to the first partition according to a pre-stored transition probability matrix, and performing data pre-reading on the second partition to obtain at least one query result, elements in the transition probability matrix are used for representing pre-reading probabilities among data in the storage system; and summarizing the at least one query result to obtain a target query result corresponding to the to-be-queried data. The range query performance is obviously optimized and the random query delay is reduced by indexing at least one partition, and in addition, the data pre-reading is realized through the transition probability matrix, so that the query efficiency can be effectively improved and the access waiting time is reduced.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Compilation optimization method and device for accelerating Attention calculation

The invention discloses a compilation optimization method and device for accelerating Attention calculation, and belongs to the field of compilation optimizing.The method comprises the following steps that multiple scheduling strategies are generated according to the number, shape and Attention type of input matrixes and a target hardware platform; defining a cost function, evaluating time and space overhead of each scheduling strategy, and selecting an optimal strategy based on hardware platform characteristics and user requirements; the method comprises the following steps of: performing multi-core parallel division on input data in sequence length and Head dimension, and reducing delay through asynchronous communication overlapping calculation and communication; in a training stage, a random floating point matrix generated by hardware and an attention weight are used for carrying out Hadamard product operation, and a Dropout function is realized. According to the method, the characteristics of different hardware platforms can be flexibly adapted through the generated Attention operator template, excellent performance can be achieved under various shapes, and efficient data division is achieved by optimizing balance between calculation and communication.
Owner:BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD

Data attribute-aware storage system and data management method for key-value database

The present application relates to the technical field of storage, and discloses a data attribute-aware storage system and data management method for a key-value database. The method comprises two parts, i.e., data hotness attribute-based writing and data density attribute-based merging. The data hotness attribute-based writing comprises classifying key-value pairs into cold key-value pairs and hot key-value pairs, to realize separate storage of cold data and hot data by means of a data hotness attribute-based writing strategy. The data density attribute-based merging comprises: when key-value pairs in an SSTable are extracted during a merging operation and form a sorted key-value pair sequence, extracting the key-value pairs one by one, and calculating a binary difference between every two adjacent key-value pairs; and if the difference is greater than a set threshold, stopping filling the current SSTable, and creating a new SSTable. In the present invention, by analyzing data hotness and data density, data having different attributes are classified to different data partitions, thereby reducing I / O amplification caused by repeated reading and writing in a compression process.
Owner:SHANDONG UNIV +1