Compression storage method, decompression method and related device for population data
By generating personal vector representations and identifying associations during compression coding, using deep learning models and multi-head attention mechanisms, the distortion problem in crowd data compression is solved, and efficient data compression and accurate data restoration are achieved.
Patent Information
- Application Number
- CN202510107938.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art has severe compression distortion in the process of crowd data compression, resulting in a large amount of information loss, especially in large-scale, high-dimensional and complex-related crowd data compression and storage. Both traditional methods and deep learning-based methods have failed to effectively capture and utilize the association relationship between data.
By generating personal vector representations and identifying the association relationship between multiple personal vector representations during the compression coding process, deep learning models such as the Transformer model and the multi-head attention mechanism are used to generate coded data containing association relationships, and dimensionality reduction is performed in combination with a linear transformation layer.
It weakens compression distortion, improves data compression efficiency and decompressed data restoration quality, and retains key information and complex relationships in crowd data.
Smart Images

Figure CN119543956B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in the present application relate to the technical field of data compression and storage, and particularly to a method for compressing and storing population data, a decompression method, and related devices. Background Art
[0002] With the rapid development of big data technology and the continuous growth of people's demand for data mining, in the field of e-commerce marketing, the acquisition of population data has become a research hotspot. Population data has the characteristics of high collection frequency and large data scale. Since storing a large amount of population data will occupy a large amount of hardware resources, and a high bandwidth resource will be consumed during the transmission of population data, data compression processing is usually required before storing and transmitting population data.
[0003] Traditional data compression methods such as Huffman coding, arithmetic coding, and LZ compression algorithm mainly perform data compression based on the statistical characteristics of data to achieve efficient compression of population data. However, after compressing population data in the prior art, there is relatively serious compression distortion, resulting in a large amount of information loss in the compressed population data. Summary of the Invention
[0004] In view of this, multiple embodiments of the present application are dedicated to providing a method for compressing and storing population data, a decompression method, and related devices, which can weaken compression distortion to a certain extent.
[0005] An embodiment of the present application provides a method for compressing and storing population data, where the population data includes multiple individual data, and the method includes: generating an individual vector representation corresponding to the individual data; performing compression encoding on the individual vector representation to obtain encoded data; wherein, during the process of performing compression encoding on the individual vector representation, identifying the association relationship between multiple individual vector representations, and enabling the information contained in the encoded data to express the association relationship; storing the encoded data corresponding to the population data.
[0006] An embodiment of the present application further provides a method for decompressing population data, including: receiving the encoded data as described in the method for compressing and storing population data as described above; during the process of decoding the encoded data, decoding to obtain the multiple individual vector representations according to the association relationship of the multiple individual vector representations expressed by the encoded data.
[0007] One embodiment of the present application also provides a compression storage device for population data. The population data includes multiple personal data. The compression storage device includes: a vector generation module for generating a personal vector representation corresponding to the personal data; an encoding module for performing compression encoding on the personal vector representation to obtain encoded data; wherein, during the process of performing compression encoding on the personal vector representation, the association relationship between multiple personal vector representations is identified, and the information contained in the encoded data can express the association relationship; a storage module for storing the encoded data corresponding to the population data.
[0008] One embodiment of the present application also provides a decompression device for population data. The decompression device includes: a receiving module for receiving the encoded data as described in the compression storage method of the population data as mentioned above; a decoding module for decoding the multiple personal vector representations according to the association relationship of the multiple personal vector representations expressed by the encoded data during the process of decoding the encoded data.
[0009] One embodiment of the present application also provides a computer device. The computer device includes a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method as described above.
[0010] One embodiment of the present application also provides a computer-readable storage medium. At least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, it can implement the method as described above.
[0011] One embodiment of the present application also provides a computer program product for implementing the method as described above.
[0012] In multiple embodiments provided by the present application, by generating a personal vector representation corresponding to the personal data and identifying the association relationship between multiple personal vectors during the compression encoding process, the encoded data can effectively express these association relationships, thereby reducing the compression distortion of the population data. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the application scenario of the compression storage method of population data provided by one embodiment of the present application.
[0014] Figure 2 It is a flowchart of the compression storage method of population data provided by one embodiment of the present application.
[0015] Figure 3 It is a schematic diagram of the generation of personal vector representation provided by one embodiment of the present application.
[0016] Figure 4 Functional schematic diagram of the multi-head attention mechanism provided for an embodiment of the present application.
[0017] Figure 5 Functional schematic diagram of the linear transformation layer provided for an embodiment of the present application.
[0018] Figure 6 Flowchart of the decompression method for population data provided for an embodiment of the present application.
[0019] Figure 7 Functional schematic diagram of the decompression device for population data provided for an embodiment of the present application.
[0020] Figure 8 Flowchart of the training method for the specified encoding model provided for an embodiment of the present application.
[0021] Figure 9 Module schematic diagram of the compression storage device for population data provided for an embodiment of the present application.
[0022] Figure 10 Module schematic diagram of the decompression device for population data provided for an embodiment of the present application.
[0023] Figure 11 Schematic diagram of a computer device provided for an embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0025] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0026] In the related art, in the field of crowd data compression and storage, traditional data compression methods usually rely on the statistical characteristics of data and reduce the data volume by eliminating redundant information. Specifically, traditional data compression methods such as Huffman coding, arithmetic coding, and LZ compression algorithm mainly achieve compression by analyzing the frequency distribution and repeated patterns of data. Although these methods can reduce the data storage space and optimize the transmission bandwidth to a certain extent, due to their mainly being based on the statistical characteristics of data, they often ignore the semantics and context relationships contained in crowd data.
[0027] With the rapid development of big data technology and the diversification of data collection means, the scale and complexity of crowd data have shown an explosive growth. When traditional data compression methods face massive and diverse crowd data, usually in order to reduce the occupation of storage space, a relatively high compression ratio will be adopted for crowd data. However, such a compression and storage method will cause some key information to be lost during the compression process. Further, there will also be a certain amount of information loss during the decompression and recovery process, resulting in a large amount of key information being lost in the finally decompressed crowd data.
[0028] In the related art, deep learning technology can be used to perform data compression. In particular, data compression methods based on neural networks have shown better effects in capturing complex relationships in data and performing efficient feature extraction. However, data compression methods based on deep learning mostly focus on specific data types such as images and texts, and the research on the compression of crowd data, which is a multi-dimensional and multi-attribute data type, is not yet sufficient. Especially in terms of how to effectively capture and utilize the correlation relationships between data to improve the compression efficiency and data restoration accuracy, there are still deficiencies.
[0029] In summary, in the related art, there is a problem of compression distortion of crowd data in dealing with the compression and storage of large-scale, high-dimensional, and complexly correlated crowd data.
[0030] In multiple embodiments provided in the present application, the compression and storage method of crowd data can be applied to a compression and storage device for crowd data. The compression and storage device for crowd data can be an electronic device with certain computing power and network access capabilities. The electronic device can be a desktop computer, a laptop computer, a tablet computer, a smart phone, or a server. The server can also be a distributed server, including multiple processors, memories, network communication modules, etc., which cooperate to achieve various functions. Or, the server can also be a server cluster formed by several servers, which has higher computing and data processing capabilities. With the development of science and technology, the server can also be realized by new form of technical means, such as a new type of "server" based on quantum computing. Of course, in some embodiments, the compression and storage device for crowd data can also be a program module running in an electronic device.
[0031] Please refer to Figure 1 An application scenario example of a method for compressing and storing population data is provided in an embodiment of the present application. The compression and storage method is applied to a device for compressing and storing population data to achieve efficient compression and storage of user data on an e-commerce platform, and can reduce compression distortion during the compression and storage of population data.
[0032] For example, in a large e-commerce platform, when users conduct shopping activities through a client or website, a large amount of personal data will be generated. These personal data include users' basic information (such as age, gender, region), consumption behaviors (such as browsing history, purchase records, shopping cart content), and interest preferences (such as categories of favorite products, evaluation content, etc.). These data need to be effectively stored and analyzed to support precise marketing strategies and personalized recommendation systems.
[0033] In this scenario example, the large e-commerce platform can manage the large amount of personal data generated through a marketing data management system. Specifically, the server of the marketing data management system can run a data collection service for automatically collecting personal data, and this data collection service can call the compression and storage device to process the large amount of personal data collected. After receiving the personal data, the compression and storage device generates a personal vector representation, identifies the correlation relationships between the personal vector representations during the compression encoding process, and finally generates encoded data and stores it in a database for subsequent data analysis, user segmentation, and the formulation of precise marketing strategies.
[0034] In this scenario example, by calling the compression and storage device to compress a large amount of personal data of e-commerce platform users and storing the compressed encoded data in a database, it is possible to identify complex correlation relationships in personal data during the data compression process, reduce compression distortion during the compression process, and retain more information of the compressed population data.
[0035] Please refer to Figure 2 An embodiment of the present application provides a method for compressing and storing population data. The method for compressing and storing population data may include the following steps:
[0036] Step S110: Generate a personal vector representation corresponding to personal data.
[0037] Step S120: Perform compression encoding on the personal vector representation to obtain encoded data; wherein, during the process of performing compression encoding on the personal vector representation, identify the correlation relationships between multiple personal vector representations, and enable the information contained in the encoded data to express the correlation relationships.
[0038] Step S130: Store the encoded data corresponding to the population data.
[0039] In this embodiment, the population data may include personal data of multiple users, and the personal data of a user may include information and statistical data related to the user. For example, the user's basic information, consumption behavior, and interest preferences, etc. The personal data can be used to analyze, understand, and segment target customers. Alternatively, the population data can also generate samples for training a machine learning model, which are used to train the machine learning model. After generating a personal vector representation corresponding to the personal data, the personal vector representation can effectively represent the relevant information of each user.
[0040] Specifically, the compression storage device can convert high-dimensional and sparse personal data into low-dimensional and dense personal vector representations through a deep learning model, so as to facilitate subsequent compression coding processing. The compression storage device can identify the correlation relationships between multiple personal vector representations through the deep learning model and perform compression coding on the personal vector representations to obtain coded data, so that while the coded data occupies less storage space, it can also retain the complex correlation relationships and semantic messages between the personal data.
[0041] In this embodiment, there will be certain correlation information between multiple personal data, and this correlation information plays an important role in the subsequent usage process. In this embodiment, taking the Transformer model as an example of the deep learning model, during the process of compressing and coding the personal vector representations of multiple personal data, the Transformer model will learn the correlation relationships between the personal vector representations of multiple personal data and can form corresponding coded data according to the learned knowledge. As a result, the finally obtained coded data can carry the said correlation relationships, realizing a reduction in data compression distortion. Further, after decoding the coded data, the vector representation of the personal data can be restored more accurately. In some embodiments, for example, the personal data may include, but is not limited to, the user browsing product information, purchasing products, and evaluating products through the client or website of an e-commerce platform.
[0042] In this embodiment, the compression storage device can obtain personal data such as the basic information, browsing records, purchase records, and product evaluations of multiple users. Further, the compression storage device can generate personal vector representations corresponding to the personal data. Specifically, the compression storage device can generate personal vector representations corresponding to the personal data based on machine learning algorithms. For example, the machine learning algorithms may include, but are not limited to, the BERT algorithm or the Word2Vec algorithm, etc.
[0043] In some embodiments, the encoded data can be stored in a distributed database. Specifically, the compression storage device can obtain the personal information of e-commerce platform users through the marketing data management system of the e-commerce platform, and send the encoded data generated corresponding to the user personal information to the marketing data management system. The marketing data management system can store the encoded data in the distributed databases established by the e-commerce platform in different regions, so as to store the data close to the users, reduce access latency, and improve the user experience. The e-commerce platform can transfer the encoded data in the distributed databases in different regions to the same marketing database for further data analysis, user segmentation, and formulation of precise marketing strategies.
[0044] In some embodiments, the encoded data can also be stored in a cloud storage system. Specifically, the compression storage device can call a remotely deployed deep learning model server or a cloud computing platform to generate vector representations and perform compression encoding on the personal data of users in the cloud, and finally store the encoded data in the cloud storage system. By performing data compression and storage through the cloud platform, not only can the pressure on local storage be significantly reduced, but also the computing power of the cloud can be utilized to accelerate the data compression process.
[0045] In multiple embodiments of the present application, by generating personal vector representations and identifying the association relationships between personal vectors during the compression encoding process, the encoded data can effectively express these association relationships, thereby improving the efficiency of data compression and the quality of data restoration after decompression.
[0046] In some embodiments, the compression storage device for crowd data can generate a primary vector representation of personal data; assign a position vector representation indicating the position order to the primary vector representation; and combine the primary vector representation and the position vector representation to obtain the personal vector representation.
[0047] Please refer to Figure 3 . In this embodiment, the compression storage device can receive the personal data of multiple users in the crowd data, and convert the multiple personal data (such as the consumption behaviors and interest preferences of users) included in the crowd data into primary vector representations through an embedding matrix. Specifically, the embedding matrix can map each word in the personal data to a low-dimensional continuous vector space, thereby capturing the user information contained in the personal data. By using the embedding matrix, the compression storage device can convert the high-dimensional and sparse personal data into low-dimensional and dense primary vector representations.
[0048] Next, the compression storage device can arrange the order of these primary vector representations according to the similarity between multiple primary vector representations. Specifically, the measurement of similarity can be based on the similarity of user basic information, the commonality of behavior patterns, or other relevant features. In some embodiments, for example, the ages of user A and user B are both in the age range of 20 to 25 years old, while the age of user C is in the age range of 30 to 35 years old. Then, among the primary vector representations corresponding to the personal data of user A, user B, and user C, the similarity between the primary vector representations of user A and user B representing age information is higher than the similarity between the primary vector representations of user A and user C or user B and user C representing age information. According to the arranged order, the compression storage device can assign a position vector representation indicating its relative position to each primary vector representation. Specifically, the compression storage device can assign the corresponding position vector representation to each primary vector representation by means of position encoding. The position vector representation contains the relative position information of different primary vector representations, which can help the deep learning model understand the order and structure of the data.
[0049] Finally, the compression storage device can combine the primary vector representation with the corresponding position vector representation by means of vector addition or concatenation to generate the final personal vector representation. By combining the primary vector representation and the corresponding position vector representation, the generated personal vector representation not only contains the personal information of each user but also reflects its relative position relationship in the dataset, providing richer and more ordered data information for the subsequent compression encoding process.
[0050] By adopting an embedding matrix to generate the primary vector representation of personal data, assigning a position vector representation indicating its relative position to each primary vector representation, and finally combining the primary vector representation and the corresponding position vector representation to obtain the personal vector representation, this embodiment can extract the rich information contained in the user's personal data into the personal vector representation to improve the integrity of user information during the data compression process.
[0051] In some embodiments, the compression storage device for crowd data can arrange the order of multiple primary vector representations according to the similarity between multiple primary vector representations; and assign a position vector representation expressing the relative position to the primary vector representations according to the arranged order of the multiple primary vector representations.
[0052] In this embodiment, the compression storage device can first arrange the order of these primary vector representations according to the similarity between multiple primary vector representations. Specifically, the compression storage device can calculate the similarity between multiple primary vector representations based on the cosine similarity between vectors, Euclidean distance, or other appropriate similarity measurement methods. In some embodiments, for example, the primary vector representations generated from the personal data of user A and user B are located in a three-dimensional vector space, and the corresponding coordinates are and , where and can be used to represent the user's basic information such as age, and can be used to represent the user's consumption behavior such as the total consumption amount of the user on the e-commerce platform within a week, and can be used to represent the user's interest preferences such as the number of times the user searches for products of a specific category within a month. The similarity between the primary vector representations of User A and User B can be obtained through the Euclidean distance calculation formula:
[0053]
[0054] By measuring the similarity between multiple primary vector representations, the compression storage device can identify personal data with high correlation or similar behavior patterns in the feature space, making the primary vector representations with high similarity closer in the arrangement order.
[0055] Secondly, the compression storage device can assign a unique position vector representation to each primary vector representation through the method of position encoding, and this position vector representation reflects its relative position in the arrangement order. Specifically, the compression storage device can generate a corresponding position vector representation for each primary vector representation according to the order of multiple primary vector representations arranged, using fixed position encoding or trainable position encoding methods generated by sine and cosine functions. By sorting and position encoding the primary vector representations, the compression storage device can assign accurate and ordered position vector representations to each primary vector representation, and then generate a final personal vector representation containing rich position information. Specifically, the sorting based on similarity can ensure the rationality and relevance of the arrangement of primary vector representations, and the position vector representation can provide necessary position information support, so that the final personal vector representation not only contains the user's personal information but also reflects its relative position relationship in the dataset. This structured vector representation method can provide a more ordered and more relevant data basis for the subsequent compression encoding process, thereby improving the efficiency of data compression and the quality of data restoration after decompression.
[0056] This embodiment arranges the order according to the similarity between primary vector representations, and assigns a position vector representation representing the relative position to the primary vector representations in this order, and then generates a structured and rich-information personal vector representation, so that the population data can retain key user information and association relationships during the compression and decompression processes.
[0057] In some embodiments, the compression storage device for crowd data inputs multiple personal vector representations into a specified encoding model with a multi-head attention mechanism to identify the correlation relationships between the multiple personal vector representations through the multi-head attention mechanism.
[0058] Please refer to Figure 4 . In this embodiment, the compression storage device can input multiple personal vector representations into a specified encoding model with a multi-head attention mechanism. The specified encoding model typically adopts a deep learning architecture (such as Transformer Encoder), and its core components include a multi-head attention mechanism and a feed-forward neural network. The multi-head attention mechanism allows the specified encoding model to concurrently focus on different subspace features of the input vectors in different attention heads, thereby enabling it to capture the multi-level and multi-dimensional correlation relationships between the personal vector representations. Specifically, the compression storage device can use multiple personal vector representations as the input sequence. After being processed by the multi-head attention mechanism, each attention head can independently calculate the attention weights between the input vectors and generate the corresponding attention output. Through the multi-head attention mechanism, the specified encoding model can identify and model the complex correlation relationships between the personal vector representations, such as user similarities, commonalities in behavior patterns, or other relevant features.
[0059] Next, the compression storage device can transfer the output data processed by the multi-head attention mechanism to the feed-forward neural network for further non-linear transformation and enhancement of the features. The feed-forward neural network can improve the expressive power of the feature representation and the deep learning ability of the specified encoding model through a series of linear transformations and activation functions. The data processed by the feed-forward neural network will be transferred to the linear transformation layer for dimensionality reduction of the high-dimensional features to generate the final encoded data. The linear transformation layer uses a low-dimensional linear matrix to perform dimensionality reduction on the personal vector representations processed by the multi-head attention mechanism and the feed-forward neural network to achieve further compression of the data. This dimensionality reduction process can not only effectively reduce the storage space requirements but also enable the encoded data to contain sufficient semantic and correlation information to support the subsequent decompression and restoration process.
[0060] In this embodiment, by inputting multiple personal vector representations into a specified encoding model with a multi-head attention mechanism, using the multi-head attention mechanism to identify the correlation relationships between the personal vector representations, and implementing dimensionality reduction and compression of the data through the linear transformation layer, it not only improves the efficiency of crowd data compression but also retains the complex correlation relationships in the crowd data, thereby improving the data restoration quality after decompression.
[0061] In some embodiments, the specified encoding model includes a linear transformation layer; the linear transformation layer uses a low-dimensional linear matrix to perform dimensionality reduction on the personal vector representations processed by the multi-head attention mechanism to obtain encoded data.
[0062] Please refer to Figure 5 . In this embodiment, the compression storage device reduces the dimension of the personal vector representation processed by the multi-head attention mechanism through the linear transformation layer in the specified encoding model. The linear transformation layer may include one or more low-dimensional linear matrices, which can be used to convert the high-dimensional personal vector representation into low-dimensional encoded data. The compression storage device can implement matrix multiplication operations through the low-dimensional linear matrices to map the high-dimensional vectors output by the multi-head attention mechanism into a lower-dimensional space, thereby effectively reducing the data dimension and storage space requirements. The design of the low-dimensional linear matrices can be based on the parameters of the pre-trained specified encoding model, so as to retain as much as possible the key information and correlation relationships in the original population data.
[0063] Next, the compression storage device can further process the dimension-reduced personal vector representation through an activation function and normalization. The activation function can be a linear or non-linear activation function (such as Relu, Sigmoid, etc.) to introduce non-linear features and enhance the expressive power of the encoded data and the fitting ability of the model. The compression storage device can normalize the dimension-reduced personal vector representation through normalization, so that the encoded data has consistent scale and distribution characteristics between different dimensions, so as to improve the data stability and reliability in the subsequent storage and decompression processes.
[0064] In some embodiments, the compression storage device can adopt a trainable low-dimensional linear matrix, and continuously optimize the parameters of the low-dimensional linear matrix during the training process of the specified encoding model through the backpropagation algorithm to further improve the dimension reduction effect of the personal vector representation and the quality of the encoded data. In addition, the compression storage device can also flexibly adjust the dimension of the dimension reduction and the structure of the low-dimensional linear matrix according to different application requirements and data characteristics to improve the compression effect of the population data and the restoration accuracy during decompression.
[0065] This embodiment introduces a linear transformation layer in the specified encoding model and uses low-dimensional linear matrices to reduce the dimension of the personal vector representation processed by the multi-head attention mechanism to generate information-rich encoded data. This method can improve the efficiency of population data compression and retain the key information and correlation relationships in the original population data, so as to accurately restore the key information and correlation relationships of the original population data during the decompression and recovery process of the encoded data.
[0066] Please refer to Figure 6 . This embodiment of the present application also provides a method for decompressing population data. The method for decompressing population data may include the following steps:
[0067] Step S210: Receive the encoded data as described in any of the foregoing.
[0068] Step S220: During the process of decoding the encoded data, according to the association relationships expressed by the multiple personal vector representations in the encoded data, decode to obtain the multiple personal vector representations.
[0069] In this embodiment, the decompression method of population data can be applied to a decompression device for population data. The decompression device for population data can be an electronic device with certain computing power and network access capabilities. The electronic device can be a desktop computer, a laptop computer, a tablet computer, a smart phone, or a server. The server can also be a distributed server, including multiple processors, memories, network communication modules, etc., which cooperate to implement various functions. Or, the server can also be a server cluster formed by several servers, with higher computing and data processing capabilities. With the development of science and technology, the server can also be implemented by new forms of technical means, such as a new type of "server" based on quantum computing. Of course, in some embodiments, the decompression device for population data can also be a program module running on an electronic device.
[0070] Please refer to Figure 7 In this embodiment, the decompression device for population data can receive the encoded data stored in a distributed database or a cloud storage system. The encoded data refers to the personal vector representations after compression encoding processing, which contains the association relationship information between multiple personal vectors. This association relationship information can be used to accurately restore the personal vector representations during the decompression process. The decompression device for population data can be based on a deep learning model (such as the Decoder of the Transformer model), and according to the association relationships expressed by the multiple personal vector representations in the encoded data, perform decoding processing on the received encoded data. In some embodiments, the decompression device for population data can use the Decoder of the Transformer model to analyze the similarity and dependence between the personal vector representations in the encoded data, and through the multi-head attention mechanism and the inverse linear transformation process, restore the original high-dimensional personal vector representations, so that the decompressed personal vector representations can accurately reflect the characteristics and association relationships of the original personal data.
[0071] In some embodiments, the decompression device for population data can achieve efficient data decompression by calling a locally or remotely deployed deep learning model server. Specifically, an enterprise can use the computing resources of a cloud service provider to download the encoded data stored in the cloud to the local, use the locally deployed deep learning model to decompress the data, and finally use the restored personal vector representations for data analysis, user segmentation, and the formulation of precise marketing strategies. This method can not only make full use of the computing power of the cloud to improve the decompression efficiency of population data, but also enhance the security of population data during the decompression process.
[0072] In this embodiment, by receiving the encoded data and based on the association relationship represented by multiple personal vector representations in the encoded data, the received encoded data is decoded based on a deep learning model to restore the original personal vector representation. This method can preserve key features and association relationships during the decompression process of population data, thereby improving the accuracy of population data decompression.
[0073] Please refer to Figure 8 This embodiment of the present application also provides a training method for a specified encoding model. The specified encoding model can adopt a deep learning architecture (such as a Transformer model), and the training method of the specified encoding model can include the following steps:
[0074] Step S310: Preprocess the collected population data to construct a training set.
[0075] Step S320: Train the specified encoding model based on the training set.
[0076] In this embodiment, the training method of the specified encoding model can be applied to a training device for the specified encoding model. The training device for the specified encoding model can be an electronic device with certain computing power and network access capabilities. This electronic device can be a desktop computer, a laptop computer, a tablet computer, a smart phone, or a server. The server can also be a distributed server, including multiple processors, memories, network communication modules, etc., that cooperate to implement various functions. Alternatively, the server can also be a server cluster formed by several servers, which has higher computing and data processing capabilities. With the development of science and technology, the server can also be implemented using new forms of technical means, such as a new type of "server" based on quantum computing. Of course, in some embodiments, the training device for the specified encoding model can also be a program module running on an electronic device.
[0077] In this embodiment, the training device for the specified encoding model can widely collect a large amount of personal data generated when users shop, browse, etc. on the e-commerce platform, including the basic information of users (such as age, gender, region), consumption behaviors (such as browsing history, purchase records, shopping cart content), and interest preferences (such as categories of collected goods, evaluation content, etc.) as the original population data. Next, the training device for the specified encoding model can preprocess the collected original population data: for example, the Z-Score or IQR method can be used to clean the original population data to remove noise and outliers; then, a normalization algorithm such as Min-Max scaling or Z-score standardization can be applied so that different features in the original population data after data cleaning are within the same dimension for subsequent processing; next, strategies such as filling (such as mean, median, etc.) or deletion can be adopted to process the missing values in the normalized population data to improve the integrity of the population data. The training device for the specified encoding model can adopt data augmentation techniques, such as jittering, translation, and rotation, etc., to expand the original population data after preprocessing to obtain a high-quality training set, thereby improving the generalization ability of the specified encoding model. In addition, the training device for the specified encoding model can select a loss function (such as the cross-entropy loss function) to measure the difference between the data in the constructed training set and the original population data, and use the Adam optimizer to train the specified encoding model, adjusting the learning rate and batch size to optimize the training process.
[0078] Please refer to Figure 9 。This embodiment of the present application also provides a compression storage device for population data. The compression storage device for population data may include:
[0079] A vector generation module for generating a personal vector representation corresponding to personal data;
[0080] An encoding module for performing compression encoding on the personal vector representation to obtain encoded data; wherein, in the process of performing compression encoding on the personal vector representation, the association relationship between multiple personal vector representations is identified, and the information contained in the encoded data can express the association relationship;
[0081] A storage module for storing the encoded data corresponding to the population data.
[0082] In this embodiment, the specific functions and effects achieved by the compression storage device for population data can be explained by referring to other embodiments of the present application and will not be elaborated here.
[0083] Please refer to Figure 10 。This embodiment of the present application also provides a decompression device for population data. The decompression device for population data may include:
[0084] A receiving module, configured to receive the encoded data as described above;
[0085] A decoding module, configured to decode the encoded data to obtain the plurality of personal vector representations according to the association relationship expressed by the plurality of personal vector representations in the encoded data during the process of decoding the encoded data.
[0086] Please refer to Figure 11 。This embodiment of the present application further provides a computer device, which includes: a memory and a processor. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method as described above.
[0087] This embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method as described above.
[0088] This embodiment of the present application further provides a computer program product containing instructions. When the computer program product is executed by a processor, the method as described above is implemented.
[0089] The user information or user account information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, etc.) involved in multiple embodiments of the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws and regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0090] It can be understood that the specific examples in this article are only to help those skilled in the art better understand the embodiments of the present application, rather than limiting the scope of the present invention.
[0091] It can be understood that in various embodiments of the present application, the magnitude of the sequence numbers of each process does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0092] It can be understood that the various embodiments described in the present application can be implemented alone or in combination, and the embodiments of the present application do not limit this.
[0093] Unless otherwise specified, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the technical field of this application. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the scope of this application. The term "and / or" used in this application includes any and all combinations of one or more of the related listed items. The singular forms "a", "above-mentioned", and "the" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0094] It can be understood that the processor in the embodiments of this application can be an integrated circuit chip with the ability to process signals. During implementation, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0095] It can be understood that the memory in the embodiments of this application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0096] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0097] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated herein.
[0098] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0099] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0101] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0102] As described above, the above are only specific embodiments of this application, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for compressed storage of population data, characterized in that, The population data includes a plurality of individual data, and the method includes: Generating an individual vector representation corresponding to the individual data; Performing compression encoding on the individual vector representation to obtain encoded data, including: inputting a plurality of individual vector representations into a specified encoding model with a multi-head attention mechanism to identify the correlation relationships between the plurality of individual vector representations through the multi-head attention mechanism; wherein, during the process of performing compression encoding on the individual vector representation, the correlation relationships between the plurality of individual vector representations are identified, and the information included in the encoded data can express the correlation relationships; wherein, the specified encoding model includes a linear transformation layer; the linear transformation layer uses a low-dimensional linear matrix to perform dimensionality reduction processing on the individual vector representation after being processed by the multi-head attention mechanism to obtain the encoded data; wherein, the structure of the low-dimensional linear matrix and the dimensionality reduction dimension are adjusted based on application requirements and data characteristics; Storing the encoded data corresponding to the population data; wherein, the encoded data is stored in a distributed database established by an e-commerce platform in different regions, the encoded data is stored at a location close to the user, and the encoded data is used for formulating marketing strategies.
2. The method according to claim 1, wherein The step of generating an individual vector representation corresponding to the individual data includes: Generating a primary vector representation of the individual data; Allocating a position vector representation indicating the position order to the primary vector representation; Combining the primary vector representation and the position vector representation to obtain the individual vector representation.
3. The method according to claim 2, wherein The step of allocating a position vector representation indicating the position order to the primary vector representation includes: Arranging the order of a plurality of primary vector representations according to the similarity between the plurality of primary vector representations; Allocating a position vector representation expressing the relative position to the primary vector representation according to the arrangement order of the plurality of primary vector representations.
4. A method for decompressing population data, characterized in that, The method includes: Receiving the encoded data as described in any one of claims 1 to 3; During the process of decoding the encoded data, decoding the plurality of individual vector representations according to the correlation relationships between the plurality of individual vector representations expressed by the encoded data.
5. A compression storage device for population data, characterized in that, The population data includes a plurality of individual data, and the compression storage device includes: A vector generation module for generating an individual vector representation corresponding to the individual data; An encoding module for performing compression encoding on the individual vector representation to obtain encoded data, including: inputting a plurality of individual vector representations into a specified encoding model with a multi-head attention mechanism to identify the correlation relationships between the plurality of individual vector representations through the multi-head attention mechanism; wherein, during the process of performing compression encoding on the individual vector representation, the correlation relationships between the plurality of individual vector representations are identified, and the information included in the encoded data can express the correlation relationships; wherein, the specified encoding model includes a linear transformation layer; the linear transformation layer uses a low-dimensional linear matrix to perform dimensionality reduction processing on the individual vector representation after being processed by the multi-head attention mechanism to obtain the encoded data; wherein, the structure of the low-dimensional linear matrix and the dimensionality reduction dimension are adjusted based on application requirements and data characteristics; A storage module for storing encoded data corresponding to the population data; wherein the encoded data is stored in a distributed database established by an e-commerce platform in different regions, the encoded data is stored close to users, and the encoded data is used for formulating marketing strategies.
6. A decompression device for population data, characterized in that, The decompression device includes: A receiving module for receiving the encoded data as described in any one of claims 1 to 3; A decoding module for decoding the encoded data to obtain the plurality of personal vector representations according to the association relationship expressed by the plurality of personal vector representations in the process of decoding the encoded data.
7. A computer device, characterized in that, The computer device includes a memory and a processor, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the method as described in any one of claims 1 to 3.
8. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and when the at least one computer program is executed by a processor, it can implement the method as described in any one of claims 1 to 3.
9. A computer program product, characterized in that, A computer program product is used to implement the method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Transform-based storage data similarity measurement method
CN118095288A