Archive information retrieval and distributed storage method and computer equipment

By generating archive information request vectors and retrieval attribute data, using archive information vector tree and Bloom filter algorithms, matching archive information is filtered out, solving the problem of low accuracy caused by non-standard keyword search, and achieving efficient and accurate archive information retrieval.

CN120371878APending Publication Date: 2025-07-25HEILONGJIANG COMM POLYTECHNIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503212.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the keyword search method can easily filter out the correct file information when the user input is not standard, resulting in low accuracy in file information retrieval.

Method used

By generating archive information request vectors and searching attribute data, using the pre-constructed archive information vector tree for clustering, combining the Bloom filter and the Beam Search algorithm, the set of archive information vectors under the target leaf node matching the search request is selected, and further filtering is performed based on the description attribute data to obtain the final search results.

Benefits of technology

It improves the efficiency and accuracy of archival information retrieval, reduces the amount of invalid data, and improves the overall retrieval efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371878A_ABST
    Figure CN120371878A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of archive information retrieval and distributed storage, in particular to an archive information retrieval and distributed storage method and computer equipment. A specific embodiment of the method comprises the steps of generating an archive information request vector in response to a received archive information retrieval request, and generating archive information retrieval attribute data according to a retrieval condition in the archive information retrieval request; searching an archive information vector set under a target leaf node matched with the archive information request vector from an archive information vector tree according to the archive information request vector; according to the description attribute data of the archive information, retrieving an alternative archive information set matched with the archive information retrieval attribute data from the archive information indicated by the archive information vector set; and according to the archive information retrieval request, performing retrieval processing on the alternative archive information in the alternative archive information set to obtain an archive information retrieval result. According to the embodiment, the efficiency and accuracy of archive information retrieval are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of archival information retrieval and distributed storage, and particularly to archival information retrieval and distributed storage methods and computer devices. Background Art

[0002] Archives are original records of various forms with preservation value directly formed in people's various social activities. Archives include personnel archives, equipment archives, enterprise archives, etc. However, in the face of a vast amount of archival information, how to achieve fast and accurate retrieval has become an urgent problem to be solved. Currently, for retrieving archival information, it is mainly to retrieve based on keywords input by users. Keyword retrieval is based on the principle of text matching. By matching the keywords input by users with the keywords appearing in the archives, the fast positioning of archives is realized. However, this method usually has the following technical problems: when the keywords input by users are not standard, it is easy to filter out correct archival information, resulting in the problem of low accuracy of archival information retrieval. Summary of the Invention

[0003] This section of the present application is used to briefly introduce concepts, which will be described in detail in the following detailed implementation section. This section of the present application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] Some embodiments of the present application propose archival information retrieval and distributed storage methods, computer devices, and computer-readable storage media to solve one or more of the technical problems mentioned in the above background art section.

[0005] In a first aspect, some embodiments of the present application provide an archival information retrieval and distributed storage method, which includes: in response to receiving an archival information retrieval request, generating an archival information request vector, and generating archival information retrieval attribute data according to the retrieval conditions in the above archival information retrieval request, where the above archival information request vector and the above archival information retrieval attribute data represent different archival data information; according to the above archival information request vector, retrieving, from a pre-constructed archival information vector tree, an archival information vector set under a target leaf node that matches the above archival information request vector, where the above archival information vector tree is obtained by clustering archival information vectors of archival information; according to the description attribute data of archival information, retrieving, from the archival information indicated by the above archival information vector set, an alternative archival information set that matches the above archival information retrieval attribute data; and performing a retrieval process on the alternative archival information in the above alternative archival information set according to the above archival information retrieval request to obtain an archival information retrieval result corresponding to the above archival information retrieval request.

[0006] In a second aspect, the present application further provides a computer device. The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the method described in any implementation manner of the first aspect is implemented.

[0007] In a third aspect, the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0008] The above various embodiments of the present application have the following beneficial effects: By pre-constructing an archive information vector tree, rapid retrieval of archive information can be achieved according to the request vector of the archive information retrieval request. Subsequently, according to the retrieval attribute data of the archive information retrieval request and the description attribute data of the archive information, further retrieval and screening can be performed on the retrieved archive information. In this way, invalid archive information that does not meet the user's requirements can be filtered out in advance during the retrieval process. Thereby reducing the amount of subsequent invalid data and resource occupation, which is beneficial to improving the overall retrieval efficiency. Additionally, by performing constraint filtering in advance during the retrieval process, the amount of valid archive information in the vest archive information can be ensured, thus improving the retrieval efficiency and accuracy. First, in response to receiving an archive information retrieval request, an archive information request vector is generated, and according to the retrieval conditions in the archive information retrieval request, archive information retrieval attribute data is generated, where the archive information request vector and the archive information retrieval attribute data represent different archive data information; then, according to the archive information request vector, from the pre-constructed archive information vector tree, an archive information vector set under the target leaf node that matches the archive information request vector is retrieved, where the archive information vector tree is obtained by clustering the archive information vectors of the archive information; then, according to the description attribute data of the archive information, from the archive information indicated by the archive information vector set, an alternative archive information set that matches the archive information retrieval attribute data is retrieved; finally, according to the archive information retrieval request, retrieval processing is performed on the alternative archive information in the alternative archive information set to obtain an archive information retrieval result corresponding to the archive information retrieval request. Thus, the efficiency and accuracy of archive information retrieval are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present application will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0010] Figure 1 is a flowchart of some embodiments of a file information retrieval and distributed storage method according to the present application; Figure 2 is a schematic structural diagram of a computer device suitable for implementing some embodiments of the present application. Detailed implementation manners

[0011] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not used to limit the protection scope of the present application.

[0012] In addition, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0013] It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0014] It should be noted that the modifications of "one" and "plural" mentioned in the present application are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0015] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0016] The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0017] Figure 1 Flow 100 of some embodiments of a file information retrieval and distributed storage method according to the present application is shown. The file information retrieval and distributed storage method includes the following steps: Step 101, in response to receiving a file information retrieval request, generate a file information request vector, and generate file information retrieval attribute data according to the retrieval conditions in the above file information retrieval request.

[0018] In some embodiments, the execution entity (e.g., a computing device) of the file information retrieval and distributed storage method may, in response to receiving a file information retrieval request, generate a file information request vector and generate file information retrieval attribute data according to the retrieval conditions in the above-mentioned file information retrieval request. Among them, the above-mentioned file information request vector and the above-mentioned file information retrieval attribute data represent different file data information. The file information request vector may represent the file information retrieval request data initially input by the user. The file information retrieval attribute data may represent the additional retrieval conditions of the user. For example, the user enters "Type A files" in the search box. In addition, the user also sets a time limit in the filtering conditions. At this time, the execution entity may convert the two words "Type A files" into a vector representation as the file information request vector. And the set time limit can be used as the file information retrieval attribute data of this file information request.

[0019] In a practical application scenario, the above-mentioned execution entity may generate file information retrieval attribute data through the following steps: First step, generate a file retrieval formula according to the retrieval conditions in the above-mentioned file information retrieval request. The constraint condition information carried in the file information request can be parsed, and a Boolean expression (file retrieval formula) can be constructed according to the AND / OR relationship between different types of conditions. Among them, if each condition expression contains multiple conditions, the conditions can be in an AND (&&) relationship. And each condition may contain multiple attribute values, and the attribute values can be in an OR (||) relationship.

[0020] Second step, for each retrieval condition in the above-mentioned file retrieval formula, calculate the hash value of the retrieval condition according to the hash algorithm of the Bloom filter to obtain the bit array of the retrieval condition. For example, K hash algorithms can be used to calculate the K hash values corresponding to each attribute value (attr_value1-5) in each retrieval condition (cond1, 2, 3). These hash values are used to construct the corresponding bit array. For example, when K = 2, the calculated hash values are 2 and 8 respectively, then the second and eighth bits in the corresponding bit array are set to 1, and the rest of the bits are set to 0.

[0021] Third step, use the bit arrays of each retrieval condition in the above-mentioned file retrieval formula as the file information retrieval attribute data of the above-mentioned file information retrieval request.

[0022] Step 102, retrieve the file information vector set under the target leaf node that matches the above-mentioned file information request vector from the pre-constructed file information vector tree according to the above-mentioned file information request vector.

[0023] In some embodiments, the above-mentioned execution entity may retrieve the set of archive information vectors under the target leaf node that matches the above-mentioned archive information request vector from the pre-constructed archive information vector tree according to the above-mentioned archive information request vector. Among them, the above-mentioned archive information vector tree is obtained by clustering the archive information vectors of the archive information. The classic BeamSearch algorithm can be used to retrieve layer by layer on the archive information vector tree. That is, starting from the root node of the tree, calculate the distance between the archive information request vector and the node vectors of the nodes in the current layer, and take out the top K nodes with the closest distance after sorting. Then, use the child nodes under these K nodes as the new candidate set, and calculate the top K nodes with the closest distance to the archive information request vector from them. Repeat this process for each layer until the top K target leaf nodes with the closest distance to the archive information request vector are found. The quantity K and the distance algorithm here can be set according to actual needs. For example, the distance can use the dot product or cosine distance.

[0024] When the number of items is not enough, the execution entity may use all the item vectors under these target leaf nodes as the retrieved item vector set. When the number of items is sufficient, for all the item vectors under these target leaf nodes, the execution entity may calculate the distance from the request vector again, so as to use the top K item vectors with the closest distance to the request vector as the retrieved item vector set.

[0025] Among them, the archive information vector tree can be obtained through the following training steps: The first step is to obtain the archive-related information of each archive to be stored and generate each archive information vector. Among them, the above-mentioned archive-related information includes: archive book title information, archive identification information, and archive pictures.

[0026] The second step is to perform clustering analysis on each archive information vector until the clustering end condition is met to obtain the archive information vector tree. The clustering end condition is not limited. For example, the number of layers of the tree reaches the set value, or the number of items in the clustering cluster reaches the set value, etc.

[0027] The third step is that for each node in the above-mentioned archive information vector tree, use the average value of the archive information vectors of each archive information belonging to the node as the node vector of the node, and store the archive identification information of the archive information belonging to the leaf node (i.e., the item node) under each leaf node. In the retrieval system, an item usually refers to a single object or record that can be retrieved, and its physical meaning depends on the specific retrieval scenario.

[0028] Step 103: Retrieve the set of alternative archive information that matches the above-mentioned archive information retrieval attribute data from the archive information indicated by the above-mentioned archive information vector set according to the description attribute data of the archive information.

[0029] In some embodiments, the above-mentioned execution entity may retrieve an alternative archive information set that matches the above-mentioned archive information retrieval attribute data from the archive information indicated by the above-mentioned archive information vector set according to the description attribute data of the archive information. The archive information description attribute data is usually data used to describe the attributes of archive information and can be obtained in various ways. The attribute (constraint condition) descriptions of different archive information may be different, and the number of attributes may also be different. To reduce or avoid retrieval errors caused by attribute descriptions, the attribute information of the archive information can first be standardized. For example, a standard attribute table (i.e., an algorithm vocabulary) can be used to unify the attribute descriptions of the archive information. Another example is to find the intersection of the standard attribute table and the archive information attributes, and use the attributes that exist in both as the attributes of the archive information after standardization processing. According to the standardized attributes and the corresponding attribute values, the archive information description attribute data can be obtained.

[0030] In an actual application scenario, the above-mentioned execution entity may retrieve an alternative archive information set that matches the above-mentioned archive information retrieval attribute data through the following steps: In the first step, obtain the bit arrays of each piece of archive information indicated by the above-mentioned archive information vector set, and match the bit arrays of each piece of archive information with the bit arrays of each retrieval condition in the above-mentioned archive information retrieval request.

[0031] In the second step, determine the archive information indicated by the bit array of the archive information that matches the bit arrays of each retrieval condition as the alternative archive information, and obtain the alternative archive information set.

[0032] That is to say, during the process of vector recall in the engine layer, first, the archive information vector tree index can be queried according to the archive information request vector to find the archive information under the leaf node. Then, business constraint filtering can be performed according to the scalar conditions of the archive information (i.e., the Bloom filter forward index). That is, match the Bloom filter information corresponding to the archive information with the bit arrays in the Boolean expression of the retrieval conditions. If any one of the conditions does not meet the match, the current archive information is filtered out, and finally, a valid alternative archive information set is retained.

[0033] It can be understood that the above-constructed archive information vector tree and the descriptive attribute data of the archive information, these archive information data, may change with the increase, decrease, or adjustment of the archive information on the retrieval system. Therefore, when the execution entity detects a change in the archive information data, such as receiving a change notification message from an offline indexing system, it can obtain the data in the Bloom filter that stores the latest descriptive attribute data of the archive information, and obtain the latest archive information vector tree for storage in memory. For the latest archive information vector tree, for the archive information indicated by the archive information vector under each leaf node, the information of the Bloom filter object corresponding to the archive information can be determined, and this information can be stored under the leaf node to which the archive information belongs.

[0034] Step 104: According to the above archive information retrieval request, perform a retrieval process on the alternative archive information in the above alternative archive information set to obtain an archive information retrieval result corresponding to the above archive information retrieval request.

[0035] In some embodiments, the above execution entity can perform a retrieval process on the alternative archive information in the above alternative archive information set according to the above archive information retrieval request to obtain an archive information retrieval result corresponding to the above archive information retrieval request. The alternative archive information in the alternative archive information set can be screened according to the archive information retrieval request. That is, according to the retrieval conditions in the archive information retrieval request, the specific business filtering rules of the archive information retrieval request at the business layer can be determined. That is, the parameters of the business rule filter. And the processing result can be used as the retrieval result of the archive information retrieval request.

[0036] For another example, each keyword field of the above archive information retrieval request can be extracted, and the alternative archive information with the highest relevance can be found from the alternative archive information set as the archive information retrieval result of the above archive information retrieval request.

[0037] Furthermore, according to the preset standard attributes, the attribute information of the archive information is standardized.

[0038] In some embodiments, the above execution entity can standardize the attribute information of the archive information according to the preset standard attributes. The attribute information of the archive information can be standardized. For example, a standard attribute table (i.e., an algorithm word table) can be used to unify the attribute descriptions of the archive information. For another example, the intersection of the standard attribute table and the archive information attributes can be taken, and the attributes that exist in both can be used as the attributes of the archive information after standardization processing. According to the standardized attributes and the corresponding attribute values, the descriptive attribute data of the archive information can be obtained.

[0039] Furthermore, according to the standardized attributes and the corresponding attribute values, archive description attribute information in a preset format is generated.

[0040] In some embodiments, the above-mentioned execution entity may generate archive description attribute information in a preset format according to the standardized attributes and corresponding attribute values. The attributes (i.e., constraints) of the archive information can be processed. The type of the attribute (i.e., constraint) can be represented by attr_type, and the specific value of the attribute, i.e., the attribute value, can be denoted as attr_value. In this way, they can be uniformly concatenated into a string in the format of "attr_type^attr_value". This processing method is applicable to condition values of any type, such as integer type, floating-point type, string, etc.

[0041] Furthermore, a Bloom filter corresponding to the archive information is constructed.

[0042] In some embodiments, the above-mentioned execution entity may construct a Bloom filter (itembloom filter) corresponding to the archive information. Establish the corresponding relationship between the archive information (item_id) and the Bloom filter.

[0043] Furthermore, calculate the hash value of the archive description attribute information of the archive information, and store the calculated hash value into the corresponding Bloom filter to obtain the bit array of the archive information.

[0044] In some embodiments, the above-mentioned execution entity may calculate the hash value of the archive description attribute information of the archive information, and store the calculated hash value into the corresponding Bloom filter to obtain the bit array of the archive information.

[0045] When constructing the Bloom filter of the archive information, different attributes under the archive information are standardized and converted into hash values. Each attribute will have its "attribute type - attr_type" and "attribute value attr_value". When implemented, the hash value will be calculated uniformly in the form of attr_type^attr_value, and then the hash value is written into the bit positions in the Bloom filter. Moreover, the size of the Bloom filter for each archive information is usually fixed.

[0046] A Bloom filter is generally a probabilistic data structure with extremely high space and query efficiency. It stores data using a bit array at the bottom layer and can quickly determine whether an element is in a set. Its core principle is as follows: 1. The length of the bit array is usually represented by the parameter m. When initializing, all m bit positions in the array are set to 0.

[0047] 2. When adding an element, use k hash functions to hash the element to generate k hash values, which are the subscripts of the bit array, and then set the bits at these positions to 1.

[0048] 3. When querying an element, the same k hash functions are needed to hash the input element to obtain k subscripts, and then check k positions in the bit array. If any of these positions is 0, it is determined that the element must not be in the set; if all positions are 1, the element may exist in the set.

[0049] By using a Bloom filter to store the scalar attribute fields of the archive information, that is, the business rule constraint condition information, a forward index data of "item_id -> Bloom filter" is constructed. The Bloom filter maps multiple elements to a finite bit array through k hash functions. With the parameters determined, the storage space occupied by the bit array is fixed, usually much smaller than the actual space occupied by the element set. And it does not increase linearly with the increase of inserted elements, which is suitable for storing a large number of scalar condition fields and has good scalability. In addition, in terms of matching efficiency, the time complexity of judging whether a certain condition field matches is O(k), which has nothing to do with the number of elements stored in the Bloom filter.

[0050] Data can also be stored in the way of a common bitmap. Among them, each bit represents a specific value of the constraint condition. This method is simpler in data structure, but usually only suitable for storing elements with a limited value space. If a hash function is introduced to map the condition value to an integer, that is, the position in the bitmap, at this time the bitmap is actually equivalent to a Bloom filter.

[0051] Furthermore, in response to detecting a change in the archive information data, obtain the latest version of the index file.

[0052] In some embodiments, the above-mentioned execution entity can, in response to detecting a change in the archive information data, obtain the latest version of the index file. Among them, the above-mentioned index file is obtained by serializing with a preset data description language after the construction of the archive information vector tree and the Bloom filter is completed.

[0053] Furthermore, use the above-mentioned preset data description language to deserialize the obtained index file to obtain the Bloom filter data and the latest archive information vector tree in the file.

[0054] In some embodiments, the above-mentioned execution entity can use the above-mentioned preset data description language to deserialize the obtained index file to obtain the Bloom filter data and the latest archive information vector tree in the file. Among them, protobuf, that is, Protocol Buffers, is usually a data description language, similar to XML, which can serialize structured data and can be used in aspects such as data storage and communication protocols.

[0055] Further, for the file information vectors under each leaf node in the above-mentioned latest file information vector tree, determine the information of the Bloom filter object corresponding to the file information indicated by the file information vector, and store the information under the leaf node to which the file information belongs.

[0056] In some embodiments, the above-mentioned execution entity may, for the file information indicated by the file information vectors under each leaf node in the above-mentioned latest file information vector tree, determine the information of the Bloom filter object corresponding to the file information, and store the information under the leaf node to which the file information belongs.

[0057] First, the forward index of the Bloom filter can be loaded and stored in a hash_map in memory. Then, the index file of the file information vector tree can be loaded, and a tree index storage structure can be constructed in memory to store the index information. When loading the file information of the leaf node, the pointer to the Bloom filter object corresponding to the current file information can be retrieved from the Bloom filter hash_map according to the file information identifier (such as item_id), and saved in memory together with other information of the item. That is, during the update process of the file information data, the execution entity can establish a method for mapping the file information identifier to the Bloom filter information in advance. In this way, during the subsequent online retrieval process, when the file information under the leaf node is queried, the corresponding Bloom filter data can be directly retrieved from it, avoiding the performance hotspot problem caused by excessive querying of the hash_map. It can not only reduce the number of queries to the hash_map, but also improve the overall query and retrieval efficiency.

[0058] Further, in response to receiving a file information storage task, determine whether there is a file information storage node corresponding to the file information storage task in the local file information storage node cluster.

[0059] In some embodiments, the above-mentioned execution entity may, in response to receiving a file information storage task, determine whether there is a file information storage node corresponding to the file information storage task in the local file information storage node cluster. Among them, the above-mentioned file information storage task includes: the file information to be stored, the type of file information, and one file information storage node corresponds to one type of file information. The file information storage task may represent a task for storing new file information. The file information storage node cluster may be a pre-established storage node for storing file information of each type of file information. That is, it can be determined whether there is a file information storage node in the local file information storage node cluster that has the same type of file information as that included in the file information storage task.

[0060] Further, in response to determining that there is an archival information storage node corresponding to the above archival information storage task, vectorize the above archival information to be stored to generate a vector of the archival information to be stored.

[0061] In some embodiments, the above execution subject may, in response to determining that there is an archival information storage node corresponding to the above archival information storage task, vectorize the above archival information to be stored to generate a vector of the archival information to be stored. For example, the above archival information to be stored may be vectorized by means of One-Hot encoding / Bag of Words (BOW) / TF-IDF, etc. to generate a vector of the archival information to be stored.

[0062] Further, store the above archival information to be stored in the corresponding archival information storage node, and add the vector of the archival information to be stored to the relevant archival information vector tree.

[0063] In some embodiments, the above execution subject may store the above archival information to be stored in the corresponding archival information storage node, and add the vector of the archival information to be stored to the relevant archival information vector tree.

[0064] Further, in response to receiving at least one archival information change data, for each archival information change data in the at least one archival information change data, determine at least one archival information data table corresponding to the archival information change data. The archival information change data may be change data indicating how the archival information corresponding to the file changes. The archival information data table may be a storage object storing the corresponding archival information data. The number of archival information data tables corresponding to the archival information change data is at least one.

[0065] In a practical application scenario, the above execution subject may determine at least one archival information data table corresponding to the archival information change data through the following: In the first step, determine the archival information type, at least one relationship attribute, and at least one relationship attribute range corresponding to the above archival information change data. The archival information type may be the corresponding type of the archival information. The at least one relationship attribute may be at least one attribute in the archival information data used to indicate that the archival information has changed. That is, the relationship attribute may reflect the attribute of the information change corresponding to the archival information. There is a one-to-one correspondence between the relationship attributes in the at least one relationship attribute and the relationship attribute ranges in the at least one relationship attribute range. The relationship attribute range may be the value range of the attribute corresponding to the relationship attribute. In practice, the archival information change data may include: archival information type, at least one relationship attribute, and at least one relationship attribute range.

[0066] In the second step, determine at least one candidate archive information data table corresponding to the above archive information type. Wherein, the at least one candidate archive information data table is at least one archive information data table including each piece of archive information data under the archive information type.

[0067] As an example, the above execution entity can determine at least one candidate archive information data table corresponding to the above archive information by querying the archive information types of the archive information data corresponding to each archive information data table.

[0068] In the third step, screen out the candidate archive information data tables having the same attributes as the above at least one relationship attribute from the above at least one candidate archive information data table to obtain a set of screened archive information data tables. Wherein, having the same attributes as the above at least one relationship attribute can indicate that the attribute values of at least one relationship attribute of the archive information data stored in the corresponding candidate archive information data table are the same as those of the archive information change data.

[0069] In the fourth step, screen out the archive information data tables having a range attribution relationship with the above at least one relationship attribute range from the above set of screened archive information data tables to obtain the above at least one archive information data table. Wherein, having a range attribution relationship with the above at least one relationship attribute range can indicate that the at least one relationship attribute range corresponding to the archive information change data is a sub-range of at least one relationship attribute range corresponding to the archive information data stored in the candidate archive information data table.

[0070] Further, for each archive information data table in the set of archive information data tables involved, perform the following processing steps: In the first step, generate a result of change of archive information data corresponding to the above archive information data table according to the correspondence between the archive information change data and at least one archive information data table. Wherein, the result of change of archive information data can indicate the result of how to change and adjust the archive information data in the archive information data table. The result of change of archive information data can also reflect the archive information change data that needs to be adjusted corresponding to the archive information data table. The set of archive information data tables can be each archive information data table having a corresponding association relationship with at least one archive information change data. In practice, the at least one archive information data table corresponding to at least one archive information change data can be de-duplicated to obtain the set of archive information data tables. For example, the above execution entity can determine the set of archive information change data corresponding to the archive information data table. Then, for each archive information change data in the set of archive information change data, generate a corresponding database statement according to the change operation corresponding to the archive information change data. Next, fuse the contents of each database statement in the obtained set of database statements to generate the result of change of archive information data.

[0071] In the second step, send the above-mentioned result of the change of the file information data to the storage node where the above-mentioned data set of file information data tables is located, so as to change the file information data in the above-mentioned file information data tables.

[0072] Thus, the confirmation of the correspondence between the file information change data and at least one file information data table and the generation of the change result of the file information data corresponding to the file information data table can be effectively carried out. Not only can the situation where the file information data needs to be changed in a timely manner be obtained in real time, but also the data adjustment of the corresponding file information data table can be realized automatically and efficiently. This application also provides a computer device 200. As Figure 2 shown, the computer device 200 includes: a bus 201, a processor 202, a memory 203, and a communication interface 204. The processor 202, the memory 203, and the communication interface 204 communicate with each other through the bus 201. The computer device 200 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computer device 200.

[0073] The bus 201 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 2 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 201 can include a path for transmitting information between various components (for example, the memory 203, the processor 202, the communication interface 204) of the computer device 200.

[0074] The processor 202 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.

[0075] The memory 203 may include volatile memory, such as random access memory (RAM). The memory 203 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0076] Executable program code is stored in the memory 203, and the processor 202 executes the executable program code to respectively implement the functions of the foregoing acquisition module, sampling module, determination module, and mixing module, so as to implement the foregoing method for file information retrieval and distributed storage. That is, instructions for executing the foregoing method for file information retrieval and distributed storage are stored on the memory 203.

[0077] The communication interface 204 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computer device 200 and other devices or a communication network.

[0078] An embodiment of this application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface to execute the foregoing method for file information retrieval and distributed storage.

[0079] An embodiment of this application also provides a computer-readable storage medium. The foregoing computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The foregoing available medium may be a magnetic medium (for example, a floppy disk, a hard disk, or a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state drive), etc. The computer-readable storage medium includes instructions, and the foregoing instructions instruct a computing device to execute the foregoing method for file information retrieval and distributed storage.

[0080] The technical features of the foregoing embodiments may be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the foregoing embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0081] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. An archive information retrieval and distributed storage method, comprising: Responding to receiving an archive information retrieval request, generating an archive information request vector, and generating archive information retrieval attribute data according to the retrieval conditions in the archive information retrieval request, wherein the archive information request vector and the archive information retrieval attribute data represent different archive data information; According to the archive information request vector, retrieving, from a pre-constructed archive information vector tree, an archive information vector set under a target leaf node that matches the archive information request vector, wherein the archive information vector tree is obtained by clustering archive information vectors of archive information; According to the description attribute data of the archive information, retrieving, from the archive information indicated by the archive information vector set, an alternative archive information set that matches the archive information retrieval attribute data; According to the archive information retrieval request, performing a retrieval process on the alternative archive information in the alternative archive information set to obtain an archive information retrieval result corresponding to the archive information retrieval request.

2. The file information retrieval and distributed storage method according to claim 1, wherein, The method further includes: Performing standardization processing on the attribute information of the archive information according to a preset standard attribute; Generating archive description attribute information in a preset format according to the standardized attribute and the corresponding attribute value; Constructing a Bloom filter corresponding to the archive information; Calculating the hash value of the archive description attribute information of the archive information, and storing the calculated hash value into the corresponding Bloom filter to obtain the bit array of the archive information.

3. The file information retrieval and distributed storage method according to claim 1, wherein, Before retrieving, from a pre-constructed archive information vector tree, an archive information vector set under a target leaf node that matches the archive information request vector according to the archive information request vector, the method further includes: Obtaining archive-related information of each archive to be stored, and generating each archive information vector, wherein the archive-related information includes: archive title information, archive identification information, archive pictures; Performing clustering analysis on each archive information vector until a clustering end condition is met to obtain an archive information vector tree; For each node in the archive information vector tree, taking the average value of the archive information vectors of the archives belonging to the node as the node vector of the node, and storing the archive identification information of the archives belonging to the leaf node under each leaf node.

4. The file information retrieval and distributed storage method according to claim 1, wherein, The method further includes: Responding to receiving an archive information storage task, determining whether there is an archive information storage node corresponding to the archive information storage task in the local archive information storage node cluster, wherein the archive information storage task includes: the archive information to be stored, the archive information type, and one archive information storage node corresponds to one archive information type; Responding to determining that there is an archive information storage node corresponding to the archive information storage task, performing vectorization processing on the archive information to be stored to generate an archive information vector to be stored; Storing the archive information to be stored into the corresponding archive information storage node, and adding the archive information vector to be stored to the related archive information vector tree.

5. A computer device, wherein, The computer device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the method according to any one of claims 1-4 are implemented.

6. A computer-readable storage medium, wherein, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the method according to any one of claims 1-4 are implemented.