Middle station data index positioning and retrieving method
By adjusting the data storage status of the middle platform and determining the search priority, the problem of inefficiency of the middle platform system in data query is solved, and orderly precise retrieval and efficient data acquisition are achieved.
Patent Information
- Application Number
- CN202510756424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
When facing data query requests from individual users, the middle platform system consumes computing resources and is difficult to conduct orderly and precise retrieval, resulting in data omissions and repeated queries, reducing data retrieval efficiency.
Based on the upload characteristics of the data source end and the status of the storage interval in the middle platform, the data storage state is adjusted and the search priority is determined. The storage interval is marked by querying the path to achieve orderly retrieval; when the target data is not successfully retrieved, the predicted data storage change information is searched again.
It improves the traceability and efficiency of middle-end data retrieval, avoids duplicate invalid searches, and ensures comprehensive acquisition of middle-end data.
Smart Images

Figure CN120256480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method for indexing, locating and retrieving data from a central station. Background Art
[0002] With the development of the industry, industrial data is also increasing and expanding, thus forming industrial big data. Industrial big data plays a vital role in industrial research. In order to realize the comprehensive and effective use of industrial big data, industrial big data can be saved through the middle-office system, which can not only classify and save industrial big data, but also facilitate the sharing and transmission of industrial big data in the front-end and back-end connected to the middle-office system, thereby improving the utilization efficiency of industrial big data. The middle-office system has the advantages of flexibility and efficiency in batch data query, but when individual users initiate data query requests to the middle-office system, the middle-office system usually needs to search for massive internal industrial big data, which not only consumes the computing power resources of the middle-office system, but also reduces the efficiency of data query, and is prone to problems such as missing data query and repeated query. It is impossible to search and search the big data in the middle-office system in an orderly and accurate manner, which reduces the data retrieval efficiency of the middle-office system. Summary of the invention
[0003] In view of the defects existing in the prior art, the present invention provides a method for indexing, locating and retrieving data on the middle platform. Based on the characteristics of the uploaded data from all connected data source ends to the middle platform, the storage status of the uploaded data on the middle platform is adjusted, and based on the data inventory status of all storage intervals in the middle platform, the retrieval priority is determined to achieve orderly retrieval of the middle platform; based on the target data attribute information and retrieval priority that the user end needs to query, the query path is determined to mark the storage interval for data browsing, thereby improving the traceability of the data retrieval on the middle platform; when the target data is successfully retrieved, based on the distribution status of the target data in the middle platform, the retrieval result data packet is returned to the user end to ensure comprehensive acquisition of the data in the middle platform; when the target data is not successfully retrieved, the data storage change attribute information of the middle platform is predicted, and data retrieval is performed again, which can avoid repeated and invalid retrieval of the middle platform, and orderly and accurate retrieval and search of the big data in the middle platform are performed, thereby improving the data retrieval efficiency of the middle platform.
[0004] The present invention provides a method for indexing and locating data in a middle platform, comprising the following steps: Step S1, based on the communication log of the middle station, determine the upload data characteristics of all connected data source ends to the middle station; based on the upload data characteristics, adjust the storage status of the uploaded data by the middle station; based on the data inventory status of all storage intervals in the middle station, determine the retrieval priority of all storage intervals; Step S2: Based on the query request of the client, determine the target data attribute information that the client needs to query; based on the target data attribute information and the retrieval priority, determine the query path for the corresponding storage range, and use this to mark the data browsing of the storage range. Step S3: Based on the data browsing mark, determine whether the target data that the client needs to query is successfully retrieved; when the retrieval is successful, generate a retrieval result data packet based on the distribution status of the target data in the middle platform and return it to the client. Step S4: When the retrieval fails, predict the data storage change attribute information of the middle platform based on the uploaded data characteristics; based on the data storage change attribute information and the target data attribute information, perform data retrieval again, and return the retrieval result to the client.
[0005] In an embodiment disclosed in the present application, in the step S1, based on the communication log of the middle platform, determining the uploaded data characteristics of all data source ends connected to the middle platform includes: Based on the gateway traffic change information in the network where the middle platform is located, determine the active gateways in the network; monitor the active gateways to determine the identity information of all data source ends accessing the active gateways. Based on the identity information, screen the data upload records of the data source ends to the middle platform from the communication log of the middle platform; based on the data upload records, determine the uploaded data types and quantities of the data source ends to the middle platform, and use this as the uploaded data characteristics.
[0006] In an embodiment disclosed in the present application, in the step S1, based on the uploaded data characteristics, adjusting the storage status of the uploaded data by the middle platform includes: Obtain the data storage history records of all storage ranges in the middle platform, analyze the data storage history records, and obtain the stored data types and quantities of each storage range in the middle platform. Compare the uploaded data types and quantities of the data source ends to the middle platform with the stored data types and quantities of each storage range to determine the storage range that matches the uploaded data of the data source ends. Based on the data upload rate of the data source ends, adjust the data storage transmission bandwidth allocated by the middle platform to the matching storage range.
[0007] In an embodiment disclosed in the present application, in the step S1, based on the data stock status of all storage ranges in the middle platform, determining the retrieval priority of all storage ranges includes: Obtain the data storage amount and data semantic information of each storage range in the middle platform, and determine the content duplication characteristics of the stored data in each storage range based on the data storage amount and the data semantic information; wherein, the content duplication characteristic refers to the proportion of data with the same or similar content in each storage range. Determine the retrieval priority for all storage ranges based on the content duplication characteristics; wherein, the retrieval priority includes the retrieval order for all storage ranges and the retrieval duration for each storage range.
[0008] In an embodiment disclosed in the present application, in step S2, based on the query request of the user terminal, determine the target data attribute information that the user terminal needs to query, including: Based on the identity information of the user terminal requesting access to the middle platform, determine the data retrieval history record of the user terminal for the middle platform; based on the data retrieval history record, determine the retrieval occupancy time information of the user terminal in the middle platform, and thereby determine whether the user terminal has the retrieval permission. When the user terminal has the retrieval permission, parse the query request of the user terminal to obtain the target data index information that the user terminal needs to query, and use this as the target data attribute information; wherein, the target data index information includes the keyword semantic information of the target data that the user terminal needs to query.
[0009] In an embodiment disclosed in the present application, in step S2, based on the target data attribute information and the retrieval priority, determine the query path for the corresponding storage range, and thereby perform data browsing marking on the storage range, including: Compare the target data index information with the storage directories of all storage ranges in the middle platform, and screen out several target storage ranges; based on the retrieval priority, determine the retrieval order for all target storage ranges and the retrieval duration for each target storage range. Determine the query path for all target storage ranges based on the retrieval order for all target storage ranges and the retrieval duration for each target storage range; based on the query path, identify the data browsed in the target storage range, and thereby add an identification code to the data that matches the keyword semantic information.
[0010] In an embodiment disclosed in the present application, in step S3, based on the data browsing marking, determine whether the target data that the user terminal needs to query is successfully retrieved, including: Compare the creation time of the data with the identification code added in the storage interval with a preset time range. If the creation time is within the preset time range, it is determined that the target data required to be queried by the client is successfully retrieved; otherwise, it is determined that the target data required to be queried by the client is not successfully retrieved.
[0011] In an embodiment disclosed in the present application, in the step S3, when the retrieval is successful, a retrieval result data packet is generated based on the distribution state of the target data in the middle platform and returned to the client, including: When the retrieval is successful, based on the distribution positions and distribution quantities of the target data in all storage intervals in the middle platform, data replication extraction and integration compression are performed on the corresponding storage intervals to generate a retrieval result data packet, and the retrieval result data packet is returned to the client.
[0012] In an embodiment disclosed in the present application, in the step S4, when the retrieval fails, the data storage change attribute information of the middle platform is predicted based on the uploaded data characteristics, including: When the retrieval fails, based on the uploaded data characteristics, predict the types and data volumes of the data to be uploaded by the data source end to the middle platform in the future; based on the types and data volumes of the future uploaded data, predict the data storage change attribute information of all storage intervals in the middle platform; wherein, the data storage change attribute information includes the increased data volumes of the corresponding type data in all storage intervals in the middle platform in the future time period.
[0013] In an embodiment disclosed in the present application, in the step S4, based on the data storage change attribute information and the target data attribute information, data retrieval is performed again, and the retrieval result is returned to the client accordingly, including: Compare the increased data volume with a preset data volume threshold. If the increased data volume exceeds the preset quantity threshold, it is determined that data retrieval is allowed to be performed again; otherwise, it is determined that data retrieval is not allowed to be performed again; When data retrieval is allowed to be performed again, based on the target data attribute information, retrieve the newly added data in all storage intervals in the middle platform, and return the retrieval result to the client accordingly.
[0014] Compared with the existing technology, the data index positioning and retrieval method of the middle platform is based on the uploaded data characteristics of all connected data source ends to the middle platform, adjusts the storage status of the uploaded data by the middle platform, and determines the retrieval priority based on the data inventory status of all storage intervals in the middle platform to achieve orderly retrieval of the middle platform; based on the target data attribute information and retrieval priority that the user end needs to query, the query path is determined to mark the storage interval for data browsing, thereby improving the traceability of the data retrieval of the middle platform; when the target data is successfully retrieved, the retrieval result data packet is returned to the user end based on the distribution status of the target data in the middle platform to ensure comprehensive acquisition of the data in the middle platform; when the target data is not successfully retrieved, the data storage change attribute information of the middle platform is predicted, and data retrieval is performed again, which can avoid repeated and invalid retrieval of the middle platform, and orderly and accurate retrieval and search of the big data in the middle platform are performed, thereby improving the data retrieval efficiency of the middle platform.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0016] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A flow chart of the middle platform data index positioning and retrieval method provided by the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] See also Figure 1 , is a flow chart of a method for locating and retrieving data indexes in a middle platform provided in an embodiment of the present invention. The method for locating and retrieving data indexes in a middle platform includes: Step S1: Based on the communication logs of the middleware platform, determine the upload data characteristics of all connected data source ends to the middleware platform; based on the upload data characteristics, adjust the storage status of the upload data by the middleware platform; based on the data stock status of all storage intervals within the middleware platform, determine the retrieval priorities of all storage intervals; Step S2: Based on the query request from the user end, determine the target data attribute information that the user end needs to query; based on the target data attribute information and the retrieval priorities, determine the query path for the corresponding storage interval, and thereby mark the data browsing of the storage interval; Step S3: Based on the data browsing mark, determine whether the target data that the user end needs to query is successfully retrieved; when the retrieval is successful, generate a retrieval result data packet based on the distribution status of the target data within the middleware platform and return it to the user end; Step S4: When the retrieval fails, predict the data storage change attribute information of the middleware platform based on the upload data characteristics; based on the data storage change attribute information and the target data attribute information, perform data retrieval again, and thereby return the retrieval result to the user end.
[0021] The beneficial effects of the above technical solution are as follows: This middleware data indexing, positioning and retrieval method adjusts the storage status of the upload data by the middleware platform based on the upload data characteristics of all connected data source ends to the middleware platform, and determines the retrieval priorities based on the data stock status of all storage intervals within the middleware platform, so as to achieve an orderly retrieval of the middleware platform; based on the target data attribute information that the user end needs to query and the retrieval priorities, determine the query path, and thereby mark the data browsing of the storage interval, improving the traceability of middleware data retrieval; when the target data is successfully retrieved, generate a retrieval result data packet based on the distribution status of the target data within the middleware platform and return it to the user end to ensure a comprehensive acquisition of the data within the middleware platform; when the target data is not successfully retrieved, predict the data storage change attribute information of the middleware platform and perform data retrieval again, which can avoid repeated and ineffective retrievals of the middleware platform, perform an orderly and accurate retrieval and search of the big data within the middleware platform, and improve the data retrieval efficiency of the middleware platform.
[0022] Preferably, in Step S1, based on the communication logs of the middleware platform, determining the upload data characteristics of all connected data source ends to the middleware platform includes: Based on the gateway traffic change information within the network where the middleware platform is located, determine the active gateways within the network; monitor the active gateways to determine the identity information of all data source ends accessing the active gateways; Based on the identity information, filter the data upload records of the data source ends to the middleware platform from the communication logs of the middleware platform; based on the data upload records, determine the upload data types and quantities of the data source ends to the middleware platform, and thereby use them as the upload data characteristics.
[0023] The beneficial effects of the above technical solution are as follows: The middle platform serves as a storage platform for big data, which stores different types of data internally. In order to store different types of data in a partitioned manner, each type of data will be independently stored in one or more storage intervals within the middle platform. To ensure that the middle platform can receive different types of data in a timely and comprehensive manner, the middle platform is also connected to different data source ends through a network, so that the data source ends can upload the data they generate or receive from the outside to the middle platform. Whenever a data source end performs a data upload to the middle platform, the communication log of the middle platform will generate a corresponding data upload record, accurately and comprehensively recording the current uploaded data type and data volume. In the actual data interaction process between the middle platform and the data source ends, the data source ends do not continuously maintain the state of uploading data to the middle platform, that is, not all data source ends will update and increase the data inside the middle platform. In order to accurately grasp the data upload situation of the corresponding data source ends to the middle platform with the least resources, monitor the change information of the gateway traffic within the network where the middle platform is located (such as the uplink traffic change information of each gateway within the network), obtain the average traffic rate of each gateway uploading to the middle platform. If the average traffic rate is greater than the preset traffic threshold, it is determined that the corresponding gateway belongs to an active gateway. Then, identify and calibrate the identity information of all data source ends connected to the active gateway, and based on the identity information of the data source ends, screen the data types and data volumes uploaded by the corresponding data source ends to the middle platform from the communication log of the middle platform, providing a reliable basis for subsequent adjustment of the data storage state of the middle platform and prediction of the future data upload trend of the corresponding data source ends.
[0024] Preferably, in step S1, based on the characteristics of the uploaded data, adjusting the storage state of the uploaded data by the middle platform includes: Obtain the data storage history records of all storage intervals within the middle platform, analyze the data storage history records, and obtain the stored data types and quantities of each storage interval within the middle platform; Compare the data types and quantities of the data uploaded by the data source end to the middle platform with the stored data types and quantities of each storage interval, and determine the storage interval that matches the data uploaded by the data source end; Based on the data upload rate of the data source end, adjust the data storage transmission bandwidth allocated by the middle platform to the matching storage interval.
[0025] The beneficial effects of the above technical solution are as follows: The middle platform includes multiple relatively independent storage intervals. Each storage interval can store data of the same type or data uploaded from the same data source end. Whenever a data storage is performed in a storage interval, the data storage history record of the middle platform regarding the storage interval will be updated accordingly. By analyzing the data storage history record, the types and amounts of stored data of all storage intervals in the middle platform can be obtained, so as to grasp the data storage status of each storage interval in real time. In order to avoid data storage chaos and crosstalk inside the middle platform, it is also necessary to adjust the storage status of the uploaded data by the middle platform in a targeted manner according to the data situation uploaded from the data source end. Specifically, compare the types and quantities of data uploaded from the data source end to the middle platform with the types and quantities of stored data of all storage intervals respectively, judge the matching degree between the data type uploaded from the data source end to the middle platform and the stored data type of each storage interval, and determine the storage intervals that are allowed to store the data uploaded from the data source end to the middle platform as the storage intervals with a matching degree exceeding the preset matching degree threshold; then, based on the remaining storage space of all the stored storage intervals, determine the storage intervals that match the uploaded data from the data source end. Usually, select the storage interval with the largest remaining storage space as the matching storage interval. Additionally, based on the data upload rate from the data source end to the middle platform, adaptively adjust the size of the data storage transmission bandwidth allocated by the middle platform to the matching storage interval to ensure that the data uploaded from the data source end to the middle platform is stored in the matching storage interval in real time and avoid data storage omission.
[0026] Preferably, in step S1, based on the data stock status of all storage intervals in the middle platform, determine the retrieval priorities of all storage intervals, including Obtain the data storage amounts and data semantic information of all storage intervals in the middle platform respectively. Based on the data storage amounts and data semantic information, determine the content repetition characteristics of the stored data in each storage interval; wherein, the content repetition characteristic refers to the proportion of data with the same or similar content in each storage interval; Based on the content repetition characteristics, determine the retrieval priorities of all storage intervals; wherein, the retrieval priorities include the retrieval order of all storage intervals and the retrieval duration of each storage interval.
[0027] The beneficial effects of the above technical solution are as follows: After the data source end uploads data to the middle platform, the middle platform will directly store the received data in the storage area without screening and sorting the data, resulting in duplicate storage of data in the storage area. When the proportion of duplicate data in the storage area is relatively high, the number of substantially retrievable data in the storage area is correspondingly low. If the traversal retrieval method is still used for the storage area with a relatively high proportion of duplicate data, it will not only increase the time-consuming of data retrieval but also cause waste of retrieval computing power resources. To ensure effective and accurate retrieval of all storage areas within a limited time, based on the data storage volume and data semantic information of each storage area in the middle platform, determine the proportion of data with the same or similar content among the stored data in each storage area. The larger the proportion of data with the same or similar content among the stored data in the storage area, the later its retrieval order among all storage areas, and the smaller the retrieval duration of the corresponding storage area, avoiding spending too much time retrieving the storage area with a relatively large proportion of data with the same or similar content.
[0028] Preferably, in step S2, based on the query request of the user end, determine the target data attribute information that the user end needs to query, including: Based on the identity information of the user end requesting to access the middle platform, determine the data retrieval history record of the user end on the middle platform; based on the data retrieval history record, determine the retrieval occupancy time information of the user end on the middle platform, so as to judge whether the user end has the retrieval permission; When the user end has the retrieval permission, parse the query request of the user end to obtain the target data index information that the user end needs to query, and use this as the target data attribute information; among them, the target data index information includes the keyword semantic information of the target data that the user end needs to query.
[0029] The beneficial effects of the above technical solution are as follows: To avoid the user end occupying the access and retrieval time of other user ends on the middle platform due to frequent access and retrieval of the middle platform, based on the identity information of the user end requesting to access the middle platform, obtain the data retrieval history record of the user end on the middle platform, so as to determine the historical retrieval occupancy time information of the user end on the middle platform (i.e., the occupancy time interval during the historical retrieval process of the user end on the middle platform). If the time interval between the most recent historical retrieval time interval of the user end on the middle platform and the current time point is less than the preset time interval threshold, it is judged that the user end does not have the retrieval permission; otherwise, it is judged that the user end has the retrieval permission. When the user end has the retrieval permission, parse the query request of the user end to obtain the keyword semantic information of the target data that the user end needs to query, and use this as the basis for subsequent data screening and identification of the storage area.
[0030] Preferably, in step S2, based on the target data attribute information and the retrieval priority, determine the query path for the corresponding storage interval, and use this to mark the data browsing of the storage interval, including: Compare the target data index information with the storage directories of all storage intervals in the middle platform, and filter out several target storage intervals; based on the retrieval priority, determine the retrieval order for all target storage intervals and the retrieval duration for each target storage interval; Based on the retrieval order for all target storage intervals and the retrieval duration for each target storage interval, determine the query path for all target storage intervals; based on the query path, identify the data browsed in the target storage interval, and thus add an identification code to the data that matches the keyword semantic information.
[0031] The beneficial effects of the above technical solution are as follows: Compare the target data index information with the storage directories of all storage intervals in the middle platform, determine the semantic similarity between the target data index information and the storage directories of all storage intervals respectively, and use the storage intervals with a semantic similarity exceeding the preset similarity threshold as the target storage intervals; then, in combination with the retrieval priorities corresponding to all target storage intervals, determine the retrieval order for all target storage intervals and the retrieval duration for each target storage interval, and thus determine the query path for all target storage intervals (i.e., the query path for all target storage intervals in the time domain). Then, based on the above query path, along the corresponding time axis, identify the data browsed in the corresponding target storage interval, and thus add an identification code to the data that matches the keyword semantic information, and perform a preliminary screening of the data in the storage interval.
[0032] Preferably, in step S3, based on the data browsing mark, determine whether the target data required by the user terminal is successfully retrieved, including: Compare the creation time of the data with the identification code added in the storage interval with the preset time range. If the creation time is within the preset time range, it is determined that the target data required by the user terminal is successfully retrieved; otherwise, it is determined that the target data required by the user terminal is not successfully retrieved.
[0033] The beneficial effects of the above technical solution are as follows: Compare the creation time of the data with the identification code added in the storage interval with the preset time range to determine whether the target data required by the user terminal is successfully retrieved, and ensure the accurate retrieval of data matching the user terminal from the storage interval.
[0034] Preferably, in step S3, when the retrieval is successful, based on the distribution state of the target data in the middle platform, generate a retrieval result data packet and return it to the user terminal, including: When the retrieval is successful, based on the distribution locations and quantities of the target data in all storage ranges within the middle platform, data replication extraction and integration compression are performed on the corresponding storage ranges to generate a retrieval result data packet, and the retrieval result data packet is returned to the user side.
[0035] The beneficial effects of the above technical solution are as follows: When the retrieval is successful, based on the distribution locations and quantities of the target data in all storage ranges within the middle platform, data replication extraction and integration compression are performed on the corresponding storage ranges to generate a retrieval result data packet, ensuring that the retrieved data is completely returned to the user side.
[0036] Preferably, in step S4, when the retrieval fails, based on the characteristics of the uploaded data, the data storage change attribute information of the middle platform is predicted, including: When the retrieval fails, based on the characteristics of the uploaded data, the types and quantities of data to be uploaded from the data source side to the middle platform in the future are predicted; based on the types and quantities of the future uploaded data, the data storage change attribute information of all storage ranges within the middle platform is predicted; among them, the data storage change attribute information includes the increased data quantities of the corresponding type of data in all storage ranges within the middle platform during the future time period.
[0037] The beneficial effects of the above technical solution are as follows: When the retrieval fails, based on the characteristics of the uploaded data, the types and quantities of data to be uploaded from the data source side to the middle platform in the future are predicted, and based on this, the increased data quantities of the corresponding type of data in all storage ranges within the middle platform during the future time period are predicted, delimiting the scope for retrieving the middle platform again later and avoiding partial duplicate retrieval of the data in the storage ranges that have been retrieved last time.
[0038] Preferably, in step S4, based on the data storage change attribute information and the target data attribute information, data retrieval is performed again, and the retrieval result is returned to the user side accordingly, including: The increased data quantity is compared with a preset data quantity threshold. If the increased data quantity exceeds the preset quantity threshold, it is determined that data retrieval is allowed to be performed again; otherwise, it is determined that data retrieval is not allowed to be performed again; When data retrieval is allowed to be performed again, based on the target data attribute information, the newly added data in all storage ranges within the middle platform is retrieved, and the retrieval result is returned to the user side accordingly.
[0039] The beneficial effects of the above technical solution are as follows: Compare the increased data volume with the preset data volume threshold. If the increased data volume exceeds the preset quantity threshold, it indicates that the newly added data volume in the middle platform can ensure that the middle platform can effectively retrieve data by fully utilizing its own retrieval computing power resources. At this time, it is determined that data retrieval is allowed again; when data retrieval is allowed again, based on the target data attribute information, retrieve the newly added data in all storage intervals in the middle platform, and return the retrieval result to the user terminal, avoiding repeated and ineffective retrieval of the middle platform, and performing an orderly and accurate retrieval of the large data in the middle platform to improve the data retrieval efficiency of the middle platform.
[0040] As can be seen from the content of the above embodiments, the middle platform data index positioning and retrieval method adjusts the storage state of the uploaded data by the middle platform based on the uploaded data characteristics of all data source ends connected to the middle platform, and determines the retrieval priority based on the data stock status of all storage intervals in the middle platform to achieve an orderly retrieval of the middle platform; based on the target data attribute information and retrieval priority required by the user terminal, determine the query path, and mark the data browsing of the storage interval accordingly to improve the traceability of the middle platform data retrieval; when the target data is successfully retrieved, based on the distribution state of the target data in the middle platform, return the retrieval result data packet to the user terminal to ensure a comprehensive acquisition of the data in the middle platform; when the target data is not successfully retrieved, predict the data storage change attribute information of the middle platform and perform data retrieval again, which can avoid repeated and ineffective retrieval of the middle platform, and perform an orderly and accurate retrieval of the large data in the middle platform to improve the data retrieval efficiency of the middle platform.
[0041] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. Middle - platform data index positioning and retrieval method, characterized in that It includes the following steps: Step S1: Based on the communication logs of the middleware, determine the upload data characteristics of all connected data source ends to the middleware; based on the upload data characteristics, adjust the storage status of the upload data by the middleware; based on the data stock status of all storage intervals in the middleware, determine the retrieval priorities of all storage intervals; Step S2: Based on the query request of the user end, determine the target data attribute information that the user end needs to query; based on the target data attribute information and the retrieval priorities, determine the query path for the corresponding storage interval, and mark the data browsing of the storage interval accordingly; Step S3: Based on the data browsing mark, determine whether the target data that the user end needs to query is successfully retrieved; When the retrieval is successful, generate a retrieval result data packet based on the distribution status of the target data in the middleware and return it to the user end; Step S4: When the retrieval fails, predict the data storage change attribute information of the middleware based on the upload data characteristics; based on the data storage change attribute information and the target data attribute information, perform data retrieval again, and return the retrieval result to the user end accordingly.
2. The middleware data index positioning and retrieval method according to claim 1, wherein: In the step S1, based on the communication logs of the middleware, determining the upload data characteristics of all connected data source ends to the middleware includes: Based on the gateway traffic change information in the network where the middleware is located, determine the active gateways in the network; monitor the active gateways to determine the identity information of all data source ends accessing the active gateways; Based on the identity information, screen the data upload records of the data source ends to the middleware from the communication logs of the middleware; based on the data upload records, determine the upload data types and quantities of the data source ends to the middleware, and use them as the upload data characteristics.
3. The middleware data index positioning and retrieval method according to claim 2, wherein: In the step S1, based on the upload data characteristics, adjusting the storage status of the upload data by the middleware includes: Obtain the data storage history records of all storage intervals in the middleware, analyze the data storage history records, and obtain the stored data types and quantities of each storage interval in the middleware; Compare the upload data types and quantities of the data source ends to the middleware with the stored data types and quantities of each storage interval to determine the storage intervals that match the upload data of the data source ends; Based on the upload data rate of the data source end, adjust the data storage transmission bandwidth allocated by the middleware to the matching storage interval.
4. The middleware data index positioning and retrieval method according to claim 3, wherein: In the step S1, based on the data stock status of all storage intervals in the middleware, determining the retrieval priorities of all storage intervals includes: Obtain the data storage amount and data semantic information of each storage range in the middle platform, and determine the content duplication characteristics of the stored data in each storage range based on the data storage amount and the data semantic information; wherein, the content duplication characteristic refers to the proportion of data with the same or similar content in each storage range. Determine the retrieval priority for all storage ranges based on the content duplication characteristics; wherein, the retrieval priority includes the retrieval order for all storage ranges and the retrieval duration for each storage range.
5. The middle platform data index positioning and retrieval method according to claim 1, characterized in that: In the step S2, based on the query request of the user terminal, determine the target data attribute information required by the user terminal, including: Based on the identity information of the user terminal requesting access to the middle platform, determine the data retrieval history record of the user terminal for the middle platform; based on the data retrieval history record, determine the retrieval occupation time information of the user terminal in the middle platform, and thereby judge whether the user terminal has the retrieval permission. When the user terminal has the retrieval permission, parse the query request of the user terminal to obtain the target data index information required by the user terminal, and use this as the target data attribute information; wherein, the target data index information includes the keyword semantic information of the target data required by the user terminal.
6. The middle platform data index positioning and retrieval method according to claim 5, characterized in that: In the step S2, based on the target data attribute information and the retrieval priority, determine the query path for the corresponding storage range, and thereby perform data browsing marking on the storage range, including: Compare the target data index information with the storage directories of all storage ranges in the middle platform, and screen out several target storage ranges; based on the retrieval priority, determine the retrieval order for all target storage ranges and the retrieval duration for each target storage range. Based on the retrieval order for all target storage ranges and the retrieval duration for each target storage range, determine the query path for all target storage ranges; based on the query path, identify the data browsed in the target storage range, and thereby add an identification code to the data that matches the keyword semantic information.
7. The middle platform data index positioning and retrieval method according to claim 1, characterized in that: In the step S3, based on the data browsing marking, judge whether the target data required by the user terminal is successfully retrieved, including: Compare the creation time of the data with the identification code added in the storage range with a preset time range. If the creation time is within the preset time range, judge that the target data required by the user terminal is successfully retrieved; otherwise, judge that the target data required by the user terminal is not successfully retrieved.
8. The middle platform data index positioning and retrieval method according to claim 7, characterized in that: In the step S3, when the retrieval is successful, generate a retrieval result data packet based on the distribution state of the target data in the middle platform and return it to the user terminal, including: When the retrieval is successful, based on the distribution locations and quantities of the target data in all storage intervals within the middleware, data replication extraction and integration compression are performed on the corresponding storage intervals to generate a retrieval result data packet, and the retrieval result data packet is returned to the client.
9. The middleware data indexing and positioning and retrieval method according to claim 1, wherein: In step S4, when the retrieval fails, based on the uploaded data characteristics, predict the data storage change attribute information of the middleware, including: When the retrieval fails, based on the uploaded data characteristics, predict the types and data volumes of the data to be uploaded to the middleware by the data source end in the future; based on the types and data volumes of the future uploaded data, predict the data storage change attribute information of all storage intervals within the middleware; wherein, the data storage change attribute information includes the increased data volumes of the corresponding type of data in all storage intervals within the middleware during the future time period.
10. The middleware data indexing and positioning and retrieval method according to claim 9, wherein: In step S4, based on the data storage change attribute information and the target data attribute information, perform data retrieval again, and return the retrieval result to the client, including: Compare the increased data volume with a preset data volume threshold. If the increased data volume exceeds the preset quantity threshold, it is determined that data retrieval is allowed to be performed again; otherwise, it is determined that data retrieval is not allowed to be performed again; When data retrieval is allowed to be performed again, based on the target data attribute information, retrieve the newly added data in all storage intervals within the middleware, and return the retrieval result to the client.
Citation Information
Patent Citations
Middle-station data query method, device and equipment and storage medium
CN114090832A
Power data flow direction monitoring and analyzing method based on data center
CN117041313A
Search method
WO2021042564A1