Data pushing method and device based on data blood relationship, equipment and medium
By acquiring user data through the data management platform and building information entity associations based on data lineage relationships, the problem of low efficiency in acquiring exploration and development data is solved, and efficient data push and management are achieved.
Patent Information
- Application Number
- CN202410319604.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2025-09-26
AI Technical Summary
The efficiency of data acquisition in the exploration and development stage is low, offline data copying poses security risks and data dissemination is blocked, and the existing data management platform lacks unified standards, making it difficult to effectively query target data.
After detecting that the user has logged in successfully through the data management platform, the user attribute data, user behavior data and historical usage data are obtained, and information entity associations are built based on data lineage relationships. The push data set is determined and merged to generate push results.
It improves the management efficiency and push hit rate of exploration and development data, and improves the work efficiency of exploration and development personnel.
Smart Images

Figure CN120705186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a data push method, device, electronic device and storage medium based on data lineage relationship. Background Art
[0002] In the field of oil exploration, researchers in the exploration and development stages need to collect various types of basic data for research.
[0003] Currently, business researchers need to contact the data source and copy the data offline for research and application. For data exchange between upstream and downstream businesses, researchers need to contact and copy the data offline for research and application. Alternatively, after a third-party unit generates basic research data, it is reviewed by relevant personnel and uploaded to an existing data management platform for data aggregation. Once the data is released, researchers can log in to the data management platform to query the data and use the required data after their application is approved.
[0004] However, offline data copying poses data security risks. Furthermore, data dissemination is often fragmented, making it difficult for researchers to identify the source of the data they need, making it extremely difficult to obtain. Existing data management platforms often lack unified data standards, making it difficult for users to effectively query their target data. Therefore, a data push method is urgently needed to improve data acquisition efficiency for exploration and development personnel. Summary of the Invention
[0005] The present invention provides a data push method, device, equipment and storage medium based on data lineage to solve the problem of low efficiency in acquiring exploration and development data. It can effectively improve the management efficiency of exploration and development data, increase the push hit rate of exploration and development data, and greatly improve the work efficiency of exploration and development personnel.
[0006] According to one aspect of the present invention, a data push method based on data lineage is provided, the method being executed by a data management platform, the method comprising:
[0007] If a successful login is detected, user attribute data, user behavior data, and at least one piece of historical usage data of the user are obtained;
[0008] Determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user;
[0009] A data push result is determined according to the first push data set, the second push data set, and the third push data set.
[0010] According to another aspect of the present invention, a data push device based on data lineage is provided, the device being configured on a data management platform and comprising:
[0011] A user data acquisition module, configured to acquire user attribute data, user behavior data, and at least one piece of historical usage data of the user if successful login information of the user is detected;
[0012] a push data set determination module, configured to determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user;
[0013] The push result determination module is configured to determine a data push result according to the first push data set, the second push data set, and the third push data set.
[0014] According to another aspect of the present invention, an electronic device is provided, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data push method based on data lineage relationship described in any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data push method based on data lineage relationship described in any embodiment of the present invention when executed.
[0019] The technical solution of an embodiment of the present invention obtains the user's user attribute data, user behavior data, and at least one piece of historical usage data upon detecting a successful user login; determines a first push data set based on each piece of historical usage data; determines a second push data set based on the first push data set; determines a third push data set based on the user's user attribute data, user behavior data, and at least one piece of historical usage data; and determines a data push result based on the first push data set, the second push data set, and the third push data set. This technical solution solves the problem of low efficiency in acquiring exploration and development data, effectively improves the management efficiency of exploration and development data, increases the push hit rate of exploration and development data, and significantly enhances the work efficiency of exploration and development personnel.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a flow chart of a data push method based on data lineage relationship provided according to the first embodiment of the present invention;
[0023] Figure 2 This is a flow chart of a data push method based on data lineage relationship provided according to the second embodiment of the present invention;
[0024] Figure 3A This is a schematic diagram of intelligent data push provided according to a specific applicable scenario 1 of the present invention;
[0025] Figure 3B This is a schematic diagram of user high-frequency usage data push provided according to a specific applicable scenario 1 of the present invention;
[0026] Figure 3C This is a schematic diagram of similar data push of user's frequently used data provided according to the specific applicable scenario 1 of the present invention;
[0027] Figure 3D This is a schematic diagram of pushing high-frequency usage data of similar users provided according to the first specific application scenario of the present invention;
[0028] Figure 4This is a structural diagram of a data push device based on data lineage relationship provided according to the third embodiment of the present invention;
[0029] Figure 5 It is a structural diagram of an electronic device that implements the data push method based on data lineage relationship according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0032] Example 1
[0033] Figure 1 The present invention provides a flowchart of a data push method based on data lineage relationship for the first embodiment. This embodiment is applicable to the retrieval scenario of exploration and development data. The method can be executed by a data push device based on data lineage relationship. The device can be implemented in the form of hardware and / or software, and the device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0034] S110: If successful login information of the user is detected, user attribute data, user behavior data, and at least one piece of historical usage data of the user are obtained.
[0035] This solution can be executed by a data management platform, which can obtain user attribute information and user behavior information of registered users, as well as exploration and development data recorded in the exploration and development database, and application tool information provided by the data management platform. User attribute information may include information such as the user's research area, business type, project group, and projects participated in. User behavior information may include information such as the exploration and development data viewed by the user, the application tools used, and the functional types of the application tools used. The exploration and development data may include exploration and development data uploaded by users, exploration and development data obtained by the data management platform from public data sharing platforms, or exploration and development data generated autonomously by the data management platform based on data simulation models. Application tool information may include information such as the application tool's name, function, input and output data, and development team.
[0036] The data management platform can construct data lineage relationships between information entities based on user attribute information, user behavior information, exploration and development data, and application tool information, and store these data lineage relationships to facilitate exploration and development data push based on these data lineage relationships. Specifically, these information entities may include users, projects, research areas, exploration and development data, and application tools. These data lineage relationships may include relationships between users and exploration and development data, between users and projects, between users and application tools, between exploration and development data and application tools, between projects and exploration and development data, and between projects and research areas.
[0037] If a successful user login is detected, the data management platform can determine the user's identity based on the user's login information and obtain the user's user attribute data, user behavior data, and at least one piece of historical usage data. The historical usage data can be the exploration and development data used by the user within a preset period, such as the exploration and development data accessed by user A within 30 days.
[0038] S120. Determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data and at least one piece of historical usage data of the user.
[0039] By performing statistical analysis on the usage information of each piece of historical usage data of the user, the data management platform can determine a first push data set. It is understandable that the usage information may include information such as the usage time and number of times of the historical usage data. The data management platform can perform statistics on the usage information of each piece of historical usage data, such as calculating the usage time ratio of each piece of historical usage data. The first push data set can be a set of historical usage data whose usage information statistics meet preset data screening conditions, such as a set of historical usage data whose usage time ratio exceeds a preset ratio threshold within a preset period, a set of historical usage data whose usage times exceed a preset number threshold within a preset period, etc.
[0040] It is understood that the data management platform can, based on data lineage relationships, determine exploration and development data associated with each historical usage data item in the first push data set, and then determine the second push data set based on the exploration and development data. The data management platform can also generate user characteristics for the user based on the user attribute data, user behavior data, and at least one piece of historical usage data, and, based on the data lineage relationships, determine target users that match the user characteristics, and then determine the third push data set based on the target user's historical usage data.
[0041] S130: Determine a data push result according to the first push data set, the second push data set, and the third push data set.
[0042] After obtaining the first, second, and third push data sets, the data management platform can merge the push data sets to obtain a data push result. Specifically, the data management platform can sequentially arrange the exploration and development data in each push data set according to a preset sorting order to generate a data push list. The data management platform can also filter the exploration and development data in each push data set according to preset filtering rules and generate a data push list based on the filtering results. For example, the data management platform can select one of two exploration and development data sets whose similarity exceeds a preset similarity threshold, thereby providing users with diverse push data.
[0043] The technical solution of an embodiment of the present invention obtains the user's user attribute data, user behavior data, and at least one piece of historical usage data upon detecting a successful user login; determines a first push data set based on each piece of historical usage data; determines a second push data set based on the first push data set; determines a third push data set based on the user's user attribute data, user behavior data, and at least one piece of historical usage data; and determines a data push result based on the first push data set, the second push data set, and the third push data set. This technical solution solves the problem of low efficiency in acquiring exploration and development data, effectively improves the management efficiency of exploration and development data, increases the push hit rate of exploration and development data, and significantly enhances the work efficiency of exploration and development personnel.
[0044] Example 2
[0045] Figure 2 This is a flow chart of a data push method based on data lineage relationship provided by the second embodiment of the present invention. This embodiment is based on the above embodiment and is refined. Figure 2 As shown, the method includes:
[0046] S210: If successful login information of the user is detected, user attribute data, user behavior data, and at least one piece of historical usage data of the user are obtained.
[0047] S220: Determine a first push data set based on each piece of historical usage data.
[0048] In this solution, optionally, determining the first push data set based on each piece of historical usage data includes:
[0049] Determine the usage frequency of each piece of historical usage data within a preset period;
[0050] A first push data set is determined according to the usage frequency.
[0051] In this solution, the data management platform can count the usage frequencies of each piece of historical usage data within a preset period, and can add historical usage data with a usage frequency higher than a preset frequency threshold to the first push data set based on the usage frequency. It can also sort the usage frequencies of each piece of historical usage data, and select a preset number of historical usage data based on the sorting results to add to the first push data set.
[0052] S230: Determine a second push data set based on the first push data set.
[0053] On the basis of the above solution, determining the second push data set according to the first push data set includes:
[0054] Taking each piece of historical usage data in the first pushed data set as target data in turn, and determining at least one candidate data having a data lineage relationship with the target data;
[0055] Determine the similarity between the target data and each piece of candidate data, and determine a second push data set based on the similarity.
[0056] As will be readily understood, the data management platform can sequentially use each piece of historical usage data in the first pushed data set as target data, identify exploration and development data that has a data lineage relationship with the target data in the exploration and development database of the data management platform, and use the associated exploration and development data as candidate data. Specifically, the candidate data that has a data lineage relationship with the target data can be data from the same source as the target data.
[0057] The data management platform can extract the data features of the target data and each candidate data, and calculate the similarity between the target data and each candidate data based on the data features of the target data and the data features of each candidate data. The data management platform can perform feature extraction on the target data and each candidate data to obtain a feature vector that matches each data. The similarity can be determined based on the distance between the feature vector of the target data and the feature vector of each candidate data, such as Euclidean distance, Mahalanobis distance, etc. The data management platform can also calculate the cosine similarity between the feature vector of the target data and the feature vector of each candidate data, and use the cosine similarity as the similarity between the target data and each candidate data. Based on the similarity between the target data and each candidate data, the data management platform can add the candidate data whose similarity with the target data is greater than a preset similarity threshold to the second push data set, or can sort the similarity of each candidate data match, and select a preset number of candidate data according to the sorting result to add to the second push data set.
[0058] S240: Determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user.
[0059] In a feasible solution, determining the third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user includes:
[0060] Determining user characteristics of the user based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user;
[0061] Determining, based on the user characteristics, at least one candidate user who has a data lineage relationship with the user;
[0062] Obtain user attribute data, user behavior data, and at least one piece of historical usage data for each candidate user, and determine user characteristics for each candidate user;
[0063] Determining the similarity between the user and each candidate user based on the user characteristics of the user and the user characteristics of each candidate user;
[0064] Determining at least one target user from among the candidate users based on similarities between the user and the candidate users;
[0065] A third push data set is determined based on the historical usage data of each target user.
[0066] The data management platform can generate user characteristics of the user based on the user attribute data, user behavior data, and at least one historical usage data of the user. At the same time, the data management platform can determine the user characteristics of other registered users. Based on the user characteristics of the user, one or more candidate users with whom the user has a data lineage relationship can be determined. Similar to the method of generating the user characteristics of the user, the data management platform can obtain the user attribute data, user behavior data, and at least one historical usage data of each candidate user, and determine the user characteristics of each candidate user. Based on the user characteristics of the user and the user characteristics of each candidate user, the data management platform can calculate the similarity between the user and each candidate user.
[0067] Based on the similarity between the user and each candidate user, the data management platform may determine as target users the candidate users whose similarity with the user exceeds a preset similarity threshold, or may determine a preset number of target users from each candidate user based on the similarity ranking results of the candidate users.
[0068] The data management platform can perform statistical analysis on the usage information of the historical usage data of each target user to obtain a third push data set. Specifically, the usage information may include information such as the usage time and number of times of the historical usage data. The data management platform can perform statistics on the usage information of each piece of historical usage data, such as calculating the usage time ratio of each piece of historical usage data. The third push data set can be a set of historical usage data of the target user whose usage information statistics meet the preset data screening conditions, such as a set of historical usage data whose usage time ratio exceeds a preset ratio threshold within a preset period, a set of historical usage data whose usage times exceed a preset number threshold within a preset period, etc.
[0069] In a preferred solution, the user characteristics include at least one characteristic indicator;
[0070] The determining, based on the user characteristics of the user and the user characteristics of the candidate users, the similarity between the user and the candidate users includes:
[0071] Determining the blood relationship evaluation value of each characteristic indicator in the user characteristics of the user and each candidate user in sequence;
[0072] Determining the similarity between the user and each candidate user based on the blood relationship evaluation value of each characteristic indicator and the predetermined weight coefficient of each characteristic indicator;
[0073] The weight coefficient is determined based on the statistical results of historical data and the random forest algorithm.
[0074] It is easy to understand that the user characteristics may include one or more characteristic indicators such as business type relevance, research area overlap, historical usage data similarity, and application tool usage similarity. The blood relationship evaluation value of business type relevance can be determined based on a pre-set evaluation standard. For example, business A and business B have a strong correlation, and the blood relationship evaluation value is 10. Business A and business C have a poor correlation, and the blood relationship evaluation value is 1. The blood relationship evaluation value of business type relevance can also be determined based on the degree of business overlap associated with the two business types, or based on the parent-child relationship between the research areas. The data management platform can use the ratio of the number of historical usage data that the user and the candidate user have in common to the total amount of historical usage data of the user as the historical usage data similarity. Similarly, the data management platform can use the ratio of the number of application tools that the user and the candidate user have in common to the total amount of application tools used by the user as the application tool usage similarity.
[0075] It is understandable that the data management platform can match a weight coefficient for each feature indicator, wherein the weight coefficient can be the degree of influence of each feature indicator on the user's similarity determined by the data management platform based on the statistical results of historical data and the random forest algorithm. After obtaining the blood relationship evaluation value of each feature indicator and the weight coefficient of each feature indicator, the data management platform can multiply the blood relationship evaluation value of each feature indicator by the weight coefficient, and then add the weighted evaluation values of each feature indicator to obtain the similarity between the user and each candidate user. Specifically, the similarity calculation formula between the user and each candidate user can be expressed as:
[0076]
[0077] Among them, i, j represent user indexes, sim i,j represents the similarity between user i and user j, k represents the feature index index, n represents the number of feature indicators, W krepresents the weight coefficient of the characteristic index k, Represents the evaluation value of the kinship relationship between user i and user j based on feature index k.
[0078] In this embodiment, optionally, determining the third push data set based on the historical usage data of each target user includes:
[0079] A third push data set is determined based on the usage frequency of the historical usage data of each target user within a preset period.
[0080] After obtaining the historical usage data of each target user, the data management platform can count the usage frequency of the historical usage data of each target user within a preset period, and add the historical usage data with a usage frequency higher than the preset frequency threshold to the third push data set. It can also sort the usage frequency of each historical usage data, and select a preset number of historical usage data based on the sorting results to add to the third push data set.
[0081] In a specific example, the calculation formula for the usage frequency can be expressed as:
[0082] Among them, m represents the historical usage data index, f m Indicates the usage frequency of historical usage data m, I m represents the number of times the historical usage data m is used, and I represents the total number of times each historical usage data of each target user is used.
[0083] The above solution can push exploration and development data that meets user needs to users in a targeted manner by determining the exploration and development data that users use frequently, the exploration and development data that is similar to the exploration and development data that users use frequently, and the exploration and development data that is used frequently by similar users, so as to improve the efficiency of users in obtaining exploration and development data.
[0084] S250: Determine a data push result according to the first push data set, the second push data set, the third push data set, and a predetermined interest level of each push data set.
[0085] If the sum of the number of exploration and development data in the first, second, and third push data sets is less than or equal to the maximum number of push data, the data management platform may push all the exploration and development data in each push data set. If the sum of the number of exploration and development data in the first, second, and third push data sets is greater than the maximum number of push data, the data management platform may filter the exploration and development data in each push data set and select the exploration and development data that is most likely to meet the user's needs for push based on the interest level of each push data set. Specifically, the data management platform may determine the user's level of interest in each push data set based on historical click records of the pushed data. The interest level is determined based on the historical hit probability of each push data set.
[0086] For example, in the historical click history of pushed data, the historical hit probability of exploration and development data in the first pushed data set is 40%, the historical hit probability of exploration and development data in the second pushed data set is 80%, and the historical hit probability of exploration and development data in the third pushed data set is 80%. The user's interest level in each pushed data set can be 0.2, 0.4, and 0.4, respectively. The data push results allow a maximum of 10 pushed data items to be displayed. The data management platform can select 2 exploration and development data items from the first pushed data set and 4 exploration and development data items from both the second and third pushed data sets to form a data push list for the exploration and development data push.
[0087] The calculation formula for data push results can be expressed as:
[0088] in, t represents the push data set index, n=3, W t Indicates the interest level of the pushed data set t, Rec t Represents the push data set t.
[0089] The technical solution of an embodiment of the present invention obtains the user's user attribute data, user behavior data, and at least one piece of historical usage data upon detecting a successful user login; determines a first push data set based on each piece of historical usage data; determines a second push data set based on the first push data set; determines a third push data set based on the user's user attribute data, user behavior data, and at least one piece of historical usage data; and determines a data push result based on the first push data set, the second push data set, and the third push data set. This technical solution solves the problem of low efficiency in acquiring exploration and development data, effectively improves the management efficiency of exploration and development data, increases the push hit rate of exploration and development data, and significantly enhances the work efficiency of exploration and development personnel.
[0090] Specific application scenario 1
[0091] This embodiment is a specific embodiment based on the above embodiment. In this solution, the research area of the data management platform is managed at three levels: basin, block, and tectonic belt. The overlap of the user's research area can be determined based on the parent-child relationship between the research areas. For example, the research area of the target user's project is the A11 tectonic belt and the B1 block, the research area of the project of user A is the A11 tectonic belt and the B2 block, the research area of the project of user B is the A12 tectonic belt, and the research area of the project of user C is the C1 block. According to the weight coefficients of the same research area, the same parent block, and the same parent basin, the weight coefficient of the same research area is reduced in sequence. This description assumes that the weight coefficient of the same research area is W1=0.5, the weight coefficient of the same parent block is W2=0.3, and the weight coefficient of the same parent basin is W3=0.2. Then the overlap of the research area of the target user and user A can be expressed as sim 目标,A =0.5×1+0.3×0+0.2×1=0.7, the overlap between the target user and user B’s research area can be expressed as sim 目标,B =0.5×0+0.3×1+0.2×0=0.3, the overlap between the target user and user C’s research area can be expressed as sim 目标,C =0.5×0+0.3×0+0.2×0=0.
[0092] Figure 3A This is a schematic diagram of intelligent data push provided according to a specific applicable scenario 1 of the present invention, such as Figure 3A As shown, user A logs in to the data management platform to obtain the user's historical work habits. The software a used by the user, such as LandMark, belongs to the business type of "seismic interpretation". The research area included in the user's project is "Basin A". The user uses "Work Area A" in the "LandMark" software to carry out research work on "Basin A".
[0093] When data is stored in the data management platform, there will be relevant tags for research area, professional software and business type. The system uses the tags to perform correlation analysis on the data and obtain the relevant data of user A, that is, the intersection of data with the tags "A Basin", "Seismic Interpretation" and "LandMark", and push it to user A for selection and use.
[0094] Figure 3B This is a schematic diagram of pushing user high-frequency data provided by a specific applicable scenario 1 of the present invention. For pushing user high-frequency data, such as Figure 3B As shown, the target user logs into the data management platform historically and performs four types of tasks.
[0095] (1) Enter the collected well data into the platform, including well head data, well deviation data, logging curve data, layer data, lithology data, well related maps and well related documents, and the subsequent use, viewing and downloading of various types of data are as follows: Figure 3B As shown;
[0096] (2) Use software a, such as professional software GeoEast, to call the input wellhead data, well deviation data, logging curve data and layer data, and conduct research on the survey lines and survey areas in the software work area. During the research process, synthetic record data is generated, and the number of times each type of data is used is as follows: Figure 3B As shown;
[0097] (3) The intermediate data generated during the research process are archived, and the number of subsequent views and downloads of various types of data is as follows: Figure 3B As shown;
[0098] (4) During the research process, the results of other projects need to be referenced and used, and the number of subsequent viewing and downloading of various types of data is as follows: Figure 3B shown.
[0099] When the target user logs in to the platform again, a collection of frequently used data is obtained based on the frequency of use of their historical data, and data is intelligently pushed to the target user.
[0100] Figure 3C This is a schematic diagram of pushing similar data of user's frequently used data provided by the specific applicable scenario 1 of the present invention. For pushing similar data of user's frequently used data, such as Figure 3C As shown, the target user and the other four users log in to the platform respectively and use different professional software to obtain the same original data for research. The data A produced by the target user is the target data, and the data BCDE produced by the other four users are all homologous data to be analyzed.
[0101] (1) The target user uses software a to perform seismic interpretation based on target data A, such as the professional software GeoEast. The output data are the layer data related to layers H1-H6, interpretation reports and cross-sectional drawings;
[0102] (2) Other users use software a to perform seismic interpretation based on the same source data B, such as the professional software GeoEast, and the output data are fault data, interpretation reports and cross-sectional drawings related to faults F1-F6;
[0103] (3) Other users use software b to perform seismic interpretation based on the same source data C, such as the professional software LandMark, and the output data are the relevant layer data of layers H1-H6, interpretation reports and cross-section drawings;
[0104] (4) Other users use software C to perform geological modeling based on the same source data D, such as the professional software Petrel, and the output data are polygon data related to polygons P1-P8, interpretation reports and stratigraphic models;
[0105] (5) Other users use software d to carry out geological mapping based on the same source data E, such as the professional software GeoMap, and the output data is a series of plan views related to the layers H1-H6.
[0106] The similarity ranking of the above homologous data can be as follows Figure 3C When the target user logs into the data management platform again, the target user will be pushed homologous data with high similarity for selection and use.
[0107] Through the user's historical high-frequency use data set, traverse each data in the set ( Figure 3C ) and find the homologous data of the target data in the data lineage relationship (homologous data means that the data is converted from the same data), and use the following steps to obtain the similar data set of the data.
[0108] Step 1: Filter out homologous data with the same software or business-related software source as the target data. If the software source is different, the similarity value is set to 0;
[0109] Step 2: Calculate the cosine similarity between the target data and the homologous data from the same software source, sort them from high to low according to the similarity value, and select the top K data with the highest similarity (the K value can be set flexibly, for example, 3).
[0110] Figure 3D This is a schematic diagram of pushing high-frequency usage data of similar users provided according to a specific applicable scenario 1 of the present invention. For pushing high-frequency usage data of similar users, as shown in FIG. Figure 3D As shown in the figure, the target user and other users A, B and C log in to the platform and perform a series of operations on the platform, such as "using software", "conducting research work on the software work area", and "archiving project research results". The target objects of each operation are recorded by the system with labels (professional software, research area, business type).
[0111] (1) Target users: software a, study area X, reservoir evaluation;
[0112] (2) Other user A: b software, Y study area, reservoir evaluation;
[0113] (3) Other user B: c software, Z study area, trap evaluation;
[0114] (4) Other users C: d software, X study area, reservoir evaluation;
[0115] Specifically, software a may be GeoEast, which performs reservoir evaluation in the study area through seismic processing or seismic interpretation; software b may be LandMark, which performs oil reservoir evaluation in the study area through seismic processing or seismic interpretation; software c may be PetroMod, which performs trap evaluation in the study area through basin simulation; and software d may be Petrel, which performs reservoir evaluation in the study area through geological modeling.
[0116] Among the above users, user A and user C are similar users, and the high-frequency data they generate are sorted as follows: Figure 3D When the target user logs into the data management platform again, a collection of high-frequency data of similar users will be pushed to him / her for selection and use.
[0117] Example 3
[0118] Figure 4 This is a structural diagram of a data push device based on data lineage relationship provided by the third embodiment of the present invention. Figure 4 As shown, the device includes:
[0119] The user data acquisition module 310 is configured to acquire user attribute data, user behavior data, and at least one piece of historical usage data of the user if successful login information of the user is detected;
[0120] A push data set determination module 320 is configured to determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user;
[0121] The push result determination module 330 is configured to determine a data push result according to the first push data set, the second push data set, and the third push data set.
[0122] In this solution, optionally, the push set determination module 320 includes a first push set determination unit, and the first push set determination unit is configured to:
[0123] Determine the usage frequency of each piece of historical usage data within a preset period;
[0124] A first push data set is determined according to the usage frequency.
[0125] Based on the above solution, optionally, the push set determination module 320 includes a second push set determination unit, and the second push set determination unit is configured to:
[0126] Taking each piece of historical usage data in the first pushed data set as target data in turn, and determining at least one candidate data having a data lineage relationship with the target data;
[0127] Determine the similarity between the target data and each piece of candidate data, and determine a second push data set based on the similarity.
[0128] In a feasible solution, the push set determination module 320 includes a third push set determination unit, and the third push set determination unit includes:
[0129] A user feature determination subunit, configured to determine a user feature of the user based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user;
[0130] a candidate user determination subunit, configured to determine, based on the user characteristics, at least one candidate user having a data lineage relationship with the user;
[0131] A candidate user feature determination subunit is configured to obtain user attribute data, user behavior data, and at least one piece of historical usage data of each candidate user, and determine user features of each candidate user;
[0132] a similarity determination subunit, configured to determine the similarity between the user and each candidate user based on the user characteristics of the user and the user characteristics of each candidate user;
[0133] a target user determination subunit, configured to determine at least one target user from among the candidate users based on similarities between the user and the candidate users;
[0134] The third push data set determining subunit is configured to determine a third push data set based on the historical usage data of each target user.
[0135] Based on the above solution, optionally, the user characteristics include at least one characteristic indicator;
[0136] The similarity determination subunit is specifically used to:
[0137] Determining the blood relationship evaluation value of each characteristic indicator in the user characteristics of the user and each candidate user in sequence;
[0138] Determining the similarity between the user and each candidate user based on the blood relationship evaluation value of each characteristic indicator and the predetermined weight coefficient of each characteristic indicator;
[0139] The weight coefficient is determined based on the statistical results of historical data and the random forest algorithm.
[0140] In a feasible solution, the third push set determination subunit is specifically configured to:
[0141] A third push data set is determined based on the usage frequency of the historical usage data of each target user within a preset period.
[0142] In a preferred solution, the push result determination module 330 is specifically configured to:
[0143] determining a data push result according to the first push data set, the second push data set, the third push data set, and a predetermined degree of interest of matching among the push data sets;
[0144] The interest level is determined based on the historical hit probability of each push data set.
[0145] The data push device based on data lineage provided by the embodiment of the present invention can execute the data push method based on data lineage provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0146] Example 4
[0147] Figure 5 A schematic diagram of the structure of an electronic device 410 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0148] like Figure 5As shown, the electronic device 410 includes at least one processor 411, and a memory connected to the at least one processor 411, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 411 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 412 or the computer program loaded from the storage unit 418 to the random access memory (RAM) 413. Various programs and data required for the operation of the electronic device 410 can also be stored in the RAM 413. The processor 411, ROM 412 and RAM 413 are connected to each other via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.
[0149] Multiple components in electronic device 410 are connected to I / O interface 415, including an input unit 416, such as a keyboard, mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a magnetic disk, optical disk, etc.; and a communication unit 419, such as a network card, modem, wireless communication transceiver, etc. The communication unit 419 allows electronic device 410 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0150] Processor 411 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Processor 411 executes the various methods and processes described above, such as the data push method based on data lineage.
[0151] In some embodiments, the data push method based on data lineage relationship can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded into the RAM 413 and executed by the processor 411, one or more steps of the data push method based on data lineage relationship described above can be performed. Alternatively, in other embodiments, the processor 411 can be configured to execute the data push method based on data lineage relationship by any other appropriate means (for example, by means of firmware).
[0152] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] Computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data lineage-based data push device, so that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0156] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0157] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0158] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0159] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A data push method based on data lineage relationship, characterized in that: The method is performed by a data management platform, and includes: If a successful login is detected, user attribute data, user behavior data, and at least one piece of historical usage data of the user are obtained; Determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user; A data push result is determined according to the first push data set, the second push data set, and the third push data set.
2. The method according to claim 1, characterized in that The determining of the first push data set according to each piece of historical usage data includes: Determine the usage frequency of each piece of historical usage data within a preset period; A first push data set is determined according to the usage frequency.
3. The method according to claim 2, characterized in that The determining, based on the first push data set, a second push data set includes: Taking each piece of historical usage data in the first pushed data set as target data in turn, and determining at least one candidate data having a data lineage relationship with the target data; Determine the similarity between the target data and each piece of candidate data, and determine a second push data set based on the similarity.
4. The method according to claim 3, characterized in that The determining of the third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user includes: Determining user characteristics of the user based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user; Determining, based on the user characteristics, at least one candidate user who has a data lineage relationship with the user; Obtain user attribute data, user behavior data, and at least one piece of historical usage data for each candidate user, and determine user characteristics for each candidate user; Determining the similarity between the user and each candidate user based on the user characteristics of the user and the user characteristics of each candidate user; Determining at least one target user from among the candidate users based on similarities between the user and the candidate users; A third push data set is determined based on the historical usage data of each target user.
5. The method according to claim 4, characterized in that The user characteristics include at least one characteristic indicator; The determining, based on the user characteristics of the user and the user characteristics of the candidate users, the similarity between the user and the candidate users includes: Determining the blood relationship evaluation value of each characteristic indicator in the user characteristics of the user and each candidate user in sequence; Determining the similarity between the user and each candidate user based on the blood relationship evaluation value of each characteristic indicator and the predetermined weight coefficient of each characteristic indicator; The weight coefficient is determined based on the statistical results of historical data and the random forest algorithm.
6. The method according to claim 4, characterized in that The determining of the third push data set based on the historical usage data of each target user includes: A third push data set is determined based on the usage frequency of the historical usage data of each target user within a preset period.
7. The method according to claim 1, characterized in that The determining of a data push result according to the first push data set, the second push data set, and the third push data set includes: determining a data push result according to the first push data set, the second push data set, the third push data set, and a predetermined degree of interest of matching among the push data sets; The interest level is determined based on the historical hit probability of each push data set.
8. A data push device based on data lineage relationship, characterized in that: The device is configured on a data management platform and includes: A user data acquisition module, configured to acquire user attribute data, user behavior data, and at least one piece of historical usage data of the user if successful login information of the user is detected; a push data set determination module, configured to determine a first push data set based on each piece of historical usage data, determine a second push data set based on the first push data set, and determine a third push data set based on the user attribute data, user behavior data, and at least one piece of historical usage data of the user; The push result determination module is configured to determine a data push result according to the first push data set, the second push data set, and the third push data set.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data push method based on data lineage relationship according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data push method based on data lineage relationship according to any one of claims 1 to 7 when executed.