Key data mining method and device based on big data

By obtaining mining scenario information and obtaining key data information based on this, the problem that existing data mining methods are prone to missing indirectly affecting data is solved, and a more comprehensive and accurate data mining effect is achieved.

CN120216567AActive Publication Date: 2025-06-27BEIJING ZHONGTIAN CHUANGYU MARKET CONSULTING CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510695204.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing data mining methods usually only focus on data that can directly affect the target, resulting in the easy missed data that can indirectly affect the target.

Method used

By obtaining mining scenario information, including mining projects and target objects, clarifying business goals and scope, and defining the boundaries and directions of data analysis. Then, based on this information, information reflecting the key data is obtained, thereby mining data that can indirectly affect the target object.

Benefits of technology

It improves the indirect data impact problems that are easily missed by existing data mining methods, and can more comprehensively mine key data and improve the accuracy and completeness of data mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216567A_ABST
    Figure CN120216567A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of key data mining, and particularly relates to a key data mining method and device based on big data, and the method comprises the steps: obtaining mining scene information; wherein the mining scene information comprises a mining item and a target object; obtaining data analysis information based on the mining scene information; wherein the data analysis information comprises at least one piece of information reflecting key data. According to the big data-based key data mining method and equipment provided by the embodiment of the invention, the problem that some data which can indirectly generate direct influence on a target and indirectly generate a result is easily missed in an existing data mining method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of key data mining, and particularly relates to a key data mining method and device based on big data. Background Art

[0002] In different fields, such as business, healthcare, finance, etc., there is a strong demand for mining key data. However, existing data mining methods usually only mine data that can directly affect the target and directly produce results, resulting in the easy omission of data that can indirectly affect the target and indirectly produce results. Summary of the Invention

[0003] The embodiments of this application provide a key data mining method and device based on big data, which can improve the problem that existing data mining methods are prone to missing data that can indirectly affect the target and indirectly produce results.

[0004] In a first aspect, the embodiments of this application provide a key data mining method based on big data, including: Obtaining mining scenario information; wherein, the mining scenario information includes a mining project and a target object; Obtaining data analysis information based on the mining scenario information; wherein, the data analysis information includes at least one piece of information reflecting key data.

[0005] The above technical solutions in the embodiments of this application have at least the following technical effects: The key data mining method based on big data provided by the embodiments of this application first obtains mining scenario information including a mining project and a target object to clarify the business objective and scope and define the boundary and direction of data analysis. Then, based on the mining scenario information, data analysis information including at least one piece of information reflecting key data is obtained to mine data that can indirectly affect the target object, improving the problem that existing data mining methods are prone to missing data that can indirectly affect the target and indirectly produce results.

[0006] In a second aspect, the embodiments of this application provide a key data mining system based on big data, including: An obtaining unit, configured to obtain mining scenario information; wherein, the mining scenario information includes a mining project and a target object; An analysis unit, configured to obtain data analysis information based on the mining scenario information; wherein, the data analysis information includes at least one piece of information reflecting key data.

[0007] In a third aspect, an embodiment of the present application provides a key data mining device based on big data, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above first aspects is implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in any one of the above first aspects is implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a key data mining device based on big data, the key data mining device based on big data is enabled to execute the key data mining method based on big data described in any one of the above first aspects.

[0010] It can be understood that the beneficial effects of the above second to fifth aspects can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 is a schematic flowchart of a key data mining method based on big data provided by an embodiment of the present application; Figure 2 is a schematic flowchart of step S200 in the key data mining method based on big data provided by an embodiment of the present application; Figure 3 is a schematic flowchart of step S2324 in the key data mining method based on big data provided by an embodiment of the present application; Figure 4 is a schematic structural diagram of a key data mining system based on big data provided by an embodiment of the present application; Figure 5 is a schematic structural diagram of a key data mining device based on big data provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0014] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0015] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0016] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0017] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0018] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0019] In different fields, such as business, medical, finance, etc., there is a strong demand for mining key data. However, existing data mining methods usually only mine data that can have a direct impact on the target and directly produce results, resulting in the easy omission of some data that can indirectly have a direct impact on the target and indirectly produce results.

[0020] To solve the above problems, the embodiments of the present application provide a key data mining method and device based on big data. In this method, by first obtaining mining scenario information including mining projects and target objects, the business objectives and scope are clarified, and the boundaries and directions of data analysis are defined. Then, based on the mining scenario information, data analysis information including at least one piece of information reflecting key data is obtained, and data that can indirectly affect the target object is mined, improving the problem that existing data mining methods are prone to missing some data that can indirectly have a direct impact on the target and indirectly produce results.

[0021] The key data mining method based on big data provided by the embodiments of the present application can be applied to a key data mining device based on big data. At this time, the key data mining device based on big data is the execution subject of the key data mining method provided by the embodiments of the present application, and the specific type of the key data mining device based on big data is not limited in any way by the embodiments of the present application.

[0022] For example, the key data mining device based on big data can be a mobile phone, a tablet computer, a laptop computer, a handheld device with wireless communication functions, a computing device, or other processing devices connected to a wireless modem, etc., but not limited thereto.

[0023] To better understand the key data mining method based on big data provided by the embodiments of the present application, the following provides an exemplary introduction to the specific implementation process of the key data mining method based on big data provided by the embodiments of the present application.

[0024] Figure 1 The schematic flowchart of the key data mining method based on big data provided by the embodiments of the present application is shown. The key data mining method based on big data includes: S100, obtaining mining scenario information; wherein, the mining scenario information includes mining projects and target objects.

[0025] It can be understood that the way to obtain mining scenario information can be to receive data transmitted by the user, or to analyze historical data, etc., but not limited thereto. Obtaining mining scenario information can clarify the business objectives and scope, define the boundaries and directions of data analysis, and provide a basis for subsequent steps.

[0026] S200, obtain data analysis information based on mined scenario information; wherein, the data analysis information includes at least one keyword reflecting key data.

[0027] It can be understood that In a possible implementation manner, please refer to Figure 2 , S100, obtain mined scenario information, including: S210, obtain keyword information based on a mining project; wherein, the keyword information includes multiple keywords related to the mining project.

[0028] It can be understood that the manner of obtaining keyword information based on a mining project can be to receive data transmitted by a user, or to input the mining project into a preset database for query to obtain at least one keyword corresponding to the mining project, etc., but not limited thereto. Obtaining keyword information based on a mining project can provide a basis for subsequent steps.

[0029] Exemplarily, assuming the mining project is "online shopping", the keyword information related to "online shopping" may include product name, brand, model, specification, material, color, size, price, promotion, etc., but not limited thereto.

[0030] S220, obtain browsing information based on a target object; wherein, the browsing information includes multiple browsed records after data cleaning and the corresponding daily timestamps, and the browsed records include all search keywords of the target object within a single day and the corresponding search times, and all web page addresses visited by the target object within a single day, the corresponding content keywords, access times, and stay durations of each web page address.

[0031] It can be understood that the manner of obtaining browsing information based on a target object can be to receive data transmitted by a user, or to receive data transmitted by an APP or website authorized by the target object, etc., but not limited thereto. Obtaining browsing information based on a target object can provide a basis for subsequent steps.

[0032] S230, obtain data analysis information based on the keyword information and the browsing information.

[0033] It can be understood that the manner of obtaining data analysis information based on the keyword information and the browsing information can be to send the keyword information and the browsing information to a user and then receive the data transmitted by the user, or to analyze the number of data belonging to the keywords in the keyword information in the browsing information, etc., but not limited thereto. Obtaining data analysis information based on the keyword information and the browsing information can analyze the data that has an indirect impact on the target.

[0034] In a possible implementation manner, please refer to Figure 2, S230, obtain data analysis information based on keyword information and browsing information, including: S231, obtain interest duration information and interest frequency information based on browsing information; wherein, the interest duration information includes at least one interest duration, and the interest duration reflects the sum of the residence durations of the web page addresses corresponding to a certain content keyword in the browsing record, and the interest frequency information includes at least one interest frequency, and the interest frequency reflects the ratio of the number of web page addresses corresponding to a certain content keyword in the browsing record to the total number of web page addresses in the browsing record.

[0035] It can be understood that the way to obtain interest duration information based on browsing information can be to receive the data transmitted by the user, or to add up the residence durations of the web page addresses corresponding to the content keywords whose content keywords corresponding to the residence durations belong to the keyword information, etc., but not limited to this. The way to obtain interest frequency information based on browsing information can be to receive the data transmitted by the user, or to count the number of content keywords that are the same as the keywords in the keyword information in the browsing record, etc., but not limited to this. Obtaining interest duration information and interest frequency information based on browsing information can provide a basis for subsequent steps.

[0036] In a possible implementation manner, please refer to Figure 2 , S231, obtain interest duration information and interest frequency information based on browsing information, including: S2311, analyze each browsing record respectively, confirm the search time corresponding to each search keyword in the browsing record as the retrieval node, confirm the web page address adjacent to and before the retrieval node as the interest web page address, and confirm the web page address adjacent to and after the retrieval node as the verification web page address.

[0037] It can be understood that the search keyword refers to the text content in the search box when the target object performs a search operation in the APP or web page. The search time refers to the time when the target object performs a search operation in the APP or web page. The web page address adjacent to and before the retrieval node is the web page address that is less than the retrieval node in time and closest to the retrieval node in time (or the web page directly after browsing this web page and then performing a search operation). The web page address adjacent to and after the retrieval node is the web page address that is greater than the retrieval node in time and closest to the retrieval node in time (or the web page directly browsed after performing a search operation). Confirming the retrieval node, and confirming the web page address adjacent to and before the retrieval node as the interest web page address, and confirming the web page address adjacent to and after the retrieval node as the verification web page address can provide a basis for subsequent steps.

[0038] S2312. If the content keyword corresponding to the interest website is related to the content keyword corresponding to the verification website adjacent to the retrieval node adjacent to the interest website, after adding an interest label to the interest website and the verification website, starting from the verification website as the retrieval starting point, add interest labels to the web addresses in the chronological order reflected by the access time in sequence until the content keyword corresponding to the web address is not related to the content keyword corresponding to the interest website; wherein, the interest label is the content keyword corresponding to the interest website.

[0039] It can be understood that the way to determine whether the content keyword corresponding to the interest website is related to the content keyword corresponding to the verification website adjacent to the retrieval node adjacent to the interest website can be exact match (the two content keywords are exactly the same), fuzzy match (the two content keywords are partially the same or the semantics expressed by the two content keywords are the same, etc.), but not limited to this. The content keyword corresponding to the web address is not related to the content keyword corresponding to the interest website means that the two content keywords are completely different or the semantics expressed by the content keywords are different. If the content keyword corresponding to the interest website is related to the content keyword corresponding to the verification website adjacent to the retrieval node adjacent to the interest website, it means that the target object's interest remains consistent before and after performing the search operation. Adding interest labels to the interest website and the verification website, and adding interest labels to the web addresses in the chronological order reflected by the access time in sequence until the content keyword corresponding to the web address is not related to the content keyword corresponding to the interest website can provide a basis for subsequent steps.

[0040] S2313. After adding up the residence times corresponding to the web addresses with the same interest label of the content keyword respectively to obtain at least one interest duration corresponding to the content keyword corresponding to the interest label, confirm all the interest durations as interest duration information.

[0041] It can be understood that adding up the residence times corresponding to the web addresses with the same interest label of the content keyword respectively to obtain at least one interest duration corresponding to the content keyword corresponding to the interest label, and confirming all the interest durations as interest duration information can provide a basis for subsequent steps.

[0042] Exemplarily, assume that the content keywords corresponding to 3 interest labels are all "apple", and the residence times corresponding to the web addresses of these 3 interest labels are 10 seconds, 15 seconds, and 5 seconds respectively. Then the interest duration corresponding to "apple" = 10 + 15 + 5 = 30 seconds.

[0043] S2314. Analyze each browsing record respectively. After adding up the number of web addresses with the same content keyword in the browsing record and dividing by the number of all web addresses in the browsing record, the obtained value is confirmed as the interest frequency, and all the interest frequencies are confirmed as interest frequency information.

[0044] It can be understood that after adding up the number of web addresses with the same content keywords in the browsing record respectively and then dividing by the total number of web addresses in the browsing record to obtain the value of the interest frequency, and then regarding all the interest frequencies as the interest frequency information can provide a basis for the subsequent steps.

[0045] Exemplarily, assume that there are 10 web addresses in the browsing record, among which 3 web addresses have the content keyword "apple", and the other 7 web addresses have the content keyword "clothes". Then the interest frequency corresponding to "apple" = 3 / 10 = 0.3, and the interest frequency corresponding to "clothes" = 7 / 10 = 0.7.

[0046] S232. Obtain preselected data analysis information based on the interest duration information and the keyword information; wherein, the preselected data analysis information includes at least one first analysis word and at least one second analysis word, the first analysis word reflects a content keyword, and the second analysis word reflects a content keyword.

[0047] It can be understood that the way to obtain the preselected data analysis information based on the interest duration information and the keyword information can be to receive the data transmitted by the user, or to regard the content keywords corresponding to the interest durations whose time lengths are in the top 30% of all the interest durations and the corresponding content keywords belong to the keyword information as the preselected data analysis information, etc., but not limited to this. Obtaining the preselected data analysis information based on the interest duration information and the keyword information can provide a basis for the subsequent steps.

[0048] In a possible implementation manner, please refer to Figure 2 , S232. Obtain preselected data analysis information based on the interest duration information and the keyword information, including: S2321. Regard the interest durations corresponding to the content keywords that belong to the keyword information in the interest duration information as preference information, and regard the interest durations corresponding to the content keywords that do not belong to the keyword information in the interest duration information as preference comparison information.

[0049] It can be understood that the way to determine whether the content keyword belongs to the keyword information can be to compare each keyword in the keyword information with the content keyword by exact match or fuzzy match, etc., but not limited to this. Regarding the interest durations corresponding to the content keywords that belong to the keyword information in the interest duration information as preference information, and regarding the interest durations corresponding to the content keywords that do not belong to the keyword information in the interest duration information as preference comparison information can provide a basis for the subsequent steps.

[0050] S2322. Divide each interest duration in the preference information by the sum of all the residence times in the corresponding browsing records respectively to obtain multiple preference ratios. Divide each interest duration in the preference comparison information by the sum of all the residence times in the corresponding browsing records respectively to obtain multiple preference comparison ratios.

[0051] It can be understood that obtaining multiple preference ratios and multiple preference comparison ratios can provide a basis for subsequent steps.

[0052] Exemplarily, assume that there are 2 interest durations numbered 1 and 2 in the preference information. The interest duration numbered 1 is 10 seconds, the interest duration numbered 2 is 20 seconds, and the sum of all the residence times in the browsing records corresponding to the interest durations numbered 1 and 2 is 600 seconds. Then, the preference ratio corresponding to the interest duration numbered 1 = 10 / 600 = 0.017 (rounded to three decimal places), and the preference ratio corresponding to the interest duration numbered 2 = 20 / 600 = 0.033 (rounded to three decimal places).

[0053] S2323. Group each preference ratio and each preference comparison ratio according to the corresponding content keywords. After making the content keywords corresponding to the preference ratios or preference comparison ratios within each group the same, sort the preference ratios or preference comparison ratios within each group in descending order according to the daily time stamps of the corresponding browsing records to obtain a sorted table.

[0054] It can be understood that the sorted table can be an Excel table, a database table, etc., but is not limited thereto. Grouping each preference ratio and each preference comparison ratio according to the corresponding content keywords, making the content keywords corresponding to the preference ratios or preference comparison ratios within each group the same, and sorting the preference ratios or preference comparison ratios within each group in descending order according to the daily time stamps of the corresponding browsing records can improve the efficiency of subsequent data processing. Obtaining the sorted table can provide a basis for subsequent steps.

[0055] S2324. Obtain analysis word information based on the sorted table.

[0056] It can be understood that the method of obtaining analysis word information based on the sorted table can be to confirm the content keywords corresponding to the preference ratios or preference comparison ratios whose values gradually increase with time in the sorted table as the analysis word information, or to send the sorted table to the user and then receive the data transmitted by the user, etc., but is not limited thereto. Obtaining analysis word information based on the sorted table can provide a basis for subsequent steps.

[0057] In a possible implementation manner, please refer to Figure 2 , S2324. Obtain analysis word information based on the sorted table, including: S23241. Analyze each group in the sorted table respectively.

[0058] It can be understood that analyzing each group in the sorting table separately can ensure the reliability of the method.

[0059] S23242, Step 1: Confirm the nth preference ratio or preference comparison ratio within the group as the first value. If n + 1 is less than or equal to the number of preference ratios or preference comparison ratios within the group, confirm the (n + 1)th preference ratio or preference comparison ratio within the group as the second value. If n + 1 is greater than the number of preference ratios or preference comparison ratios within the group, then execute Step 3; where the initial value of n is 1.

[0060] It can be understood that if n + 1 is less than or equal to the number of preference ratios or preference comparison ratios within the group, it means that there are still preference ratios or preference comparison ratios within the group that have not been processed. If n + 1 is greater than the number of preference ratios or preference comparison ratios within the group, it means that all the preference ratios or preference comparison ratios within the group have been processed. Confirming the first value and the second value can provide a basis for subsequent steps.

[0061] Exemplarily, assume that there are 3 preference ratios within the group, and the 3 preference ratios are 0.1, 0.2, and 0.3 in sequence. Then when n = 1, the first value = 0.1, the second value = 0.2. When n = 2, the first value = 0.2, the second value = 0.3.

[0062] S23243, Step 2: Confirm the value obtained by subtracting the first value from the second value and then dividing by the first value as the change rate, and substitute n + 1 into n in Step 1 and repeat Step 1.

[0063] It can be understood that if the change rate is positive, it means that the interest of the target object in the analysis word corresponding to this preference ratio or preference comparison ratio is in an upward state during the time corresponding to the first value to the time corresponding to the second value. If the change rate is negative, it means that the interest of the target object in the analysis word corresponding to this preference ratio or preference comparison ratio is in a downward state during the time corresponding to the first value to the time corresponding to the second value. Confirming the change rate can provide a basis for subsequent steps.

[0064] Exemplarily, assume that the first value is 0.1 and the second value is 0.2. Then the change rate = (0.2 - 0.1) / 0.1 = 1.

[0065] S23244, Step 3: If all the data within the group are preference ratios, confirm all the obtained change rates as the preference change trend. If all the data within the group are preference comparison ratios, confirm all the obtained change rates as the comparison change trend.

[0066] It can be understood that confirming all the obtained change rates as the preference change trend or the comparison change trend can provide a basis for subsequent steps.

[0067] After all groups in the sorting table have been analyzed, the intervals formed by the values obtained by multiplying the change rates in each preference change trend by 1.05 and by 0.95 respectively are confirmed as change intervals.

[0068] It can be understood that confirming the intervals formed by the values obtained by multiplying the change rates in each preference change trend by 1.05 and by 0.95 respectively as change intervals can provide a basis for subsequent steps.

[0069] Exemplarily, assume that there are 3 change rates in the preference change trend, and the 3 change rates are 0.1, 0.11, and 0.09 in sequence. Then the change intervals corresponding to these 3 change rates are (0.1 * 0.95, 0.1 * 1.05) = (0.095, 0.105), (0.11 * 0.95, 0.11 * 1.05) = (0.1045, 0.1155), and (0.09 * 0.95, 0.09 * 1.05) = (0.0855, 0.0945) in sequence.

[0070] S23246. Analyze each comparison change trend respectively, compare the comparison change trend with each preference change trend, and confirm the content keywords corresponding to the comparison change trend in which each change rate is within the change interval corresponding to the same preference change trend according to time as analysis words.

[0071] It can be understood that confirming the content keywords corresponding to the comparison change trend in which each change rate is within the change interval corresponding to the same preference change trend according to time as analysis words can provide a basis for subsequent steps.

[0072] Exemplarily, assume that there are 3 change rates in a comparison change trend, and the 3 change rates are 0.1, 0.11, and 0.09 in sequence according to time. Assume that there are 3 change intervals in a preference change trend, and the 3 change intervals are (0.095, 0.105), (0.1045, 0.1155), and (0.0855, 0.0945) in sequence according to time. Then each change rate in this comparison change trend is within the change intervals corresponding to this preference change trend according to time, and the content keywords corresponding to this comparison change trend are confirmed as analysis words.

[0073] S23247. Confirm all the analysis words as analysis word information.

[0074] It can be understood that confirming all the analysis words as analysis word information can provide a basis for subsequent steps.

[0075] S2325. Analyze each browsing record respectively based on the analysis word information to obtain preselected data analysis information.

[0076] It can be understood that the method of analyzing each browsing record based on the analysis word information to obtain the preselected data analysis information can be to receive the data transmitted by the user, or to analyze the proportion of each analysis word in the analysis word information in each browsing record, etc., but not limited to this. Analyzing each browsing record based on the analysis word information to obtain the preselected data analysis information can ensure the reliability of the preselected data analysis information and provide a basis for subsequent steps.

[0077] In a possible implementation manner, please refer to Figure 3 , S2325, analyzing each browsing record based on the analysis word information to obtain the preselected data analysis information, including: S23251, step a, confirming an analysis word in the analysis word information as a separator word.

[0078] It can be understood that the method of confirming an analysis word in the analysis word information as a separator word can be to sequentially confirm the analysis words in the analysis word information as separator words in order, or to confirm a random analysis word in the analysis word information as a separator word, etc., but not limited to this. Confirming an analysis word in the analysis word information as a separator word can provide a basis for subsequent steps.

[0079] S23252, step b, respectively confirming the search time corresponding to the first search keyword that is the same as the separator word or the access time corresponding to the content keyword that is the same as the separator word that appears in each browsing record in chronological order as the separator time.

[0080] It can be understood that confirming the search time corresponding to the first search keyword that is the same as the separator word or the access time corresponding to the content keyword that is the same as the separator word that appears in each browsing record in chronological order as the separator time can provide a basis for subsequent steps.

[0081] S23253, step c, respectively adding up the residence times corresponding to the web page addresses with the same corresponding content keyword before the separator time in each browsing record to obtain at least one pre-separator duration, and adding up the residence times corresponding to the web page addresses with the same corresponding content keyword after the separator time to obtain at least one post-separator duration.

[0082] It can be understood that obtaining the pre-separator duration and the post-separator duration can provide a basis for subsequent steps.

[0083] Exemplarily, assuming that in the browsing record, there are 3 web page addresses with the same corresponding content keyword before the separator time, and the residence times corresponding to these 3 web page addresses are 10 seconds, 51 seconds, and 20 seconds respectively, then the pre-separator duration corresponding to this content keyword = 10 + 15 + 20 = 45 seconds.

[0084] S23254, step d: Determine whether all the parsing words in the parsing word information are confirmed as delimiter words. If all the parsing words in the parsing word information are confirmed as delimiter words, then confirm all the pre-delimiter durations and post-delimiter durations as delimiter time information. If there is a parsing word in the parsing word information that is not confirmed as a delimiter word, then confirm one of the parsing words that is not confirmed as a delimiter word as a delimiter word and repeat step b, step c, and step d; wherein, the delimiter time information includes at least one pre-delimiter duration, at least one post-delimiter duration, and at least one delimiter word corresponding to the pre-delimiter duration and the post-delimiter duration for each browsing record.

[0085] It can be understood that determining whether all the parsing words in the parsing word information are confirmed as delimiter words can ensure that all the parsing words in the parsing word information are confirmed as delimiter words, ensuring the reliability of the method.

[0086] S23255, Obtain preselected data analysis information based on the delimiter time information.

[0087] It can be understood that the way to obtain preselected data analysis information based on the delimiter time information can be to receive the data transmitted by the user, or to judge the magnitudes of the pre-delimiter duration and the post-delimiter duration corresponding to the same content keyword in the same browsing record, etc., but not limited to this. Obtaining preselected data analysis information based on the delimiter time information can provide a basis for the subsequent steps.

[0088] In a possible implementation, please refer to Figure 3 , S23255, Obtain preselected data analysis information based on the delimiter time information, including: S232551, step e: Confirm the content keyword corresponding to a post-delimiter duration as a comparison word.

[0089] It can be understood that confirming the content keyword corresponding to a post-delimiter duration as a comparison word can provide a basis for the subsequent steps.

[0090] S232552, step f: Determine whether there is a content keyword identical to the comparison word among the content keywords corresponding to the pre-delimiter duration of the browsing record corresponding to the comparison word. If there is, then execute step g. If not, then confirm the delimiter word corresponding to the post-delimiter duration as the first parsing word and then execute step h.

[0091] It can be understood that if there is no content keyword identical to the comparison word among the content keywords corresponding to the pre-delimiter duration of the browsing record corresponding to the comparison word, it means that the target object obtains the content related to the comparison word only after obtaining the delimiter word. Confirming the delimiter word corresponding to the post-delimiter duration as the first parsing word can provide a basis for the subsequent steps.

[0092] S232553, step g, determine the size of the front separation duration and the rear separation duration. If the front separation duration is greater than or equal to the rear separation duration, execute step h. If the front separation duration is less than the rear separation duration, confirm the separation word corresponding to the rear separation duration as the second analysis word.

[0093] It can be understood that if the pre-separation duration is greater than or equal to the post-separation duration, it means that the target object has browsed the content related to the comparison word for a sufficient time before obtaining the separation word. If the pre-separation duration is less than the post-separation duration, it means that after obtaining the separation word, the target object has increased the time spent browsing the web page with content related to the comparison word. Confirming the separation word corresponding to the post-separation duration as the second analysis word can provide a basis for subsequent steps.

[0094] S232554, step h, confirm the content keyword corresponding to another post-separation time length as a comparison word and then repeat steps f and g until all content keywords are confirmed as comparison words.

[0095] It can be understood that after confirming the content keyword corresponding to another post-separation time length as a comparison word, repeating steps f and g until all content keywords are confirmed as comparison words can ensure the reliability of the method.

[0096] S232555, confirm all the first analysis words and the second analysis words as pre-selected data analysis information.

[0097] It can be understood that confirming all the first analysis words and the second analysis words as pre-selected data analysis information can provide a basis for subsequent steps.

[0098] S233, obtaining data analysis information based on the pre-selected data analysis information and the interest frequency information.

[0099] It can be understood that the method of obtaining data analysis information based on pre-selected data analysis information and interest frequency information can be to confirm the first analysis word and the second analysis word with larger values ​​in the interest frequency information as data analysis information, or to send the pre-selected data analysis information and interest frequency information to the user and then receive the data transmitted by the user, etc., but it is not limited to this. Obtaining data analysis information based on pre-selected data analysis information and interest frequency information can mine data that indirectly affects the target object, improving the problem that existing data mining methods easily miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0100] In one possible implementation, see Figure 3 , S233, obtaining data analysis information based on the pre-selected data analysis information and the interest frequency information, including: S2331, sort the interest frequencies in the interest frequency information from large to small to obtain a frequency table.

[0101] It can be understood that sorting the interest frequencies in the interest frequency information from largest to smallest can improve the processing speed of subsequent steps. Obtaining the frequency table can provide a basis for subsequent steps.

[0102] S2332: Confirm the interest frequencies in the top 30% of the frequency table as the first analysis frequency information, and confirm the interest frequencies in the top 50% of the frequency table as the second analysis frequency information.

[0103] It can be understood that confirming the first analysis frequency information and the second analysis frequency information can provide a basis for subsequent steps.

[0104] S2333: Confirm the first analysis words with the same content keywords corresponding to the interest frequencies in the second analysis frequency information among all the first analysis words as the first data, and confirm the second analysis words with the same content keywords corresponding to the interest frequencies in the first analysis frequency information among all the second analysis words as the second data.

[0105] It can be understood that since the first analysis words are those for which the target object has obtained the content related to the comparison words after obtaining the first analysis words, and the possibility of a potential connection relationship between the first analysis words and the target object is relatively large, the first analysis words are compared with the second analysis frequency information with a larger range. Since the second analysis words are those for which the target object has obtained the content related to the comparison words before obtaining the second analysis words, and the possibility of an implicit connection between the second analysis words and the target object is relatively small, the second analysis words are compared with the first analysis frequency information with a smaller range to improve the accuracy of the second data.

[0106] S2334: Confirm the first data and the second data as the data analysis information.

[0107] It can be understood that confirming the first data and the second data as the data analysis information can mine the data that indirectly affects the target object, and improve the problem that the existing data mining methods are prone to miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0108] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0109] Corresponding to the key data mining method based on big data described in the above embodiments, the embodiments of the present application further provide a key data mining system based on big data, and each unit of the system can implement each step of the key data mining method based on big data. Figure 4The structural block diagram of the key data mining system based on big data provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0110] Referring to Figure 4 , the system includes: An acquisition unit, configured to acquire mining scenario information; wherein, the mining scenario information includes a mining project and a target object.

[0111] An analysis unit, configured to obtain data analysis information based on the mining scenario information; wherein, the data analysis information includes at least one keyword reflecting key data.

[0112] It should be noted that for the information interaction, execution process, etc. between the above units, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0113] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used for illustration. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0114] The embodiments of the present application also provide a key data mining device based on big data. Figure 5 It is a structural schematic diagram of the key data mining device based on big data provided by an embodiment of the present application. As Figure 5 shown, the key data mining device based on big data in this embodiment includes a control device 6. Among them, the control device 6 includes: at least one processor 60 ( Figure 5 only one is shown in Figure 5only one is shown) and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the big data-based key data mining device is caused to implement the steps in any of the above-described big data-based key data mining method embodiments, or the functions of the various units in the above-described system embodiments are implemented in the big data-based key data mining device.

[0115] Exemplarily, the computer program 62 may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the control device 6.

[0116] The control device 6 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The control device 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 5 merely an example of a big data-based key data mining device, which does not constitute a limitation on the big data-based key data mining device, and may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, buses, etc.

[0117] The processor 60 may be a central processing unit (CPU), and the processor 60 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The memory 61 may be an internal storage unit of the control device 6 in some embodiments, such as a hard disk or memory of the control device 6. The memory 61 may also be an external storage device of the control device 6 in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the control device 6. Further, the memory 61 may also include both the internal storage unit of the control device 6 and external storage devices. The memory 61 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of the computer program. The memory 61 may also be used to temporarily store data that has been output or will be output.

[0119] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0120] An embodiment of the present application provides a computer program product, and when the computer program product runs on a big data-based key data mining device, the big data-based key data mining device implements the steps in any of the above method embodiments.

[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program may be used to instruct relevant hardware to complete. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the big data-based key data mining device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.

[0122] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0123] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0124] In the embodiments provided in this application, it should be understood that the disclosed key data mining system, device, and method based on big data can be implemented in other ways. For example, the key data mining system and device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A key data mining method based on big data, characterized in that, Including: Obtaining mining scenario information; wherein, the mining scenario information includes a mining project and a target object; Obtaining data analysis information based on the mining scenario information; wherein, the data analysis information includes at least one piece of information reflecting key data.

2. The key data mining method based on big data according to claim 1, characterized in that The obtaining of the data analysis information based on the mining scenario information includes: Obtaining keyword information based on the mining project; wherein, the keyword information includes a plurality of keywords related to the mining project; Obtaining browsing information based on the target object; wherein, the browsing information includes a plurality of browsed records after data cleaning and the daily timestamps corresponding to the browsed records, the browsed records include all the search keywords of the target object within a single day and the search times corresponding to each of the search keywords, and all the web page addresses accessed by the target object within a single day, the content keywords, access times and stay durations corresponding to each of the web page addresses; Obtaining the data analysis information based on the keyword information and the browsing information.

3. The key data mining method based on big data according to claim 2, characterized in that, The obtaining of the data analysis information based on the keyword information and the browsing information includes: Obtaining interest duration information and interest frequency information based on the browsing information; wherein, the interest duration information includes at least one interest duration, the interest duration reflects the sum of the stay durations of the web page addresses corresponding to a certain content keyword in the browsed records, the interest frequency information includes at least one interest frequency, the interest frequency reflects the ratio of the number of web page addresses corresponding to a certain content keyword in the browsed records to the number of all the web page addresses in the browsed records; Obtaining preliminary data analysis information based on the interest duration information and the keyword information; wherein, the preliminary data analysis information includes at least one first analysis word and at least one second analysis word, the first analysis word reflects a content keyword, and the second analysis word reflects a content keyword; Obtaining the data analysis information based on the preliminary data analysis information and the interest frequency information.

4. The key data mining method based on big data according to claim 3, wherein, The obtaining of the interest duration information and the interest frequency information based on the browsing information includes: Analyzing each of the browsed records respectively, confirming the search times corresponding to each of the search keywords in the browsed records as retrieval nodes, confirming the web page addresses adjacent to and before the retrieval nodes as interest web pages, and confirming the web page addresses adjacent to and after the retrieval nodes as verification web pages; If the content keyword corresponding to the interest website is related to the content keyword corresponding to the verification website adjacent to the retrieval node adjacent to the interest website, after adding an interest tag to the interest website and the verification website, starting from the verification website as the retrieval starting point, the interest tag is sequentially added to the web address backward in the time order reflected by the access time until the content keyword corresponding to the web address is not related to the content keyword corresponding to the interest website; wherein, the interest tag is the content keyword corresponding to the interest website. After adding up the residence durations corresponding to the web addresses with the same content keyword among the interest tags respectively to obtain at least one interest duration corresponding to the content keyword corresponding to the interest tag, all the interest durations are confirmed as interest duration information. By analyzing each browsing record respectively, the value obtained by adding up the numbers of the web addresses with the same content keyword corresponding to the browsing record and then dividing by the number of all the web addresses in the browsing record is confirmed as the interest frequency, and all the interest frequencies are confirmed as the interest frequency information.

5. The key data mining method based on big data according to claim 3, characterized in that The obtaining of the preselected data analysis information based on the interest duration information and the keyword information includes: The interest duration corresponding to the content keyword belonging to the keyword information in the interest duration information is confirmed as the preference information, and the interest duration corresponding to the content keyword not belonging to the keyword information in the interest duration information is confirmed as the preference comparison information. Each interest duration in the preference information is divided by the sum of all the residence times in the corresponding browsing record respectively to obtain a plurality of preference ratios, and each interest duration in the preference comparison information is divided by the sum of all the residence times in the corresponding browsing record respectively to obtain a plurality of preference comparison ratios. The respective preference ratios and the respective preference comparison ratios are grouped according to the corresponding content keyword, so that the content keywords corresponding to the preference ratios or the preference comparison ratios within each group are the same, and then the preference ratios or the preference comparison ratios within each group are sorted in descending order according to the daily time stamps of the corresponding browsing records to obtain a sorting table. Analysis word information is obtained based on the sorting table. Based on the analysis word information, each browsing record is analyzed respectively to obtain the preselected data analysis information.

6. The key data mining method based on big data according to claim 5, characterized in that The obtaining of the analysis word information based on the sorting table includes: Each group in the sorting table is analyzed respectively. Step 1, the nth preference ratio or preference comparison ratio within the group is confirmed as the first value. If n + 1 is less than or equal to the number of the preference ratios or preference comparison ratios within the group, the (n + 1)th preference ratio or preference comparison ratio within the group is confirmed as the second value. If n + 1 is greater than the number of the preference ratios or preference comparison ratios within the group, then step 3 is executed; wherein, the initial value of n is 1. Step 2: Confirm the value obtained by subtracting the first value from the second value, dividing the result by the first value, and taking it as the change rate. Then substitute n + 1 for n in Step 1 and repeat Step 1; Step 3: If all the data in the group are the preference ratios, confirm all the obtained change rates as the preference change trend. If all the data in the group are the preference comparison ratios, confirm all the obtained change rates as the comparison change trend; After analyzing each group in the sorting table, respectively confirm the interval formed by the values obtained by multiplying the change rates in each preference change trend by 1.05 and by 0.95 as the change interval; Analyze each comparison change trend respectively, compare the comparison change trend with each preference change trend, and confirm the content keywords corresponding to the comparison change trend in which each change rate is within the corresponding change interval according to time in the same preference change trend as the analysis words; Confirm all the analysis words as the analysis word information.

7. The key data mining method based on big data according to claim 5, characterized in that, Analyze each browsing record respectively based on the analysis word information to obtain the preselected data analysis information, including: Step a: Confirm one of the analysis words in the analysis word information as the separator word; Step b: Respectively confirm the search time corresponding to the search keyword or the access time corresponding to the content keyword that is the same as the separator word and appears first in time order in each browsing record as the separation time; Step c: Respectively add up the residence times corresponding to the web page addresses with the same content keyword before the separation time in each browsing record to obtain at least one pre-separation duration, and add up the residence times corresponding to the web page addresses with the same content keyword after the separation time to obtain at least one post-separation duration; Step d: Judge whether all the analysis words in the analysis word information have been confirmed as the separator words. If all the analysis words in the analysis word information have been confirmed as the separator words, confirm all the pre-separation durations and post-separation durations as the separation time information. If there are analysis words in the analysis word information that have not been confirmed as the separator words, confirm one of the analysis words that have not been confirmed as the separator words as the separator word and then repeat Step b, Step c, and Step d; wherein, the separation time information includes at least one pre-separation duration, at least one post-separation duration, and at least one separator word corresponding to the pre-separation duration and the post-separation duration corresponding to each browsing record; Obtain the preselected data analysis information based on the separation time information.

8. The key data mining method based on big data according to claim 7, characterized in that The obtaining of the preselected data analysis information based on the separation time information includes: Step e: Confirm the content keyword corresponding to one post-separation duration as the comparison word; Step f: Determine whether there is a content keyword in the content keywords corresponding to the browsing record corresponding to the comparison word that is the same as the comparison word. If there is, execute Step g; if not, confirm the separator word corresponding to the post-separation duration as the first analysis word and then execute Step h; Step g: Determine the magnitudes of the pre-separation duration and the post-separation duration. If the pre-separation duration is greater than or equal to the post-separation duration, execute Step h; if the pre-separation duration is less than the post-separation duration, confirm the separator word corresponding to the post-separation duration as the second analysis word; Step h: Confirm the content keyword corresponding to the other post-separation duration as the comparison word and then repeat Steps f and g until all content keywords are confirmed as comparison words; Confirm all the first analysis words and the second analysis words as the preselected data analysis information.

9. The key data mining method based on big data according to claim 3, characterized in that Obtaining the data analysis information based on the preselected data analysis information and the interest frequency information includes: Sort the interest frequencies in the interest frequency information from largest to smallest to obtain a frequency table; Confirm the interest frequencies in the top 30% of the frequency table as the first analysis frequency information, and confirm the interest frequencies in the top 50% of the frequency table as the second analysis frequency information; Confirm the first analysis words that are the same as the content keywords corresponding to the interest frequencies in the second analysis frequency information among all the first analysis words as the first data, and confirm the second analysis words that are the same as the content keywords corresponding to the interest frequencies in the first analysis frequency information among all the second analysis words as the second data; Confirm the first data and the second data as the data analysis information.

10. A key data mining device based on big data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Software application data mining method based on big data positioning and software service platform

    CN112600893A

  • Intelligent data mining method and system based on big data

    CN117874103A

  • User portrait construction method and system based on big data

    CN118098584A

  • AI-based user behavior preference data analysis method

    CN118520164A

  • Electrolyte additives for secondary battery, non-aqueous electrolyte for secondary battery comprising same and secondary battery

    KR1020250033028A