Key data mining method and equipment based on big data

By obtaining mining scenario information and analyzing key data, the problem of missing indirectly affecting data in the existing technology is solved, and comprehensive data mining of target objects is achieved.

CN120216567BActive Publication Date: 2025-08-12BEIJING ZHONGTIAN CHUANGYU MARKET CONSULTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510695204.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-12
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing data mining methods usually only focus on data that can have a direct impact on the target, resulting in missing data that can have a direct impact on the target.

Method used

By obtaining mining scenario information, clarifying business goals and scope, defining the boundaries and directions of data analysis, and obtaining information that reflects key data based on mining scenario information, mining data that can indirectly affect the target object.

Benefits of technology

It improves the shortcomings of existing data mining methods, can mine data that indirectly has a direct impact on the target, and improves the comprehensiveness and accuracy of data mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216567B_ABST
    Figure CN120216567B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of key data mining technology, and in particular relates to a key data mining method and device based on big data, the method comprising: obtaining mining scenario information; wherein the mining scenario information includes a mining project and a target object; obtaining data analysis information based on the mining scenario information; wherein the data analysis information includes at least one information reflecting key data. The key data mining method and device based on big data provided by the embodiments of the present application can improve the problem that existing data mining methods are prone to miss some data that can indirectly have a direct impact on the target and indirectly produce results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of key data mining technology, and in particular relates to a key data mining method and device based on big data. Background Art

[0002] Across diverse fields, such as business, healthcare, and finance, there's a strong demand for key data mining. However, existing data mining methods typically only target data that directly impacts the target and produces direct results, leading to the potential for missing data that could indirectly impact the target and produce results. Summary of the Invention

[0003] The embodiments of the present application provide a key data mining method and device based on big data, which can improve the problem that existing data mining methods easily miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0004] In a first aspect, an embodiment of the present application provides a key data mining method based on big data, comprising:

[0005] Acquire mining scenario information; wherein the mining scenario information includes mining items and target objects;

[0006] Data analysis information is obtained based on the mining scenario information; wherein the data analysis information includes at least one information reflecting key data.

[0007] The above technical solutions in the embodiments of the present application have at least the following technical effects:

[0008] The key data mining method based on big data provided in the embodiments of the present application first obtains mining scenario information including mining projects and target objects, clarifies business objectives and scope, and defines the boundaries and direction of data analysis. Then, based on the mining scenario information, data analysis information including at least one piece of information reflecting key data is obtained, and data that can indirectly affect the target object is mined, thereby improving the problem that existing data mining methods are prone to miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0009] In a second aspect, an embodiment of the present application provides a key data mining system based on big data, comprising:

[0010] An acquisition unit, configured to acquire mining scenario information; wherein the mining scenario information includes mining items and target objects;

[0011] An analysis unit is used to obtain data analysis information based on the mining scenario information; wherein the data analysis information includes at least one information reflecting key data.

[0012] In a third aspect, an embodiment of the present application provides a key data mining device based on big data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the methods described in the first aspect above when executing the computer program.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the above-mentioned first aspects is implemented.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a key data mining device based on big data, the key data mining device based on big data executes the key data mining method based on big data described in any one of the first aspects above.

[0015] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a flowchart of a key data mining method based on big data provided by an embodiment of the present application;

[0018] Figure 2 This is a flow chart of step S200 in the key data mining method based on big data provided by an embodiment of the present application;

[0019] Figure 3 This is a flow chart of step S2324 in the key data mining method based on big data provided by one embodiment of the present application;

[0020] Figure 4 This is a schematic diagram of the structure of a key data mining system based on big data provided by an embodiment of the present application;

[0021] Figure 5 This is a schematic diagram of the structure of a key data mining device based on big data provided by an embodiment of the present application. DETAILED DESCRIPTION

[0022] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0023] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0024] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0025] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0026] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0027] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0028] Across diverse fields, such as business, healthcare, and finance, there's a strong demand for key data mining. However, existing data mining methods typically only target data that directly impacts the target and produces direct results, leading to the potential for missing data that could indirectly impact the target and produce results.

[0029] To address the aforementioned issues, embodiments of the present application provide a method and apparatus for mining key data based on big data. This method first obtains mining scenario information, including mining projects and target objects, to clarify business objectives and scope, and define the boundaries and direction of data analysis. Then, based on the mining scenario information, data analysis information, including at least one piece of information reflecting key data, is obtained. Data that can indirectly affect the target object is mined, addressing the problem with existing data mining methods, which tend to miss data that can indirectly affect the target and indirectly produce results.

[0030] The key data mining method based on big data provided in the embodiment of the present application can be applied to the key data mining device based on big data. At this time, the key data mining device based on big data is the execution entity of the key data mining method based on big data provided in the embodiment of the present application. The embodiment of the present application does not impose any restrictions on the specific type of the key data mining device based on big data.

[0031] For example, a key data mining device based on big data may be a mobile phone, a tablet computer, a laptop computer, a handheld device with wireless communication capabilities, a computing device, or other processing device connected to a wireless modem, but is not limited thereto.

[0032] In order to better understand the key data mining method based on big data provided in the embodiment of the present application, the specific implementation process of the key data mining method based on big data provided in the embodiment of the present application is exemplarily introduced below.

[0033] Figure 1 A schematic flow chart of a key data mining method based on big data provided by an embodiment of the present application is shown. The key data mining method based on big data includes:

[0034] S100, obtaining mining scenario information; wherein the mining scenario information includes mining items and target objects.

[0035] It is understood that mining scenario information can be obtained by receiving data transmitted by users or analyzing historical data, but is not limited to these methods. Obtaining mining scenario information can clarify business goals and scope, define the boundaries and direction of data analysis, and provide a basis for subsequent steps.

[0036] S200, obtaining data analysis information based on the mining scenario information; wherein the data analysis information includes at least one keyword reflecting key data.

[0037] I understand.

[0038] In one possible implementation, see Figure 2 , S100, obtain mining scene information, including:

[0039] S210 , obtaining keyword information based on the mining project; wherein the keyword information includes a plurality of keywords related to the mining project.

[0040] It is understood that the method for obtaining keyword information based on the mining project can be to receive data transmitted by the user, or to input the mining project into a preset database for query to obtain at least one keyword corresponding to the mining project, etc., but is not limited thereto. Obtaining keyword information based on the mining project can provide a basis for subsequent steps.

[0041] For example, assuming that the mining item is "online shopping", the keyword information related to "online shopping" may include product name, brand, model, specification, material, color, size, price, promotion, etc., but is not limited thereto.

[0042] S220, obtaining browsing information based on the target object; wherein the browsing information includes multiple browsing records after data cleaning and daily timestamps corresponding to the browsing records, the browsing records include all search keywords of the target object in a single day and the search time corresponding to each search keyword, and all web page addresses visited by the target object in a single day, the content keywords corresponding to each web page address, the access time and the length of stay.

[0043] It is understood that the method of obtaining browsing information based on the target object can be to receive data transmitted by the user, or to receive data transmitted by an APP or website authorized by the target object, etc., but is not limited thereto. Obtaining browsing information based on the target object can provide a basis for subsequent steps.

[0044] S230: Obtain data analysis information based on the keyword information and the browsing information.

[0045] It is understood that the method for obtaining data analysis information based on keyword information and browsing information can be to send the keyword information and browsing information to the user and then receive the data transmitted by the user, or to analyze the number of data items in the browsing information that belong to the keywords in the keyword information, etc., but is not limited to this. Obtaining data analysis information based on keyword information and browsing information can analyze data that indirectly affects the target.

[0046] In one possible implementation, see Figure 2, S230, obtaining data analysis information based on keyword information and browsing information, including:

[0047] S231, obtaining interest duration information and interest frequency information based on browsing information; wherein, the interest duration information includes at least one interest duration, and the interest duration reflects the sum of the stay times of the web page addresses corresponding to a certain content keyword in the browsing record; the interest frequency information includes at least one interest frequency, and the interest frequency reflects the ratio of the number of web page addresses corresponding to a certain content keyword in the browsing record to the total number of web page addresses in the browsing record.

[0048] It is understood that the method for obtaining interest duration information based on browsing information can be to receive data transmitted by the user, or to add the dwell times corresponding to the web page addresses corresponding to the content keywords in the keyword information, etc., but is not limited to this. The method for obtaining interest frequency information based on browsing information can be to receive data transmitted by the user, or to count the number of content keywords in the browsing history that are identical to the keywords in the keyword information, etc., but is not limited to this. Obtaining interest duration information and interest frequency information based on browsing information can provide a basis for subsequent steps.

[0049] In one possible implementation, see Figure 2 S231: Obtaining interest duration information and interest frequency information based on the browsing information, including:

[0050] S2311, analyze each browsing record separately, confirm the search time corresponding to each search keyword in the browsing record as the retrieval node, confirm the web page address adjacent to the retrieval node and located before the retrieval node as the interest URL, and confirm the web page address adjacent to the retrieval node and located after the retrieval node as the verification URL.

[0051] It can be understood that the search keyword refers to the text content in the search box when the target object performs a search operation in the APP or web page. The search time refers to the time when the target object performs a search operation in the APP or web page. The web page address adjacent to the retrieval node and located before the retrieval node is the web page address that is less than the retrieval node in time and closest to the retrieval node in time (or the web page that performs the search operation directly after browsing the web page). The web page address adjacent to the retrieval node and located after the retrieval node is the web page address that is greater than the retrieval node in time and closest to the retrieval node in time (or the web page that is browsed directly after performing the search operation). Confirming the retrieval node, and confirming the web page address adjacent to the retrieval node and located before the retrieval node as the interest URL, and confirming the web page address adjacent to the retrieval node and located after the retrieval node as the verification URL can provide a basis for subsequent steps.

[0052] S2312, if the content keyword corresponding to the interest URL is related to the content keyword corresponding to the verification URL adjacent to the search node adjacent to the interest URL, then after adding interest tags to the interest URL and the verification URL, take the verification URL as the search starting point, and add interest tags to the web page addresses in sequence according to the time sequence reflected by the access time until the content keyword corresponding to the web page address is no longer related to the content keyword corresponding to the interest URL; wherein the interest tag is the content keyword corresponding to the interest URL.

[0053] It is understood that the method for determining whether the content keywords corresponding to the interest URL and the content keywords corresponding to the verification URL adjacent to the search node adjacent to the interest URL are related can be an exact match (the two content keywords are exactly the same), a fuzzy match (the two content keywords are partially the same or the semantics expressed by the two content keywords are the same, etc.), but is not limited to this. The content keywords corresponding to the web page address are not related to the content keywords corresponding to the interest URL, which means that the two content keywords are completely different or the semantics expressed by the content keywords are different. If the content keywords corresponding to the interest URL are related to the content keywords corresponding to the verification URL adjacent to the search node adjacent to the interest URL, it means that the target object's interests are consistent before and after the search operation is performed. Adding interest tags to the interest URL and the verification URL, and adding interest tags to the web page addresses in the chronological order reflected by the access time until the content keywords corresponding to the web page address are no longer related to the content keywords corresponding to the interest URL, can provide a basis for subsequent steps.

[0054] S2313, after adding up the dwell times corresponding to the web page addresses corresponding to the interest tags with the same content keywords to obtain at least one interest duration corresponding to the content keyword corresponding to the interest tag, all the interest durations are confirmed as interest duration information.

[0055] It can be understood that by adding the dwell times corresponding to the web page addresses corresponding to the interest tags with the same content keywords, at least one interest duration corresponding to the content keyword corresponding to the interest tag can be obtained, and confirming all the interest durations as interest duration information can provide a basis for subsequent steps.

[0056] For example, assuming that there are three interest tags and the corresponding content keywords are all "apple", and the corresponding dwell times of the web page addresses corresponding to these three interest tags are 10 seconds, 15 seconds, and 5 seconds respectively, then the interest duration corresponding to "apple" = 10+15+5=30 seconds.

[0057] S2314, analyze each browsing record separately, add up the number of web page addresses with the same content keyword in the browsing record and divide it by the total number of web page addresses in the browsing record to obtain the value as the interest frequency, and then confirm all the interest frequencies as interest frequency information.

[0058] It can be understood that after the value obtained by adding the number of web page addresses with the same content keyword in the browsing history and dividing it by the total number of web page addresses in the browsing history is confirmed as the interest frequency, confirming all interest frequencies as interest frequency information can provide a basis for subsequent steps.

[0059] For example, assuming that there are 10 web page addresses in the browsing history, and the content keywords corresponding to 3 of these 10 web page addresses are all "apple", and the content keywords corresponding to the other 7 web page addresses are all "clothes", then the interest frequency corresponding to "apple" = 3 / 10 = 0.3, and the interest frequency corresponding to "clothes" = 7 / 10 = 0.7.

[0060] S232, obtaining pre-selected data analysis information based on the interest duration information and keyword information; wherein the pre-selected data analysis information includes at least one first analysis word and at least one second analysis word, the first analysis word reflects a content keyword, and the second analysis word reflects a content keyword.

[0061] It is understood that the method for obtaining pre-selected data analysis information based on the interest duration information and keyword information may be to receive data transmitted by the user, or to identify the content keywords corresponding to the interest durations whose durations are in the top 30% of all interest durations and whose corresponding content keywords belong to the keyword information as pre-selected data analysis information, etc., but is not limited thereto. Obtaining pre-selected data analysis information based on the interest duration information and keyword information can provide a basis for subsequent steps.

[0062] In one possible implementation, see Figure 2 , S232, based on the interest duration information and keyword information, obtain pre-selected data analysis information, including:

[0063] S2321: Confirm the interest durations whose corresponding content keywords in the interest duration information belong to the keyword information as preference information, and confirm the interest durations whose corresponding content keywords in the interest duration information do not belong to the keyword information as preference comparison information.

[0064] It is understood that the method for determining whether a content keyword belongs to the keyword information may be to compare the content keyword with each keyword in the keyword information through exact matching or fuzzy matching, but is not limited thereto. Confirming the interest duration of the content keyword corresponding to the interest duration information that belongs to the keyword information as preference information, and confirming the interest duration of the content keyword corresponding to the interest duration information that does not belong to the keyword information as preference comparison information can provide a basis for subsequent steps.

[0065] S2322, divide each interest duration in the preference information by the sum of all stay durations in the corresponding browsing records to obtain multiple preference ratios, and divide each interest duration in the preference comparison information by the sum of all stay durations in the corresponding browsing records to obtain multiple preference comparison ratios.

[0066] It can be understood that obtaining multiple preference ratios and multiple preference comparison ratios can provide a basis for subsequent steps.

[0067] For example, assuming that there are two interest durations numbered 1 and 2 in the preference information, the interest duration numbered 1 is 10 seconds, and the interest duration numbered 2 is 20 seconds, and the sum of all stay durations in the browsing records corresponding to the interest durations numbered 1 and 2 is 600 seconds, then the preference ratio corresponding to the interest duration numbered 1 = 10 / 600 = 0.017 (keep three decimal places), and the preference ratio corresponding to the interest duration numbered 2 = 20 / 600 = 0.033 (keep three decimal places).

[0068] S2323, grouping each preference ratio and each preference comparison ratio according to the corresponding content keyword, so that the content keywords corresponding to the preference ratio or preference comparison ratio in each group are the same, and then sorting the preference ratio or preference comparison ratio in each group from far to near according to the corresponding browsing record time stamp to obtain a sorting table.

[0069] It is understood that the ranking table can be an Excel spreadsheet, a database table, etc., but is not limited thereto. Grouping the preference ratios and preference comparison ratios by corresponding content keywords, ensuring that the preference ratios or preference comparison ratios within each group correspond to the same content keywords, and sorting the preference ratios or preference comparison ratios within each group from earliest to latest by the corresponding browsing record time stamps can improve the efficiency of subsequent data processing. The resulting ranking table can provide a basis for subsequent steps.

[0070] S2324, obtain analysis word information based on the sorting table.

[0071] It is understood that the method for obtaining analysis word information based on the ranking table may be to identify as analysis word information the content keywords corresponding to preference ratios or preference comparison ratios whose values in the ranking table gradually increase over time, or to send the ranking table to the user and then receive data transmitted by the user, etc., but is not limited thereto. Obtaining analysis word information based on the ranking table can provide a basis for subsequent steps.

[0072] In one possible implementation, see Figure 2 , S2324, obtain analysis word information based on the sorting table, including:

[0073] S23241, analyze each group in the sorted table separately.

[0074] It can be understood that analyzing each group in the ranking table separately can ensure the reliability of the method.

[0075] S23242, step 1, confirm the nth preference ratio or preference comparison ratio in the group as the first value; if n+1 is less than or equal to the number of preference ratios or preference comparison ratios in the group, confirm the n+1th preference ratio or preference comparison ratio in the group as the second value; if n+1 is greater than the number of preference ratios or preference comparison ratios in the group, execute step 3; wherein, the initial value of n is 1.

[0076] It can be understood that if n+1 is less than or equal to the number of preference ratios or preference comparison ratios within the group, it means that there are still preference ratios or preference comparison ratios within the group that have not been processed. If n+1 is greater than the number of preference ratios or preference comparison ratios within the group, it means that all preference ratios or preference comparison ratios within the group have been processed. Confirming the first and second values can provide a basis for subsequent steps.

[0077] For example, assuming that there are three preference ratios in the group, and the three preference ratios are 0.1, 0.2, and 0.3 in order, then when n=1, the first value=0.1, the second value=0.2, and when n=2, the first value=0.2, the second value=0.3.

[0078] S23243, step 2, subtract the first value from the second value and divide the result by the first value to obtain a value as the rate of change, substitute n+1 into n in step 1 and repeat step 1.

[0079] It can be understood that if the rate of change is a positive number, it means that the target subject's interest in the analysis word corresponding to the preference ratio or preference comparison ratio is increasing from the time corresponding to the first value to the time corresponding to the second value. If the rate of change is a negative number, it means that the target subject's interest in the analysis word corresponding to the preference ratio or preference comparison ratio is decreasing from the time corresponding to the first value to the time corresponding to the second value. Confirming the rate of change can provide a basis for subsequent steps.

[0080] For example, assuming that the first value is 0.1 and the second value is 0.2, the change rate=(0.2−0.1) / 0.1=1.

[0081] S23244, step three, if the data in the group are all preference ratios, then all the obtained change rates will be confirmed as preference change trends; if the data in the group are all preference comparison ratios, then all the obtained change rates will be confirmed as comparison change trends.

[0082] It can be understood that confirming all obtained change rates as preference change trends or comparative change trends can provide a basis for subsequent steps.

[0083] S23245, after all groups in the ranking table are analyzed, the interval formed by multiplying the change rate of each preference change trend by 1.05 and 0.95 respectively is confirmed as the change interval.

[0084] It can be understood that confirming the interval formed by the values obtained by multiplying the change rate in each preference change trend by 1.05 and 0.95 as the change interval can provide a basis for subsequent steps.

[0085] For example, assuming that there are three change rates in the preference change trend, and the three change rates are 0.1, 0.11, and 0.09 in order, then the change intervals corresponding to these three change rates are (0.1*0.95, 0.1*1.05) = (0.095, 0.105), (0.11*0.95, 0.11*1.05) = (0.1045, 0.1155), and (0.09*0.95, 0.09*1.05) = (0.0855, 0.0945) in order.

[0086] S23246, analyze each comparative change trend separately, compare the comparative change trend with each preference change trend, and confirm the content keywords corresponding to the comparative change trends whose change rates in the comparative change trends are within the change interval corresponding to the time in the same preference change trend as analysis words.

[0087] It can be understood that identifying the content keywords corresponding to the comparative change trends whose change rates in the comparative change trends are all within the change interval corresponding to time in the same preference change trend as analysis words can provide a basis for subsequent steps.

[0088] For example, suppose there are three change rates in a comparative change trend, and these three change rates are 0.1, 0.11, and 0.09 in chronological order; there are three change intervals in a preference change trend, and these three change intervals are (0.095, 0.105), (0.1045, 0.1155), and (0.0855, 0.0945) in chronological order. Then, each change rate in the comparative change trend is within the change intervals corresponding to the preference change trend in time, and the content keyword corresponding to the comparative change trend is confirmed as the analysis word.

[0089] S23247: Confirm all analysis words as analysis word information.

[0090] It can be understood that confirming all analysis words as analysis word information can provide a basis for subsequent steps.

[0091] S2325: Analyze each browsing record based on the analysis word information to obtain pre-selected data analysis information.

[0092] It is understood that the method for analyzing each browsing record based on the analysis word information to obtain the pre-selected data analysis information may be, but is not limited to, receiving data transmitted by the user or analyzing the proportion of each analysis word in the analysis word information in each browsing record. Analyzing each browsing record based on the analysis word information to obtain the pre-selected data analysis information can ensure the reliability of the pre-selected data analysis information and provide a basis for subsequent steps.

[0093] In one possible implementation, see Figure 3 S2325: Analyze each browsing record based on the analysis word information to obtain pre-selected data analysis information, including:

[0094] S23251, step a, confirming an analysis word in the analysis word information as a separation word.

[0095] It is understood that the method of confirming an analysis word in the analysis word information as a separator word can be to confirm the analysis words in the analysis word information as separator words in sequence, or to confirm a random analysis word in the analysis word information as a separator word, etc., but is not limited thereto. Confirming an analysis word in the analysis word information as a separator word can provide a basis for subsequent steps.

[0096] S23252, step b, respectively confirming the search time corresponding to the first search keyword identical to the separator word or the access time corresponding to the content keyword that appears in chronological order in each browsing record as the separator time.

[0097] It can be understood that confirming the search time corresponding to the first search keyword with the same separator word appearing in chronological order or the access time corresponding to the first content keyword with the same separator word appearing in chronological order in each browsing record as the separation time can provide a basis for subsequent steps.

[0098] S23253, step c, respectively add up the dwell times corresponding to the web page addresses with the same content keywords before the separation time in each browsing record to obtain at least one before-separation time, and add up the dwell times corresponding to the web page addresses with the same content keywords after the separation time to obtain at least one after-separation time.

[0099] It can be understood that obtaining the pre-separation duration and the post-separation duration can provide a basis for subsequent steps.

[0100] For example, assuming that in the browsing history, there are three web page addresses with the same content keyword before the separation time, and the corresponding stay times of these three web page addresses are 10 seconds, 51 seconds, and 20 seconds respectively, then the previous separation time corresponding to the content keyword = 10+15+20=45 seconds.

[0101] S23254, step d, determines whether all the analysis words in the analysis word information are confirmed as separator words. If all the analysis words in the analysis word information are confirmed as separator words, all the front separation durations and back separation durations are confirmed as separation time information. If there are analysis words in the analysis word information that have not been confirmed as separator words, one of the analysis words that has not been confirmed as separator words is confirmed as a separator word and then steps b, c and d are repeated; wherein, the separation time information includes at least one front separation duration, at least one back separation duration and at least one separator word corresponding to both the front separation duration and the back separation duration corresponding to each browsing record.

[0102] It can be understood that determining whether all analysis words in the analysis word information are confirmed as separation words can ensure that all analysis words in the analysis word information are confirmed as separation words, thereby ensuring the reliability of the method.

[0103] S23255, obtaining pre-selected data analysis information based on the separation time information.

[0104] It is understood that the method for obtaining the pre-selected data analysis information based on the separation time information may be to receive data transmitted by the user, or to determine the difference between the pre-separation time and the post-separation time of the same content keyword in the same browsing history, but is not limited to this. The pre-selected data analysis information obtained based on the separation time information can provide a basis for subsequent steps.

[0105] In one possible implementation, see Figure 3 , S23255, obtain preselected data analysis information based on the separation time information, including:

[0106] S232551, step e, confirming a content keyword corresponding to a post-separation duration as a comparison word.

[0107] It can be understood that identifying a content keyword corresponding to a post-separation duration as a contrasting word can provide a basis for subsequent steps.

[0108] S232552, step f, determines whether there is a content keyword identical to the comparison word in the content keyword corresponding to the front separation time length of the browsing record corresponding to the comparison word. If so, execute step g; if not, confirm the separation word corresponding to the rear separation time length as the first analysis word and then execute step h.

[0109] It can be understood that if the content keyword corresponding to the preceding delimiting duration of the browsing record corresponding to the comparison word does not contain the same content keyword as the comparison word, it means that the target object only obtained content related to the comparison word after obtaining the delimiting duration. Determining the delimiting duration corresponding to the following delimiting duration as the first analysis word can provide a basis for subsequent steps.

[0110] S232553, step g, determine the size of the front separation duration and the rear separation duration. If the front separation duration is greater than or equal to the rear separation duration, execute step h. If the front separation duration is less than the rear separation duration, confirm the separation word corresponding to the rear separation duration as the second analysis word.

[0111] It can be understood that if the pre-separation duration is greater than or equal to the post-separation duration, it means that the target subject has browsed content related to the comparison word for a sufficient time before obtaining the separation word. If the pre-separation duration is less than the post-separation duration, it means that after obtaining the separation word, the target subject has spent more time browsing web pages with content related to the comparison word. Identifying the separation word corresponding to the post-separation duration as the second analysis word can provide a basis for subsequent steps.

[0112] S232554, step h, confirming the content keyword corresponding to another post-separation time length as a comparison word and then repeating steps f and g until all content keywords are confirmed as comparison words.

[0113] It can be understood that after confirming the content keyword corresponding to another post-separation time length as a comparison word, repeating steps f and g until all content keywords are confirmed as comparison words can ensure the reliability of the method.

[0114] S232555: Confirm all the first analysis words and the second analysis words as pre-selected data analysis information.

[0115] It can be understood that confirming all the first analysis words and the second analysis words as pre-selected data analysis information can provide a basis for subsequent steps.

[0116] S233, obtaining data analysis information based on the pre-selected data analysis information and the interest frequency information.

[0117] It is understood that the method for obtaining data analysis information based on the preselected data analysis information and interest frequency information may be to identify the first analysis word and the second analysis word with the larger value in the interest frequency information as the data analysis information, or to send the preselected data analysis information and interest frequency information to the user and then receive the data transmitted by the user, etc., but is not limited to this. Obtaining data analysis information based on the preselected data analysis information and interest frequency information can mine data that indirectly affects the target object, improving the problem that existing data mining methods often miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0118] In one possible implementation, see Figure 3 S233: obtaining data analysis information based on the pre-selected data analysis information and the interest frequency information, including:

[0119] S2331 , sort the interest frequencies in the interest frequency information from largest to smallest to obtain a frequency table.

[0120] It is understandable that sorting the interest frequencies in the interest frequency information from largest to smallest can improve the processing speed of subsequent steps. Obtaining the frequency table can provide a basis for subsequent steps.

[0121] S2332: Confirm the interest frequencies in the top 30% in the frequency table as first analysis frequency information, and confirm the interest frequencies in the top 50% in the frequency table as second analysis frequency information.

[0122] It can be understood that confirming the first analysis frequency information and the second analysis frequency information can provide a basis for subsequent steps.

[0123] S2333, confirm the first analysis words that are identical to the content keywords corresponding to the interest frequency in the second analysis frequency information among all the first analysis words as the first data, and confirm the second analysis words that are identical to the content keywords corresponding to the interest frequency in the first analysis frequency information among all the second analysis words as the second data.

[0124] It is understandable that because the first analysis word clearly indicates that the target object only obtains content related to the comparison word after obtaining the first analysis word, there is a greater possibility that the first analysis word and the target object have a potential connection, so the first analysis word is compared with the second analysis frequency information with a larger scope. Because the second analysis word indicates that the target object has already obtained content related to the comparison word before obtaining the second analysis word, there is a relatively small possibility that the second analysis word and the target object have an implicit connection, so the second analysis word is compared with the first analysis frequency information with a smaller scope to improve the accuracy of the second data.

[0125] S2334: Confirm the first data and the second data as data analysis information.

[0126] It can be understood that identifying the first data and the second data as data analysis information can mine data that indirectly affects the target object, improving the problem that existing data mining methods easily miss some data that can indirectly have a direct impact on the target and indirectly produce results.

[0127] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0128] Corresponding to the key data mining method based on big data described in the above embodiment, the embodiment of the present application also provides a key data mining system based on big data, and each unit of the system can implement each step of the key data mining method based on big data. Figure 4 A structural block diagram of a key data mining system based on big data provided by an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0129] Reference Figure 4 , the system comprises:

[0130] The acquisition unit is used to acquire mining scenario information; wherein the mining scenario information includes mining items and target objects.

[0131] The analysis unit is used to obtain data analysis information based on the mining scenario information; wherein the data analysis information includes at least one keyword reflecting key data.

[0132] It should be noted that the information interaction, execution process, etc. between the above-mentioned units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0133] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0134] The present application also provides a key data mining device based on big data. Figure 5 This is a schematic diagram of the structure of a key data mining device based on big data provided by an embodiment of the present application. Figure 5 As shown, the key data mining device based on big data of this embodiment includes a control device 6. The control device 6 includes: at least one processor 60 ( Figure 5 Only one is shown), at least one memory 61 ( Figure 5Only one is shown in the figure) and a computer program 62 stored in the at least one memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the key data mining device based on big data implements the steps of any of the above-mentioned key data mining method embodiments based on big data, or implements the functions of each unit in the above-mentioned system embodiments.

[0135] For example, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 62 in the control device 6.

[0136] The control device 6 can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The control device 6 can include, but is not limited to, a processor 60 and a memory 61. It will be understood by those skilled in the art that Figure 5 The illustrations are merely examples of key data mining equipment based on big data and do not constitute a limitation on the key data mining equipment based on big data. The illustrations may include more or fewer components than shown in the illustrations, or a combination of certain components, or different components. For example, the illustrations may also include input and output devices, network access devices, buses, etc.

[0137] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0138] In some embodiments, the memory 61 may be an internal storage unit of the control device 6, such as a hard drive or memory of the control device 6. In other embodiments, the memory 61 may also be an external storage device of the control device 6, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the control device 6. Furthermore, the memory 61 may include both the internal storage unit of the control device 6 and an external storage device. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or is about to be output.

[0139] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0140] An embodiment of the present application provides a computer program product. When the computer program product is run on a key data mining device based on big data, the key data mining device based on big data implements the steps of any of the above method embodiments.

[0141] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a key data mining device based on big data, a recording medium, computer memory, read-only memory (ROM), random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. Examples include a USB flash drive, a removable hard drive, a magnetic disk, or an optical disk. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunications signals.

[0142] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0143] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0144] In the embodiments provided in the present application, it should be understood that the disclosed key data mining system, device and method based on big data can be implemented in other ways. For example, the key data mining system and device embodiments based on big data described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0145] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0146] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A key data mining method based on big data, characterized in that: include: Acquire mining scenario information; wherein the mining scenario information includes mining items and target objects; Obtaining data analysis information based on the mining scenario information; wherein the data analysis information includes at least one information reflecting key data; The obtaining of data analysis information based on the mining scenario information includes: Obtaining keyword information based on the mining project; wherein the keyword information includes a plurality of keywords related to the mining project; Obtaining browsing information based on the target object; wherein the browsing information includes a plurality of browsing records after data cleaning and daily timestamps corresponding to the browsing records, the browsing records including all search keywords of the target object in a single day and the search time corresponding to each search keyword, and all web page addresses visited by the target object in a single day, the content keywords corresponding to each web page address, the visit time, and the length of stay; Obtaining the data analysis information based on the keyword information and the browsing information; The obtaining of the data analysis information based on the keyword information and the browsing information includes: Obtaining interest duration information and interest frequency information based on the browsing information; wherein the interest duration information includes at least one interest duration, the interest duration reflecting the sum of the dwell times of the web page address corresponding to a certain content keyword in the browsing record; and the interest frequency information includes at least one interest frequency, the interest frequency reflecting the ratio of the number of the web page addresses corresponding to a certain content keyword in the browsing record to the total number of the web page addresses in the browsing record; Obtaining preselected data analysis information based on the interest duration information and the keyword information; wherein the preselected data analysis information includes at least one first analysis word and at least one second analysis word, the first analysis word reflects one of the content keywords, and the second analysis word reflects one of the content keywords; The data analysis information is obtained based on the preselected data analysis information and the interest frequency information.

2. The key data mining method based on big data according to claim 1, characterized in that: The obtaining of interest duration information and interest frequency information based on the browsing information includes: Analyze each of the browsing records separately, identify the search time corresponding to each search keyword in the browsing record as a search node, identify the web page address adjacent to and before the search node as an interest URL, and identify the web page address adjacent to and after the search node as a verification URL; If the content keyword corresponding to the interest website is related to the content keyword corresponding to the verification website adjacent to the search node adjacent to the interest website, then after adding interest tags to the interest website and the verification website, taking the verification website as the search starting point, the interest tags are added to the web page addresses in sequence in the time order reflected by the access time until the content keyword corresponding to the web page address is no longer related to the content keyword corresponding to the interest website; wherein the interest tag is the content keyword corresponding to the interest website; After adding the dwell times corresponding to the web page addresses corresponding to the interest tags with the same content keywords to obtain at least one interest duration corresponding to the content keyword corresponding to the interest tag, all the interest durations are confirmed as interest duration information; Each browsing record is analyzed separately, and the value obtained by adding up the number of web page addresses with the same content keyword in the browsing record and dividing it by the total number of web page addresses in the browsing record is confirmed as the interest frequency, and then all the interest frequencies are confirmed as the interest frequency information.

3. The key data mining method based on big data according to claim 1, characterized in that: The obtaining of pre-selected data analysis information based on the interest duration information and the keyword information includes: Confirming the interest duration in which the content keyword corresponding to the interest duration information belongs to the keyword information as preference information, and confirming the interest duration in which the content keyword corresponding to the interest duration information does not belong to the keyword information as preference comparison information; Dividing each interest duration in the preference information by the sum of all the stay durations in the corresponding browsing history to obtain multiple preference ratios, and dividing each interest duration in the preference comparison information by the sum of all the stay durations in the corresponding browsing history to obtain multiple preference comparison ratios; grouping the preference ratios and the preference comparison ratios according to the corresponding content keywords, so that the preference ratios or the preference comparison ratios in each group correspond to the same content keywords, and then sorting the preference ratios or the preference comparison ratios in each group from earliest to latest according to the corresponding time stamps of the browsing records to obtain a sorting table; Obtaining analysis word information based on the sorting table; The preselected data analysis information is obtained by analyzing each of the browsing records based on the analysis word information.

4. The key data mining method based on big data according to claim 3, characterized in that: The obtaining of analysis word information based on the sorting table includes: Analyze each group in the ranking table separately; Step 1: Determine the nth preference ratio or the preference comparison ratio in the group as the first value. If n+1 is less than or equal to the number of the preference ratios or the preference comparison ratios in the group, determine the n+1th preference ratio or the preference comparison ratio in the group as the second value. If n+1 is greater than the number of the preference ratios or the preference comparison ratios in the group, execute step 3; wherein, the initial value of n is 1; Step 2: Determine the value obtained by subtracting the first value from the second value and dividing the result by the first value as the rate of change, and substitute n+1 into n in step 1 and repeat step 1; Step 3: If the data in the group are all the preference ratios, all the obtained change rates are confirmed as the preference change trend; if the data in the group are all the preference comparison ratios, all the obtained change rates are confirmed as the comparison change trend; After all groups in the ranking table are analyzed, the interval formed by multiplying the change rate in each preference change trend by 1.05 and 0.95 is determined as the change interval; Analyze each of the comparative change trends respectively, compare the comparative change trends with each of the preference change trends, and identify the content keywords corresponding to the comparative change trends whose change rates in the comparative change trends are within the change interval corresponding to the time in the same preference change trend as analysis words; All of the analysis words are confirmed as the analysis word information.

5. The key data mining method based on big data according to claim 3, characterized in that: The analyzing each of the browsing records based on the analysis word information to obtain the pre-selected data analysis information includes: Step a, confirming one of the analysis words in the analysis word information as a separator word; Step b, respectively determining the search time corresponding to the first search keyword identical to the separator word or the access time corresponding to the content keyword in each browsing record that appears in chronological order as the separator time; Step c, summing the dwell times corresponding to the web page addresses with the same content keyword in each browsing record before the separation time to obtain at least one pre-separation time, and summing the dwell times corresponding to the web page addresses with the same content keyword after the separation time to obtain at least one post-separation time; Step d, determining whether all the analysis words in the analysis word information are confirmed as the separation words; if all the analysis words in the analysis word information are confirmed as the separation words, confirming all the pre-separation durations and post-separation durations as separation time information; if there are analysis words in the analysis word information that are not confirmed as the separation words, confirming one of the analysis words that are not confirmed as the separation words as the separation word and then repeating steps b, c, and d; wherein the separation time information includes at least one pre-separation duration, at least one post-separation duration, and at least one separation word corresponding to both the pre-separation duration and the post-separation duration corresponding to each browsing record; The preselected data analysis information is obtained based on the separation time information.

6. The key data mining method based on big data according to claim 5, characterized in that: The obtaining of the preselected data analysis information based on the separation time information includes: Step e, identifying the content keyword corresponding to the post-separation duration as a comparison word; Step f, determining whether the content keyword corresponding to the preceding separation duration of the browsing record corresponding to the comparison word contains the same content keyword as the comparison word; if so, executing step g; if not, confirming the separation word corresponding to the following separation duration as the first analysis word and then executing step h; Step g, determining the difference between the pre-separation duration and the post-separation duration; if the pre-separation duration is greater than or equal to the post-separation duration, executing step h; if the pre-separation duration is less than the post-separation duration, determining the separator word corresponding to the post-separation duration as the second analysis word; Step h, confirming the content keyword corresponding to another of the post-separation durations as the comparison word and then repeating steps f and g until all the content keywords are confirmed as the comparison words; All of the first analysis words and the second analysis words are confirmed as the pre-selected data analysis information.

7. The key data mining method based on big data according to claim 1, characterized in that: The obtaining the data analysis information based on the pre-selected data analysis information and the interest frequency information includes: Sort the interest frequencies in the interest frequency information from largest to smallest to obtain a frequency table; Confirming the interest frequencies in the top 30% of the frequency table as first analysis frequency information, and confirming the interest frequencies in the top 50% of the frequency table as second analysis frequency information; Confirming the first analysis words that are identical to the content keyword corresponding to the interest frequency in the second analysis frequency information among all the first analysis words as first data, and confirming the second analysis words that are identical to the content keyword corresponding to the interest frequency in the first analysis frequency information among all the second analysis words as second data; The first data and the second data are confirmed as the data analysis information.

8. A key data mining device based on big data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Software application data mining method based on big data positioning and software service platform

    CN112600893A

  • User portrait construction method and system based on big data

    CN118098584A