Internet open source information data cleaning method and system

By analyzing users' travel route data and calculating travel preferences, the problem that existing methods cannot effectively clean up travel information data of non-travel enthusiasts is solved, and the construction quality of knowledge graphs and the accuracy of recommendations is improved.

CN120216876APending Publication Date: 2025-06-27SHANDONG CENTURY SAIFU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312866.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing methods fail to effectively clean travel information data from non-travel enthusiasts, resulting in a decline in the construction quality of knowledge graphs.

Method used

By obtaining users' travel route data, analyzing travel time clustering, travel possibility, popular index and vehicle speed similarity, filtering out travel routes and calculating travel preferences, and then cleaning the travel information data of non-travel enthusiasts.

Benefits of technology

Effectively remove travel information data from non-travel enthusiasts, improve the quality of the construction of knowledge graphs, and ensure that the recommended travel information is more in line with the needs of travel enthusiasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216876A_ABST
    Figure CN120216876A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data cleaning, in particular to an internet open source information data cleaning method and system. The method comprises the following steps: firstly, acquiring travel start time and travel end time of each travel route selected by different users, and driving speed data of the users at each moment on each travel route, and acquiring travel time aggregation degree of the users according to distribution of the travel start time of the travel routes of the users; the method comprises the steps of obtaining travel time aggregation degree and quantity of users, screening out travel routes from travel routes of each user, obtaining travel preference degree of each user according to travel time aggregation degree and quantity of the users selecting the same travel route and in combination with difference of travel speed data of the users at each moment on the same travel route, and obtaining travel preference degree of each user on the basis of the travel preference degree. And cleaning the travel information data of each user, and constructing a knowledge graph. According to the method, the travel information data of the non-travel enthusiasts can be effectively cleaned and removed, and the construction quality of the knowledge graph is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data cleaning, and particularly to an Internet open-source information data cleaning method and system. Background Art

[0002] There is a large amount of open-source information data on the Internet. In the field of tourism information recommendation, a knowledge graph related to travel for each user can be constructed, and various personalized travel information can be recommended to specific travel users through the knowledge graph, improving the travel experience of users.

[0003] Since there are some non-travel lovers among a large number of users, the travel information data of non-travel lovers will reduce the quality of the constructed knowledge graph. In related technologies, usually, the travel information data of non-travel lovers is cleaned and removed, and the travel information of each user after cleaning is used to construct the corresponding knowledge graph. However, due to the unclear characteristics of the travel information data of non-travel lovers, the existing methods cannot effectively clean and remove the travel information data of non-travel lovers, thus reducing the construction quality of the knowledge graph. Summary of the Invention

[0004] In order to solve the technical problem that the existing methods cannot effectively clean and remove the travel information data of non-travel lovers, thus reducing the construction quality of the knowledge graph, the purpose of the present invention is to provide an Internet open-source information data cleaning method and system, and the specific technical solutions adopted are as follows:

[0005] The present invention proposes an Internet open-source information data cleaning method, and the method includes:

[0006] Obtain the travel start time and travel end time of each travel route selected by different users within a preset time period, and the driving speed data of each user at each moment on each travel route;

[0007] Obtain the travel time aggregation degree of each user according to the distribution of the travel start time of each travel route of each user;

[0008] Take any user as the target user, and obtain the travel possibility of each travel route of the target user according to the differences in the travel start time and the travel end time between the target user and other users except the target user on the same travel route; based on the travel possibility, screen out multiple travel routes of the target user from all the travel routes of the target user;

[0009] Take any one of the target user's travel routes as the target travel route of the target user, and obtain the popularity index of the target user's target travel route according to the travel time concentration degree and quantity of each user who selects the target travel route; obtain the vehicle speed similarity of the target user's target travel route according to the difference in the driving speed data of the target user and other users at each moment on the target travel route; obtain the travel preference degree of the target user according to the vehicle speed similarity and the popularity index of all the travel routes of the target user.

[0010] Based on the travel preference degree of each user, clean the travel information data of each user and construct a knowledge graph.

[0011] Further, the obtaining of the travel time concentration degree of each user includes:

[0012] Take any one user as the user to be measured. Based on the travel start time of each travel route of the user to be measured, divide all the travel routes of the user to be measured into four categories, and the categories are spring category, summer category, autumn category and winter category respectively;

[0013] Use the number of all travel routes of the user to be measured in each category as the numerator, and the number of all travel routes of the user to be measured as the denominator, and use the ratio as the proportion of the travel quantity of each category of the user to be measured.

[0014] Analyze the dispersion degree of the proportion of the travel quantity of all categories of the user to be measured to obtain the travel time concentration degree of the user to be measured.

[0015] Further, the obtaining of the travel possibility of each travel route of the target user includes:

[0016] Take any one of the target user's travel routes as the target travel route of the target user, and select multiple first reference users of the target user regarding the target travel route from all other users except the target user, where each of the first reference users' travel routes includes the target user's target travel route;

[0017] Take the absolute value of the difference in the travel start time between the target user and each first reference user on the target travel route as the travel start time difference between the target user and each first reference user regarding the target travel route, and take the absolute value of the difference in the travel end time between the target user and each first reference user on the target travel route as the travel end time difference between the target user and each first reference user regarding the target travel route;

[0018] Take the sum value of the travel start time difference and the travel end time difference as the comprehensive time difference between the target user and each first reference user regarding the target travel route.

[0019] Perform a negative correlation normalization on the average value of the comprehensive time difference between the target user and all first reference users regarding the target travel route to obtain the travel possibility of the target travel route of the target user.

[0020] Furthermore, screening out multiple travel routes of the target user from all travel routes of the target user includes:

[0021] Among all travel routes of the target user, use the travel routes with a travel possibility greater than a preset possibility threshold as the travel routes of the target user.

[0022] Furthermore, obtaining the popularity index of the target travel route of the target user includes:

[0023] Select multiple audience users of the target travel route from all users, where the travel route of each said audience user includes the target travel route of the target user;

[0024] Use the cumulative value of the travel time concentration of all audience users of the target travel route as the popularity evaluation value of the target travel route;

[0025] Integrate the number of all audience users of the target travel route and the popularity evaluation value of the target travel route to obtain the popularity index of the target travel route of the target user.

[0026] Furthermore, obtaining the vehicle speed similarity of the target travel route of the target user includes:

[0027] Obtain the traffic flow data of each user at each moment on each travel route;

[0028] According to the correlation between the traffic flow data and the driving speed data of each user at each moment on each travel route, adjust the driving speed data of each user at each moment on each travel route to obtain the adjusted driving speed of each user at each moment on each travel route;

[0029] Sort the adjusted driving speeds of each user at each moment on each travel route in chronological order to obtain the adjusted driving speed sequence of each user on each travel route;

[0030] Select multiple second reference users of the target user regarding the target travel route from all other users except the target user, where the travel route of each said second reference user includes the target travel route of the target user;

[0031] Using the dynamic time warping algorithm, process the adjusted driving speed sequences of the target user and each second reference user on the target travel route to obtain the vehicle speed difference degree between the target user and each second reference user regarding the target travel route;

[0032] Perform a negative correlation mapping on the average value of the vehicle speed difference degrees between the target user and all second reference users regarding the target travel route to obtain the vehicle speed similarity of the target travel route of the target user.

[0033] Further, the obtaining of the adjusted driving speed of each user at each moment on each travel route includes:

[0034] Sort the driving speed data of the target user at each moment on the target travel route in chronological order to obtain the driving speed data sequence of the target user on the target travel route, and sort the traffic flow data of the target user at each moment on the target travel route to obtain the traffic flow data sequence of the target user on the target travel route;

[0035] Perform a negative correlation normalization on the Pearson correlation coefficient between the driving speed data sequence and the traffic flow data sequence to obtain the vehicle speed adjustment weight of the target user on the target travel route;

[0036] Take the product value of the vehicle speed adjustment weight of the target user on the target travel route and the driving speed data of the target user at each moment on the target travel route as the vehicle speed adjustment amount of the target user at each moment on the target travel route;

[0037] Take the sum value of the driving speed data of the target user at each moment on the target travel route and the vehicle speed adjustment amount as the adjusted driving speed of the target user at each moment on the target travel route.

[0038] Further, the obtaining of the travel preference degree of the target user includes:

[0039] Integrate the popularity index and the vehicle speed similarity of each travel route of the target user to obtain the travel status coefficient of each travel route of the target user;

[0040] Perform a normalization on the average value of the travel status coefficients of all travel routes of the target user to obtain the travel preference degree of the target user.

[0041] Further, the cleaning of the travel information data of each user based on the travel preference degree of each user and the construction of the knowledge graph include:

[0042] Remove the users whose travel preference level is less than the preset preference threshold and their travel information data, and retain the remaining users and their travel information data;

[0043] Perform negative correlation normalization on the travel preference level of each retained user to obtain the distance coefficient of each retained user;

[0044] Based on all the retained users and their travel information data, construct a knowledge graph, where each retained user serves as an entity in the knowledge graph. The knowledge graph includes a node for the group of travel enthusiasts and nodes corresponding to each retained user, and the distance between the node corresponding to each retained user and the node for the group of travel enthusiasts is equal to the distance coefficient of each retained user.

[0045] The present invention also proposes an Internet open-source information data cleaning system, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of any one of the Internet open-source information data cleaning methods.

[0046] The present invention has the following beneficial effects:

[0047] In view of the fact that the existing methods cannot effectively clean and remove the travel information data of non-travel enthusiasts, thus reducing the construction quality of the knowledge graph, the present invention first obtains the departure time and arrival time of each travel route selected by different users within a preset time period, as well as the driving speed data of each user at each moment on each travel route. Subsequently, based on various data, the degree of preference of each user for travel can be analyzed, so as to clean and remove the travel information data of non-travel enthusiasts, and construct a knowledge graph with better quality. Considering that the multiple travel activities of each user are aggregated in terms of time and season, the degree of aggregation of each user's travel activities in time can be reflected by the travel time aggregation degree. Considering that not all of the travel routes of each user are the routes selected by the user during travel, and the departure time and arrival time between the travel routes selected by a certain user during travel and the same travel routes selected by other users are usually relatively close, the possibility that the target user's travel activity belongs to each travel route can be reflected by the travel possibility, and then the travel routes of the target user can be screened out, and the popularity index obtained can reflect the popularity of the target travel routes of the target user. Considering that the driving speeds of users in the travel state on the same travel route are relatively similar, the similarity degree of the driving speeds of the target user and other users on the target travel route can be reflected by the obtained vehicle speed similarity, and then the degree of preference of the target user for travel can be reflected by the travel preference degree. Based on the travel preference degree of each user, the travel information data of each user is cleaned, the travel information data of non-travel enthusiasts is effectively cleaned and removed, and a knowledge graph with better quality is constructed. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 Flowchart of an Internet open-source information data cleaning method provided by an embodiment of the present invention;

[0050] Figure 2 Flowchart of a method for obtaining the vehicle speed similarity of the target travel route of the target user provided by an embodiment of the present invention;

[0051] Figure 3 Structural schematic diagram of the knowledge graph constructed by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method and system for cleaning Internet open-source information data proposed according to the present invention, including its specific implementation manners, structures, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0054] The following specifically describes the specific solutions of a method and system for cleaning Internet open-source information data provided by the present invention with reference to the accompanying drawings.

[0055] Please refer to Figure 1 , which shows a flowchart of a method for cleaning Internet open-source information data provided by an embodiment of the present invention. The method includes:

[0056] Step S1: Obtain the departure time and arrival time of each travel route selected by different users within a preset time period, as well as the driving speed data of the user at each moment on each travel route.

[0057] In an embodiment of the present invention, first, through a cloud retrieval system, the departure time and arrival time of each travel route selected by different users within a preset time period, as well as the driving speed data of the user at each moment on each travel route, are collected from the navigation recorders of each user. Among them, the preset time period is set to 5 years, and the specific value of the preset time period can also be set by the implementer according to the specific implementation scenario and is not limited herein. It can be understood that when collecting or acquiring data in an embodiment of the present invention, the consent of the relevant users is obtained, and the process does not violate relevant laws and regulations and does not violate public order and good customs.

[0058] It should be noted that a travel route of a certain user can represent a travel activity of the user. The departure time of a certain travel route of the user represents the time when starting from the starting point of this travel route selected by the user during this going-out activity of the user. The arrival time of a certain travel route of the user represents the corresponding end time when arriving at the end point of this travel route selected by the user during this going-out activity of the user. And different users may select the same travel route.

[0059] Step S2: Obtain the travel time aggregation degree of each user according to the distribution of the departure times of each travel route of each user.

[0060] Since the start times of the travel routes selected by users traveling for the purpose of travel are characterized by aggregation, the distribution of the start times of each travel route of each user can be analyzed, and the aggregation degree of travel time obtained can reflect the degree of aggregation of each user's travel activities in time. The greater the time aggregation degree of a certain user, the more obvious the aggregation characteristics of the user's travel activities in time, which further indicates that the user has more travel activities for the purpose of travel, and the greater the possibility that the user is a travel enthusiast.

[0061] Preferably, in an embodiment of the present invention, the method for obtaining the travel time aggregation degree of each user specifically includes:

[0062] Taking any user as the user to be measured, based on the start time of each travel route of the user to be measured, all travel routes of the user to be measured are divided into four categories, where the four categories are the spring category, the summer category, the autumn category, and the winter category. For example, if the start time of a certain travel route of the user to be measured belongs to January to March, then this travel route of the user to be measured is divided into the spring category; if the start time of a certain travel route of the user to be measured belongs to April to June, then this travel route of the user to be measured is divided into the summer category, and so on.

[0063] Taking the number of all travel routes of the user to be measured in each category as the numerator and the number of all travel routes of the user to be measured as the denominator, and taking the ratio as the proportion of the travel quantity of each category of the user to be measured.

[0064] Since travel activities are characterized by aggregation in seasons, for example, the climate in spring and autumn is suitable and more suitable for travel than other seasons. Therefore, if the degree of dispersion of the distribution of the proportion of travel quantities in each category of the user to be measured is greater, it indicates that the travel activities of the user to be measured have aggregation characteristics in time or seasons, which further indicates that most of the travel routes of the user to be measured are selected for the purpose of travel. Therefore, the degree of dispersion of the proportion of travel quantities in all categories of the user to be measured can be analyzed to obtain the travel time aggregation degree of the user to be measured.

[0065] In an embodiment of the present invention, the standard deviation or variance of the proportion of travel quantities in all categories of the user to be measured can be used as the travel time aggregation degree of the user to be measured to realize the analysis of the degree of dispersion of the proportion of travel quantities in all categories of the user to be measured, and no limitation is made here.

[0066] The travel time aggregation degree of each user can be obtained by the above same method.

[0067] Step S3: Take any user as the target user, and obtain the travel possibility of each travel route of the target user based on the differences in the start time and end time of the same travel route between the target user and other users except the target user; based on the travel possibility, filter out multiple travel routes of the target user from all the travel routes of the target user.

[0068] Since a user does not choose travel routes solely for travel purposes. For example, a user may also choose travel routes for other purposes such as business trips. That is to say, not all of a user's travel routes are the routes chosen by the user during travel. Since users usually have the habit of traveling in groups, and the travel activities of each user are clustered in time, the start time of the travel route chosen by a user during travel is usually close to the start time of the same travel route chosen by other users during travel, and the end time is also close. Therefore, any user can be taken as the target user first, and the differences in the start time and end time of the same travel route between the target user and other users except the target user are analyzed. The obtained travel possibility reflects the possibility that the target user's travel on each travel route belongs to a travel activity. The greater the travel possibility, the more likely it is that the target user chose the travel route for travel purposes. Subsequently, multiple travel routes of the target user can be filtered out based on the travel possibility.

[0069] Preferably, in an embodiment of the present invention, the method for obtaining the travel possibility of each travel route of the target user specifically includes:

[0070] Take any travel route of the target user as the target travel route of the target user, and select multiple first reference users of the target user for the target travel route from all other users except the target user. Among them, the travel routes of each first reference user include the target travel route of the target user.

[0071] Take the absolute value of the difference in the start time of the target travel route between the target user and each first reference user as the start time difference of the target travel route between the target user and each first reference user, and take the absolute value of the difference in the end time of the target travel route between the target user and each first reference user as the end time difference of the target travel route between the target user and each first reference user.

[0072] Take the sum of the start time difference and the end time difference as the comprehensive time difference of the target travel route between the target user and each first reference user.

[0073] The smaller the comprehensive time difference, the smaller the difference in the start time of the target travel route selected by the target user and other users, and the smaller the difference in the end time of the travel. This indicates that there is a characteristic of group travel between the target user and other users, and the travel activities between the target user and other users have an aggregation characteristic in time. Furthermore, it shows that the target user is more likely to choose the target travel route for the purpose of travel. Therefore, a negative correlation normalization process can be performed on the average value of the comprehensive time difference between the target user and all first reference users regarding the target travel route, and the calculation result is limited to within the range, so as to obtain the travel possibility of the target travel route of the target user.

[0074] In an embodiment of the present invention, a function form of can be used to implement the negative correlation normalization process, and the negative correlation normalization in subsequent steps can all be processed using this function form. Among them, represents the normalization function. In an embodiment of the present invention, the normalization process can specifically be, for example, the maximum-minimum normalization process, and the normalization in subsequent steps can all use the maximum-minimum normalization process. In other embodiments of the present invention, other normalization methods can be selected according to the specific range of values, which will not be elaborated here.

[0075] As an example, in an embodiment of the present invention, the expression of the travel possibility of the target travel route of the target user can specifically be, for example:

[0076]

[0077] Among them, represents the travel possibility of the target travel route of the target user; represents the start time of the target travel route of the target user; represents the th start time of the target travel route of the first reference user of the target user; represents the th difference in the start time of the target travel route between the target user and the first reference user; represents the end time of the target travel route of the target user; represents the th end time of the target travel route of the first reference user of the target user; represents the th difference in the end time of the target travel route between the target user and the first reference user; represents the th comprehensive time difference of the target travel route between the target user and the first reference user; represents the number of all first reference users of the target user; represents a normalization function.

[0078] By the same method as above, the travel possibility of each travel route of the target user can be obtained. The greater the travel possibility of a certain travel route of the target user, the more likely it is that this travel route is selected by the target user for travel purposes, and further indicates that this travel route is more likely to be the travel route of the target user. Therefore, based on the travel possibility, multiple travel routes of the target user can be screened out from all the travel routes of the target user.

[0079] Preferably, in an embodiment of the present invention, the method for obtaining multiple travel routes of the target user specifically includes:

[0080] Among all the travel routes of the target user, the travel routes with travel possibility greater than the preset possibility threshold are used as the travel routes of the target user, where the value range of the preset possibility threshold is set to , in an embodiment of the present invention, the preset possibility threshold is set to 0.7. The specific value of the preset possibility threshold can also be set by the implementer according to the specific implementation scenario and is not limited herein.

[0081] By the same method as above, multiple travel routes of each user can be obtained.

[0082] Step S4: Take any one of the travel routes of the target user as the target travel route of the target user, and obtain the popularity index of the target travel route of the target user according to the travel time concentration degree and quantity of the users who select the target travel route; obtain the vehicle speed similarity of the target travel route of the target user according to the difference in the vehicle speed data of each moment on the target travel route between the target user and other users; obtain the travel preference degree of the target user according to the vehicle speed similarity and popularity index of all the travel routes of the target user.

[0083] A certain travel route of the target user may also be selected by other users. For a certain travel route of the target user, the more users who select this travel route and the greater the travel time concentration degree of the users who select this travel route, it indicates that the popularity of this travel route of the target user is higher and it is more popular among travel enthusiasts. Further, it indicates that the target user who selects this travel route is more likely to be a travel enthusiast. Therefore, first take any one of the travel routes of the target user as the target travel route of the target user, and obtain the popularity index of the target travel route of the target user according to the travel time concentration degree and quantity of the users who select the target travel route. Subsequently, based on the popularity index of each travel route of the target user, the travel preference degree of the target user can be analyzed.

[0084] Preferably, in an embodiment of the present invention, the method for obtaining the popularity index of the target travel route of the target user specifically includes:

[0085] Select multiple audience users of the target travel route from all users. Among them, the travel routes of each audience user include the target travel route of the target user. Therefore, the audience users necessarily include the target user.

[0086] Take the cumulative value of the travel time concentration of all audience users of the target travel route as the popularity evaluation value of the target travel route.

[0087] Comprehensively consider the number of all audience users of the target travel route and the popularity evaluation value of the target travel route to obtain the popularity index of the target travel route of the target user.

[0088] In the embodiment of the present invention, the sum value or product value of the number of all audience users of the target travel route and the popularity evaluation value of the target travel route can be used as the popularity index of the target travel route of the target user, so as to achieve the comprehensive consideration of the two, and there is no limitation here.

[0089] As an example, in an embodiment of the present invention, the expression of the popularity index of the target travel route of the target user can be specifically, for example:

[0090]

[0091] Among them, represents the popularity index of the target travel route of the target user; represents the number of all audience users of the target travel route; represents the th audience user's travel time concentration of the target travel route; represents the popularity evaluation value of the target travel route.

[0092] By the same method as above, the popularity index of each travel route of the target user, as well as the popularity index of each travel route of each user, can be obtained, and the popularity indexes of the same travel route of different users are the same.

[0093] Since travel enthusiasts are in a travel and leisure state when driving on a travel route, and the road conditions of the same travel route selected by different users are similar, the driving speeds of different travel enthusiasts on the same travel route are relatively close. Therefore, the differences in the driving speed data of the target user and other users at each moment on the target travel route can be analyzed, and the similarity of vehicle speeds obtained can reflect the degree of closeness of the driving speeds of the target user and other users on the target travel route. The greater the similarity of vehicle speeds, the more likely it is that the target user is in a travel state on the target travel route, and further indicates that the target user who selects the target travel route is more likely to be a travel enthusiast. Subsequently, the degree of preference of the target user for travel can be accurately calculated and analyzed by combining the similarity of vehicle speeds and the popularity index of each travel route of the target user.

[0094] Preferably, in an embodiment of the present invention, the method for obtaining the similarity of vehicle speeds of the target travel route of the target user specifically includes:

[0095] Please refer to Figure 2 , which shows a flowchart of the method for obtaining the similarity of vehicle speeds of the target travel route of the target user provided by an embodiment of the present invention.

[0096] Step S401: Obtain the traffic flow data of each user at each moment on each travel route; according to the correlation between the traffic flow data and the driving speed data of each user at each moment on each travel route, adjust the driving speed data of each user at each moment on each travel route to obtain the adjusted driving speed of each user at each moment on each travel route.

[0097] First, use the cloud collection system to collect the traffic flow data of each user at each moment on each travel route from the navigation recorders of each user. Among them, the collection frequencies of the traffic flow data and the driving speed data of a specific user on the travel route are the same, that is, each moment on the travel route corresponds to a traffic flow data and a driving speed data. Since in the actual driving process, the traffic flow on the road will affect the driving speed, the greater the traffic flow, the smaller the driving speed. Therefore, in order to remove the interference of traffic flow on the driving speed, it is necessary to adjust the driving speed data of each user at each moment on each travel route according to the correlation between the traffic flow data and the driving speed data of each user at each moment on each travel route to obtain the adjusted driving speed of each user at each moment on each travel route. Subsequently, based on the differences in the adjusted driving speeds of the target user and other users at each moment on the same travel route, the similarity of vehicle speeds of the target travel route of the target user can be accurately calculated and analyzed.

[0098] Preferably, in an embodiment of the present invention, the method for obtaining the adjusted driving speed of each user at each moment on each travel route specifically includes:

[0099] Sort the driving speed data of the target user at each moment on the target travel route in chronological order to obtain the driving speed data sequence of the target user on the target travel route, and sort the traffic flow data of the target user at each moment on the target travel route to obtain the traffic flow data sequence of the target user on the target travel route.

[0100] Since the traffic flow data and the driving speed data usually show a negative correlation, the stronger the negative correlation between the two, the greater the degree of interference of the driving speed data by the traffic flow, and the greater the degree of adjustment required for the driving speed data. Therefore, the Pearson correlation coefficient between the driving speed data sequence and the traffic flow data sequence can be normalized for negative correlation, and the calculation result is limited to within the range, so as to obtain the vehicle speed adjustment weight of the target user on the target travel route. The greater the vehicle speed adjustment weight, the greater the degree of adjustment required for the driving speed data of the target user at each moment on the target travel route.

[0101] Furthermore, the product value of the vehicle speed adjustment weight of the target user on the target travel route and the driving speed data of the target user at each moment on the target travel route can be used as the vehicle speed adjustment amount of the target user at each moment on the target travel route. Since a large traffic flow will reduce the driving speed, the sum value of the driving speed data of the target user at each moment on the target travel route and the vehicle speed adjustment amount can be used as the adjusted driving speed of the target user at each moment on the target travel route.

[0102] As an example, in an embodiment of the present invention, the expression of the adjusted driving speed of the target user at each moment on the target travel route can be specifically, for example:

[0103]

[0104]

[0105] Among them, represents the adjusted driving speed of the target user at the th moment on the target travel route; represents the driving speed data of the target user at the th moment on the target travel route; represents the vehicle speed adjustment weight of the target user on the target travel route; represents the vehicle speed adjustment amount of the target user at the th moment on the target travel route; represents the Pearson correlation coefficient between the driving speed data sequence and the traffic flow data sequence of the target user on the target travel route; represents a normalization function.

[0106] By the same method as above, the adjusted driving speed of the target user at each moment on each travel route, and the adjusted driving speed of each user at each moment on each travel route can be obtained.

[0107] Step S402: Sort the adjusted driving speeds of each user at each moment on each travel route in chronological order to obtain the adjusted driving speed sequence of each user on each travel route.

[0108] Subsequently, based on the adjusted driving speed sequences of the target user and other users on the same travel route, the proximity of the driving speeds between the target user and other users on the same travel route can be analyzed.

[0109] Step S403: Select multiple second reference users of the target user regarding the target travel route from all other users except the target user, where the travel route of each second reference user includes the target travel route of the target user; use the dynamic time warping algorithm to process the adjusted driving speed sequences of the target user and each second reference user on the target travel route to obtain the vehicle speed difference degree between the target user and each second reference user regarding the target travel route.

[0110] Then select multiple second reference users of the target user regarding the target travel route from all other users except the target user, where the travel route of each second reference user includes the target travel route of the target user. Since the lengths of the adjusted driving speed sequences of the target user and other users on the target travel route may be different, the dynamic time warping algorithm can be used to process the adjusted driving speed sequences of the target user and each second reference user on the target travel route to obtain the vehicle speed difference degree between the target user and each second reference user regarding the target travel route. The dynamic time warping algorithm can evaluate the difference between two time series sequences, which is a well-known technical means in the art and will not be elaborated here.

[0111] Step S404: Perform a negative correlation mapping on the average value of the vehicle speed difference degrees between the target user and all second reference users regarding the target travel route to obtain the vehicle speed similarity of the target travel route of the target user.

[0112] The smaller the speed difference degree between the target user and the second reference user regarding the target travel route, the closer the driving speeds of the target user and the second reference user on the same travel route, which further indicates that the target user who selects the target travel route is more likely to be a travel enthusiast. Therefore, a negative correlation mapping can be performed on the average value of the speed difference degrees between the target user and all second reference users regarding the target travel route to obtain the speed similarity of the target travel route of the target user.

[0113] As an example, in an embodiment of the present invention, the expression of the speed similarity of the target travel route of the target user can be specifically, for example:

[0114]

[0115] Wherein, represents the speed similarity of the target travel route of the target user; represents the speed difference degree between the target user and the th second reference user regarding the target travel route; represents the number of second reference users of the target user regarding the target travel route; represents the exponential function with the natural constant as the base, which is used for negative correlation mapping.

[0116] It should be noted that in other embodiments of the present invention, negative correlation mapping can also be achieved through other basic mathematical operations, which will not be elaborated here.

[0117] Through the above same method, the speed similarity of each travel route of the target user can be obtained. The greater the speed similarity and popularity index of each travel route of the target user, the more likely the target user is a travel enthusiast, and the greater the degree of preference of the target user for travel. Therefore, the degree of travel preference of the target user can be obtained based on the speed similarity and popularity index of all travel routes of the target user. Subsequently, based on the degree of travel preference of each user, some users and their travel information data can be cleaned and removed, so as to construct a knowledge graph with better quality.

[0118] Preferably, in an embodiment of the present invention, the method for obtaining the degree of travel preference of the target user specifically includes:

[0119] Integrate the popularity index and speed similarity of each travel route of the target user to obtain the travel status coefficient of each travel route of the target user. The greater the travel status coefficient, the more likely the target user is in a travel state on each travel route, and the greater the degree of preference of the target user for travel. Then, perform normalization processing on the average value of the travel status coefficients of all travel routes of the target user, and limit the calculation result within within a certain range to obtain the degree of travel preference of the target user.

[0120] In an embodiment of the present invention, the sum value or product value of the popularity index and the vehicle speed similarity of each travel route of the target user can be used as the travel status coefficient of each travel route of the target user, so as to achieve the integration of the two, and no limitation is made here.

[0121] As an example, in an embodiment of the present invention, the expression of the degree of travel preference of the target user can be specifically, for example:

[0122]

[0123] wherein, represents the degree of travel preference of the target user; represents the th popularity index of the travel route of the target user; represents the th vehicle speed similarity of the travel route of the target user; represents the th travel status coefficient of the travel route of the target user; represents the number of travel routes of the target user; represents the normalization function.

[0124] The degree of travel preference of each user can be obtained by the same method as above.

[0125] Step S5: Based on the degree of travel preference of each user, clean the travel information data of each user and construct a knowledge graph.

[0126] The greater the degree of travel preference of a certain user, the more likely it is that the user is a travel enthusiast, and further indicates that the user and his travel information data are more important in the process of constructing the knowledge graph. On the contrary, the smaller the degree of travel preference of a certain user, the more likely it is that the user is not a travel enthusiast, and further indicates that the user and his travel information data are less important in the process of constructing the knowledge graph. In order to construct a better-quality knowledge graph, it is necessary to clean and remove the user and his travel information data. Therefore, based on the degree of travel preference of each user, the travel information data of each user can be cleaned and a knowledge graph can be constructed. Among them, the travel information data of the user can be collected through the cloud retrieval system in step S1. The travel information data of the user includes, for example, the departure time and end time information when the user travels, the route information selected during the travel, and various information such as accommodation, scenic spots and food during the travel.

[0127] Preferably, in an embodiment of the present invention, the method for cleaning the travel information data of each user and constructing a knowledge graph specifically includes:

[0128] Remove users whose travel preference level is less than the preset preference threshold and their travel information data, and retain the remaining users and their travel information data, so as to effectively clean the travel information data of non-travel enthusiasts and improve the quality of the subsequent constructed knowledge graph. Among them, the value range of the preset preference threshold is , in an embodiment of the present invention, the preset preference threshold is set to 0.6. The specific value of the preset preference threshold can also be set by the implementer according to the specific implementation scenario, and is not limited herein.

[0129] Perform negative correlation normalization processing on the travel preference level of each retained user, and limit the calculation result within range, so as to obtain the distance coefficient of each retained user. The smaller the distance coefficient of a certain retained user, the greater the user's preference for travel. Then, when constructing the knowledge graph subsequently, the distance between the node corresponding to the user and the node of the travel enthusiast group is closer. Therefore, during the process of travel information recommendation, the travel information of this user can be preferentially recommended to specific users.

[0130] As an example, in an embodiment of the present invention, the expression of the distance coefficient of each retained user can be specifically, for example:

[0131]

[0132] Among them, represents the distance coefficient of the th retained user; represents the travel preference level of the th retained user; represents the normalization function.

[0133] Then, based on all retained users and their travel information data, construct a knowledge graph. Among them, each retained user serves as an entity of the knowledge graph. The Neo4j graph database is used in the construction process of the knowledge graph. The knowledge graph includes a node of the travel enthusiast group and nodes corresponding to each retained user. And the distance between the node corresponding to each retained user and the node of the travel enthusiast group is equal to the distance coefficient of each retained user. It should be noted that the construction process of the knowledge graph is a well-known technical means in the art and will not be elaborated here. Please refer to Figure 3 , which shows the structural schematic diagram of the constructed knowledge graph provided by an embodiment of the present invention. Among them, Q represents the node of the travel enthusiast group in the knowledge graph, and P1~P5 represent the nodes corresponding to each retained user.

[0134] An embodiment of the present invention provides an Internet open-source information data cleaning system, which includes a memory, a processor, and a computer program. The memory is used to store the corresponding computer program, the processor is used to run the corresponding computer program, and when the computer program runs in the processor, it can implement the methods described in steps S1 to S5.

[0135] It should be noted that: the above sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0136] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for cleaning Internet open source information data, characterized in that: The method comprises: Obtain the travel start time and travel end time of each travel route selected by different users within a preset time period, as well as the driving speed data of the user at each moment on each travel route; Obtaining the travel time concentration of each user according to the distribution of the travel start time of each travel route of each user; Taking any user as a target user, obtaining the travel possibility of each travel route of the target user according to the difference in the travel start time and the difference in the travel end time of the same travel route between the target user and other users except the target user; based on the travel possibility, screening multiple travel routes of the target user from all travel routes of the target user; Taking any travel route of the target user as the target travel route of the target user, obtaining the popularity index of the target travel route of the target user according to the travel time concentration and the number of users who select the target travel route; obtaining the vehicle speed similarity of the target travel route of the target user according to the difference in driving speed data at each moment between the target user and other users on the target travel route; obtaining the travel preference degree of the target user according to the vehicle speed similarity and the popularity index of all travel routes of the target user; Based on each user's travel preferences, each user's travel information data is cleaned and a knowledge graph is constructed.

2. The method for cleaning Internet open source information data according to claim 1, characterized in that: The obtaining of the travel time concentration of each user comprises: Taking any user as the user to be tested, and dividing all the travel routes of the user to be tested into four categories based on the travel start time of each travel route of the user to be tested, the categories being spring category, summer category, autumn category and winter category; The number of all travel routes of the user to be tested in each category is used as the numerator, the number of all travel routes of the user to be tested is used as the denominator, and the ratio is used as the proportion of the number of travels of each category of the user to be tested; The discrete degree of the travel quantity proportions of all categories of the users to be tested is analyzed to obtain the travel time concentration of the users to be tested.

3. The method for cleaning Internet open source information data according to claim 1, characterized in that: The obtaining of the travel possibility of each travel route of the target user includes: Taking any travel route of the target user as the target travel route of the target user, selecting multiple first reference users of the target user with respect to the target travel route from all other users except the target user, wherein the travel route of each of the first reference users contains the target travel route of the target user; The absolute value of the difference between the travel start time of the target travel route between the target user and each first reference user is used as the travel start time difference between the target user and each first reference user with respect to the target travel route, and the absolute value of the difference between the travel end time of the target travel route between the target user and each first reference user is used as the travel end time difference between the target user and each first reference user with respect to the target travel route; The sum of the travel start time difference and the travel end time difference is used as the comprehensive time difference between the target user and each first reference user with respect to the target travel route; A negative correlation normalization process is performed on the average value of the comprehensive time differences between the target user and all first reference users regarding the target travel route to obtain the travel possibility of the target user's target travel route.

4. The method for cleaning Internet open source information data according to claim 1, characterized in that: The step of selecting multiple travel routes of the target user from all travel routes of the target user includes: Among all the travel routes of the target user, the travel route whose travel possibility is greater than a preset possibility threshold is used as the travel route of the target user.

5. The method for cleaning Internet open source information data according to claim 1, characterized in that: The hot index of the target travel route of the target user includes: Selecting multiple audience users of a target travel route from all users, wherein the travel route of each of the audience users includes the target travel route of the target user; The accumulated value of the travel time concentration of all audience users of the target travel route is used as the popularity evaluation value of the target travel route; The number of all audience users of the target travel route and the popularity evaluation value of the target travel route are integrated to obtain the popularity index of the target travel route of the target user.

6. The method for cleaning Internet open source information data according to claim 1, characterized in that: The obtaining of the vehicle speed similarity of the target travel route of the target user comprises: Obtain traffic flow data for each user at each time on each travel route; According to the correlation between the traffic flow data and the driving speed data of each user at each moment on each travel route, the driving speed data of each user at each moment on each travel route is adjusted to obtain the adjusted driving speed of each user at each moment on each travel route; Sorting the adjusted driving speeds of each user at each time on each travel route in chronological order to obtain a sequence of adjusted driving speeds of each user on each travel route; Selecting a plurality of second reference users of the target user with respect to the target travel route from all other users except the target user, wherein the travel route of each of the second reference users includes the target travel route of the target user; Using a dynamic time warping algorithm, the adjusted driving speed sequences of the target user and each second reference user on the target travel route are processed to obtain a speed difference between the target user and each second reference user on the target travel route; A negative correlation mapping is performed on the average value of the vehicle speed differences between the target user and all second reference users with respect to the target travel route to obtain the vehicle speed similarity of the target travel route of the target user.

7. The method for cleaning Internet open source information data according to claim 6, characterized in that: The step of obtaining the adjusted driving speed of each user at each time on each travel route comprises: Sorting the driving speed data of the target user at each moment on the target travel route in chronological order to obtain a driving speed data sequence of the target user on the target travel route, and sorting the traffic flow data of the target user at each moment on the target travel route to obtain a traffic flow data sequence of the target user on the target travel route; Performing negative correlation normalization processing on the Pearson correlation coefficient between the driving speed data sequence and the traffic flow data sequence to obtain a speed adjustment weight of the target user on the target travel route; The product value of the speed adjustment weight of the target user on the target travel route and the driving speed data of the target user at each moment on the target travel route is used as the speed adjustment amount of the target user at each moment on the target travel route; The sum of the target user's driving speed data at each moment on the target travel route and the speed adjustment amount is used as the adjusted driving speed of the target user at each moment on the target travel route.

8. The method for cleaning Internet open source information data according to claim 1, characterized in that: The obtaining of the target user's travel preference degree includes: The popularity index and the vehicle speed similarity of each travel route of the target user are combined to obtain a travel status coefficient of each travel route of the target user; The average values ​​of the travel status coefficients of all travel routes of the target user are normalized to obtain the travel preference degree of the target user.

9. The method for cleaning Internet open source information data according to claim 1, characterized in that: Based on each user's travel preference, the travel information data of each user is cleaned and a knowledge graph is constructed, including: Remove users whose travel preference is less than a preset preference threshold and their travel information data, and retain the remaining users and their travel information data; Performing negative correlation normalization processing on the travel preference degree of each retained user to obtain a distance coefficient of each retained user; A knowledge graph is constructed based on all retained users and their travel information data, wherein each retained user serves as an entity of the knowledge graph, and the knowledge graph contains a travel enthusiast group node and a node corresponding to each retained user, and the distance between the node corresponding to each retained user and the travel enthusiast group node is equal to the distance coefficient of each retained user.

10. An Internet open source information data cleaning system, the system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.