A Method for Visualizing Traffic Control Information
Through the method of matching regular expressions and participle participle, multi-dimensional geochemical display of road control information is combined with geographical data, which solves the problem of unsatisfactory information accuracy in the existing technology, and realizes refined road control information display and travel decision support.
Patent Information
- Application Number
- CN202310093834.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-02
AI Technical Summary
In the prior art, the way of publishing road control information cannot fully reflect the control situation of urban roads, and the information published in text is difficult to associate with the map, making it difficult for travelers to accurately and quickly determine the specific geographical location of the control event, affecting travel decisions.
The method of regular expression matching and part-of-word part-time matching is used to extract and structure road control information, and multi-dimensional geochemical display is performed in combination with open source geographical data, including refined display at road level, location level, regional level and lane level.
It improves the accuracy and flexibility of identification of road control information, realizes the fine and reasonable division of road control information and multi-dimensional geochemical display, and enhances the accuracy and convenience of travelers to obtain information.
Smart Images

Figure CN116150255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual analysis, and in particular to a method for visualizing traffic control information. Background Art
[0002] Fully obtaining various traffic events occurring on the road and obtaining control information affecting traffic travel has a very wide impact on fields such as public travel, driverless driving, and intelligent transportation. In the prior art, road control information is mainly released by government platforms (such as transportation authorities) for road maintenance plans. However, on the one hand, road emergencies are accidental, sudden, and dissipative, and the road control information released only by the government platform cannot fully reflect the control situation of urban roads; on the other hand, the road control information of the government platform is generally released in text form, with problems such as complex formats and lack of association with maps. It is difficult for travelers to obtain all relevant information at the same time and accurately and quickly determine the specific geographical location where the control event occurs, so as to make decisions, resulting in the underutilization of information.
[0003] Crowdsourced road control information is collected on network platforms with the help of group wisdom, and has advantages such as strong timeliness, wide sources, and rich content. It can effectively guide traffic travel and has important data value. Therefore, it is necessary to extract the fine road control information contained in the text semantics, fuse it with spatial data and perform geographical display, so as to facilitate the route planning and decision-making of travelers, and then improve the traffic conditions.
[0004] In current research on traffic information extraction and geographicalization based on text data, usually various word segmentation algorithms are used to extract geographical locations contained in natural language texts, and are mapped to specific coordinate points through road network matching for visual display. These methods ignore the standardization of geographical location semantic expression and the multidimensionality of spatial expression, and the accuracy of extraction and positioning needs to be further improved.
[0005] In summary, there is currently a lack of a method for visualizing traffic control information to solve the problem of unsatisfactory accuracy of the traditional method of using various word segmentation algorithms to extract geographical locations contained in natural language texts. Summary of the Invention
[0006] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for visualizing traffic control information, which realizes the content parsing of complex texts based on regular expression matching and part-of-speech matching of word segmentation, effectively improves the accuracy of information extraction and positioning, and solves or partially solves the problem of unsatisfactory accuracy of the traditional method of using various word segmentation algorithms to extract geographical locations contained in natural language texts.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] The present invention provides a method for visualizing traffic control information, comprising the following steps:
[0009] Obtain the original road control data and open-source geographic data and perform preprocessing on them respectively to obtain control text information and road vector layer information;
[0010] For the control text information, through information extraction and structured processing, obtain refined control information. Based on multiple refined control information from different sources, obtain standardized control records through fusion processing;
[0011] Based on the standardized control records, construct a control information knowledge graph. Based on the control information knowledge graph and the road vector layer information, realize multi-dimensional geographical display of control information,
[0012] Among them, the information extraction includes the following steps:
[0013] For information with a standard format, match it with the text content through a pre-set regular expression to achieve information extraction;
[0014] For information containing a preset keyword, achieve information extraction through part-of-speech matching;
[0015] For information that cannot be directly obtained from the text, perform attribute matching query through the unified location names in the control text information and the open-source geographic data to achieve information extraction.
[0016] As a preferred technical solution, the preprocessing includes the following steps:
[0017] For the urban road network dataset and point-of-interest dataset in the open-source geographic data, based on the urban road network dataset, obtain road line layer information, section line layer information, and intersection point layer information. Based on the point-of-interest dataset, obtain point-of-interest layer information. Based on the road line layer information, section line layer information, intersection point layer information, and point-of-interest layer information, obtain the road vector layer information;
[0018] Based on the open-source geographic data, obtain a list of geographical location names including road names and POI names, and update the preset word segmentation dictionary;
[0019] For the original road control data, obtain the control text information based on the word segmentation dictionary.
[0020] As a preferred technical solution, obtaining the control text information based on the word segmentation dictionary includes the following steps:
[0021] Perform word segmentation using the said word segmentation dictionary, locate the text containing specific geographical location information through part-of-speech matching, classify and filter out the high-frequency irrelevant words in the text through word frequency analysis, and delete the redundant information in the text content through data cleaning to obtain the said controlled text information.
[0022] As a preferred technical solution, the said structured processing includes the following steps:
[0023] Convert the said controlled standardized record into a preset information description structure including controlled location information, controlled time information, and controlled type information.
[0024] As a preferred technical solution, the said controlled location information includes at least one of location name, starting section, ending section, location form, lane information, display level, and specific longitude and latitude. The said controlled time information includes starting time and ending time.
[0025] As a preferred technical solution, the acquisition of the said controlled information knowledge graph includes the following steps:
[0026] Input the said controlled standardized record into a preset graph database, and obtain the said controlled information knowledge graph by creating nodes and relationships.
[0027] As a preferred technical solution, the said graph database is a Neo4j graph database.
[0028] As a preferred technical solution, the implementation of the said multi-dimensional geographical display includes the following steps:
[0029] For the road control information in different location forms, determine the display level, and use the controlled location information extracted from the control data to retrieve the specific longitude and latitude information in the corresponding layer attribute table of the said vector layer information to achieve positioning and multi-dimensional geographical display, where the said display level includes road level, location level, area level, and lane level.
[0030] As a preferred technical solution, the said part-of-speech matching is specifically:
[0031] Based on a preset word segmentation dictionary, implement information extraction through this line matching.
[0032] As a preferred technical solution, obtain the said original road control data and / or the said open-source geographical data through web crawlers for open-source web pages and / or radio stations and / or map platforms.
[0033] Compared with the prior art, the present invention has the following advantages:
[0034] (1) The content parsing of complex texts is realized based on regular expression matching and part-of-speech matching of word segmentation. It is not limited to the recognition of specific location names, but fully utilizes the normativity of geographical location semantic expressions, incorporates more location forms into consideration, and divides the types of road control locations more precisely and reasonably, greatly improving the accuracy and flexibility of information recognition, and solving or partially solving the problem of unsatisfactory accuracy when traditional methods extract geographical locations contained in natural language texts.
[0035] (2) The further parsing of complex texts is realized based on part-of-speech matching of word segmentation. By fully utilizing the similarity of geographical location semantic expressions and cooperating with regular expression matching, higher recognition accuracy and flexibility can be achieved.
[0036] (3) The method provided by the present invention defines four display levels, namely road level, location level, region level, and lane level, based on different location forms. Combining the extracted road control locations, specific longitude and latitude information is further obtained through feature attribute matching in the corresponding geographical data, and the road control information is geographically displayed with different levels of fineness according to the text richness. Description of the Drawings
[0037] Figure 1 It is a schematic flow chart of the traffic control information visualization method in Embodiment 1;
[0038] Figure 2 It is a schematic diagram of data screening and cleaning processing;
[0039] Figure 3 It is a schematic diagram of extracting road fine control information;
[0040] Figure 4 It is a schematic diagram of constructing a knowledge graph of road control information;
[0041] Figure 5 It is a schematic diagram of geographical visualization of road control information;
[0042] Figure 6 It is a schematic diagram of two extraction methods, namely based on a standard format and based on specific keywords, in the extraction of road fine control information;
[0043] Figure 7 It is a schematic diagram of the geographical method of road control information. Detailed Embodiment
[0044] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0045] Embodiment 1
[0046] As Figure 1 described, this embodiment provides a method for visualizing traffic control information. Based on web crawlers, this method first obtains crowdsourced urban road control text data and related open-source geographic data, and screens and cleans the two types of data; secondly, extracts fine-grained road control information from the text corpus, performs structured processing and data fusion; finally, constructs a knowledge graph, and realizes the geographicalization of road control information based on the longitude and latitude information extracted by matching the attributes of geographical elements. The specific steps are as follows:
[0047] Step S101, use web crawler technology to obtain crowdsourced text data and open-source geographic data.
[0048] Use a web crawler program to obtain in real time the road control information released open source, including but not limited to government official microblogs, traffic-related websites, radio stations, etc.; at the same time, obtain the urban road network dataset and the POI interest point dataset from open-source map websites and map open platforms respectively.
[0049] Step S102, respectively perform screening and cleaning work on the text data and geographic data, and extract the text records related to road control and the vector layers related to roads.
[0050] As Figure 2 described, first, perform preprocessing on the open-source geographic data, screen the road network line data, retain roads of types such as "urban arterial roads", "urban secondary arterial roads", "urban expressways", etc., and perform element fusion processing according to the road names to form a road line layer, and then realize road network segmentation and intersection extraction through operations such as "breaking intersecting lines" and "converting element fold points to points" to form a section line layer and an intersection point layer respectively; screen the POI point data, retain POI elements of types such as "road ancillary facilities", "place name and address information", "traffic facility services", etc., to form a POI point layer.
[0051] Secondly, integrate the screened geographic element attribute tables, extract the road names and POI names to form a list of urban geographical location names, import a custom word segmentation dictionary, and assign part-of-speech according to the location type.
[0052] Then, preprocess the text data. Traverse each text record and perform word segmentation on it using a custom dictionary; through part-of-speech matching, accurately locate the text containing specific geographical location information; through word frequency analysis, classify and uniformly screen out the irrelevant words that appear frequently in the text.
[0053] Finally, clean the text data related to road control after screening, and delete the redundant information in the text content, such as "[[Subtitle]]", "@Username", "#Topic Name#", etc.
[0054] Step S103, extract the fine road control information from the text data based on the specification format matching, specific keyword matching, and geographical element attribute matching methods, and perform structured and fusion processing.
[0055] As Figure 3 described, first, construct a control information description structure including control location, control time, control type, and current status; analyze the text content and define the types of road control-related information that can be extracted, mainly including control location (location name, starting section, ending section, location form, lane information, display level, specific longitude and latitude), control time (starting time, ending time), control type, and current situation, etc.
[0056] Secondly, analyze the semantic representation forms of different control information types and propose corresponding information extraction methods:
[0057] (1) Information with a standard format, such as the section form and road intersection / entrance and exit form in the control time and control location, is matched with the text content by defining regular expressions to achieve information extraction.
[0058] (2) Information containing specific keywords, such as control type, current situation, lane information, single road form and POI point / area form in the control location, is extracted through part-of-speech matching after text segmentation.
[0059] (3) Information that cannot be directly obtained from the text, such as the specific longitude and latitude of the control location, is extracted by performing attribute matching queries on the unified location name in the text data and geographical data.
[0060] Finally, perform structured and fusion processing on the road control information extracted from multi-source unstructured text data simultaneously.
[0061] Figure 6 For the specific explanations of the two extraction methods based on the specification format and specific keywords, taking the extraction of control location information as an example, the specific steps are as follows:
[0062] (1) Analyze the representation forms of road control positions in the text. It can be seen that road segments and intersections / entrances / exits are two position forms with standard formats, and regular expressions are defined for matching.
[0063] (2) If no content is matched, it indicates that the text contains other non-standard position types. Import a custom dictionary for text word segmentation, and identify geographical locations in the text through part-of-speech matching.
[0064] (3) Extract geographical locations in the text through the above two methods, classify them into corresponding position types, and identify and determine other control information based on the position types to achieve refined extraction of road control information.
[0065] Step S104: Traverse the structured standardized records of road control, and construct a knowledge graph of road control information.
[0066] As Figure 4 described, after structured processing, different road control information has different numbers of fields, and some record fields will have null values. For example, in the records of POI point types, the "starting road segment" and "ending road segment" fields are null, in the records of single time types, the "ending time" field is null, and in the records without lane descriptions, the "lane information" field is null.
[0067] Use the graph database Neo4j to store urban road control information; connect to the Neo4j graph database, create nodes and relationships respectively, and obtain the knowledge graph of each road control information.
[0068] Step S105: Based on the specific longitude and latitude information of the control positions obtained by matching geographical element attributes, perform multi-dimensional geographical display for different road control information.
[0069] As Figure 5 described, based on road control information in different position forms, starting from the three dimensions of point, line, and surface, define four display levels of road level, location level, area level, and lane level respectively. Use the road control positions extracted from the text data to retrieve specific longitude and latitude information in the corresponding geographical layer attribute table to achieve precise positioning and multi-dimensional geographical display.
[0070] Figure 7 The above is a schematic diagram of the geographical method for road control information. The specific steps are as follows:
[0071] (1) Display level division: First, divide road control information into three categories of road level, location level, and area level according to five position forms. If there is lane-related description information, it is classified into the lane level type with a higher level of detail.
[0072] (2) Geographic display: Match the information of different location types to the corresponding geographic feature layers, and perform a matching query through the name of the controlled location in the text information and the corresponding attributes of the geographic data to obtain the specific longitude and latitude information of the controlled location, and perform geographic display in three dimensions of points, lines, and surfaces.
[0073] Compared with the traditional method, this method realizes the content parsing of complex texts based on regular expression matching and part-of-speech matching of word segmentation. It is not limited to the recognition of specific place names, but fully utilizes the normativity and similarity of the semantic expression of geographical locations, incorporates more location forms into consideration, and divides the types of road control locations more finely and reasonably, greatly improving the accuracy and flexibility of information recognition. Based on different location forms, this method defines four display levels. Combining the extracted road control locations, further obtain the specific longitude and latitude information through feature attribute matching in the corresponding geographic data, and perform geographic display of road control information with different levels of fineness according to the text richness.
[0074] Embodiment 2
[0075] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the traffic control information visualization method as described in Embodiment 1.
[0076] Embodiment 3
[0077] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the traffic control information visualization method as described in Embodiment 1.
[0078] As mentioned above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for visualizing traffic control information, characterized in that, It includes the following steps: Obtain the original road control data and open-source geographic data and perform preprocessing on them respectively to obtain control text information and road vector layer information; For the control text information, through information extraction and structured processing, obtain refined control information. Based on multiple refined control information from different sources, obtain standardized control records through fusion processing; Based on the standardized control records, construct a control information knowledge graph. Based on the control information knowledge graph and the road vector layer information, realize multi-dimensional geographical display of control information, Among them, the information extraction includes the following steps: For information with a standard format, match it with the text content through a pre-set regular expression to realize information extraction; For information containing a preset keyword, realize information extraction through part-of-speech matching; For information that cannot be directly obtained in the text, realize information extraction through attribute matching query of the control text information and the unified location name in the open-source geographic data, The preprocessing includes the following steps: For the urban road network dataset and point-of-interest dataset in the open-source geographic data, based on the urban road network dataset, obtain road line layer information, section line layer information, and intersection point layer information. Based on the point-of-interest dataset, obtain point-of-interest layer information. Based on the road line layer information, section line layer information, intersection point layer information, and point-of-interest layer information, obtain the road vector layer information; Based on the open-source geographic data, obtain a list of geographical location names including road names and POI names, and update the pre-set word segmentation dictionary; For the original road control data, obtain the control text information based on the word segmentation dictionary, Obtaining the control text information based on the word segmentation dictionary includes the following steps: Use the word segmentation dictionary for word segmentation, locate the text containing specific geographical location information through part-of-speech matching, classify and screen out irrelevant words with high text frequency through word frequency analysis, and delete redundant information in the text content through data cleaning to obtain the control text information.
2. The visualization method of traffic control information according to claim 1, wherein The structured processing includes the following steps: Convert the standardized control records into a pre-set information description structure including control location information, control time information, and control type information.
3. A traffic control information visualization method according to claim 2, characterized in that, The control location information includes at least one of location name, starting section, ending section, location form, lane information, display level, and specific longitude and latitude. The control time information includes starting time and ending time.
4. A traffic control information visualization method according to claim 1, characterized in that, The acquisition of the control information knowledge graph includes the following steps: Input the standardized control records into a pre-set graph database, and obtain the control information knowledge graph by creating nodes and relationships.
5. A traffic control information visualization method according to claim 4, characterized in that, The graph database is a Neo4j graph database.
6. A traffic control information visualization method according to claim 1, characterized in that The realization of the multi-dimensional geographical display includes the following steps: For road control information in different location forms, determine the display level, and use the control location information extracted from the control data to retrieve specific longitude and latitude information in the corresponding layer attribute table of the vector layer information, so as to achieve positioning and multi-dimensional geographical display, where the display level includes road level, location level, area level and lane level.
7. A traffic control information visualization method according to claim 1, characterized in that, The specific part-of-speech matching is as follows: Based on a preset word segmentation dictionary, information extraction is achieved through this line matching.
8. A traffic control information visualization method according to claim 1, characterized in that, The original road control data and / or the open-source geographical data are obtained through web crawlers from open-source web pages and / or radio stations and / or map platforms.
Citation Information
Patent Citations
Visual analysis method combining real-time traffic road condition text data with traffic volume
CN111552772A
Method for improving accuracy of positioning traffic accident site
WO2020192123A1