An AI vision-based landscape sign design method, device and medium
By using AI vision technology to quantitatively evaluate the legibility of landscape signs, the problem of quantifying the impact of occlusion and glare in landscape sign design has been solved, achieving closed-loop consistency from design to on-site installation and reducing review costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG COLLEGE OF CONSTR
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing landscape signage design methods are difficult to quantify and assess legibility in actual traffic flow, and the lack of a unified spatial verification loop for location and orientation from design to on-site installation makes it difficult to ensure consistency and results in high costs.
AI vision technology is used to collect first-person view video, restore the camera pose to obtain unified coordinates, identify key guidance road sections and extract a list of candidate installation surfaces and obstructions, calculate the readable level through sign detection and text direction recognition, generate sign configuration and conduct on-site verification, locate the causes of obstruction and glare and make directional corrections.
It enables quantitative characterization of the effects of occlusion, glare, and imaging blur under continuous travel field of view, supports candidate sign ranking, ensures closed-loop consistency from design to on-site installation and verification, and forms a sign configuration and verification record that can be directly implemented.
Smart Images

Figure CN122116095A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision image processing technology, and in particular to a landscape signage design method, device and medium based on AI vision. Background Technology
[0002] Landscape signage design is an important component of environmental guidance information expression and spatial cognition support. It is widely used in scenic spots, parks, and public open spaces for wayfinding organization, node prompts, and unified visual identification. Existing technologies typically rely on planning drawings and access routes, combined with wayfinding specifications, information hierarchy, and visual identification systems, to complete the configuration of signage types, location layout, and layout design. Computer-aided design, 3D visualization, and photorealistic rendering are used to present the solutions intuitively and conduct cross-disciplinary collaborative verification. These methods are highly mature and adaptable, and can achieve consistent expression and rapid iteration of wayfinding information while meeting landscape and management requirements. This facilitates the formation of feasible signage solutions and construction documents under different scene scales and spatial structures.
[0003] Existing technologies still have the problem of not being consistent with the actual travel experience. The verification of the solution relies heavily on static views or limited perspectives, making it difficult to quantify the combined impact of occlusion, glare and image blur on the reading results in a continuous field of view. When converting the design points and orientations to on-site installation, there is a lack of unified spatial expression and verification loop, resulting in high costs for point reproduction and orientation correction and difficulty in ensuring consistency. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a landscape sign design method based on AI vision to solve the problems that the legibility of landscape signs is difficult to quantify and evaluate in actual traffic flow and that the location and orientation lack a unified spatial verification loop from design to on-site installation.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a landscape sign design method based on AI vision, which includes: acquiring first-view video and recording the acquisition time period and weather; restoring the camera pose to obtain unified coordinates; determining key guidance road sections and extracting a list of candidate installation surfaces and obstructions; and obtaining candidate signs through the list of candidate installation surfaces.
[0008] Based on key guidance sections, a sequence of real background frames is obtained and occluders are segmented. Counterfactual sample frames are generated by perspective coverage of candidate sign projection areas. Sign detection and text direction recognition are performed, and the readable level is calculated.
[0009] Based on the readability level, the worst critical guidance road segment is determined, and the final location, final orientation, and unified format element template are determined to obtain the sign configuration and output the sign list.
[0010] Perform trial installations according to the label list and collect first-person perspective video again. Conduct on-site verification based on the readability level, locate the causes of obstruction and glare, and make directional corrections, forming a verification record.
[0011] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0012] The process of acquiring first-person perspective video and recording the acquisition time and weather, and restoring camera pose to obtain unified coordinates includes: acquiring first-person perspective video along the main passageway of the landscape area and recording the acquisition time and weather; recording timestamps frame by frame and simultaneously acquiring camera pose, camera intrinsic parameters, and distance map for each frame; aligning the timestamps, camera pose, camera intrinsic parameters, and distance map with the video frames to form a frame sequence record; performing smooth interpolation to complete frames with missing camera poses based on adjacent valid frames; obtaining unified coordinates by using the camera position of the first frame as the unified coordinate origin and the camera orientation of the first frame as the unified coordinate reference direction; and writing the observation position and observation direction of each frame.
[0013] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0014] The specific steps for determining key guidance road segments and extracting candidate mounting surfaces and obstructions, and obtaining candidate identifiers from the candidate mounting surface list, are as follows: Under a unified coordinate system, observe directional changes based on adjacent frames to obtain an orientation angle sequence; determine the turning center frame and determine the candidate road segment boundary based on the angle change to form a candidate road segment set; perform walkable region segmentation and connected component analysis on the turning center frame of the candidate road segments; confirm key guidance road segments and write them into the key guidance road segment table; perform semantic segmentation on the turning center frame of the key guidance road segments to extract candidate mounting surfaces and obstructions; extract boundary contours from the candidate mounting surfaces and combine them with the camera's intra-camera participation distance map to map pixels within the boundary contours into a spatial point set; perform plane fitting based on the spatial point set to obtain the spatial plane of the candidate mounting surfaces; map the boundary contours onto the spatial plane to form an installable boundary and write it into the candidate mounting surface list.
[0015] Within the installable boundary, shrink inward and generate candidate points by equidistant grid points. Generate candidate orientations for the candidate points and bind them to layout element templates. Determine the combination of candidate points, candidate orientations, and layout element templates as candidate identifiers and write them into the candidate identifier list.
[0016] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0017] The steps for obtaining the real background frame sequence and segmenting the occluders, and generating counterfactual sample frames by perspective coverage of the candidate identifier projection area are as follows: read the start and end frame range and the turning center frame index from the key guidance road segment table, take all video frames within the start and end frame range as the real background frame sequence, and perform semantic segmentation on each frame to obtain the occluder mask.
[0018] On a sequence of real background frames, mark detection and text direction recognition are performed, and consistent recognition frames are selected. The mark regions detected by the consistent recognition frames are cropped to form a reference sample set. The median of the contrast quantity, viewpoint scale quantity, and local sharpness quantity are taken as reference statistics. The candidate points, candidate orientations, and layout element templates are rendered into two-dimensional mark images. The mark boundaries are determined according to the outer boundary of the layout element template rendering result. The mark boundaries are projected onto the real background frames according to the camera pose and camera intrinsic parameters to obtain the candidate mark projection areas. After perspective transformation of the two-dimensional mark images, counterfactual sample frames are generated and the candidate mark projection area mask is saved.
[0019] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0020] The specific steps for performing identifier detection and text direction recognition, and calculating the readable level, are as follows: perform identifier detection and text direction recognition within the candidate identifier projection area of the counterfactual sample frame and write the recognition record; then, filter and mark consistent and inconsistent recognition frames.
[0021] For each counterfactual sample frame, the occlusion ratio is calculated based on the contrast ratio, viewpoint scale, local sharpness measurement, candidate sign projection area mask, occlusion mask, distance map, and reference statistics. Combined with the consistent readability frame marker, the readability level of a single frame is obtained. The distance difference between the observation positions of adjacent frames under the same coordinate system is used as the travel distance increment. The readability levels of each frame in the same key guidance road segment are merged according to the travel distance increment to obtain the readability level of the candidate sign in the key guidance road segment.
[0022] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0023] The specific steps for determining the worst critical guidance road segment based on the readability level are as follows: for each critical guidance road segment, the candidate identifiers are sorted from high to low according to their readability level in the critical guidance road segment to generate a candidate identifier sorting list.
[0024] For each critical guidance road segment, the highest readable level in the candidate identifier ranking list is taken as the optimal readable level of the critical guidance road segment, and the critical guidance road segment with the smallest value among all the optimal readable levels of critical guidance road segments is selected as the worst critical guidance road segment.
[0025] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0026] The process involves determining the final location, final orientation, and unified format element template to obtain the identifier configuration and output the identifier list. Specifically, the process involves selecting the candidate identifier with the highest readability level from the candidate identifier ranking list of the worst critical guidance road segment as the identifier configuration for the critical guidance road segment, and determining the candidate location, candidate orientation, and format element template bound to the candidate identifier as the final location, final orientation, and unified format element template. When there are ties, the unique identifier configuration is determined sequentially according to the longest continuous consistent read segment length, the total number of consistent read frames, and the order in which the candidate identifier list is written.
[0027] For each key guidance road segment except the worst key guidance road segment, candidate points and candidate orientations are combined one by one using a unified format element template to generate candidate identifiers. Counterfactual sample frames are generated on the real background frame sequence to calculate the readability level. The combination with the highest readability level is selected as the final point and final orientation of the key guidance road segment. When there are ties, the unique configuration is determined sequentially according to the longest continuous consistent readability segment length, the total number of consistent readability frames, and the order in which the candidate identifier list is written. The unified format element template is used as the final format element template for each key guidance road segment. The final point, final orientation, and final format element template of each key guidance road segment are written into the identifier list.
[0028] As a preferred embodiment of the AI vision-based landscape signage design method described in this invention, wherein:
[0029] The process involves trial installation according to the identification list and re-capturing first-view video. On-site verification is performed based on the readability level to locate and correct the causes of obstruction and glare, forming a verification record. Specifically, the steps are as follows: Mark the installable boundary on the candidate installation surface; project the final point onto the spatial plane of the candidate installation surface to obtain the landing point; reproduce the landing point on-site; align the landing point with the geometric center of the physical identifier's outer boundary and align the final orientation with the directional arrow; re-shoot along the main traffic route to form a frame sequence record; generate an obstruction mask and an identifier area mask, and perform identifier detection and text direction recognition; calculate the readability level of a single frame and aggregate it according to the travel distance increment, writing the re-shoot readability level into the verification record; locate the obstruction source area and glare source area; sequentially change the candidate orientation, replace the nearest candidate point, switch the candidate installation surface list, and re-shoot and verify; and write the results of each re-shoot and verification into the verification record.
[0030] In a second aspect, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the AI vision-based landscape signage design method as described in the first aspect of the present invention.
[0031] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the AI vision-based landscape signage design method as described in the first aspect of the present invention.
[0032] The beneficial effects of this invention are as follows: By generating counterfactual sample frames by perspective overlaying candidate signs according to the projection area of the candidate signs on the real background frame sequence, and combining the occlusion mask, the candidate sign projection area mask, the distance map and the reference statistics to calculate the readable level, a quantitative characterization of the effects of occlusion, glare and imaging blur under continuous travel field of view is achieved, which is used to support the sorting of candidate signs. By determining the worst key guidance road segment and based on this, the final location, final orientation and unified format element template are determined to output the sign list, a closed-loop consistency is achieved from design to on-site installation and re-examination and verification, forming a configuration and verification record that can be directly implemented. Attached Figure Description
[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a flowchart of a landscape signage design method based on AI vision.
[0035] Figure 2 A flowchart for extracting key guidance road sections and candidate installation surfaces.
[0036] Figure 3 Flowchart for generating counterfactual sample frames and calculating readable levels.
[0037] Figure 4 A flowchart illustrating the causes of location obstruction and glare during on-site verification.
[0038] Figure 5 This is a graph showing the changes in readable level and continuous consistent readability stability under repeated scanning and verification iterations.
[0039] Figure 6 This is an optimal readable horizontal heat map of key guidance road sections and unified layout element templates. Detailed Implementation
[0040] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0041] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0042] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0043] Reference Figures 1-6 This is one embodiment of the present invention, which provides a landscape signage design method based on AI vision, including the following steps:
[0044] S1. Collect first-person perspective video and record the collection time and weather, restore the camera pose to obtain unified coordinates, determine key guidance road sections and extract a list of candidate installation surfaces and obstructions, and obtain candidate identifiers through the list of candidate installation surfaces.
[0045] First-person perspective video was collected along the main access route of the scenic area. During the collection process, the camera height and line of sight were kept natural and stable. The video was recorded continuously. The collection time period and weather information were recorded at the beginning and end of the collection. The time period, weather information and video file were bound to the same collection record. During the collection process, the timestamp was recorded frame by frame, and the camera pose, camera intrinsic parameters and distance map aligned with the video frame were recorded simultaneously for each frame.
[0046] A distance map refers to the depth or distance measurement results output with each frame. The depth or distance measurement results are stored as a distance map with the same resolution as the video frame, indexed by pixel position.
[0047] The timestamp, camera pose, camera intrinsic parameters, and distance map are aligned with the video frames to form a frame sequence record. If individual frames are missing camera poses, the missing frames are smoothly interpolated using the camera poses of adjacent valid frames, while maintaining the time order. The camera position of the first frame is used as the unified coordinate origin, and the camera orientation of the first frame is used as the unified coordinate reference direction to obtain unified coordinates. The observation position, observation direction, and frame index of each frame are written together into the same record to obtain the observation position and observation direction corresponding to each frame.
[0048] Smooth interpolation uses linear interpolation of camera position and continuous interpolation of camera orientation on the time axis to give missing frames a camera pose that changes continuously with adjacent frames.
[0049] Based on the change in observation direction between adjacent frames under a unified coordinate system, the angle between the orientations of adjacent frames is calculated and formed into a sequence that changes with the frame. Frames with an angle greater than both the previous and next frames are defined as turning center frames. The starting frame is determined by tracing back from the turning center frame to where the angle no longer increases continuously, and the ending frame is determined by tracing back to where the angle no longer decreases continuously. The candidate road segment boundary corresponding to the turning center frame is determined, resulting in a set of candidate road segments. The walkable area is segmented for the turning center frame of each candidate road segment. Connectivity analysis is performed on the walkable area within the forward field of view of the turning center frame. The number of connected components in the walkable area within the forward field of view is used as the criterion. When the number of connected components is greater than one, the candidate road segment is confirmed as a key guidance road segment. The start and end frame ranges of the key guidance road segment and the index of the turning center frame are written into the key guidance road segment table.
[0050] Walkable region segmentation is achieved through pixel-level semantic segmentation. The walkable region pixel set is output for the turning center frame, and morphological denoising and hole filling are performed on the walkable region pixel set to obtain a binary mask, which is the walkable region mask. It indicates whether the region in the turning center frame image is a walkable region or an inwalkable region.
[0051] Semantic segmentation is performed on the turning center frame of each key guidance road segment to extract candidate mounting surfaces and occlusions. Candidate mounting surfaces include walls, pillars, railings, and ground mounting areas, while occlusions include tree canopies, pergolas, and billboards. Boundary contours are extracted from the candidate mounting surfaces. Within the pixel area covered by the boundary contours of the candidate mounting surfaces, pixels are mapped to a set of spatial points in a unified coordinate system using the camera intrinsics and distance map of the current frame. Plane fitting is then performed on the spatial point set to obtain the spatial plane of the candidate mounting surfaces. The boundary contours of the candidate mounting surfaces are mapped onto the spatial plane of the candidate mounting surfaces to form mounting boundaries. The spatial plane and mounting boundaries are then written into the candidate mounting surface list.
[0052] Plane fitting, through least squares plane fitting, minimizes the sum of squared distances from the set of spatial points to the fitted plane, thus obtaining a unique spatial plane.
[0053] Mapping pixels to a set of spatial points in a unified coordinate system means querying the distance map for the pixel positions within the boundary contour of the candidate mounting surface, obtaining the distance values of the pixels, back-projecting the pixels to the camera coordinates based on the camera intrinsic parameters to obtain spatial points, and transforming the spatial points to a unified coordinate system based on the camera pose of the current frame.
[0054] Mapping the candidate mounting surface boundary contour to a spatial plane to form a mountable boundary refers to obtaining a boundary space point set by back-projecting and transforming the pixels on the candidate mounting surface boundary contour, and then projecting the boundary space point set onto the fitting plane to obtain the in-plane polygon boundary, which serves as the mountable boundary.
[0055] Within the installable boundary of each candidate installation surface, inward shrinkage and equidistant grid points are arranged to generate candidate points. For each candidate point, a candidate orientation is generated. The candidate orientation is determined by the direction of the line connecting the observation positions of adjacent frames before and after the turning center frame of the key guide segment as the main direction of travel. A layout element template is bound to each candidate point and candidate orientation (the layout element template includes character height, character thickness, white space, arrow ratio, background color, font color combination, and text and directional arrow content that can be read by text direction recognition). The combination of candidate points, candidate orientations, and layout element templates is determined as candidate identifiers and written into the candidate identifier list.
[0056] Inward contraction refers to generating an inner boundary along the inner normal direction of the installable boundary and using the inner boundary as the range for point placement, so as to avoid candidate points falling outside the installable boundary.
[0057] S2. Based on key guidance road sections, obtain the real background frame sequence and segment the occluders. Generate counterfactual sample frames according to the perspective coverage of the candidate sign projection area, perform sign detection and text direction recognition, and calculate the readable level.
[0058] Read the start and end frame range and turning center frame index of each key guidance segment from the key guidance segment table one by one, and take all video frames within the start and end frame range from the frame sequence record as the real background frame sequence. For each frame in the real background frame sequence, read the current frame image and perform semantic segmentation on the current frame to obtain the occlusion pixel set, i.e., the occlusion mask, which represents the set of pixel positions of all occlusion areas in the current frame.
[0059] Identifier detection and text direction recognition are performed on a sequence of real background frames to obtain the recognized text and direction result for each frame. Consistent recognition frames are selected based on consecutive frames (e.g., 3 frames) where the recognized text and direction result are the same. For each consistent recognition frame, the detected identifier region is cropped from the real background frame into a reference sample image fragment. The reference sample image fragments obtained from each consistent recognition frame are summarized by frame index to obtain a reference sample set. The contrast, viewpoint scale, and local sharpness measure corresponding to the reference sample image fragments are saved. The median of the contrast, viewpoint scale, and local sharpness measure of the reference sample set are taken as reference statistics.
[0060] Contrast measure refers to the set of foreground pixels and the set of background pixels within a reference sample image segment. The contrast measure is characterized by the difference between the brightness of the foreground and the brightness of the background, and is normalized by the dispersion of the brightness of the foreground and the background.
[0061] The foreground pixel set is taken from the pixel positions of the text and directional arrow content within the reference sample image fragment, while the background pixel set is taken from the pixel positions within the reference sample image fragment excluding the foreground pixel set.
[0062] The visual scale measure refers to reading the character height from the layout element template as the physical size representation of the character, reading the observation distance corresponding to the reference sample image fragment from the distance map, and using the ratio of character height to observation distance to represent the scale of the character in the observer's field of vision.
[0063] Local sharpness measurement refers to the steepness of brightness change within a reference sample image segment, and the average magnitude of the brightness gradient is used to characterize local sharpness.
[0064] The candidate sign list is processed item by item. The candidate point, candidate orientation, and layout element template bound to the candidate sign are rendered into a two-dimensional sign image. The sign boundary is determined based on the outer boundary of the layout element template rendering result. The spatial position of the sign boundary is determined under a unified coordinate system by combining the candidate point and candidate orientation. In the real background frame sequence of the key guidance road section, the sign boundary is projected onto the image of each frame using the current frame's camera pose and camera intrinsic parameters to obtain the candidate sign projection area. The two-dimensional sign image is then overlaid onto the real background frame after perspective transformation according to the candidate sign projection area to obtain a counterfactual sample frame. At the same time, the candidate sign projection area mask is saved. The candidate sign projection area mask is generated by binarizing the pixels of the candidate sign projection area to produce a binary image with the same resolution as the real background frame. The candidate sign projection area is set to 1, and the rest of the area is set to 0, indicating the coverage range of the candidate sign. If the candidate sign projection area falls completely outside the image, no counterfactual sample frame is generated for the current frame. If the candidate sign projection area partially falls within the image, only the area within the image is overlaid and the in-image mask is saved.
[0065] The outer boundary refers to the smallest enclosing boundary of all the content (text, arrows, and background blocks, etc.) that will actually be displayed in the image after the layout element template is rendered as a two-dimensional icon image.
[0066] Before covering the area within the image, a background ring immediately surrounding the candidate identifier projection area is taken as the adjacent background area in the current frame. The steepness of the brightness change within the background ring is calculated to obtain the local background sharpness measure. The larger the local background sharpness measure, the sharper the edges and the clearer the image. The smaller the local background sharpness measure, the smoother the edges and the blurrier the image. The local background sharpness measure is used as the characterization of the imaging blur level of the current frame. The same blurring process is applied to the two-dimensional identifier image before perspective transformation and coverage are performed.
[0067] Equal blurring refers to calculating the local sharpness measure of the background on the set of background pixels immediately adjacent to the outside of the candidate label projection area, applying blur intensity processing to the two-dimensional label image in sequence, and selecting the blur intensity that makes the local sharpness measure of the processed two-dimensional label image closest to the local sharpness measure of the background as the equal blur intensity.
[0068] For each counterfactual sample frame, sign detection is performed within the candidate sign projection area to obtain the sign region. Text direction recognition is performed on the detected sign region, and the recognition text and direction result are output. The frame index, recognition text, direction result, and recognition record of whether the sign was detected are written. For adjacent counterfactual sample frames within the key guidance road segment, three consecutive frames with the same recognition text and the same direction result are marked as consistent recognition frames and inconsistent recognition frames. When no sign is detected in the counterfactual sample frame or the text direction recognition output is empty, the current frame is marked as an inconsistent recognition frame.
[0069] For each counterfactual sample frame within each key guidance road segment, the occlusion percentage is calculated based on contrast ratio, viewpoint scale, local sharpness metric, candidate sign projection area mask, occlusion mask, distance map, and reference statistics. The single-frame readable level is then calculated using consistent readability markers. The single-frame readable increments of each frame within the key guidance road segment are merged to obtain the readable level of the candidate sign within the key guidance road segment, expressed as:
[0070] ;
[0071] ;
[0072] in, This indicates the readable level of a single counterfactual sample frame. This indicates the occlusion percentage of the candidate identifier projection area, which is the ratio of the overlapping area between the occluding mask and the candidate identifier projection area mask to the total area of the candidate identifier projection area mask. This represents a consistency flag; it is set to one if the current frame is a consistency frame, and zero otherwise. This represents the single-frame normalization consistency term. Indicates the comparison amount of the current frame. This represents the median of the comparison values in the reference sample set. This represents the viewpoint scale of the current frame. This represents the median of the viewpoint scale in the reference sample set. This indicates the local sharpness measurement of the current frame. This represents the median of the locally clear measure in the reference sample set.
[0073] Using the distance difference between the observation positions of adjacent frames under the same coordinate system as the travel distance increment, the single-frame readable level of each frame in the same key guidance road segment is merged according to the travel distance increment to obtain the readable level of the candidate identifier in the key guidance road segment.
[0074] Inference is performed collaboratively by a trained semantic segmentation model, a marker detection model, and a text direction recognition model. The training data consists of first-person view video frames of the main traffic routes in the landscape area, covering imaging conditions such as sunny days, cloudy days, rain and fog, front lighting, back lighting, day and night, and different pedestrian densities. The training frames are annotated at both the pixel and target levels. The pixel-level annotations include walkable areas, candidate installation surface categories, and occlusion categories. The target-level annotations include the bounding boxes of markers, the text content within the markers, and the directional arrow labels. The data is divided into training, validation, and test sets. Data augmentation is applied to the training set using brightness perturbation, blur perturbation, glare-saturation perturbation, and random occlusion perturbation to improve the model's robustness to occlusion, glare, and image blur.
[0075] The semantic segmentation model is an encoder-decoder structure. It takes video frames as input and outputs a pixel category probability map with the same resolution as the input. It uses a combination of pixel cross-entropy loss and region overlap class loss as the training objective. The label detection model is a single-stage object detection network that outputs candidate label boxes and confidence scores. It is trained using a combination of classification loss and bounding box regression loss. The text direction recognition model is a cascaded structure of a text sequence recognition sub-model and a direction classification sub-model. The text sequence recognition sub-model outputs a text sequence for the image within the detection box, and the direction classification sub-model outputs the direction category for the direction arrow. During training, sequence recognition loss and multi-class cross-entropy loss are used respectively.
[0076] S3. Determine the worst critical guidance road segment based on the readability level, and determine the final location, final orientation, and unified format element template to obtain the sign configuration and output the sign list.
[0077] For the same key guidance road segment, all candidate signs are sorted from highest to lowest readability level in the key guidance road segment to generate a sorted list of candidate signs for the key guidance road segment.
[0078] For each critical guidance road segment, the candidate identifier with the highest readability level is selected from the candidate identifier ranking list of critical guidance road segments, and denoted as the optimal readability level of the critical guidance road segment. The optimal readability levels of all critical guidance road segments are compared, and the critical guidance road segment corresponding to the lowest optimal readability level is selected as the worst critical guidance road segment. The expression is:
[0079] ;
[0080] in, This represents the index of the worst critical guidance segment in the critical guidance segment table. Indicates the index of key guiding routes. Indicates the candidate identifier index. This indicates the readability level of candidate signs on key guidance road sections.
[0081] In the candidate sign ranking list corresponding to the worst critical guidance road segment, the candidate sign with the highest readability level is directly selected as the sign configuration for the worst critical guidance road segment. The candidate point, candidate orientation, and layout element template bound to the candidate sign with the highest readability level are determined as the final point, final orientation, and final layout element template for the critical guidance road segment. If there are candidate signs with the same highest readability level, the counterfactual sample frame reading records of the candidate signs in the critical guidance road segment are read one by one. The length of the continuous segment formed by the consistent reading frames is calculated and the maximum continuous segment length is taken. The candidate sign with the larger maximum continuous segment length is selected as the final sign configuration. If the maximum continuous segment length is still the same, the total number of consistent reading frames of the candidate signs in the critical guidance road segment is compared. The candidate sign with more consistent reading frames is selected as the final sign configuration. If they are still the same, the candidate sign written earlier in the candidate sign list is selected as the final sign configuration. The layout element template in the final sign configuration of the worst critical guidance road segment is recorded as the unified layout element template.
[0082] For each critical guidance road segment except the worst critical guidance road segment, a unified format element template is used to sequentially read the candidate point and candidate orientation of each candidate identifier in the candidate identifier ranking list of the critical guidance road segment. The candidate point, candidate orientation and unified format element template are combined to form a candidate identifier for evaluation. Counterfactual sample frames are generated on the real background frame sequence of the critical guidance road segment, and identifier detection and text direction recognition are performed to calculate the readable level and obtain the readable level of the candidate point and candidate orientation under the constraints of the unified format element template.
[0083] When calculating the readable level, the fixed reference statistics are used, and the occlusion mask and candidate sign projection area mask are re-observed for each frame in the real background frame sequence of the key guidance road section. Then, the occlusion ratio, contrast, viewing scale and local sharpness measure are calculated and combined with the consistent readable mark to obtain the readable level of a single frame.
[0084] After calculating the readable level of all candidate points and candidate orientations of the key guidance road segment, the combination with the highest readable level is selected as the final point and final orientation of the key guidance road segment, and the unified format element template is used as the final format element template of the key guidance road segment. If there are ties for the highest readable level, the length of the longest continuous consistent readable segment is compared in sequence, then the total number of consistent readable frames is compared, and finally the unique final identifier configuration is determined according to the order in which the candidate identifier list is written.
[0085] The final determined sign configuration for each key guidance road segment is written into the sign list. Each record in the sign list includes the key guidance road segment index, the candidate installation surface list information corresponding to the final point and the description of the installable boundary of the candidate installation surface, the location description of the final point under unified coordinates, the direction description of the final orientation under unified coordinates, and the description of the final layout element template.
[0086] The final candidate installation surface list information refers to the process of searching the installable boundaries of the candidate installation surface list for each candidate point under a unified coordinate system, projecting the candidate point onto the spatial plane of the corresponding candidate installation surface, and determining whether it falls within the polygonal range enclosed by the installable boundary. The candidate installation surface list that falls within the range is used as the candidate installation surface list information corresponding to the candidate point. When multiple candidate installation surface lists simultaneously meet the condition of falling within the range, the candidate installation surface list with the smaller polygonal range in the plane is taken as the candidate installation surface list information corresponding to the candidate point.
[0087] Figure 6 Using the key guidance road segment number as the vertical axis and the unified template element as the horizontal axis (corresponding to templates T1, T2, and T3 respectively), the number in each color block represents the optimal readability level that the key guidance road segment can achieve after evaluating each candidate point, candidate orientation, and corresponding template element combination in the candidate sign list. The warmer the color, the larger the value, indicating that after generating counterfactual sample frames through perspective coverage of the candidate sign projection area in the real background frame sequence, the higher the readability stability obtained by combining the occlusion mask, candidate sign projection area mask, distance map, and reference statistics, the more reliably it can resist occlusion, glare, and image blurring. The grid marked with the worst point in the figure represents the most difficult key guidance road segment to obtain a stable consistent readability frame under a certain template constraint. This supports the comparison of the optimal readability levels of all key guidance road segments and the selection of the worst key guidance road segment, quantifying and ranking the readability level in the design stage to determine the final point, final orientation, and unified template element.
[0088] S4. Perform trial installation according to the label list and collect first-person perspective video again. Based on the readability level, conduct on-site verification, locate the causes of obstruction and glare, and make directional corrections, forming a verification record.
[0089] According to the identification list, prepare physical identifications that are consistent with the final template and install them. Mark the installable boundary on the corresponding candidate installation surface. Project the final points in the identification list onto the spatial plane of the candidate installation surface under a unified coordinate system to obtain the landing points. Check whether the landing points fall within the polygonal area enclosed by the installable boundary. Use the landing points as the on-site installation layout positions. Measure the in-plane distances from the landing points to the three non-collinear vertices of the installable boundary polygon and reproduce the three distances on-site. Take the intersection point that satisfies the three distances and falls within the polygonal area as the final landing point of the identification list.
[0090] During installation, the geometric center of the outer boundary of the physical marker should coincide with the final point in the marker list on the candidate installation surface space plane. The direction of the arrow in the layout element template should be used as the positive direction of the physical marker, so that the positive direction is consistent with the final orientation in the candidate installation surface space plane. After installation, the marker list item, the corresponding candidate installation surface list information, the actual center point and the actual orientation should be written into the trial installation record.
[0091] After the trial setup was completed, first-person perspective video was captured again along the main route. During the capture process, timestamps were recorded frame by frame, and the camera pose, camera intrinsic parameters, and distance map aligned with the video frame were recorded simultaneously for each frame, forming a reshot frame sequence record.
[0092] The distance map and camera pose are obtained synchronously from a first-view acquisition device with depth measurement capabilities. The depth measurement capabilities include any one of structured light, binocular parallax, or time-of-flight ranging. The distance map is a pixel-by-pixel distance value map indexed by pixel position. It is bound to the corresponding video frame with the same timestamp and completed frame-level alignment. Based on the calibration mapping relationship between depth imaging and video imaging, the distance map is resampled to the same resolution as the video frame. The camera pose is obtained by performing pose estimation on the frame sequence of the first-view video. The pose estimation is based on the image feature matching results between adjacent frames. The position and orientation of the camera in each frame are output in a unified coordinate system, and the output results are bound to the frame timestamp for storage. The scale of the camera position is constrained by the synchronous distance map. When the acquisition device synchronously provides inertial measurement data, the inertial measurement data and the image feature matching results are fused to obtain the camera pose.
[0093] The reshot frame sequence record and the frame sequence record use the same unified coordinate origin and reference direction, so that the reshot video can directly reuse the start and end frame range of the key guidance segment table and the turning center frame index to locate the same key guidance segment.
[0094] Read the start and end frame range and turning center frame index of each key guidance segment from the key guidance segment table one by one, and extract all video frames within the start and end frame range from the reshot frame sequence record as the reshot real background frame sequence. For each frame in the reshot real background frame sequence, read the current frame image and perform semantic segmentation to obtain the occlusion pixel set, forming an occlusion mask.
[0095] Perform identifier detection and text direction recognition on the sequence of re-shot real background frames, output the recognized text and direction result of each frame and write it into the recognition record, and mark consistent recognition frames and inconsistent recognition frames according to three consecutive frames with the same recognized text and the same direction result.
[0096] For each frame of the detected identifier region, the identifier region is rasterized into a binary image with the same resolution as the current frame as the identifier region mask. The identifier region is set to 1 and the rest is set to zero. The identifier region mask is then used as the candidate identifier projection region mask for the reshooting stage.
[0097] For each frame of the real background frame sequence in each key guidance road segment, the occlusion ratio, contrast ratio, viewing angle scale, and local sharpness measurement are calculated based on the marker area mask, occlusion object mask, distance map, and reference statistics. The single-frame readable level is obtained by combining the consistent reading mark. Then, the single-frame readable levels of each frame in the key guidance road segment are merged according to the travel distance increment to obtain the re-enabled level of the key guidance road segment. The key guidance road segment index, marker list items, and re-enabled level are written into the review record.
[0098] For the identification list entries in the review records that are empty after filtering out three consecutive frames with the same text and direction results within the key guidance road segment, or frames in the identification records that have no detected identification or empty text direction identification output, the reason for location and orientation correction will be performed.
[0099] To pinpoint the cause of occlusion, in the sequence of repeated real background frames for the corresponding key guidance road section, the ratio of the overlapping area of the occluding object mask and the marking area mask to the area of the marking area mask is calculated frame by frame. The connected components of the overlapping area are marked as the source region of occlusion in the frame. At the same time, the source of occlusion is determined by combining the semantic segmentation category, such as tree canopy, pergola or billboard, and written into the review record.
[0100] To locate the cause of glare, the current frame image is converted into a luminance map. Then, pixels whose values in the luminance map are equal to the maximum possible values in the luminance map (taken from the upper limit of the data format of the luminance map) are taken as saturated pixels, and all saturated pixels are combined into a saturated pixel set.
[0101] A luminance map is a single-channel luminance value map obtained by proportionally combining the red, green, and blue components of the current frame.
[0102] Within the pixels covered by the mask of the marked area, the current frame image is converted into a luminance map, and then the luminance map is binarized to obtain the foreground pixel set. The foreground pixel set is regarded as the set of pixel positions of the text and directional arrow content.
[0103] If the set of saturated pixels overlaps with the set of pixel locations within the pixels covered by the mask in the marked area, the overlapping location is marked as the glare source area and written into the verification record.
[0104] Keep the final position unchanged, and replace it with another candidate orientation that is different from the current final orientation in the candidate orientation set corresponding to the final position. After re-aligning the physical identifier orientation, take a second photo and check. If there is no candidate orientation that is different from the current final orientation in the candidate orientation set corresponding to the final position, skip the current adjustment.
[0105] If obstruction or glare still exists after reshooting and verification, keep the candidate installation surface list unchanged. Within the spatial plane of the candidate installation surface, use the in-plane distance difference between the final point and other candidate points as the nearest criterion, select the other candidate point with the smallest in-plane distance difference to replace the final point, and keep the unified layout element template unchanged. Reshoot and verify according to the replaced candidate point and candidate orientation.
[0106] If there is no other candidate point on the candidate installation surface besides the current final point, then skip the point replacement.
[0107] If the obstruction or glare cannot be eliminated after re-shooting and verification, the key guidance road section in the identification list will be switched to another candidate installation surface list that meets the requirement that the candidate point falls within the polygonal range enclosed by the installation boundary. On the candidate installation surface, the point and orientation will be re-selected according to the highest readable level of the candidate identification on the key guidance road section, and then re-shooting and verification will be performed.
[0108] If a consistent identification frame still cannot be formed, the identification list item will be marked as an item that needs to be returned to regenerate the identification list, and the corresponding occlusion source area and glare source area will be retained in the review record as constraint inputs for re-determining the location and orientation.
[0109] The results of each re-shoot and verification are written into the verification record. The verification record includes the collection period and weather information, key guidance road section index, sign list items, description of the occlusion source area, description of the glare source area, orientation correction order and specific replacement content, re-shoot readability level and consistent readability frame marking results. When all key guidance road section corresponding signs meet the requirement of consistently obtainable readability frames, the final installation status is written back to the sign list and used as the final output.
[0110] Figure 5Using the number of re-shot and verification iterations as the horizontal axis and the index value as the vertical axis, the blue curve represents the re-shot readable level calculated and merged based on the identification area mask, obstruction mask, distance map, and reference statistics after locating the same key guidance road segment according to the key guidance road segment table on the re-shot frame sequence record. The orange curve represents the relative value of the number of consecutive consistent readable frames formed by consistent readable frames within the same key guidance road segment (used to characterize whether the stable readable segment continues to lengthen). From the peak or rise of the curve, it can be seen that as the causes of the obstruction source area and glare source area in the verification record are located, and the orientation replacement, point replacement, or installation surface switching are performed according to the directional correction order, the re-shot readable level increases from low to high, and the relative value of the number of consecutive consistent readable frames increases synchronously. This shows that the present invention not only completes the quantitative sorting of readable level using counterfactual sample frames in the design stage, but also accurately attributes the problem to the cause and guides directional correction through re-shot and verification after on-site installation, realizing closed-loop consistency and solidification of verifiable verification records.
[0111] This embodiment also provides a computer device applicable to the landscape signage design method based on AI vision, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the landscape signage design method based on AI vision as proposed in the above embodiment.
[0112] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0113] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the AI vision-based landscape signage design method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0114] In summary, this invention generates counterfactual sample frames by perspective overlaying candidate signs according to their projection areas on a real background frame sequence. It then calculates the readable level by combining occlusion masks, candidate sign projection area masks, distance maps, and reference statistics. This enables a quantitative representation of the effects of occlusion, glare, and image blurring under continuous travel vision, supporting the ranking of candidate signs. By identifying the worst-case key guidance road segment and using this to determine the final location, final orientation, and output a unified template of elements for the sign list, it achieves closed-loop consistency from design to on-site installation and re-examination, forming a configuration and verification record that can be directly implemented.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A landscape signage design method based on AI vision, characterized in that: include, First-person view video is collected and the collection time and weather are recorded. The camera pose is restored to obtain unified coordinates. Key guidance road sections are identified and a list of candidate installation surfaces and obstructions is extracted. Candidate identifiers are obtained from the list of candidate installation surfaces. Based on key guidance road sections, the real background frame sequence is obtained and the occluders are segmented. Counterfactual sample frames are generated according to the perspective coverage of the candidate sign projection area. Sign detection and text direction recognition are performed, and the readable level is calculated. Based on the readability level, the worst critical guidance road segment is determined, and the final location, final orientation, and unified format element template are determined to obtain the sign configuration and output the sign list; Perform trial installations according to the label list and collect first-person perspective video again. Conduct on-site verification based on the readability level, locate the causes of obstruction and glare, and make directional corrections, forming a verification record.
2. The landscape signage design method based on AI vision as described in claim 1, characterized in that: The process of acquiring first-person perspective video, recording the acquisition time and weather, and restoring the camera pose to obtain unified coordinates includes... First-person perspective video was collected along the main access route of the landscape area, and the collection time and weather were recorded. Timestamps were recorded frame by frame, and camera pose, camera intrinsic parameters and distance map were collected simultaneously for each frame. The timestamps, camera pose, camera intrinsic parameters and distance map were aligned with the video frames to form a frame sequence record. Frames with missing camera poses were filled by smooth interpolation based on adjacent valid frames. The camera position of the first frame was used as the unified coordinate origin, and the camera orientation of the first frame was used as the unified coordinate reference direction to obtain unified coordinates, which were written into the observation position and observation direction of each frame.
3. The landscape signage design method based on AI vision as described in claim 2, characterized in that: The specific steps for identifying key guidance road sections and extracting a list of candidate installation surfaces and obstructions, and obtaining candidate identifiers from the candidate installation surface list, are as follows: Under a unified coordinate system, the orientation angle sequence is obtained by observing the change in direction between adjacent frames. The turning center frame is determined, and the candidate road segment boundary is determined based on the change in angle to form a candidate road segment set. Walkable region segmentation and connected component analysis are performed on the turning center frame of the candidate road segment to identify key guidance road segments and write them into the key guidance road segment table. Semantic segmentation is performed on the turning center frame of the key guidance road segment to extract candidate mounting surfaces and occluders. The boundary contour of the candidate mounting surface is extracted and the pixels within the boundary contour are mapped to a spatial point set by combining the in-camera participation distance map. Plane fitting is performed based on the spatial point set to obtain the spatial plane of the candidate mounting surface. The boundary contour is mapped to the spatial plane to form an installable boundary and written into the candidate mounting surface list. Within the installable boundary, shrink inward and generate candidate points by equidistant grid placement. Generate candidate orientations for the candidate points and bind them to layout element templates. Determine the combination of candidate points, candidate orientations, and layout element templates as candidate identifiers and write them into the candidate identifier list.
4. The landscape signage design method based on AI vision as described in claim 1, characterized in that: The specific steps for obtaining the real background frame sequence, segmenting the occluders, and generating counterfactual sample frames by perspective overlay of candidate identifier projection areas are as follows: Read the start and end frame range and the turning center frame index from the key guidance road segment table, take all video frames within the start and end frame range as the real background frame sequence, and perform semantic segmentation on each frame to obtain the occlusion mask. On a sequence of real background frames, mark detection and text direction recognition are performed, and consistent recognition frames are selected. The mark regions detected by the consistent recognition frames are cropped to form a reference sample set. The median of the contrast quantity, viewpoint scale quantity, and local sharpness quantity are taken as reference statistics. The candidate points, candidate orientations, and layout element templates are rendered into two-dimensional mark images. The mark boundaries are determined according to the outer boundary of the layout element template rendering result. The mark boundaries are projected onto the real background frames according to the camera pose and camera intrinsic parameters to obtain the candidate mark projection areas. After perspective transformation of the two-dimensional mark images, counterfactual sample frames are generated and the candidate mark projection area mask is saved.
5. The landscape signage design method based on AI vision as described in claim 4, characterized in that: The specific steps for performing identifier detection and text direction recognition, and calculating the readable level, are as follows: In the candidate identifier projection area of the counterfactual sample frame, perform identifier detection and text direction recognition and write the recognition record, and filter the consistent recognition frames and inconsistent recognition frames according to the selection mark; For each counterfactual sample frame, the occlusion ratio is calculated based on the contrast ratio, viewpoint scale, local sharpness measurement, candidate sign projection area mask, occlusion mask, distance map, and reference statistics. Combined with the consistent readability frame marker, the readability level of a single frame is obtained. The distance difference between the observation positions of adjacent frames under the same coordinate system is used as the travel distance increment. The readability levels of each frame in the same key guidance road segment are merged according to the travel distance increment to obtain the readability level of the candidate sign in the key guidance road segment.
6. The landscape signage design method based on AI vision as described in claim 1, characterized in that: The specific steps for determining the worst-case critical guidance route based on the readability level are as follows: For each key guidance road segment, the candidate icons are sorted from high to low in terms of their readability level in the key guidance road segment to generate a candidate icon sorting list; For each critical guidance road segment, the highest readable level in the candidate identifier ranking list is taken as the optimal readable level of the critical guidance road segment, and the critical guidance road segment with the smallest value among all the optimal readable levels of critical guidance road segments is selected as the worst critical guidance road segment.
7. The landscape signage design method based on AI vision as described in claim 6, characterized in that: The process involves determining the final location, final orientation, and unified layout element template, obtaining the identifier configuration, and outputting the identifier list. The specific steps are as follows: In the candidate sign ranking list of the worst critical guidance road segment, the candidate sign with the highest readability level is selected as the sign configuration of the critical guidance road segment. The candidate point, candidate orientation and layout element template bound to the candidate sign are determined as the final point, final orientation and unified layout element template. When there are ties, the unique sign configuration is determined in sequence according to the longest continuous consistent reading segment length, the total number of consistent reading frames and the order of writing the candidate sign list. For each key guidance road segment except the worst key guidance road segment, candidate points and candidate orientations are combined one by one using a unified format element template to generate candidate identifiers. Counterfactual sample frames are generated on the real background frame sequence to calculate the readability level. The combination with the highest readability level is selected as the final point and final orientation of the key guidance road segment. When there are ties, the unique configuration is determined sequentially according to the longest continuous consistent readability segment length, the total number of consistent readability frames, and the order in which the candidate identifier list is written. The unified format element template is used as the final format element template for each key guidance road segment. The final point, final orientation, and final format element template of each key guidance road segment are written into the identifier list.
8. The landscape signage design method based on AI vision as described in claim 1, characterized in that: The process involves trial installation according to the identification list, re-capturing first-person perspective video, conducting on-site verification based on the readability level, identifying and correcting the causes of obstruction and glare, and creating a verification record. The specific steps are as follows: Mark the installable boundary on the candidate installation surface, project the final point onto the spatial plane of the candidate installation surface to obtain the landing point, and obtain the on-site landing point by reproduction. Align the on-site landing point with the geometric center of the outer boundary of the physical marker and align the final orientation with the direction arrow. Repeat the shooting along the main traffic route to form a frame sequence record, generate the occlusion mask and the marker area mask, and perform marker detection and text direction recognition. Calculate the readable level of a single frame and merge it into the readable level of the repeated shooting according to the travel distance increment, and write it into the verification record. Locate the occlusion source area and the glare source area, and sequentially change the candidate orientation, replace the nearest candidate point, switch the candidate installation surface list, and perform repeated shooting and verification. Write the results of each repeated shooting and verification into the verification record.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the landscape signage design method based on AI vision as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the landscape signage design method based on AI vision as described in any one of claims 1 to 8.