Methods for digitally archiving house numbers and readable storage media

CN117829784BActive Publication Date: 2026-08-14BEIJING REMARKABLES UNITED TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410225967.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2026-08-14
Estimated Expiration
2044-02-29

AI Technical Summary

Technical Problem

[0003]首先,目前人工巡检建档方式效率低下

Benefits of technology

[0018]通过如上所提供的对门牌进行数字化建档的方法,本申请实施例通过门牌信息的采集,过滤与存储,能够完成对街道门牌进行数字化建档的任务。进一步,在一些实施例中,通过无人机来采集多媒体数据,可以提高建档的速度和效率。更进一步地,在一些实施例中,通过基于卡尔曼滤波算法对门牌进行跟踪和裁剪出识别区域,可以降低诸如遗漏、错误等人为因素与传统图像算法对建档质量造成影响的发生率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117829784B_ABST
    Figure CN117829784B_ABST
Patent Text Reader

Abstract

This application discloses a method and readable storage medium for digitally archiving house numbers, including collecting multimedia data containing house number images; extracting house number information from the multimedia data; filtering the house number information to generate valid house number information; and storing the valid house number information for archiving. Using this method, digital archiving of street house numbers can be completed, facilitating comparison and analysis of changes to street house numbers in numerous historical archives, identifying potential customers, and thus providing them with more precise services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to the field of artificial intelligence technology. More specifically, this application relates to a method for digitally archiving house numbers. Background Technology

[0002] With the continuous advancement of lean management and digital transformation, digital record-keeping technology has been widely applied across various industries. To improve the intelligence, precision, and efficiency of electricity marketing services, it is necessary to digitally record street address information. This involves intelligently comparing and analyzing numerous historical archives to identify new installations, changes (transfers of shops), and removals of street address information. This allows for the identification of potential customers changing their electricity usage, enabling more precise services such as account cancellation, name change, transfer of ownership, meter relocation, capacity reduction (restore), suspension (restore), temporary removal (reinstallation), and category modification.

[0003] First, the current manual inspection and record-keeping method is inefficient. Second, due to differences in factors such as the size, color, font, and installation location of street address signs, traditional image processing-based address sign recognition methods have low accuracy.

[0004] Therefore, improving the speed and efficiency of archiving and avoiding the impact of human factors such as omissions and errors, as well as traditional image algorithms, on the quality of archiving have become urgent technical problems to be solved. Summary of the Invention

[0005] In order to at least address one or more of the technical issues mentioned above, this disclosure proposes a method for digitally archiving house numbers in several aspects.

[0006] In one aspect, a method for digitally archiving house numbers includes: acquiring multimedia data containing house number images; extracting house number information from the multimedia data; filtering the house number information to generate valid house number information; and storing the valid house number information for archiving purposes.

[0007] In one embodiment, the multimedia data is collected via a drone.

[0008] In another embodiment, extracting address information from the multimedia data includes: locating the address image in the multimedia data; and extracting the text content of the address from the located address image.

[0009] In yet another embodiment, locating the doorplate image in the multimedia data includes: identifying a single-frame image containing the doorplate image; extracting the single-frame image; and cropping the doorplate image from the single-frame image.

[0010] In yet another embodiment, extracting the text content of the doorplate from the located doorplate image includes: cropping an image including a text region from the doorplate image; and recognizing the text content from the image including the text region.

[0011] In yet another embodiment, the text content includes at least one of the following: store name, brand, and contact information.

[0012] In yet another embodiment, filtering the address information to generate valid address information includes: deleting incomplete text content to retain valid text content; and merging identical valid text content.

[0013] In another embodiment, when the multimedia data is video, door number information is extracted from the multimedia data and filtered to generate valid door number information, including: assigning the same number to the same door number in the video; obtaining the text content of the door number in each frame of the video; and merging the text content of the door numbers with the same number.

[0014] In another embodiment, assigning the same number to the same doorplate in the video includes: locating the position of the doorplate image in each frame of the video; predicting the position of the doorplate image in the next frame of the video; locating the actual position of the doorplate image in the next frame of the video; and assigning the same number to the doorplate in the two frames if the overlap between the actual position and the predicted position of the doorplate image in the next frame of the video is greater than a set threshold, otherwise assigning different numbers.

[0015] In yet another embodiment, the algorithm for predicting the position of the doorplate image in the next frame of the video includes an object detection algorithm and a Kalman filter algorithm.

[0016] In another embodiment, merging the text content of the same numbered doorplates includes: obtaining the capture direction of the doorplate in the video; and incrementally merging the text content of the doorplate according to the capture direction to retain the valid text content.

[0017] In yet another embodiment, storing the valid address information for filing includes: storing the filtered valid text content; storing the address image corresponding to the valid text content; and storing the street name corresponding to the address image.

[0018] By employing the method for digitally archiving house numbers provided above, this application embodiment, through the collection, filtering, and storage of house number information, can complete the task of digitally archiving street house numbers. Furthermore, in some embodiments, using drones to collect multimedia data can improve the speed and efficiency of archiving. Even further, in some embodiments, by tracking and cropping the recognition area of ​​the house numbers based on a Kalman filter algorithm, the incidence of human factors such as omissions and errors, as well as the impact of traditional image algorithms, on the quality of archiving can be reduced. Attached Figure Description

[0019] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts, wherein: Figure 1 An exemplary flowchart of a method for digitally archiving door numbers according to an embodiment of this application is shown; Figure 2a An exemplary flowchart of a method for extracting address information according to an embodiment of this application is shown; Figure 2b An exemplary flowchart of a method for extracting address information according to yet another embodiment of this application is shown; Figure 3 An exemplary flowchart of the door number tracking algorithm according to an embodiment of this application is shown; Figure 4 An exemplary schematic diagram illustrating the integrity of a doorplate according to an embodiment of this application is shown. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] It should be understood that the terms "comprising" and "including" used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0022] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0023] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0024] The specific embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0025] Figure 1 An exemplary flowchart of a method for digitally archiving door numbers according to an embodiment of this application is shown. Figure 1 As shown, in step S101, multimedia data including doorplate images is collected.

[0026] The multimedia data mentioned here can include images, videos, etc. Images have the advantage of occupying less space, making them easy to store and transmit. However, a single image may not provide enough information to obtain accurate results. While multimedia data may occupy more space or have a larger transmission bandwidth, it contains sufficient detailed information, making it easier to obtain more accurate and comprehensive address information.

[0027] Multimedia data in different formats can be acquired in various ways. For example, in one embodiment, multimedia data can be collected using a drone. In another embodiment, the multimedia data includes video. By using a drone to collect address information of the streets to be documented, a video file containing street address information is obtained. Specifically, the drone's flight speed, altitude, angle, and shooting distance from the street addresses are first controlled to obtain high-quality video data. Then, the video file is parsed and frames are extracted.

[0028] The video frame extraction technology described above is a method for extracting still images from a video stream. This technology can identify and extract keyframes from the video stream, thereby reducing storage space and processing resource requirements. This method is typically based on image processing, computer vision, machine learning, and other technologies, and can be implemented using various algorithms and models. For example, traditional motion detection and background subtraction algorithms can be used, as well as deep learning and convolutional neural network models for keyframe identification and extraction. Furthermore, the frame extraction frequency can be controlled by setting a timed frame extraction strategy. To balance processing resources, recognition efficiency, and adaptability to different scenarios, a timed frame extraction strategy can be set. That is, the system can periodically extract keyframes from the video stream according to a set time interval. This ensures effective frame extraction of the entire video stream without excessively consuming processing resources.

[0029] exist Figure 1 In step S102, the address information is extracted from the multimedia data. In one embodiment, Figure 2a An exemplary flowchart of a method for extracting address information according to an embodiment of this application is shown, such as... Figure 2a As shown, in step S201, the doorplate image is located in the multimedia data. Then, in step S202, the text content of the doorplate is extracted from the located doorplate image.

[0030] In step S201, locating the doorplate image in the multimedia data includes the following steps: In another embodiment, Figure 2b An exemplary flowchart of a method for extracting address information according to another embodiment of this application is shown. Figure 2b In step S2011, a single-frame image containing a doorplate image is identified; in step S2012, the single-frame image is extracted; and in step S2013, the doorplate image is cropped from the single-frame image.

[0031] Specifically, recognizing a single frame containing a doorplate image requires building and training a doorplate target detection model to obtain the trained model. Traditional image processing methods often require manual annotation or threshold setting to find regions of interest. This method is not only time-consuming and labor-intensive but also easily affected by factors such as image quality and lighting conditions, leading to inaccurate results. By using the target detection model, automatic recognition of street doorplates can be achieved without manual annotation or threshold setting. Furthermore, the model can adapt to street doorplates of different styles, sizes, and lighting conditions, exhibiting good robustness and generalization ability.

[0032] Doorplate detection can be performed using a YOLOv8-based target detection algorithm to train a doorplate detection model. This approach ensures both accuracy and speed, effectively improving overall recognition accuracy. The training process includes the following steps: acquiring street doorplate images from various application scenarios; selecting samples from each scenario to construct a training set; and labeling the training set images with doorplate information. All labeled images are then divided into training, validation, and test sets in an 8:1:1 ratio, and the doorplate detection model is trained using the YOLOv8 algorithm. The learning rate, batch size, and number of iterations for training the doorplate detection model are set. The model is evaluated during training, with evaluation performed every 5 generations. The model with the best accuracy is selected after training. Finally, the best accuracy model is tested using the test set until its accuracy reaches a set threshold, resulting in the trained doorplate detection model.

[0033] By applying this model, single-frame images containing doorplates can be accurately located. Based on the doorplate location information provided by the object detection model, the position and size of the doorplate within a single frame image can be accurately determined. Image processing algorithms are then applied to separate the doorplate from the background, and target region cropping is performed to obtain an image containing only the doorplate. This also reduces the computational load of subsequent processing, improving processing efficiency.

[0034] Returning to step S202, extracting the text content of the address plate from the located address plate image includes the following steps: Figure 2b In step S2021, an image including the text area is cropped from the doorplate image; and in step S2022, the text content is identified from the image including the text area. The text content includes the store name, brand, and contact information.

[0035] Specifically, the text model employs an independent construction approach, combining a text region object detection model and a text recognition model. The text region object detection model categorizes doorplates into three types: store name, brand, and contact information. This reduces interference from non-store name text, allowing the model to focus more on the relevant text regions and improve the accuracy of store name extraction. After doorplate region detection, the text recognition model further processes each region. The text recognition model converts the text information within the doorplate region into readable text. By recognizing each character or word individually, it accurately extracts key information such as the store name. The two models, each focusing on different tasks, collaborate to improve the accuracy of doorplate information extraction. The text region object detection model, by subdividing doorplates into different categories, reduces interference from non-store name text, making subsequent text recognition more accurate. Meanwhile, the text recognition model can recognize each character within the doorplate region, further improving the accuracy of store name extraction.

[0036] In one embodiment of this application, training a text model for doorplates includes the following steps: acquiring street doorplate images in the application scenario; selecting sample images (different appearances and layouts) from various scenarios to construct a training set; labeling the training set images using a labeling tool, with labeling items including store name, brand, and contact information; dividing all labeled images into a training set, validation set, and test set in an 8:1:1 ratio, and training a text region object detection model based on the DBNet algorithm; setting the learning rate, batch size, and number of iterations for training the text region object detection model; evaluating the text region object detection model during training, with an evaluation performed every 5 generations, and selecting the model with the best accuracy after training; testing the best accuracy model using a test set until the model's accuracy reaches a set threshold, thus obtaining the training result. The text region object detection model was trained. The text regions containing store names, brands, and contact information located by the text region object detection model were cropped from the original image. Images of text (different fonts, sizes, and colors) from various scenarios were selected to construct a training set. Annotation tools were used to label the training set images, with the annotation items being the text content. All labeled images were divided into training, validation, and test sets in an 8:1:1 ratio, and a text recognition model was trained based on the SVTR algorithm. The learning rate, batch size, and number of iterations for training the text recognition model were set. The text detection model was evaluated during training, with an evaluation performed every 5 generations. The model with the best accuracy was selected after training. The best accuracy model was tested using the test set until the model's accuracy reached the set threshold, resulting in the trained text recognition model.

[0037] Back Figure 1In step S102, after obtaining the trained doorplate target detection model and text model, the doorplate location is located on each frame of the image based on the doorplate target detection model, and the store name on the doorplate is obtained based on the text model, thus obtaining the doorplate name and location for each frame of the video image. This step includes the following steps: inputting the current frame image into the trained doorplate target detection model, the model output includes doorplate detection box location information and confidence level, wherein the detection box location information includes the coordinates of the upper left corner of the detection box. of The coordinates of the top left corner of the detection box of The coordinates of the lower right corner of the detection box of and the coordinates of the lower right corner of the detection box of The result is as follows: The doorplate is cropped from the original image according to the coordinates of the detection box and input into the text model. The model output includes the text category code, the detection box position information, and the text content. The text category code includes the store name, brand, and contact information, and the detection box position information includes the coordinates of the top left corner of the detection box. of The coordinates of the top left corner of the detection box of , coordinates of the lower right corner of the detection frame of and the coordinates of the lower right corner of the detection box of The results are as follows: Text content with the store name as the category code is filtered out as the name of the doorplate in the current image frame. If there is no category code for the store name, the doorplate name is empty. The doorplate detection box position information and doorplate name are output.

[0038] In step S103, the address information is filtered to generate valid address information. In one embodiment, incomplete text content is deleted to retain valid text content, and identical valid text content is merged.

[0039] In another embodiment, when the multimedia data is video, the same number is assigned to the same doorplate in the video, the text content of the doorplate in each frame of the video is obtained, and the text content of the doorplates with the same number is merged.

[0040] Assigning the same number to the same doorplate in the video requires locating the position of the doorplate image in each frame of the video, predicting the position of the doorplate image in the next frame of the video, locating the actual position of the doorplate image in the next frame of the video, and assigning the same number to the doorplate in the two frames if the overlap between the actual position and the predicted position of the doorplate image in the next frame of the video is greater than a set threshold; otherwise, assigning different numbers.

[0041] The localization and tracking of house numbers, or house number tracking algorithm, identifies and tracks house number regions in a video by analyzing changes between consecutive frames. This algorithm utilizes image processing and pattern recognition methods to automatically extract features of the house numbers, such as color, shape, and text, thus achieving effective detection and tracking. Based on the house number tracking algorithm, house numbers in the video can be numbered and merged. By matching and comparing the position and features of the same house number in different frames, they can be marked as the same house number instance and assigned a unique identifier, i.e., a house number ID. This allows for accurate tracking of the appearance of a house number throughout the entire video and enables statistical analysis.

[0042] Furthermore, the doorplate tracking algorithm can also recognize textual information on doorplates, such as door numbers or shop names. Through text recognition technology, the text on the doorplate is extracted to obtain the doorplate's name or identification information. This provides a foundation for subsequent data processing and management. Finally, the doorplate tracking algorithm can also acquire images of the doorplates. During the identification and tracking process, images of the doorplate area are captured and saved for further image processing or manual review. These images can be used not only for doorplate verification and comparison but also for applications such as building doorplate databases.

[0043] In yet another embodiment, the algorithm for predicting the position of the doorplate image in the next frame of the video includes an object detection algorithm and a Kalman filter algorithm.

[0044] Specifically, this step includes the following sub-steps: constructing a doorplate tracking algorithm based on the doorplate target detection model and Kalman filtering; obtaining the location information and name of the doorplates in each frame of the video; assigning IDs to the doorplates in the current frame using the doorplate tracking algorithm; obtaining the ID, name, and location of the doorplates contained in each frame; merging the names of all doorplates in the video according to their IDs; selecting the optimal frame in which the doorplate appears as the doorplate image; and outputting the ID, name, and image of all doorplates in the video. Specifically, the doorplate target detection model is used as the detection box generation model for the doorplate tracking algorithm. During data association, Kalman filtering is used to predict the position of the doorplate in the current frame in the next frame. If the overlap between the detection box of the doorplate in the next frame and the predicted box of the doorplate in the previous frame is greater than a set threshold (which can be set to 0.7), then the doorplate is determined to belong to the same doorplate as the doorplate in the previous frame; otherwise, it is determined to be a new doorplate.

[0045] Figure 3An exemplary flowchart of a doorplate tracking algorithm according to an embodiment of this application is shown. In one embodiment, video frames are first read and feature extraction is performed on the doorplate image in step S301. The SURF algorithm can be used for feature extraction. After preprocessing steps such as grayscale conversion and Gaussian filtering, the SURF algorithm is used to detect feature points in the doorplate image and calculate the descriptor of each feature point to obtain a set of feature vectors for the doorplate image. Then, in step S302, the Kalman filter is initialized, and parameters such as state variables, state transition matrix, and observation matrix are set. In this algorithm, a conventional two-dimensional Kalman filter can be used for target tracking. The center coordinates of the doorplate image are used as the state variable, while the state transition matrix describes the prediction and correction process of the doorplate image position, and the observation matrix is ​​used to map the feature vectors of the doorplate image to the observation space. The Kalman filter is used to predict the doorplate image to obtain the prediction result.

[0046] In the initial stage of doorplate image tracking, since there are no actual detection results to provide a reference, a Kalman filter can be used for prediction to obtain the initial position of the doorplate image. Based on the state transition matrix and the observation matrix, the predicted position of the doorplate image in the next time step can be obtained through the prediction step S303 of the Kalman filter. At step S304, it is determined whether the actual detection result is consistent with the prediction result. If they are consistent, tracking continues; otherwise, corrections are made.

[0047] In each frame of the video, a feature point matching algorithm (such as FLANN) can be used to detect the doorplate image, and the feature vector of the doorplate image is recalculated. Then, the center coordinates of the detected doorplate image are compared with the position predicted by the Kalman filter. If the distance is small, the actual detection result is considered consistent with the prediction result, and tracking of the doorplate image continues; if the distance is too large, the target is considered to have deviated from the predicted trajectory, and correction is needed. The Kalman filter parameters, including state variables, state transition matrix, and observation matrix, are updated based on the detection results. When the detection result is inconsistent with the Kalman filter's prediction result, the Kalman filter parameters need to be corrected based on the actual detection result to improve the accuracy and stability of tracking. The state variables, state transition matrix, and observation matrix can be updated using the Kalman Update step.

[0048] Meanwhile, the above steps need to be repeated continuously during the tracking of the doorplate image until the video ends or the doorplate image tracking fails (e.g., the doorplate image is obscured).

[0049] In another embodiment, merging the text content of the same numbered doorplates includes: obtaining the capture direction of the doorplate in the video, and incrementally merging the text content of the doorplate according to the capture direction to retain the valid text content.

[0050] Specifically, merging the text content of doorplates with the same number (i.e., merging doorplate names with the same ID) and selecting the optimal frame requires adding a completeness status attribute to each doorplate to further improve the accuracy of doorplate recognition. The completeness status attribute refers to the completeness of the doorplate in the video frame, which is determined by comparing the positional relationship between the doorplate and the image boundary.

[0051] Figure 4 An exemplary schematic diagram of the integrity of a doorplate according to an embodiment of this application is shown. Specifically, an integrity status attribute is introduced for each doorplate, and the integrity status is divided into four cases based on the relationship between the doorplate's position P2 and the image boundary P2. If the left and right boundaries of the doorplate do not coincide with the left and right boundaries of the image, its integrity status is set to 0; if the left boundary of the doorplate coincides with the left boundary of the image, its integrity status is set to 1; if the right boundary of the doorplate coincides with the right boundary of the image, its integrity status is set to 2; and if both the left and right boundaries of the doorplate coincide with the left and right boundaries of the image, its integrity status is set to 3.

[0052] If multiple doorplates with a completeness state of 0 exist among all doorplate names for the current ID, the name with the highest probability of occurrence needs to be selected as the doorplate name, and the frame with the doorplate detection box closest to the image center is selected as the optimal frame. This method can avoid misidentification and missed identification due to incomplete doorplates. If no doorplates with a completeness state of 0 exist, the matching method for doorplate names needs to be determined according to the doorplate acquisition order. If the completeness state changes from 1 to 3 to 2, the doorplates are acquired from left to right; if the completeness state changes from 2 to 1 to 3, the doorplates are acquired from right to left. Based on the acquisition direction, the names of all doorplates are incrementally merged to obtain the doorplate name for the current ID. Among the doorplates with a completeness state of 3, the frame with the highest character matching score between the doorplate name and the final doorplate name needs to be selected as the optimal frame. This method can further improve the accuracy and reliability of doorplate recognition.

[0053] Returning to step S104, valid address information is stored for archiving. In one embodiment, the filtered valid text content is stored first, followed by the address image corresponding to the valid text content and the street name corresponding to the address image, to complete the archiving of addresses for the current street. Specifically, in the address recognition system, valid text content is filtered out through a series of filtering steps. This text content may include address numbers, shop names, and other related information. After filtering, only text content that conforms to the prescribed format and standards is considered valid and then stored. Next, for each valid text content, the system stores its corresponding address image. This ensures that a correspondence is established between address information and the corresponding image, facilitating subsequent queries and viewing. In addition, to improve the archiving of addresses for the current street, the system also stores the street name corresponding to the address image. Thus, when querying address information for a certain street, the system can quickly retrieve the corresponding address image and related text content based on the street name, improving the efficiency and accuracy of address archiving.

[0054] In some embodiments, various aspects of this disclosure can also be implemented as a program product comprising program code that, when executed by a processor, causes the processor to perform the methods described above. The program product may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0055] While numerous embodiments of this disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and intent of this disclosure. It should be understood that various alternatives to the embodiments of this disclosure described herein may be employed in the practice of this disclosure. The appended claims are intended to define the scope of this disclosure and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A method for digitally archiving house numbers, characterized in that, include: Collect multimedia data including images of doorplates; The address information is extracted from the multimedia data, and the address information includes text content associated with the store name, brand and contact information; The doorplate information is filtered to generate valid doorplate information, wherein the multimedia data is video. Doorplate information is extracted from the multimedia data and filtered to generate valid doorplate information, including: assigning the same number to the same doorplate in the video; obtaining the text content of the doorplate in each frame of the video; and merging the text content of doorplates with the same number, wherein merging the text content of doorplates with the same number includes: obtaining the acquisition direction of the doorplate in the video; and incrementally merging the text content according to the acquisition direction to retain valid text content; and The valid address information is stored for filing purposes.

2. The method according to claim 1, characterized in that, The multimedia data was collected using drones.

3. The method according to claim 1, characterized in that, Extracting address information from the multimedia data includes: Locate the doorplate image in the multimedia data; Extract the text content of the doorplate from the located doorplate image.

4. The method according to claim 3, characterized in that, Locating the doorplate image in the multimedia data includes: Identify a single frame image containing the doorplate image; Extract the single-frame image; and The signboard image is cropped from the single-frame image.

5. The method according to claim 4, characterized in that, The text content extracted from the located doorplate image includes: Cropping an image including the text area from the address plate image; and Identify the text content from an image that includes text regions.

6. The method according to claim 5, characterized in that, The filtering of the address information to generate valid address information includes: Delete incomplete text content to retain valid text content; and Merge identical valid text content.

7. The method according to claim 1, characterized in that, The following are examples of assigning the same number to the same doorplate in the video: Locate the position of the doorplate image in each frame of the video; Predict the position of the doorplate image in the next frame of the video; Locate the actual position of the doorplate image in the next frame of the video; and If the overlap between the actual and predicted positions of the doorplate image in the next frame of the video is greater than a set threshold, then the doorplates in the two frames are assigned the same number; otherwise, they are assigned different numbers.

8. The method according to claim 7, characterized in that, The algorithms for predicting the location of the doorplate image in the next frame of the video include object detection algorithms and Kalman filtering algorithms.

9. The method according to any one of claims 1-8, characterized in that, The storage of the valid address information for filing purposes includes: Store the filtered valid text content; Store the doorplate image corresponding to the valid text content; and Store the street name corresponding to the doorplate image.

10. A computer-readable storage medium, characterized in that, It stores computer program instructions that are executed by one or more processors to perform the operation of the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Multi-target behavior identification method and system for monitoring video

    CN110378259A

  • Mobile robot doorplate positioning method combined with deep learning

    CN110458161A

  • Traffic asset checking method and device, medium and electronic equipment

    CN115273025A