Method, device and equipment for identifying street tree species in street scene image and medium

By using the yolo supervision model and semi-supervised domain adaptation framework in street scene images, combining the global domain adaptation module and enhanced pseudo-tagged adapter to identify the street tree species, solving the problem of difficult to take into account both economic and accuracy in the existing technology, and achieving efficient and accurate tree species recognition.

CN120047822AActive Publication Date: 2025-05-27GUANGZHOU UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510028144.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-27
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The prior art is difficult to take into account both economic and accuracy in large-scale urban street tree evaluations. Traditional field surveys are costly, advanced sensor equipment is costly or individual identification accuracy is insufficient.

Method used

A method for identifying street tree species in street scene images is proposed. By obtaining street scene data sets and tree species single-plane data sets, using the yolo supervision model and semi-supervised domain adaptation framework, combining the global domain adaptation module and enhanced pseudo-label adapter, the tree species recognition model is trained.

Benefits of technology

The cost of identifying street tree species is significantly reduced, the recognition accuracy is improved, and the effect of efficient identification of tree species in street scene images is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047822A_ABST
    Figure CN120047822A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device, equipment and a medium for identifying street tree species in a street view image, and relates to the technical field of image identification, and the method comprises the steps: obtaining a street view data set of a research area; obtaining a tree species single plant data set; training a yolo supervisory model by using the tree species single plant data set; adding the global domain adaptive module and the enhanced pseudo tag adapter into a semi-supervised target detection framework to obtain a semi-supervised domain adaptive framework; inputting the street view data set, the tree species single plant data set and the trained yolk supervision model into a semi-supervised domain adaptation framework for training to obtain a tree species identification model; and identifying the street tree species by using the tree species identification model. According to the method, the streetscape data set and the tree species single-plant data set of the labeled tree species are used for training, so that the labeling cost of the data set can be reduced; the features are aligned by using the global domain self-adaptive module, and the pseudo tag is optimized by using the enhanced pseudo tag adapter, so that the training effect of the model can be improved, and the accuracy of identifying the tree species in the streetscape image is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a method, device, equipment and medium for identifying street tree species in street view images. Background Art

[0002] An up-to-date list of street tree species is crucial for urban management and planning. Different types of street trees have relatively different functions and advantages, which enables them to play specific roles in urban services and ecological benefits. For example, some tree species are outstanding in adsorbing particulate matter and carbon dioxide in the air and are commonly used to improve urban air quality, while others are more suitable for providing shade and cooling effects in high-temperature areas. In addition, the selection and distribution of street trees also affect urban ecological diversity, landscape aesthetics, and the mental health level of residents.

[0003] Currently, common methods for urban street tree assessment include field surveys, light detection and ranging (LiDAR), using satellite images, and street view recognition. Among them, traditional field surveys are often labor-intensive and time-consuming, especially in large-scale urban street tree censuses, with high costs. With the progress of computer vision technology, advanced sensors (such as LiDAR, remote sensing satellites, etc.) have made significant progress in urban tree censuses. For example, light detection and ranging (LiDAR) shows high accuracy and feasibility in tree species identification by capturing the three-dimensional structure of trees; satellite images identify trees based on the spectral radiance and surface reflectivity of different tree species and are suitable for large-scale assessments. However, due to reasons such as high equipment costs or insufficient individual recognition accuracy, these technologies are difficult to balance economy and accuracy in large-scale urban street tree assessments and lack cost-effectiveness. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method, device, equipment and medium for identifying street tree species in street view images, so as to reduce the cost of identifying street tree species and improve the recognition accuracy.

[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a method for identifying street tree species in street view images, and the method includes the following steps:

[0006] Obtain a street view data set of the study area; wherein, the street view data set includes street tree images on both sides of the street;

[0007] Obtain a single-tree data set of tree species; wherein, the single-tree data set of tree species includes single-tree images of multiple labeled tree species;

[0008] Train a YOLO supervision model using the single-tree data set of tree species to obtain the trained YOLO supervision model;

[0009] Add a global domain adaptation module and an enhanced pseudo-label adapter to the semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework; wherein, the global domain adaptation module is used to align the features of the street view dataset and the single-tree species dataset; the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression.

[0010] Input the street view dataset, the single-tree species dataset, and the trained YOLO supervised model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model.

[0011] Use the tree species recognition model to identify the tree species of street trees.

[0012] In some embodiments, obtaining the street view dataset of the study area includes the following steps:

[0013] Generate street view sampling points at a set distance threshold on the streets of the study area.

[0014] At each street view sampling point, use a camera to capture the images of the street trees on both sides of the street in the 90° and 270° directions.

[0015] The calculation formula for the set distance threshold is:

[0016] T = 2 * A × tan45° - I;

[0017] wherein, T is the set distance threshold; A is the distance between the camera and both sides of the street; 45° is the field of view range of each street view sampling point; I is the overlapping part of adjacent fields of view.

[0018] In some embodiments, obtaining the single-tree species dataset includes the following steps:

[0019] Obtain the single-tree images of multiple labeled tree species from different data open platforms as the single-tree species dataset.

[0020] In some embodiments, training the YOLO supervised model using the single-tree species dataset to obtain the trained YOLO supervised model includes the following steps:

[0021] Perform data augmentation on the single-tree species dataset to obtain the augmented single-tree species dataset; wherein, the data augmentation includes exposure, light reduction, random cropping, scaling, and mirroring of each single-tree image.

[0022] Using each of the single-tree images in the enhanced single-tree dataset of the tree species as training samples and the labeled tree species as training labels, train the YOLO supervision model to obtain the trained YOLO supervision model.

[0023] In some embodiments, adding the global domain adaptation module and the enhanced pseudo-label adapter to the semi-supervised object detection framework to obtain the semi-supervised domain adaptation framework includes the step of constructing the global domain adaptation module. The step of constructing the global domain adaptation module includes the following steps:

[0024] Determine the mixing coefficient; wherein, the mixing coefficient is used to determine the proportion of using the street tree images and the proportion of using the single-tree images when training the semi-supervised domain adaptation framework;

[0025] The calculation formula of the mixing coefficient is:

[0026]

[0027] where α is the mixing coefficient, α min is the minimum value of the mixing coefficient, D is the feature distance, and β represents a hyperparameter that controls the influence degree of the feature distance on the mixing coefficient;

[0028] The calculation formula of the feature distance is:

[0029] D = ||U S - U T ||;

[0030] where U S and U T respectively represent the feature mean vectors of the source domain and the target domain; the source domain is the single-tree dataset of the tree species, and the target domain is the street view dataset;

[0031] The calculation formula of the feature mean vector is:

[0032]

[0033] where U is the feature mean vector, N represents the number of images in the dataset, f(x) represents the feature extraction function of the YOLO supervision model, and x i represents the images in the dataset;

[0034] Generate mixed samples using the street view dataset and the single-tree dataset of prime tree species according to the mixing coefficient;

[0035] The expression of the mixed sample is:

[0036]

[0037] Among them, is the mixed sample, x S is the image of the source domain, x T is the image of the target domain;

[0038] Construct the global domain adaptation module according to the mixed sample and the domain adaptation loss function;

[0039] The expression of the domain adaptation loss function is:

[0040] L da = -∑ h,w [B log p(h, w)+(1 - B)log(1 - p(h, w))];

[0041] Among them, L da is the domain adaptation loss function, p(h, w) is the output of the global domain adaptation module, B = 0 represents the labeled image, and B = 1 represents the unlabeled image.

[0042] In some embodiments, adding the global domain adaptation module and the enhanced pseudo-label adapter to the semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework includes the step of constructing the enhanced pseudo-label adapter. The step of constructing the enhanced pseudo-label adapter includes the following steps:

[0043] Divide the pseudo-labels into reliable pseudo-labels and uncertain pseudo-labels according to the original pseudo-label adapter of the Efficient Teacher framework;

[0044] Determine the reliable pseudo-labels discarded due to regression errors as auxiliary pseudo-labels;

[0045] Determine the semi-supervised training loss function according to the reliable pseudo-labels, the auxiliary pseudo-labels and the uncertain pseudo-labels;

[0046] The expression of the semi-supervised training loss function is:

[0047]

[0048] Among them, L U is the semi-supervised training loss function, are the classification loss term, regression loss term and objectness loss term during semi-supervised training respectively;

[0049] Among them:

[0050]

[0051] Among them, RPL, APL, and UPL are the reliable pseudo-labels, the auxiliary pseudo-labels, and the uncertain pseudo-labels respectively, The sampling result of the original pseudo-label adapter representing the position (h, w) on the feature map Representing the objectivity score of the pseudo-label at (h, w) Representing the indicator function that outputs 1 when the condition is met and 0 otherwise; CE represents the cross-entropy loss function, CIoU represents the complete intersection between the predicted box and the ground truth box, and X (h,w) Is the output of the student model; cls, reg, and obj represent the classification score, regression score, and objectivity score respectively.

[0052] In some embodiments, the identifying the tree species of street trees by using the tree species identification model includes the following steps:

[0053] Using the tree species identification model to identify the street view images taken on both sides of any street, and obtaining the tree species of the street trees in the street view images.

[0054] To achieve the above object, on the other hand, an embodiment of the present application proposes a device for identifying the tree species of street trees in a street view image, the device includes:

[0055] A first data set acquisition unit, configured to acquire a street view data set of a research area; wherein, the street view data set includes images of street trees on both sides of the street;

[0056] A second data set acquisition unit, configured to acquire a single-tree data set of tree species; wherein, the single-tree data set of tree species includes images of single trees of multiple labeled tree species;

[0057] A model training unit, configured to train a yolo supervision model by using the single-tree data set of tree species to obtain the trained yolo supervision model;

[0058] A framework transformation unit, configured to add a global domain adaptation module and an enhanced pseudo-label adapter to a semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework; wherein, the global domain adaptation module is used to align the features of the street view data set and the single-tree data set of tree species; the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression;

[0059] A model improvement unit, configured to input the street view data set, the single-tree data set of tree species, and the trained yolo supervision model into the semi-supervised domain adaptation framework for training to obtain a tree species identification model;

[0060] A tree species identification unit, configured to identify the tree species of street trees by using the tree species identification model.

[0061] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above method for identifying street tree species in street view images.

[0062] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above method for identifying street tree species in street view images.

[0063] The embodiments of the present application at least include the following beneficial effects:

[0064] The present application can obtain a street view dataset of the study area, where the street view dataset includes images of street trees on both sides of the street; obtain a single-tree species dataset, where the single-tree species dataset includes images of single trees of multiple labeled tree species; use the single-tree species dataset to train a YOLO supervision model to obtain a trained YOLO supervision model; add a global domain adaptation module and an enhanced pseudo-label adapter to a semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework, where the global domain adaptation module is used to align the features of the street view dataset and the single-tree species dataset, and the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression; input the street view dataset, the single-tree species dataset, and the trained YOLO supervision model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model; use the tree species recognition model to identify street tree species. The present application jointly trains using the street view dataset and the single-tree species dataset with labeled tree species, avoiding various challenges in obtaining tree species annotation samples in street view images, and can significantly reduce the cost of sample screening and annotation. In addition, the present application uses global domain adaptation to align the features of the street view dataset and the single-tree species dataset, and uses the enhanced pseudo-label adapter to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression, which can further improve the utilization rate of pseudo-labels, improve the training effect of the tree species recognition model, and thus improve the accuracy of tree species recognition in street view images. Description of the Drawings

[0065] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0066] Figure 1Schematic flowchart of a method for identifying street tree species in street view images provided by an embodiment of the present application;

[0067] Figure 2 Example flowchart of a method for identifying street tree species in street view images provided by an embodiment of the present application;

[0068] Figure 3 Schematic diagram of street view sampling point intervals provided by an embodiment of the present application;

[0069] Figure 4 Flowchart of an improved semi-supervised object detection framework provided by an embodiment of the present application;

[0070] Figure 5 Workflow diagram of a global domain adaptation module provided by an embodiment of the present application;

[0071] Figure 6 Workflow diagram of an enhanced pseudo-label adapter provided by an embodiment of the present application;

[0072] Figure 7 Schematic diagram of the structure of a device for identifying street tree species in street view images provided by an embodiment of the present application;

[0073] Figure 8 Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0074] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application detailed in the appended claims.

[0075] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "when...", or "in response to determining".

[0076] The terms "at least one", "a plurality", "each", "any one", etc. used in this application, "at least one" includes one, two or more than two, "a plurality" includes two or more than two, "each" refers to each one of the corresponding plurality, and "any one" refers to any one of the plurality.

[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0078] Before elaborating on the embodiments of this application in detail, some related technologies that may be involved in the embodiments of this application are described as follows:

[0079] Due to reasons such as high equipment costs or insufficient individual recognition accuracy, it is difficult for the prior art to balance economy and accuracy in large-scale urban street tree assessments and lacks cost-effectiveness. In contrast, street view data, as an openly accessible resource, is undoubtedly the best choice for large-scale urban street tree assessments due to its high resolution, wide coverage, and human-like perspective. In fact, the use of street view data for large-scale urban street tree assessments has made gratifying progress. Research shows that, under the condition of having a sufficient number of high-quality labeled data sets and combined with advanced object detection algorithms, large-scale urban street tree assessments using street view data as the main data source have been successful in some individual cities.

[0080] However, the above-mentioned related technologies still have significant limitations, especially the problem of completely relying on manual screening and annotation of data, which is particularly prominent in the task of street view tree species recognition. This task faces two main challenges. First, street trees usually have a long-tailed distribution characteristic in street views (a few dominant tree species account for the majority), and at the same time, there is a low inter-class variance and a high intra-class variance (the distinction between trees is not obvious), which makes it require a large amount of resources and time to screen street tree species data to be annotated with broad representativeness. Second, the annotation of tree species usually requires professional botanical knowledge, and the background of street view data is complex, and the appearance of trees will change due to factors such as perspective, light, and obstacles, and errors are prone to occur during the annotation process, and the cost of manual annotation is high. Therefore, in the context of large-scale urban street tree assessments, there is an urgent need to propose a new solution to obtain or replace high-quality and widely representative street tree species labeled data at a lower cost.

[0081] The embodiments of the present application provide a method, apparatus, device and medium for identifying street tree species in street view images. The technical solution of the present application includes: obtaining a street view data set of a research area; wherein, the street view data set includes images of street trees on both sides of the street; obtaining a single-tree data set of tree species; wherein, the single-tree data set of tree species includes images of single trees of multiple labeled tree species; training a YOLO supervision model using the single-tree data set of tree species to obtain a trained YOLO supervision model; adding a global domain adaptation module and an enhanced pseudo-label adapter to a semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework; wherein, the global domain adaptation module is used to align the features of the street view data set and the single-tree data set of tree species; the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression; inputting the street view data set, the single-tree data set of tree species and the trained YOLO supervision model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model; using the tree species recognition model to identify street tree species. The present application jointly trains using the street view data set and the single-tree data set of labeled tree species, avoiding various challenges in obtaining tree species annotation samples in street view images, and can significantly reduce the screening and annotation cost of samples; in addition, the present application uses global domain adaptation to align the features of the street view data set and the single-tree data set of tree species, and uses the enhanced pseudo-label adapter to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression, which can further improve the utilization rate of pseudo-labels, improve the training effect of the tree species recognition model, and thus improve the accuracy of tree species recognition in street view images.

[0082] The embodiments of the present application provide a method, apparatus, device and storage medium for identifying street tree species in street view images, which relates to the technical field of image recognition. The method, apparatus, device and medium for identifying street tree species provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application for implementing the knowledge extraction method, etc., but is not limited to the above forms.

[0083] This application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0084] Referring to Figure 1 , an embodiment of this application provides a method for identifying street tree species in street view images. This method may include but is not limited to steps S100 to S150, specifically as follows:

[0085] S100: Obtain a street view data set of the research area; wherein, the street view data set includes images of street trees on both sides of the street.

[0086] Further, S100 may include the following steps S101 to S102:

[0087] S101: Generate street view sampling points on the streets of the research area at a set distance threshold;

[0088] S102: Use a camera to capture images of the street trees on both sides of the street at 90° and 270° directions at each of the street view sampling points;

[0089] The calculation formula for the set distance threshold is:

[0090] T = 2 * A × tan45° - I;

[0091] Wherein, T is the set distance threshold; A is the distance between the camera and both sides of the street; 45° is the field of view range of each street view sampling point; I is the overlapping part of adjacent fields of view.

[0092] S110: Obtain a data set of single tree species; wherein, the data set of single tree species includes images of single trees of multiple labeled tree species.

[0093] Further, S110 may include the following step S111:

[0094] S111: Obtain the single-tree images of multiple labeled tree species from different data open platforms as the single-tree dataset of tree species.

[0095] S120: Use the single-tree dataset of tree species to train the YOLO supervision model to obtain the trained YOLO supervision model.

[0096] Further, S120 may include the following steps S121 to S123:

[0097] S121: Perform data augmentation on the single-tree dataset of tree species to obtain the augmented single-tree dataset of tree species; wherein, the data augmentation includes exposure, light reduction, random cropping, scaling, and mirroring of each single-tree image.

[0098] S122: Use each single-tree image in the augmented single-tree dataset of tree species as a training sample and the labeled tree species as a training label to train the YOLO supervision model to obtain the trained YOLO supervision model.

[0099] S130: Add a global domain adaptation module and an enhanced pseudo-label adapter to the semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework; wherein, the global domain adaptation module is used to align the features of the street view dataset and the single-tree dataset of tree species; the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression.

[0100] Further, S130 includes the step of constructing the global domain adaptation module, and the step of constructing the global domain adaptation module includes the following steps S131 to S133:

[0101] S131: Determine the mixing coefficient; wherein, the mixing coefficient is used to determine the proportion of using the street tree images and the proportion of using the single-tree images when training the semi-supervised domain adaptation framework.

[0102] The calculation formula of the mixing coefficient is:

[0103]

[0104] where α is the mixing coefficient, α min is the minimum value of the mixing coefficient, D is the feature distance, and β represents the hyperparameter that controls the influence degree of the feature distance on the mixing coefficient.

[0105] The calculation formula of the feature distance is:

[0106] D = ‖U S - U T ‖;

[0107] Among them, U S and U T respectively represent the characteristic mean vectors of the source domain and the target domain; the source domain is the single-tree dataset of the tree species, and the target domain is the street view dataset;

[0108] The calculation formula of the characteristic mean vector is:

[0109]

[0110] Among them, U is the characteristic mean vector, N represents the number of images in the dataset, f(x) represents the feature extraction function of the yolo supervision model, and x i represents the image in the dataset;

[0111] S132: Generate a mixed sample according to the mixing coefficient by using the street view dataset and the single-tree dataset of prime tree species;

[0112] The expression of the mixed sample is:

[0113]

[0114] Among them, is the mixed sample, x S is the image of the source domain, and x T is the image of the target domain;

[0115] S133: Construct the global domain adaptation module according to the mixed sample and the domain adaptation loss function;

[0116] The expression of the domain adaptation loss function is:

[0117] L da =-∑ h,w [Blogp(h,w)+(1 - B)log(1 - p(h,w))];

[0118] Among them, L da is the domain adaptation loss function, p(h,w) is the output of the global domain adaptation module, B = 0 represents a labeled image, and B = 1 represents an unlabeled image.

[0119] Furthermore, S130 includes the step of constructing the enhanced pseudo-label adapter, and the step of constructing the enhanced pseudo-label adapter includes the following steps S134 to S136:

[0120] S134: Divide the pseudo-labels into reliable pseudo-labels and uncertain pseudo-labels according to the original pseudo-label adapter of the Efficient Teacher framework;

[0121] S135: Determine the reliable pseudo-labels discarded due to regression errors as auxiliary pseudo-labels;

[0122] S136: Determine a semi-supervised training loss function based on the reliable pseudo-labels, the auxiliary pseudo-labels, and the uncertain pseudo-labels;

[0123] The expression of the semi-supervised training loss function is:

[0124]

[0125] where L U is the semi-supervised training loss function, are the classification loss term, regression loss term, and objectiveness loss term during semi-supervised training, respectively;

[0126] where:

[0127]

[0128] where RPL, APL, and UPL are the reliable pseudo-labels, the auxiliary pseudo-labels, and the uncertain pseudo-labels, respectively, represents the sampling result of the original pseudo-label adapter at position (h, w) on the feature map, represents the objectiveness score of the pseudo-label at (h, w), represents an indicator function that outputs 1 when the condition is met and 0 otherwise; CE represents the cross-entropy loss function, CIoU represents the complete intersection between the predicted box and the ground truth box, and X (h,e) is the output of the student model; cls, reg, and obj represent the classification score, regression score, and objectiveness score, respectively.

[0129] S140: Input the street view dataset, the single-tree species dataset, and the trained yolo supervised model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model.

[0130] S150: Use the tree species recognition model to identify the tree species of street trees.

[0131] Further, S150 may include the following steps S151:

[0132] S151: Use the tree species recognition model to identify the street view images taken on both sides of any street to obtain the tree species of the street trees in the street view images.

[0133] Next, the solution of the embodiment of the present application will be introduced and described in detail with specific application examples.

[0134] Based on the problems existing in the prior art, this embodiment proposes a method for efficiently identifying street tree species in street view images based on semi-supervised domain adaptation. The specific improvement ideas are as follows:

[0135] First, aiming at the problem of difficult screening of street tree species data to be labeled, this embodiment proposes an innovative solution, which uses the tree species image resources of the data open platform to replace the screening of tree species in street view data. This solution not only significantly reduces the time and resource investment required in the traditional screening process of data to be labeled, but also can quickly construct a dataset with broad representativeness.

[0136] Secondly, after obtaining a tree species dataset with broad representativeness from the data open platform, in order to further reduce the labeling cost, this embodiment further screens out single-tree images, and processes the screened single-tree images through data augmentation techniques to simulate the changes that tree species may appear due to real factors such as perspective, light, and occluders in street views. These processing methods have additional benefits. The single-tree images ensure that each image in the dataset can represent the single characteristics of the tree species, thereby improving the accuracy and pertinence of the dataset; data augmentation effectively expands the scale and diversity of the dataset and enhances the generalization ability of the model.

[0137] Finally, to solve the domain differences between the above single-tree datasets of tree species and street view datasets, this embodiment performs cross-domain learning on single-tree images and street view images, and improves the ability of the model to identify tree species in street view data through semi-supervised domain adaptation.

[0138] Exemplarily, Figure 2 It is an example flowchart of a method for identifying street tree species in street view images.

[0139] Specifically, this embodiment may include the following steps:

[0140] Step 1:

[0141] Collect the street view dataset of the study area. In this embodiment, the street network of the study area is used to generate street view sampling points at a set threshold distance (T), and the trees on both sides of the road are detected in the 90° and 270° directions. Figure 3 It is a schematic diagram of the interval of street view sampling points. Since the field of view range of each street view sampling point is 45°, the calculation formula of the threshold T between street view sampling points is shown in Equation (1):

[0142] T = 2 * A × tan45° - I(1)

[0143] Where A is the distance between the camera and the roadside, and I is the overlapping part of adjacent fields of view.

[0144] By retaining the overlap of a small number of trees, ensure that the selected street view dataset can comprehensively cover all street trees and maintain a reasonable perspective.

[0145] Step 2:

[0146] Collect a single-tree dataset of publicly available tree species. By consulting government materials, this embodiment can determine the common street tree species in the research area. Subsequently, this embodiment uses a data open platform to crawl relevant tree species datasets and filters out a single-tree dataset of tree species from them.

[0147] 1) Data augmentation: Perform data augmentation on the single-tree images of tree species, including exposure, dimming, random cropping, scaling, and mirroring methods to simulate the changes of tree species in street views due to factors such as lighting, occlusion, and perspective.

[0148] 2) Train the yolov5 supervised model: Use the augmented single-tree dataset of tree species to train the yolov5 object detection model, learn the features of the trees and perform object localization to obtain the yolo supervised model.

[0149] Step 3:

[0150] Improve the semi-supervised object detection framework Efficient Teacher to obtain a semi-supervised domain adaptation framework. Efficient Teacher (ET) is an efficient semi-supervised object detection framework that focuses on improving the performance of one-stage Anchor-based detectors (yolov5), while taking into account the detection efficiency and the consistency of pseudo-label quality.

[0151] As Figure 4 shown in the flowchart, improving the semi-supervised object detection framework can be divided into two steps. One is data processing, which generates unlabeled street view data and labeled tree species pictures respectively. Among them, the labeled tree species pictures are used to train a yolo supervised model, and the specific implementation refers to Step 1 and Step 2. The other is model training, which inputs the unlabeled street view data, labeled tree species pictures, and the yolo supervised model into the semi-supervised domain adaptation framework to achieve knowledge transfer from the source domain to the target domain. Among them, the semi-supervised domain adaptation framework is optimized based on the open-source semi-supervised object detection framework Efficient Teacher by introducing a Global Domain Adaptor and an Enhanced Pseudo Label Assigner. The following is a detailed elaboration of the two improved modules. Among them, Figure 4Mosaic in it represents an augmentation technique that stitches multiple images into a new training sample; Strong represents strong data augmentation, that is, it includes multiple augmentation strategies; EMA (Exponential Moving Average) is a technique commonly used to smooth data and is used to smooth the generation process of pseudo-labels in semi-supervised object detection; in the one-stage object detection algorithm yolo, Backbone is the feature extraction network, Neck is the feature fusion module, and Head is the prediction layer that outputs the detection results.

[0152] 1) Global Domain Adaptor:

[0153] The Efficient Teacher (ET) framework introduces the Burn-in concept, that is, in the initial stage of training, the model will still undergo a period of warm-up supervised training, and semi-supervised training will only start after the warm-up ends. No pseudo-labels are generated during this stage, but the model is adapted to unlabeled data through domain adaptation, with the aim of stabilizing the training of the model in the early stage. However, the above method in the original framework is mainly local and limited domain adaptation, which is reflected in the calculation of the supervised training loss function L S The original ET framework performs domain adaptation through weighted loss, and the specific calculation formula is shown in Equation (2):

[0154]

[0155] where CE represents the cross-entropy loss function, CIoU represents the complete intersection between the predicted box and the GT box, X (h,w) is the output of the student model, Y (h,w) represents the sampling result generated by the yolo detector, cls, reg, and obj represent the classification score, regression score, and objectness score respectively, and L da is the domain adaptation loss function, λ is a hyperparameter that controls the contribution of domain adaptation, and the ET framework is set to 0.1.

[0156] The original domain adaptation scheme of the ET framework simply sums weights by fixing λ, which obviously cannot meet the cross-domain requirements of this experiment. Therefore, as Figure 5 shown in this embodiment, a global domain adaptation improvement scheme is proposed. For the convenience of description, this embodiment defines the labeled single-tree species dataset as the source domain and the unlabeled street view dataset as the target domain. In this scheme, this embodiment introduces the Mixup data mixing technology and dynamically adjusts the mixing coefficient α through the feature distance between the source domain and the target domain to smoothly achieve the global alignment of the source domain and the target domain. Among them, Figure 5Among them, Burn-In represents the preheating supervised training for a period of time at the initial stage of training of the semi-supervised domain adaptation framework; Mixup represents the enhancement technology of generating new samples by mixing data; respectively represent the classification loss term, regression loss term, and objectness loss term during supervised training; L da represents the domain adaptation loss term during supervised training. Specifically, it can include the following steps:

[0157] (1) At the beginning of training, the mixing coefficient α is fixed at 1, and only the source domain is used for training.

[0158] (2) As training progresses, the target domain is gradually introduced into training. The core idea is to dynamically adjust the value of the mixing coefficient α by calculating the distance D between the features of the source domain and the target domain, so as to achieve a smooth transition between the source domain and the target domain data.

[0159] Calculation of feature mean: In each epoch, calculate the mean vectors of the features of the source domain and the target domain respectively. The calculation formula of the mean vector U is shown in Equation (3):

[0160]

[0161] Among them, N represents the number of data sets, f(x) represents the feature extraction function of the yolo detector, and x i represents the sample.

[0162] Calculation of feature distance: Use the Euclidean distance to measure the difference between the feature means of the source domain and the target domain. The calculation formula is shown in Equation (4):

[0163] D = ‖U S -U T ‖ (4)

[0164] Among them, D represents the feature distance, and U S 、U T respectively represent the feature mean vectors of the source domain and the target domain.

[0165] Calculation of the mixing coefficient α: When D is large, the feature difference between the source domain and the target domain is large. To avoid introducing too much noise, the model should be mainly trained with the source domain. At this time, α is close to 1, and the proportion of the target domain data introduced is low. When D is small, the feature distributions of the source domain and the target domain tend to be consistent, indicating that the timing of introducing the target domain data is ripe. At this time, α gradually decreases, and the proportion of the target domain data gradually increases, enabling the model to make a smooth transition and achieve domain adaptation. The calculation formula of the mixing coefficient α is shown in Equation (5):

[0166]

[0167] Among them, β represents the hyperparameter that controls the influence degree of the feature distance D on α.

[0168] (3) Generate mixed samples: In this embodiment, the source domain and target domain samples are weighted and mixed by controlling the Mixup coefficient to generate new training samples. The purpose is to provide a smooth way to introduce target domain data globally, thereby optimizing the cross-domain learning process. The calculation formula of the mixed samples is shown in Equation (6):

[0169]

[0170] where x S and x T are the mixed samples, source domain samples, and target domain samples respectively.

[0171] (4) Calculate the domain adaptation loss function L da : The domain adaptation technology with a classifier is adopted to confuse the ability of the detector to distinguish two types of data. The calculation formula is shown in Equation (7):

[0172] L da = -∑ h,w [Blogp(h,w)+(1 - B)log(1 - p(h,w))](7)

[0173] where p(h,w) is the output of the domain classifier. B = 0 represents labeled data, and B = 1 represents unlabeled data.

[0174] (5) Through dynamic mixing and gradient reversal, the model can smoothly introduce the target domain data features (as Figure 5 shown) during the global training process, while avoiding the training instability caused by the excessive difference in feature distributions between the source domain and the target domain, and finally achieving better recognition performance on the target domain.

[0175] 2) Enhanced Pseudo Label Assigner:

[0176] The original pseudo label adapter PLA of the Efficient Teacher (ET) framework designs a soft loss to handle uncertain pseudo labels. PLA classifies pseudo labels into reliable and uncertain categories according to high and low thresholds. The uncertain pseudo labels are assigned as soft labels to participate in the calculation of regression and target loss, which enables PLA to be used to optimize the pseudo labels with accurate regression but inaccurate classification.

[0177] As Figure 6As shown in the figure, in this embodiment, an Enhanced Pseudo-Label Adapter (EPLA) is proposed based on the PLA mechanism. The pseudo-labels are more finely divided into reliable pseudo-labels, auxiliary pseudo-labels, and uncertain pseudo-labels. The auxiliary pseudo-labels are used to fallback the reliable pseudo-labels discarded due to regression errors to participate in the calculation process of the classification loss. EPLA can not only optimize the pseudo-labels with accurate regression but inaccurate classification, but also effectively handle the pseudo-labels with accurate classification but inaccurate regression.

[0178] Improvement background of EPLA: The ET framework focuses on improving the performance of yolov5. The anchor matching strategy of yolov5 changes from iou matching to shape matching. This change will regard the prediction boxes that do not match the anchor as the background and discard them. However, in PLA, this processing method obviously does not efficiently utilize the reliable pseudo-labels. In actual applications, the number of reliable pseudo-labels is already small. In the effective semi-supervised learning process, the proportion of prediction boxes that do not match the anchor is as high as 30%. This excessive discarding approach leads to a large amount of useful information being ignored, affecting the effective utilization of pseudo-labels.

[0179] (1) The semi-supervised training loss function L U in EPLA is calculated as follows:

[0180]

[0181] where respectively represent the classification loss term, regression loss term, and objectness loss term during semi-supervised training. RPL, APl, and UPL represent reliable pseudo-labels, auxiliary pseudo-labels, and uncertain pseudo-labels respectively. represents the PLA sampling result at position (h, w) on the feature map. represents the objectiveness score of the pseudo-label at (h, w). represents the indicator function, which outputs 1 when the condition is met and 0 otherwise.

[0182] (2) The differences between EPLA and PLA. By introducing auxiliary pseudo-labels, EPLA retains the reliable pseudo-labels with poor regression effects for the calculation of the classification loss during semi-supervised training, making the classification of cross-entropy during semi-supervised training also replaced by soft labels (as Figure 6 shown). It realizes the use of soft loss to process all pseudo-labels during the semi-supervised training process, and expands the function of the pseudo-label allocator. It not only retains the optimization of the pseudo-labels with accurate regression but inaccurate classification in PLA, but also further realizes the purpose of converting the pseudo-labels with good classification but insufficient regression into true positives.

[0183] Step 4:

[0184] Obtain a model suitable for identifying tree species in street view data. To verify the effectiveness of the semi-supervised domain adaptation framework, in this embodiment, some street view validation sets are manually screened and labeled to verify the accuracy of the model, and horizontal comparison and ablation experiments are conducted. The verification results are shown in Table 1:

[0185] Table 1 Horizontal Comparison of Experiments

[0186] Model mAP0.5 Study area (Choi et al.,2022) 0.564 City 1 (Branson et al.,2018) 0.581 City 2 (Liu et al,2023) 0.587 City 3 The model of this embodiment 0.613 City 4

[0187] (Choi et al., 2022), (Branson et al., 2018), (Liu et al, 2023), etc. are existing studies on the classification of street tree species using SVI as the only data source. By comparing with existing solutions for identifying street tree species in street view images, the superiority of the method in this embodiment is demonstrated. mAP@0.5 (mean Average Precision at IoU threshold 0.5) is a commonly used evaluation metric in object detection, which represents the average precision of the model across all classes when the Intersection over Union (IoU) threshold is 0.5. That is, only when the IoU between the predicted bounding box and the ground truth bounding box is greater than 0.5, the prediction is considered correct.

[0188] Table 2 Ablation Experiments

[0189]

[0190] Each method in the ablation experiments shown in Table 2 demonstrates the gradual optimization of the framework in this embodiment in different aspects. The finally combined semi-supervised domain adaptation framework achieved the highest mAP50 value of 61.2 (+2.6 improvement). mAP@50 (mean Average Precision at IoU threshold 50) is an evaluation metric in object detection, which represents the average precision (AP) across all classes when the IoU threshold is 50%. That is, only when the IoU between the predicted bounding box and the ground truth bounding box is greater than or equal to 50%, the prediction is considered correct.

[0191] The semi-supervised domain adaptation framework proposed in this embodiment, by effectively combining labeled data and unlabeled data, not only significantly reduces the sample acquisition cost of street tree species recognition, but also maintains a high accuracy rate, achieving efficient recognition of street tree species and having strong practical application potential. Therefore, the beneficial effects of this embodiment include:

[0192] (1) Reduce the screening cost of street view samples: In this embodiment, through the semi-supervised domain adaptation framework, high-quality samples can be effectively screened out from a large amount of unlabeled street view data for training. Researchers only need to focus on publicly available tree species images that are easy to obtain. This embodiment not only reduces the manual intervention in data screening, but also improves the automation of the screening process through the self-training mechanism, thus significantly reducing the screening cost of street view samples.

[0193] (2) Reduce the sample annotation cost: In this embodiment, by combining a small amount of labeled data with a large amount of unlabeled data, the framework reduces the quantity and difficulty of the annotation tasks. With the help of pseudo-label generation and self-training mechanisms, the need for manual annotation is reduced, significantly reducing the annotation cost.

[0194] (3) Improve the recognition accuracy: The semi-supervised domain adaptation framework of this embodiment improves the accuracy of street view tree species recognition through knowledge transfer between the source domain and the target domain.

[0195] (4) Identify more tree species: By enhancing the generalization ability and adaptability of the model, the framework can identify more tree species. Even if tree species that have not been seen before appear in the target domain (unlabeled street view data), the model can effectively identify them using the learned features and extend to more tree species recognition tasks.

[0196] Refer to Figure 7 , this application embodiment also provides a device for identifying street tree species in street view images, which can implement the method for identifying street tree species in street view images described above. The device includes:

[0197] The first dataset acquisition unit is used to acquire the street view dataset of the study area; wherein, the street view dataset includes the images of street trees on both sides of the street;

[0198] The second dataset acquisition unit is used to acquire the single-tree dataset of tree species; wherein, the single-tree dataset of tree species includes the images of single trees of multiple labeled tree species;

[0199] The model training unit is used to train the YOLO supervised model using the single-tree dataset of tree species to obtain the trained YOLO supervised model;

[0200] The framework transformation unit is used to add a global domain adaptation module and an enhanced pseudo-label adapter to the semi-supervised object detection framework to obtain a semi-supervised domain adaptation framework; wherein, the global domain adaptation module is used to align the features of the street view dataset and the single-tree dataset of tree species; the enhanced pseudo-label adapter is used to optimize the pseudo-labels with accurate regression but inaccurate classification and the pseudo-labels with accurate classification but inaccurate regression;

[0201] A model improvement unit for inputting the street view dataset, the single-tree species dataset, and the trained YOLO supervision model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model;

[0202] A tree species recognition unit for recognizing the tree species of street trees by using the tree species recognition model.

[0203] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0204] An embodiment of the present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above method for recognizing the tree species of street trees in a street view image. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0205] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0206] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0207] A processor 801, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0208] A memory 802, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the method for recognizing the tree species of street trees in a street view image according to an embodiment of the present application;

[0209] An input / output interface 803 for implementing information input and output;

[0210] A communication interface 804 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0211] A bus 805 for transmitting information between various components of the device (such as a processor 801, a memory 802, an input / output interface 803, and a communication interface 804);

[0212] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.

[0213] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above method for identifying tree species of roadside trees in street view images.

[0214] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0215] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0216] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0217] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0218] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0220] It should be understood that in this application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0221] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0222] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0223] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0224] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.

[0226] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A method for identifying roadside tree species in street view images, characterized in that: The method comprises the following steps: Obtain a street view dataset of the study area; wherein the street view dataset includes images of roadside trees on both sides of the street; Acquire a tree species single tree dataset; wherein the tree species single tree dataset includes single tree images of multiple labeled tree species; Using the tree species single plant data set to train a YOLO supervised model to obtain the trained YOLO supervised model; A global domain adaptation module and an enhanced pseudo-label adapter are added to a semi-supervised target detection framework to obtain a semi-supervised domain adaptation framework; wherein the global domain adaptation module is used to align the features of the street view dataset and the tree species dataset; the enhanced pseudo-label adapter is used to optimize pseudo-labels with accurate regression but inaccurate classification and pseudo-labels with accurate classification but inaccurate regression; Inputting the street view dataset, the tree species single tree dataset and the trained YOLO supervised model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model; The tree species identification model is used to identify the street tree species.

2. The method for identifying roadside tree species in street view images according to claim 1, characterized in that: The step of obtaining the street view dataset of the study area includes the following steps: Generate street view sampling points at a set distance threshold on the streets of the study area; At each of the street scene sampling points, use a camera to capture images of the roadside trees on both sides of the street at 90° and 270° directions; The calculation formula for setting the distance threshold is: T = 2*A×tan45°-I; Wherein, T is the set distance threshold; A is the distance between the camera and both sides of the street; 45° is the field of view of each street view sampling point; and I is the overlapping part of adjacent fields of view.

3. The method for identifying roadside tree species in street view images according to claim 1, characterized in that: The step of obtaining a tree species individual tree dataset comprises the following steps: The single tree images of multiple labeled tree species are obtained from different data open platforms as the single tree dataset of the tree species.

4. The method for identifying roadside tree species in street view images according to claim 1, characterized in that: The method of training the YOLO supervised model using the tree species single plant data set to obtain the trained YOLO supervised model includes the following steps: Performing data enhancement on the tree species single tree dataset to obtain an enhanced tree species single tree dataset; wherein the data enhancement includes exposing, reducing light, randomly cropping, scaling and mirroring each of the single tree images; The YOLO supervised model is trained using each of the individual tree images in the enhanced tree species individual tree data set as a training sample and the labeled tree species as a training label to obtain the trained YOLO supervised model.

5. The method for identifying roadside tree species in street view images according to claim 1, characterized in that: The step of adding the global domain adaptation module and the enhanced pseudo-label adapter to the semi-supervised target detection framework to obtain the semi-supervised domain adaptation framework includes the step of constructing the global domain adaptation module. The step of constructing the global domain adaptation module includes the following steps: Determining a mixing coefficient; wherein the mixing coefficient is used to determine the proportion of the street tree images and the proportion of the single tree images used when training the semi-supervised domain adaptation framework; The calculation formula of the mixing coefficient is: Wherein, α is the mixing coefficient, α min is the minimum value of the mixing coefficient, D is the characteristic distance, and β represents a hyperparameter that controls the influence of the characteristic distance on the mixing coefficient; The calculation formula of the characteristic distance is: D=||U S -U T ||; Among them, U S , U T Respectively represent the feature mean vectors of the source domain and the target domain; the source domain is the tree species single tree dataset, and the target domain is the street view dataset; The calculation formula of the characteristic mean vector is: Where U is the feature mean vector, N represents the number of images in the data set, f(x) represents the feature extraction function of the YOLO supervised model, and x i represents an image in the dataset; Generate a mixed sample using the street view dataset and the prime tree species single tree dataset according to the mixing coefficient; The expression of the mixed sample is: in, is the mixed sample, x S is the image in the source domain, x T is an image of the target domain; Constructing the global domain adaptation module according to the mixed sample and the domain adaptation loss function; The expression of the domain adaptation loss function is: L da =-∑ h,w [B log p(h,w)+(1-B)log(1-p(h,w))|; Among them, L da is the domain adaptation loss function, p(h, w) is the output of the global domain adaptation module, B=0 represents a labeled image, and B=1 represents an unlabeled image.

6. The method for identifying roadside tree species in street view images according to claim 1, characterized in that: The step of adding the global domain adaptation module and the enhanced pseudo-label adapter to the semi-supervised target detection framework to obtain the semi-supervised domain adaptation framework includes the step of constructing the enhanced pseudo-label adapter. The step of constructing the enhanced pseudo-label adapter includes the following steps: According to the original pseudo-label adapter of the Efficient Teacher framework, pseudo-labels are divided into reliable pseudo-labels and uncertain pseudo-labels; Determine the reliable pseudo-labels discarded due to regression errors as auxiliary pseudo-labels; Determine a semi-supervised training loss function according to the reliable pseudo-label, the auxiliary pseudo-label and the uncertain pseudo-label; The expression of the semi-supervised training loss function is: Among them, L U is the semi-supervised training loss function, They are the classification loss term, regression loss term, and object loss term during semi-supervised training; in: Among them, RPL, APL, and UPL are the reliable pseudo-label, the auxiliary pseudo-label, and the uncertain pseudo-label, respectively. represents the sampling result of the original pseudo-label adapter at position (h,w) on the feature map, represents the objectivity score of the pseudo-label at (h,w), represents the indicator function, outputting 1 when the condition is met, otherwise outputting 0; CE represents the cross entropy loss function, CIoU represents the complete intersection between the predicted box and the true box, and X (h,w) is the output of the student model; cls, reg, and obj represent the classification score, regression score, and target score, respectively.

7. A method for identifying roadside tree species in street view images according to any one of claims 1 to 6, characterized in that: The method of identifying roadside tree species using the tree species identification model comprises the following steps: The tree species recognition model is used to recognize street view images taken on both sides of any street to obtain the street tree species in the street view images.

8. A device for identifying roadside tree species in street view images, characterized in that: The device comprises: A first data set acquisition unit is used to acquire a street view data set of a research area; wherein the street view data set includes images of roadside trees on both sides of the street; A second data set acquisition unit is used to acquire a tree species single tree data set; wherein the tree species single tree data set includes single tree images of multiple labeled tree species; A model training unit, used to train a YOLO supervised model using the tree species single plant data set to obtain the trained YOLO supervised model; A framework transformation unit, used for adding a global domain adaptation module and an enhanced pseudo-label adapter to a semi-supervised target detection framework to obtain a semi-supervised domain adaptation framework; wherein the global domain adaptation module is used for aligning the features of the street view dataset and the tree species dataset; the enhanced pseudo-label adapter is used for optimizing pseudo-labels with accurate regression but inaccurate classification and pseudo-labels with accurate classification but inaccurate regression; A model improvement unit, used for inputting the street view dataset, the tree species single tree dataset and the trained YOLO supervised model into the semi-supervised domain adaptation framework for training to obtain a tree species recognition model; The tree species identification unit is used to identify the street tree species using the tree species identification model.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements a method for identifying street tree species in street view images as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a method for identifying roadside tree species in a street view image is implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Component identification and positioning method and system containing local feature constraint, and storage medium

    CN115775278A

  • Road scene semantic segmentation method for unsupervised traffic element alignment and related device

    CN117710679A

  • Semi-supervised medical image segmentation method and system based on visual language model

    CN118115516A

  • Using image pre-processing to generate a machine learning model

    US20200210769A1