Visibility value detection method, device, apparatus and storage medium
By using neural network feature extraction and mapping relationships to determine the visibility value of the image to be detected, the problem of inaccurate visibility value determination in existing technologies is solved, and fine-grained visibility detection is achieved under weather conditions such as fog and dust storms, thereby improving the accuracy of traffic safety management.
Patent Information
- Application Number
- CN202111656430.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-12-30
AI Technical Summary
Existing technologies struggle to accurately and efficiently determine visibility values in areas under testing, especially in smoggy or dusty weather conditions, leading to significant traffic safety hazards.
By using a pre-defined mapping relationship between representation quantities and visibility values, a pre-trained neural network is used to extract features from the image to be detected, and determine the scene visibility values contained in the image. This includes optimizing and training the feature extraction network, the representation quantity determination network, and the output layer to obtain fine-grained visibility values.
It can accurately determine the visibility value of the image under test without the need for specialized visibility testing instruments, providing more granular visibility information, which facilitates traffic management and early warning.
Smart Images

Figure CN114359211B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for detecting visibility values. Background Technology
[0002] Fog, haze, and dust storms can all reduce visibility, impacting transportation. For example, common patchy fog, influenced by localized microclimates, occurs within a larger fog area, resulting in even lower visibility. When patchy fog forms, visibility drops suddenly and drastically, making it difficult to predict and forecast, posing a significant threat to road safety, especially on highways, where it can easily lead to serious accidents. Therefore, identifying the visibility value of a monitored area can help determine the impact of fog, dust storms, and other similar weather phenomena, thus guiding traffic decisions. Therefore, it is necessary to provide a convenient and accurate method for determining the visibility value of a monitored area. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for detecting visibility values.
[0004] According to a first aspect of the present disclosure, a visibility value detection method is provided, the method comprising:
[0005] Acquire the image to be detected;
[0006] Feature extraction is performed on the image to be detected to obtain the representational value of the image to be detected;
[0007] Based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected, the visibility value of the image to be detected is determined; wherein, the representation quantity is used to represent the magnitude of the visibility value of the scene contained in the image.
[0008] In some embodiments, feature extraction is performed on the image to be detected, including:
[0009] Feature extraction is performed on the image to be detected using a pre-trained neural network; wherein the neural network is trained in the following manner:
[0010] Acquire sample image pairs carrying label information, wherein the label information is used to indicate the relationship between the visibility values of two frames in the sample image pair;
[0011] The preset initial neural network outputs the prediction result of the relationship between the visibility values of the two frames in the sample image pair;
[0012] Based on the difference between the prediction result and the label information, the network parameters of the initial neural network are adjusted to obtain the neural network.
[0013] In some embodiments, the initial neural network includes a feature extraction network, a representation determination network, and an output layer;
[0014] The step of outputting the prediction result of the visibility value relationship between the two frames in the sample image pair through a preset initial neural network includes:
[0015] The feature extraction network is used to extract features from two frames of the sample image pair.
[0016] The network determines the representation quantity of each of the two frames in the sample image pair based on the features extracted by the feature extraction network using the representation quantity determination network.
[0017] The output layer determines the relationship between the visibility values of the two frames in the sample image pair based on their respective representation values.
[0018] In some embodiments, the characterization determination network includes a first network branch and a second network branch;
[0019] The step of determining the representation quantities of each of the two frames in the sample image pair based on the features extracted by the feature extraction network through the representation quantity determination network includes:
[0020] The first network branch determines the representation quantity of the first image based on the features extracted from the first image in the sample image pair;
[0021] The second network branch determines the representation of the second image based on the features extracted from the second image in the sample image pair.
[0022] In some embodiments, the sample image pairs are obtained based on the following method:
[0023] At least two first image sets are obtained, wherein the visibility value range of each image in each first image set is the same, and the visibility value ranges of each image in any two first image sets do not overlap.
[0024] One frame is selected from each of the two first image sets to form the sample image pair.
[0025] In some embodiments, the representation quantity is a one-dimensional scalar, and the mapping relationship between the representation quantity and the visibility value is determined based on the following:
[0026] Obtain at least one second image set, wherein the visibility values of each image in each second image set are within the same range;
[0027] Based on the determined representation quantities of each image in each of the second image sets, the distribution range of the representation quantities of each image in each of the second image sets is obtained;
[0028] The mapping relationship is determined based on the visibility value range of each image in each of the second image sets and the distribution range of the characterization quantity; or
[0029] Obtain a third image set, wherein the visibility value of each image in the third image set is known;
[0030] The mapping relationship is determined based on the characterization of each image in the determined third image set and the visibility value of each image in the third image set.
[0031] In some embodiments, determining the mapping relationship based on the visibility value range of each image in each of the second image sets and the distribution range of the representation quantity includes:
[0032] Based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, the representation quantity corresponding to the target visibility value is determined, wherein the visibility value range of each image in two adjacent second image sets is continuous and separated by the target visibility value;
[0033] The mapping relationship is determined based on the representation quantity corresponding to the target visibility value.
[0034] In some embodiments, the representation quantity corresponding to the target visibility value is determined based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, including:
[0035] Based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, the initial representation quantity corresponding to the target visibility value is determined;
[0036] The initial representation quantity is adjusted, and the images in two adjacent second image sets are classified based on the adjusted representation quantity until the accuracy of the classification results of the two adjacent second image sets reaches the maximum value.
[0037] The adjusted representation value is used as the representation value corresponding to the target visibility value.
[0038] In some embodiments, the representation of the image to be detected is obtained by extracting features from the image to be detected through a pre-trained neural network, wherein the neural network includes a feature extraction network, a first network branch, and a second network branch.
[0039] Feature extraction is performed on the image to be detected using a pre-trained neural network to obtain the representational parameters of the image to be detected, including:
[0040] The feature extraction network extracts features from the image to be detected, and the extracted features are input into either the first network branch or the second network branch to obtain the representation of the image to be detected; or
[0041] The feature extraction network of the neural network is used to extract features from the image to be detected, and the extracted features are input into the first network branch to obtain the first representation quantity;
[0042] The extracted features are input into the second network branch to obtain the second representation quantity;
[0043] The characterization of the image to be detected is obtained based on the first characterization and the second characterization.
[0044] In some embodiments, the image to be detected includes an image of a road area, and the method further includes:
[0045] Based on the visibility values of the road areas contained in the image to be detected, the hazard level of a specific climate is determined;
[0046] Traffic within the road area is controlled according to the control strategy corresponding to the hazard level.
[0047] In some embodiments, the image to be detected includes multiple frames of road areas acquired at different times, and the method includes: predicting a specific climate change trend within the road area based on the changing trend of visibility values in the multiple frames of the road area; or
[0048] The image to be detected includes images of road areas collected at preset time intervals. The method includes: determining that a specific climate has occurred in the road area in response to the visibility value corresponding to the image exceeding a preset threshold; and determining the frequency of the specific climate in the road area based on the total number of times the specific climate has occurred in the road area within a target time period.
[0049] According to a second aspect of the present disclosure, a visibility value detection device is provided, the device comprising:
[0050] The acquisition module is used to acquire the image to be detected;
[0051] The representation module is used to extract features from the image to be detected using a pre-trained neural network to obtain the representational quantity of the image to be detected.
[0052] The visibility value determination module is used to determine the visibility value corresponding to the image to be detected based on a pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected. The representation quantity is obtained by extracting features from the image through the neural network and is used to represent the magnitude of the visibility value of the scene contained in the image.
[0053] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device including a processor, a memory, and computer instructions stored in the memory that are executable by the processor, wherein when the processor executes the computer instructions, it can implement the method mentioned in the first aspect above.
[0054] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing computer instructions that, when executed, implement the method mentioned in the first aspect above.
[0055] In this embodiment, a mapping relationship between a representation quantity used to characterize the visibility value of a scene contained in an image and the visibility value can be pre-calibrated. When determining the visibility value of a scene contained in an image to be detected, feature extraction can be performed on the image to be detected to obtain the representation quantity of the image to be detected. Then, based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected, the visibility value of the scene contained in the image to be detected is determined. In this way, the visibility value of the image to be detected can be determined without the aid of a dedicated visibility testing instrument. Compared with directly determining the range of visibility values, this embodiment can obtain a more granular visibility value, which is convenient for subsequent applications.
[0056] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0057] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.
[0058] Figure 1 This is a flowchart of a visibility value detection method according to an embodiment of the present disclosure.
[0059] Figure 2 This is a schematic diagram of a detection method for determining visibility values according to an embodiment of the present disclosure.
[0060] Figure 3 This is a schematic diagram of an application scenario according to an embodiment of this disclosure.
[0061] Figure 4 This is a schematic diagram of a neural network structure according to an embodiment of the present disclosure.
[0062] Figure 5 This is a schematic diagram of the logic structure of a road visibility value detection device according to an embodiment of the present disclosure.
[0063] Figure 6 This is a schematic diagram of the logical structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0064] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0065] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Additionally, the term “at least one” herein means any combination of at least two of any one or more of a plurality.
[0066] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0067] To enable those skilled in the art to better understand the technical solutions in the embodiments of this disclosure, and to make the above-mentioned objectives, features and advantages of the embodiments of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings.
[0068] Foggy weather, smog, and dust storms can reduce visibility, increasing the risk of traffic accidents. Typically, the level of smog or dust in a monitored area can be determined by identifying its visibility value, thus guiding traffic decisions. Currently, some methods for determining visibility in a monitored area involve directly measuring the visibility value using a visibility meter. This method requires deploying specialized detection equipment within the monitored area, which is costly and difficult to implement. Other methods determine visibility based on images of the monitored area. For example, images of the monitored area can be input into a pre-trained neural network to predict visibility. However, due to the significant difficulty in accurately calibrating the visibility values of sample images and the scarcity of a large number of images carrying visibility value labels, the calibration of sample images used to train neural networks is usually based on human experience to estimate the visibility range of the regions included in the sample images, such as less than 50m, less than 100m, or greater than 1000m. As a result, when the neural network predicts the image of the region to be detected, it can only predict a rough range of visibility values and cannot obtain precise visibility values. Furthermore, the visibility range of sample images estimated by human experience is affected by subjective factors, leading to inaccurate prediction results.
[0069] Based on this, this disclosure provides a visibility value detection method. It can pre-calibrate the mapping relationship between the representation quantity used to characterize the visibility value of the scene contained in the image and the visibility value. When it is necessary to determine the visibility value of the scene contained in the image to be detected, feature extraction can be performed on the image to be detected to obtain the representation quantity of the image to be detected. Then, based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected, the visibility value of the scene contained in the image to be detected is determined.
[0070] In this way, the visibility value of the image to be tested can be determined without the need for a specialized visibility value tester. Compared to related technologies that can only determine a rough range of visibility values, the embodiments of this disclosure can also obtain more granular visibility values, which is convenient for subsequent applications.
[0071] The method provided in this disclosure can be executed by various devices, such as cloud servers and user terminals. This method can be used to predict visibility values under weather conditions such as fog, haze, and dust storms.
[0072] Specifically, such as Figure 1 As shown, the method for determining visibility values in this embodiment of the disclosure may include the following steps:
[0073] S102. Obtain the image to be detected;
[0074] In step S102, an image to be detected can be acquired. The image to be detected can be an image of the area whose visibility value needs to be detected. The image to be detected may include scenes such as fog, dust, and haze.
[0075] S104. Perform feature extraction on the image to be detected to obtain the representational quantity of the image to be detected;
[0076] In step S104, feature extraction can be performed on the image to be detected to obtain a representation quantity that characterizes the visibility value of the scene contained in the image. When extracting features from the image to be detected, a pre-trained neural network can be used, or other methods can be employed, as long as the representation quantity used to characterize the visibility value of the scene contained in the image can be determined. For example, the neural network can be trained using sample images, enabling it to accurately extract features related to the visibility value of the image and output a representation quantity that characterizes the visibility value. This representation quantity can be various information representing the visibility value, such as vectors, matrices, or specific numerical values. In some embodiments, to facilitate the mapping between the representation quantity and the visibility value, the representation quantity can be a one-dimensional scalar.
[0077] S106. Determine the visibility value of the image to be detected based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected, wherein the representation quantity is used to represent the magnitude of the visibility value of the scene contained in the image.
[0078] In step S106, after obtaining the representational quantity of the image to be detected, the representational quantity of the image to be detected can be mapped to a visibility value based on a pre-calibrated mapping relationship between the representational quantity and the visibility value, thereby obtaining the visibility value of the scene contained in the image to be detected. The representational quantity mentioned in this embodiment can be used to represent the visibility value of the scene contained in the image. When calibrating the mapping relationship, features can be extracted from images with known visibility values to obtain the representational quantities of these images, and the mapping relationship can be constructed based on the correspondence between the representational quantities and the visibility values of these images. Alternatively, since it is difficult to calibrate the accurate visibility value corresponding to an image, features can be extracted from images with known visibility value ranges to determine the distribution of the representational quantities of these images, and the mapping relationship can be constructed based on the distribution range of the representational quantities and the visibility value range of these images. There are many specific ways to calibrate the mapping relationship, as long as the mapping relationship can accurately reflect the correspondence between the representational quantity and the visibility value, this embodiment does not impose any limitations.
[0079] By means of the embodiments disclosed herein, the accurate visibility value of the scene contained in the image can be predicted, the visibility of the image can be characterized in a more granular way, and subsequent applications can be facilitated.
[0080] In some embodiments, features can be extracted from the image to be detected using a pre-trained neural network to obtain the representational quantity of the image. For example... Figure 2 As shown, a mapping relationship between representation quantities and visibility values can be constructed based on a pre-trained neural network. Then, this neural network is used to determine the representation quantity of the image to be detected, and subsequently, the visibility value of the image to be detected is determined based on the representation quantity and the mapping relationship. Since accurately labeling the visibility values of sample images is difficult, it is challenging to obtain a large number of samples carrying visibility value labels for training the neural network. To ensure that the trained neural network can accurately extract features from images and output a representation quantity indicating the magnitude of the scene visibility values contained in the image, in some embodiments, sample image pairs carrying label information can be obtained during neural network training. This label information indicates the magnitude relationship of the visibility values of the two sample images in the sample image pair. The sample image pair is then input into a preset initial neural network, which outputs a prediction result of the magnitude relationship of the visibility values of the two sample images in the sample image pair. Based on the difference between this prediction result and the true result indicated by the label information, the network parameters of the initial neural network are continuously adjusted to train the neural network. For example, we can obtain sample image A and sample image B, where the visibility value of sample image A is greater than that of sample image B. Then, we use an initial neural network to output the probability that sample image A is the image with the larger visibility value. This probability is then compared with the true probability that sample image A is the image with the larger visibility value to determine the deviation. Based on this deviation, we continuously adjust the parameters of the initial neural network until it converges, resulting in a trained neural network. In this way, we can train a neural network that can accurately extract the representation of the magnitude of image visibility values without obtaining sample images labeled with visibility values.
[0081] In some embodiments, the preset initial neural network may include a feature extraction network, a representation determination network, and an output layer. When outputting the prediction result of the visibility value relationship between two frames in a sample image pair through the preset initial neural network, features can first be extracted from the two frames in the sample image pair by the feature extraction network. Then, the representation determination network determines the representation values of each of the two frames in the sample image pair based on the features extracted by the feature extraction network. Finally, the output layer determines the visibility value relationship between the two frames based on the representation values of each of the two frames in the sample image pair.
[0082] Since the representation output by the neural network needs to reflect the visibility values of the scene contained in the image as realistically as possible, the visibility value obtained based on the representation mapping will be accurate. To make the representation output by the neural network as accurate as possible, the applicant optimized the structure of the neural network to improve its performance. In some embodiments, the representation determination network may include two network branches, namely a first network branch and a second network branch. When predicting the relationship between the visibility values of two sample images in a preset initial neural network output sample image pair, features can first be extracted from the first image in the sample image pair by a feature extraction network, and the extracted features are input into the first network branch. The network branch obtains the representation of the first image based on the extracted features. Then, features are extracted from the second image in the sample image pair by a feature extraction network, and the extracted features are input into the second network branch. The second network branch obtains the representation of the second image based on the extracted features. Finally, the prediction result is determined based on the representation of the first image and the representation of the second image. After using a feature extraction network to extract features from two frames of a sample image pair, two network branches are used to further extract features from the two extracted frames to obtain a representation of the image visibility value. Compared to using only one network branch to further extract features from the two extracted frames, the result is more accurate.
[0083] Of course, in some embodiments, the representation determination network may also include only one network branch. The features extracted by the feature extraction network from the two frames in the sample image pair can be input into the network branch, and the representation values corresponding to the two frames can be determined by the network branch respectively.
[0084] In some embodiments, the first input sample image may randomly select one of the network branches, while the subsequent input sample image may select another network branch. In some embodiments, the two network branches may also correspond to the visibility values of the sample images. For example, the first network branch may be specifically used for feature extraction of sample images with higher visibility values, and the second network branch may be specifically used for feature extraction of sample images with lower visibility values.
[0085] While determining the visibility value of an image is difficult, determining the approximate visibility range is relatively simple and can be done visually based on experience. Therefore, a large number of images can be acquired, and the visibility range of each frame can be determined, resulting in a visibility range label for the image, such as 0-50m, 50-100m, 100-150m, or greater than 1000m. Images within each visibility range can constitute a first image set, thus yielding multiple first image sets carrying visibility range labels. When acquiring sample image pairs, at least two first image sets can be acquired, where the visibility ranges of images in each first image set are the same, and the visibility ranges of images in any two first image sets do not overlap. Then, one frame is acquired from each of any two first image sets to form a sample image pair. Since the visibility ranges of images in any two sample image sets are different, by randomly combining the first images from each pair of first image sets, a large number of sample image pairs with varying visibility values can be obtained for training the neural network.
[0086] In some embodiments, the representation quantity that characterizes the visibility value of the scene contained in the image can be a one-dimensional scalar, that is, a specific numerical value, thereby facilitating the establishment of a mapping relationship between the representation quantity and the visibility value.
[0087] In some embodiments, when constructing the mapping relationship between representation quantities and visibility values, the specific visibility values of some images can be determined first using a visibility meter or other methods. Then, a mapping relationship can be constructed based on the representation quantities and visibility values of these images. For example, a third image set can be obtained, which may include multiple frames of images with known visibility values. Feature extraction can then be performed on these multiple frames of images in the third image set to determine their representation quantities. Subsequently, a mapping relationship can be constructed based on the representation quantities and visibility values of these multiple frames.
[0088] In some embodiments, when constructing a mapping relationship using multiple frames of images in a third image set, in order to obtain the mapping relationship as accurately as possible, the visibility values of the multiple frames of images in the third image set can be uniformly distributed within a specified visibility range. The specified visibility range can be determined based on the approximate distribution range of the visibility values of the image to be tested. For example, assuming that the visibility values of the image to be tested are roughly distributed within 1000m, it is necessary to accurately calibrate the mapping relationship between the visibility values and the characterization quantity within this visibility range. Therefore, the specified visibility range is 0-1000m. In order to make the calibrated mapping relationship as accurate as possible, the visibility values of these multiple frames of images can be distributed as uniformly as possible within this range to cover the various visibility gradients within this range.
[0089] Of course, determining the visibility value of an image requires specialized instruments, which is quite cumbersome. To eliminate the need for specialized instruments, a mapping relationship can be constructed. In some embodiments, when establishing the mapping relationship between the representation quantity and the visibility value, at least one second image set can be obtained. Each second image set includes multiple frames of images, and the visibility value range of the images in each second image set is the same, for example, 0-50m or 50-100m. For each second image set, features can be extracted from each frame of the second image set to obtain the representation quantity of each frame. For example, each frame of the image can be input into a trained neural network, and the neural network can output the representation quantity of the image, thereby obtaining the distribution range of the representation quantities of each image in the second image set. Then, based on the visibility value range and the distribution range of the representation quantities of each image in each second image set, the mapping relationship between the representation quantity and the visibility value can be determined. Each second image set can include a large number of sample images. The more images there are, the more visibility values can cover all visibility values within the visibility value range, and the more accurate the established mapping relationship will be.
[0090] For example, suppose a second image set contains a large number of images with visibility values ranging from 0 to 50 meters. Due to the large number of images, these images can essentially cover all visibility values between 0 and 50 meters. Then, a neural network can output representations of these images, where the representation and visibility value can be positively or negatively correlated; for example, the higher the visibility value, the higher the representation. Furthermore, a mapping relationship can be determined based on the distribution range of these representations and the range of visibility values. For instance, if the distribution range of the representation is A and B, assuming a positive correlation between the representation and the visibility value, then representation A corresponds to a visibility value of 0, and representation B corresponds to a visibility value of 50, thus establishing a mapping relationship. Of course, this is just a simple example; in practice, constructing mapping relationships can be much more complex.
[0091] In some embodiments, there can be multiple second image sets. These multiple second image sets are arranged in order of visibility value range, with the visibility value ranges of images in adjacent second image sets being continuous. For example, consider six second image sets with visibility value ranges of (0m, 50m], (50m, 100m], (100m, 200m], (200m, 500m], (500m, 1000m], and (1000m, ∞), respectively. The visibility value ranges of adjacent second image sets are continuous and separated by a target visibility value. For example, two second image sets with visibility ranges of (0m, 50m] and (50m, 100m] are separated by 50m. This ensures accurate construction of the mapping relationship. In determining the mapping relationship based on the visibility value range and representation quantity distribution range corresponding to each second image set, the representation quantity corresponding to the target visibility value can be determined first based on the visibility value range and representation quantity distribution range of each image in two adjacent second image sets. Then, the mapping relationship is determined based on the representation quantity of the target visibility value. Since the target visibility value is the boundary of each visibility value range, and the determination of the mapping relationship is mainly based on the boundary of the visibility value range and the boundary of the representation quantity distribution range, accurately determining the representation quantity of the boundary is the key to accurately determining the mapping relationship.
[0092] For example, taking two second image sets with visibility ranges of (0m, 50m] and (50m, 100m] respectively, the target visibility value of 50m is a boundary of the two visibility value ranges. The representation quantity corresponding to the target visibility value of 50m can be determined based on the two second image sets. For example, taking the representation quantity and visibility value as positively correlated, the representation quantity corresponding to the target visibility value of 50m is the maximum representation quantity in the first second image set and the minimum representation quantity in the second second image set. The representation quantity of the target visibility value of 50m determined based on the two second image sets may be inconsistent. Therefore, the final representation quantity of the target visibility value can be obtained by combining the representation quantities of the target visibility value of 50m determined based on the two second image sets.
[0093] In some embodiments, when determining the representation quantity of the target visibility value based on the visibility value range and representation quantity distribution range corresponding to two adjacent second image sets, the initial representation quantity corresponding to the target visibility value can be determined based on the visibility value range and representation quantity distribution range of two adjacent second image sets. The initial representation quantity is then continuously adjusted until the accuracy of the classification results of the images in the two adjacent second image sets based on the adjusted representation quantity reaches the maximum value. Then, the adjusted representation quantity is used as the representation quantity corresponding to the target visibility value.
[0094] For example, taking two second image sets with visibility values ranging from (0m, 50m) to (50m, 100m), assuming a positive correlation between visibility values and the representation value, after inputting images from the first second image set into the neural network, the output representation value is between 0 and 20. Therefore, based on the first second image set, the representation value corresponding to 50m can be initially determined to be 20. After inputting images from the second second image set into the neural network, the output representation value is between 18 and 40. Therefore, based on the second second image set, the representation value corresponding to 50m can be initially determined to be 18. Furthermore, one of the representation values (18 or 20) or the average of the two can be taken as the initial representation value for 50m. Then, the initial representation value is adjusted within a certain range (e.g., increased or decreased by a certain step size). Each adjustment results in a new representation value based on... The adjusted representation value determines the accuracy of classification results when images in two second image sets are classified according to the adjusted representation value. For example, assuming the adjusted representation value is 19, meaning the representation value corresponding to 50m is 19, then images in the first image set with representation values between 19 and 20 will be misclassified, thus allowing for the calculation of a classification accuracy. Similarly, images in the second image set with representation values between 18 and 19 will also be misclassified, allowing for the calculation of another accuracy. The representation value with the highest classification accuracy between the two image sets can be taken as the representation value corresponding to the target visibility value. A similar method can be used for other target visibility values (e.g., 100m, 200m, 500m, and 1000m in the six second image sets mentioned above).
[0095] In some embodiments, the pre-trained neural network may include a feature extraction network, a first network branch, and a second network branch. When extracting features from the image to be detected using the pre-trained neural network to obtain the representation of the image to be detected, the feature extraction network can be used first to extract features from the image to be detected. The extracted features are then input into either the first or second network branch to obtain the representation of the image to be detected. That is, when extracting features from the image to be detected, only one network branch is used to output the representation of the image to be detected. In some embodiments, the representations of the image to be detected output from both network branches can be combined to obtain the final representation. For example, the feature extraction sub-network of the neural network can be used to extract features from the image to be detected. The extracted features are input into the first network branch to obtain the first representation, and then the extracted features are input into the second network branch to obtain the second representation. The representation of the image to be detected is obtained based on the first and second representations. For example, the representation of the image to be detected can be obtained by weighted averaging of the first and second representations, where the weights can be the same or different depending on the actual needs.
[0096] In some embodiments, the image to be detected can be an image of a road area, which can be acquired by an image acquisition device installed in the road. After determining the visibility value of the scene contained in the image to be detected, the hazard level of the current specific climate (e.g., foggy weather, dust storm, haze, etc.) can be determined based on the visibility value, and a control strategy corresponding to the hazard level can be selected to control traffic on the road. For example, multiple levels for evaluating the hazard level of foggy weather can be preset, with different levels corresponding to different visibility value ranges. Each hazard level can have a corresponding control strategy preset to control traffic on the road. For example, the control strategy can be to issue a fog weather warning to vehicles on the road, to remind vehicles on the road to control their speed below a certain level, or to close road intersections and prohibit vehicles from passing through.
[0097] In some embodiments, the image to be detected can be an image of a road area. This image can be multiple frames captured at different times by an image acquisition device installed along the road. After determining the visibility values corresponding to these multiple frames, the changing trend of a specific climate (e.g., foggy weather, dust storm weather, haze weather, etc.) can be predicted based on these visibility values. For example, an image of the road area can be acquired every 30 minutes, and the corresponding visibility value can be determined to obtain the changing trend of the road's visibility value over a period of time. This allows prediction of whether the fog concentration on the road is gradually increasing or decreasing, thus predicting the changing trend of foggy weather.
[0098] In some embodiments, the image to be detected can be an image of a road area. This image can be captured by an image acquisition device installed on the road at preset time intervals, such as capturing one or more frames every hour. Then, a neural network is used to determine the visibility value corresponding to this frame or frames to determine whether a specific weather phenomenon (e.g., fog, dust storm, haze, etc.) has occurred. For example, for each frame, it can be determined whether the visibility value exceeds a preset threshold. If the visibility value of each frame exceeds the preset threshold, or if the visibility value of more than a certain number of images exceeds the preset threshold, then it is determined that a specific weather phenomenon has occurred in the road area. Alternatively, the average visibility value of this frame or multiple frames can be determined. If the average value exceeds a preset threshold, then it is determined that a specific weather phenomenon has occurred. By statistically analyzing the total number of times a specific weather phenomenon occurs within a target time period, the frequency of a specific weather phenomenon in a certain area can be predicted, which can be used to study the climate patterns of that area.
[0099] For example, one or more frames of images of a road area can be collected every two hours. The visibility value corresponding to the one or more frames of the road area can be predicted using a neural network. If the visibility value of these images exceeds a preset threshold, it can be determined whether fog has occurred in the road area. By counting the total number of times fog occurs in a day, the frequency of fog occurrence in a day can be determined.
[0100] To further explain the method for detecting visibility values provided in the embodiments of this disclosure, the following explanation is based on a specific embodiment.
[0101] For weather conditions such as smog, sandstorms, and fog, it is often necessary to detect visibility in road areas for traffic management. This is typically achieved by using cameras installed along roads to capture images of road surfaces, and then using neural networks to predict the visibility of the road areas within the images. However, because it is difficult to accurately label the visibility values of scenes within images, it is challenging to obtain a large number of sample images labeled with visibility values to train the neural network. Currently, the neural network is trained using sample images labeled with a range of visibility values, and the network outputs only a classification result of the image's visibility value range, failing to provide accurate visibility values. This embodiment provides a method for detecting image visibility values, such as... Figure 3 The diagram illustrates one application scenario of this method, which can be used to detect visibility values in road areas and subsequently manage road traffic. The method specifically includes the following steps:
[0102] (1) Neural network training phase
[0103] A large number of sample images can be collected. For each sample image, it is not necessary to label its specific visibility value; only the visibility value range category to which it belongs needs to be determined. For example, the sample images can be divided into 6 categories, with corresponding 6 visibility value ranges as follows: (0m, 50m], (50m, 100m], (100m, 200m], (200m, 500m], (500m, 1000m], and (1000m, ∞). Since only the visibility value range of the images needs to be labeled, it is easy to implement, and thus a large number of sample images can be obtained.
[0104] Then, any two sample images from the above six categories can be paired to obtain a large number of sample image pairs, where the visibility value relationship of each sample image pair is known. These sample image pairs can then be used as training samples, and the visibility value relationship of the sample image pairs can be used as labels to train the initial neural network, resulting in the trained neural network.
[0105] The structure of the neural network is as follows: Figure 4As shown, the network includes a feature extraction network, a first network branch, a second network branch, and an output layer. The feature extraction network can be a ResNet18 network, which consists of five convolutional layers (conv1-conv5 in the diagram). After conv5, two network branches are connected: the first and second network branches. Each branch contains three fully connected layers. The fully connected layer fc1 has an output dimension of 256 and uses PreLU activation; fc2 has an output dimension of 128 and uses PreLU activation; and fc3 has an output dimension of 1 and uses Liner activation.
[0106] When a pair of sample images (sample image A and sample image B) are input into a neural network, sample image A can have its features extracted by a feature extraction network. These extracted features are then further input into the first network branch for feature extraction, ultimately outputting a one-dimensional scalar RA, which represents the visibility value of the scene included in sample image A. Similarly, sample image B can have its features extracted by a feature extraction network, and these extracted features are then further input into the second network branch for feature extraction, ultimately outputting a one-dimensional scalar RB, which also represents the visibility value of the scene included in sample image B. The one-dimensional scalar output by the neural network and the visibility value can be positively correlated; the larger the one-dimensional scalar, the larger the visibility value.
[0107] Then, RA and RB can be input into the output layer. The output layer can determine the probability p that sample image A has a higher visibility value, and the probability 1-p that sample image B has a higher visibility value, based on RA and RB. The loss function can be cross-entropy loss. For example, assuming the visibility value of sample image A is greater than that of sample image B, the true probability that sample image A has a higher visibility value is y (y is 1 or a pre-set value close to 1, such as 0.8), and the true probability that sample image B has a higher visibility value is 1-y. Then, the cross-entropy loss between the predicted probability and the true probability can be calculated. The parameters of the neural network can be continuously adjusted based on the cross-entropy loss until the loss function converges, resulting in a well-trained neural network.
[0108] (2) Calibration of the mapping relationship between the one-dimensional scalar output of the neural network and the visibility value
[0109] The images from each of the six categories mentioned above can be input into a neural network. The neural network outputs a one-dimensional scalar representing the visibility value of each image frame. Since there are many sample images, the output one-dimensional scalar can basically encompass various visibility values. Then, a mapping relationship is constructed based on the range of visibility values of images in each category and the range of the one-dimensional scalar. In order to accurately construct the mapping relationship, the one-dimensional scalar corresponding to the boundary values of each visibility value range can be determined first, for example, 0m, 50m, 100m, 200m, 500m, and 1000m. Taking 0m as an example, since the one-dimensional scalar and visibility value are positively correlated, the minimum value of the one-dimensional scalar in the image with visibility values in the range (0, 50) corresponds to the boundary value 0m. Taking 50m as an example, since 50m is the boundary value between the two categories, the maximum value of the one-dimensional scalar in the image with visibility values in the range (0, 50) corresponds to 50m, and the minimum value of the one-dimensional scalar in the image with visibility values in the range (50, 100) also corresponds to 50m. The one-dimensional scalar corresponding to 50m may not be equal based on the two categories of images, for example, one might be 10 and the other 15. To accurately determine the one-dimensional scalar corresponding to the boundary value 50m, the one-dimensional scalar corresponding to the boundary value 50m can be continuously adjusted (for example, by taking values between 10 and 15), and then the value can be determined based on... The adjusted one-dimensional scalar determines the accuracy of image classification in both categories. For example, when the one-dimensional scalar is set to 12, classifying images with visibility values in the range (0, 50) based on this scalar yields only 80% accuracy, while classifying images with visibility values in the range (50, 100) based on this scalar yields only 90% accuracy. The value of the one-dimensional scalar is then continuously adjusted until the classification accuracy for both categories reaches its maximum value. This scalar is then identified as the one-dimensional scalar corresponding to a visibility value of 50m. Similarly, a similar method can be used to determine the corresponding scalar for other visibility value range boundaries. After determining the one-dimensional scalar corresponding to the boundary values of the visibility value range, a mapping relationship can be constructed based on this scalar.
[0110] Of course, in some embodiments, the visibility values of multiple frames of images can be determined first using a visibility meter, and then a one-dimensional scalar of the visibility values of these multiple frames of images can be determined using a pre-trained neural network. A mapping relationship can then be constructed based on the one-dimensional scalar of these multiple frames of images and the visibility values. In order to make the constructed mapping relationship as accurate as possible, the visibility values of these multiple frames of images should cover the gradient of each visibility value as much as possible.
[0111] (3) Neural Network Inference Stage
[0112] When determining the visibility value of an image, the image is acquired and input into a neural network. The network outputs a one-dimensional scalar value of the image (which can be determined using a single network branch or the average of the outputs from two branches). The visibility value is then determined based on this one-dimensional scalar value and a pre-defined mapping relationship. Once the visibility value is determined, traffic management on roads can be implemented based on it.
[0113] The method provided in this embodiment allows for the training of a neural network without calibrating the visibility values of sample images, and the determination of the visibility value of the image to be detected based on the trained neural network. Compared to directly determining the range of visibility values, this method yields more accurate visibility results.
[0114] Corresponding to the above method, this disclosure also provides a visibility value detection device, such as... Figure 5 As shown, the device 50 includes:
[0115] Acquisition module 51 is used to acquire the image to be detected;
[0116] The representation module 52 is used to extract features from the image to be detected through a pre-trained neural network to obtain the representation quantity of the image to be detected.
[0117] The visibility value determination module 53 is used to determine the visibility value corresponding to the image to be detected based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected. The representation quantity is obtained by extracting features from the image through the neural network and is used to represent the magnitude of the visibility value of the scene contained in the image.
[0118] The specific implementation process of the device for determining the visibility value of the image to be detected can be referred to the description in the above method embodiments, and will not be repeated here.
[0119] Furthermore, embodiments of this disclosure also include an electronic device, such as... Figure 6 As shown, the electronic device includes a processor 61, a memory 62, and computer instructions stored in the memory 62 that can be executed by the processor 61. When the processor executes the computer instructions, it can implement the method described in any of the foregoing embodiments.
[0120] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the foregoing embodiments.
[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0122] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of this specification can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the embodiments of this specification, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this specification.
[0123] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0124] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. When implementing the embodiments of this specification, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0125] The above description is merely a specific implementation of the embodiments of this specification. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of the embodiments of this specification, and these improvements and modifications should also be considered within the protection scope of the embodiments of this specification.
Claims
1. A method for detecting visibility values, characterized in that, The method includes: Acquire the image to be detected; Feature extraction is performed on the image to be detected to obtain the representational value of the image to be detected; Based on the pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected, the visibility value of the image to be detected is determined; wherein, the representation quantity is used to characterize the magnitude of the visibility value of the scene contained in the image. Wherein, the representation quantity is a one-dimensional scalar, and the mapping relationship between the representation quantity and the visibility value is determined based on the following method: acquiring at least one second image set, wherein the visibility value range of each image in each second image set is the same; obtaining the representation quantity distribution range of each image in each second image set based on the determined representation quantity of each image in each second image set; determining the mapping relationship based on the visibility value range of each image in each second image set and the representation quantity distribution range; or A third image set is obtained, wherein the visibility value of each image in the third image set is known; the mapping relationship is determined based on the determined representation quantity of each image in the third image set and the visibility value of each image in the third image set.
2. The method according to claim 1, characterized in that, Feature extraction of the image to be detected includes: Feature extraction is performed on the image to be detected using a pre-trained neural network; wherein the neural network is trained in the following manner: Acquire sample image pairs carrying label information, wherein the label information is used to indicate the relationship between the visibility values of two frames in the sample image pair; The preset initial neural network outputs the prediction result of the relationship between the visibility values of the two frames in the sample image pair; Based on the difference between the prediction result and the label information, the network parameters of the initial neural network are adjusted to obtain the neural network.
3. The method according to claim 2, characterized in that, The initial neural network includes a feature extraction network, a representation quantity determination network, and an output layer; The prediction result of the visibility value relationship between the two frames in the sample image pair, output through a preset initial neural network, includes: The feature extraction network is used to extract features from two frames of the sample image pair. The network determines the representation quantity of each of the two frames in the sample image pair based on the features extracted by the feature extraction network using the representation quantity determination network. The output layer determines the relationship between the visibility values of the two frames in the sample image pair based on their respective representations.
4. The method according to claim 3, characterized in that, The characterization determination network includes a first network branch and a second network branch; The step of determining the representation quantities of each of the two frames in the sample image pair based on the features extracted by the feature extraction network through the representation quantity determination network includes: The first network branch determines the representation quantity of the first image based on the features extracted from the first image in the sample image pair; The second network branch determines the representation of the second image based on the features extracted from the second image in the sample image pair.
5. The method according to any one of claims 2-4, characterized in that, The sample image pairs were obtained based on the following method: At least two first image sets are obtained, wherein the visibility value range of each image in each first image set is the same, and the visibility value ranges of each image in any two first image sets do not overlap; One frame is selected from each of the two first image sets to form the sample image pair.
6. The method according to claim 1, characterized in that, Based on the visibility value range of each image in each of the second image sets and the distribution range of the characterization quantity, the mapping relationship is determined, including: Based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, the representation quantity corresponding to the target visibility value is determined, wherein the visibility value range of each image in two adjacent second image sets is continuous and separated by the target visibility value; The mapping relationship is determined based on the representation quantity corresponding to the target visibility value.
7. The method according to claim 6, characterized in that, Based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, the representation quantity corresponding to the target visibility value is determined, including: Based on the visibility value range of each image in two adjacent second image sets and the distribution range of the representation quantity, the initial representation quantity corresponding to the target visibility value is determined; The initial representation quantity is adjusted, and the images in two adjacent second image sets are classified based on the adjusted representation quantity until the accuracy of the classification results of the two adjacent second image sets reaches the maximum value. The adjusted representation value is used as the representation value corresponding to the target visibility value.
8. The method according to any one of claims 1-4, characterized in that, The representation of the image to be detected is obtained by extracting features from the image through a pre-trained neural network, which includes a feature extraction network, a first network branch, and a second network branch. Feature extraction is performed on the image to be detected using a pre-trained neural network to obtain the representational parameters of the image to be detected, including: The feature extraction network extracts features from the image to be detected, and the extracted features are input into either the first network branch or the second network branch to obtain the representation of the image to be detected; or The feature extraction network of the neural network is used to extract features from the image to be detected, and the extracted features are input into the first network branch to obtain the first representation quantity; The extracted features are input into the second network branch to obtain the second representation quantity; The characterization of the image to be detected is obtained based on the first characterization and the second characterization.
9. The method according to any one of claims 1-4, characterized in that, The image to be detected includes an image of a road area, and the method further includes: Based on the visibility value of the road area contained in the image to be detected, the hazard level of a specific climate is determined, wherein the specific climate includes any one of the following: fog, haze, or dust storm; Traffic within the road area is controlled according to the control strategy corresponding to the hazard level.
10. The method according to any one of claims 1-4, characterized in that, The image to be detected includes multiple frames of images of a road area collected at different times. The method includes: predicting the change trend of a specific climate within the road area based on the changing trend of visibility values in the multiple frames of images; or The image to be detected includes images of a road area collected at preset time intervals. The method includes: in response to the visibility value corresponding to the image exceeding a preset threshold, determining that a specific climate has occurred in the road area; and determining the frequency of the specific climate occurring in the road area based on the total number of times the specific climate has occurred in the road area within a target time period. The specific climate includes any of the following: fog, haze, or dust storm.
11. A visibility value detection device, characterized in that, The device includes: The acquisition module is used to acquire the image to be detected; The representation module is used to extract features from the image to be detected using a pre-trained neural network to obtain the representational quantity of the image to be detected. The visibility value determination module is used to determine the visibility value of the image to be detected based on a pre-calibrated mapping relationship between the representation quantity and the visibility value, and the representation quantity of the image to be detected. The representation quantity is obtained by extracting features from the image through the neural network and is used to represent the magnitude of the visibility value of the scene contained in the image. Wherein, the characterization quantity is a one-dimensional scalar, and the mapping relationship between the characterization quantity and the visibility value is determined based on the following method: Obtain at least one second image set, where the visibility value range of each image in each second image set is the same; based on the determined representation values of each image in each second image set, obtain the representation value distribution range of each image in each second image set; based on the visibility value range of each image in each second image set and the representation value distribution range, determine the mapping relationship; or A third image set is obtained, wherein the visibility value of each image in the third image set is known; the mapping relationship is determined based on the determined representation quantity of each image in the third image set and the visibility value of each image in the third image set.
12. An electronic device, characterized in that, The device includes a processor, a memory, and computer instructions stored in the memory that are executable by the processor, wherein the processor executes the computer instructions to implement the method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed, implement the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Foggy weather visibility estimation method based on deep relative learning under discrete labels
CN112989994A