Arithmetic device

The computing device improves visibility estimation accuracy by identifying useful features and incorporating gaze maps into images, addressing the challenges of reduced accuracy in conventional AI-based methods.

JP2025174673APending Publication Date: 2025-11-28HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024081175
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Conventional AI-based visibility estimation methods face accuracy issues due to the presence of large objects like buildings or aircraft near the camera, leading to reduced estimation accuracy.

Method used

A computing device that extracts useful and harmful features from multiple images with similar angles and directions, generates a gaze map indicating feature intensities, and incorporates this map into images for learning and inference to improve visibility estimation accuracy.

Benefits of technology

Enhances the accuracy of automatic visibility estimation by focusing on features highly correlated with visibility, improving robustness against occlusions and disturbances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174673000001_ABST
    Figure 2025174673000001_ABST
Patent Text Reader

Abstract

To improve robustness against occlusion and disturbance uncorrelated with a visibility range, and enhance accuracy in automatic estimation of the visibility range.SOLUTION: An arithmetic device extracts beneficial features or harmful features from a plurality of images in which at least an imaging angle of view, an imaging direction, or an estimation target object is substantially identical, calculates strengths of the beneficial features or the harmful features in the plurality of images, aggregates the strengths of the beneficial features or the harmful features, generates an attention map indicating a degree of attention of the beneficial features or the harmful features on the basis of the aggregated strengths, generates learning images or inference images by incorporating the attention map into each of the plurality of images, generates an AI model by learning the learning images, and estimates a visibility range with the beneficial features set to have a high degree of attention predominantly as estimation grounds in accordance with the attention map incorporated into the inference images by using the AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to automated weather observation. [Background technology]

[0002] In recent years, technology that evaluates visibility information by utilizing image analysis technology using AI (artificial intelligence) has been attracting attention. One example of this type of technology is "automated weather observation," which automates the observation of weather information. Observation of weather information, such as runway visual distance and the presence or absence of fog, can be performed automatically using existing meteorological observation instruments. However, observing visibility (the distance visible to the naked eye) using current meteorological observation instruments is not easy. For example, experienced observers currently conduct visual weather observations 24 hours a day. From the economic perspective of reducing the number of personnel, and from the social perspective of supporting the training of new observers, automating visibility observation is considered an urgent priority.

[0003] Here, Patent Document 1, in relation to visibility evaluation technology, discloses "a visibility evaluation device including: an image acquisition means for acquiring an image whose imaging range includes the position of a target that serves as an evaluation index for visibility; a feature amount calculation unit for calculating feature amounts from the image; a distance information acquisition unit for acquiring distance information from the imaging point of the image to the target; an ambient light information acquisition unit for acquiring ambient light information that indicates the state of ambient light at the time the image was captured; and a visibility evaluation unit for evaluating the visibility at the time the image was captured using the feature amounts, the distance information, and the ambient light information." [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-20911 Summary of the Invention [Problem to be solved by the invention]

[0005] Conventional evaluation methods using AI are thought to learn the characteristic that targets close to the camera are mapped at the bottom of the image, and distant targets are mapped at the top of the image. For example, when evaluating visibility, it is thought that visibility is estimated based on the height coordinate information of the target and the degree of blur. For this reason, there was an issue of reduced accuracy in estimating visibility when large objects such as buildings or aircraft are present near the camera. Therefore, further improvement in accuracy is essential for automatic visibility estimation in automatic weather observation.

[0006] Therefore, an object of the present invention is to improve the accuracy of automatic visibility estimation and to contribute from an economic perspective of reducing the number of observers, and from a social perspective of training observers. [Means for solving the problem]

[0007] The present invention includes multiple means for solving at least part of the above-mentioned problems, and an example thereof is as follows: That is, a computing device for estimating visibility using images includes: an attention degree calculation unit that extracts useful features or harmful features from multiple images having at least substantially the same shooting angle of view, shooting direction, or estimation target, and calculates the intensity of the useful features or harmful features in the multiple images, an aggregation unit that aggregates the intensities of the useful features or harmful features and generates a gaze map indicating the attention degree of the useful features or harmful features based on the aggregated intensities, and an image reconstruction unit that incorporates the gaze map into each of the multiple images.

[0008] Also, a method for estimating visibility using images includes extracting beneficial features or harmful features from a plurality of images having at least the same photographing angle of view, photographing direction, or estimated object, calculating the intensity of the beneficial features or harmful features in the plurality of images, aggregating the intensities of the beneficial features or harmful features, generating a gaze map indicating the degree of gaze of the beneficial features or harmful features based on the intensities after the aggregating, and incorporating the gaze map into each of the plurality of images. [Effects of the Invention]

[0009] According to the present invention, it is possible to improve the accuracy of automatic visibility estimation, and to contribute from the economic viewpoint of reducing the number of observers, and from the social viewpoint of training observers. Note that the problems, configurations, and effects other than those described above will become clear from the following description of the embodiment of the invention. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 10 is a diagram illustrating an example of a processing flow of a learning phase performed by the arithmetic device according to the embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example of a processing flow of an inference phase performed by the arithmetic device according to the embodiment. [Figure 3] FIG. 10 is a diagram illustrating an effect of the embodiment. [Figure 4] FIG. 1 is a diagram illustrating an example of the overall configuration of a computing device according to a first embodiment. [Figure 5] FIG. 2 is a diagram illustrating an example of the configuration of a gaze degree calculation unit (D100) in the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of the configuration of a consolidation unit (E100) in the first embodiment. [Figure 7] FIG. 3 is a diagram showing an example of a processing flow of a learning phase executed in the arithmetic device (1) in the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a processing flow of an inference phase executed in the arithmetic device (1) in the first embodiment. [Figure 9] FIG. 10 is a diagram showing an example of a processing flow of a learning phase executed in the arithmetic device according to the second embodiment. [Figure 10] FIG. 11 is a diagram showing an example of a processing flow of an inference phase executed in a computing device according to the second embodiment. [Figure 11] FIG. 11 is a diagram showing an example of a processing flow of a learning phase executed in a computing device according to the third embodiment. [Figure 12] FIG. 11 is a diagram showing an example of a processing flow of an inference phase executed in a computing device according to the third embodiment. [Figure 13] FIG. 2 is a diagram illustrating an example of a hardware configuration of a computing device. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment is an example for explaining the present invention, and for clarity of explanation, appropriate omissions and simplifications have been made. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.

[0012] In order to facilitate understanding of the invention, the position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings.

[0013] When there are multiple components with the same or similar functions, they may be described using the same reference numeral with different subscripts. When there is no need to distinguish between these multiple components, the subscripts may be omitted.

[0014] In the embodiments, there may be a description of processing performed by executing a program. Here, a computer executes the program using a processor (e.g., a CPU or a GPU), and performs processing defined by the program while using storage resources (e.g., a memory) and interface devices (e.g., a communication port). Therefore, the entity that executes the program and performs the processing may be the processor. Similarly, the entity that executes the program and performs the processing may be a controller, device, system, computer, or node that has a processor.

[0015] The processing performed by executing the program may be performed by a computing unit, and may include a dedicated circuit for performing specific processing. Here, the dedicated circuit may be, for example, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), or a Complex Programmable Logic Device (CPLD).

[0016] A program may be installed on a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server may include a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. In addition, in an embodiment, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0017] When estimating visibility (the maximum distance visible to the naked eye) using AI, if you try to determine the image features that serve as the basis for estimating visibility from a single arbitrary inference image, there is a problem in that the accuracy of visibility estimation decreases if the inference image contains, for example, occlusions that are uncorrelated with visibility (power lines, lightning rods, buildings near the observation point, etc.) or disturbances that cause blurring that occurs only in certain images (rainfall, icing, the appearance of moving objects, etc.). In order to achieve stable visibility estimation in automatic weather observation, it is necessary to improve robustness against the aforementioned occlusions and disturbances.

[0018] In the embodiment, an AI that estimates visibility (hereinafter, visibility estimation AI) is provided with a mechanism that identifies image features (useful features) that are useful for improving the accuracy of visibility and image features (harmful features) that are harmful for improving the accuracy of visibility in advance as a gaze map from multiple training images that have approximately the same shooting angle of view, shooting direction, and estimation target, and feeds this back from outside as domain knowledge during AI learning and inference. This enables the learning and inference of the visibility estimation AI to focus on image features that are highly correlated with visibility, improving robustness against the above-mentioned occlusion and disturbances.

[0019] First, an outline of the embodiment will be described with reference to FIGS. 1, 2 and 3. FIG.

[0020] 1 is a diagram showing an example of a processing flow of the learning phase by a computing device according to an embodiment. In FIG. 1, the camera images (A100) are made up of multiple types of camera images with different shooting angles of view, shooting directions, and estimated objects, and each type includes multiple camera images with approximately the same shooting angles of view, shooting directions, and estimated objects.

[0021] The image grouping unit (B100) takes in the above-mentioned camera images (A100) and groups a plurality of camera images that have substantially the same photographing angle of view, photographing direction, and estimated object for each type.

[0022] The data refinement unit (C100) quantifies the impact of each camera image on the learning and inference of the visibility estimation AI, and tags each camera image as either useful or harmful. Here, useful images are those that contribute to improving the accuracy of visibility estimation, while harmful images are those that hinder this improvement.

[0023] The gaze degree calculation unit (D100) calculates the intensities of the useful features and harmful features possessed by the useful image and the harmful image, respectively, in pixel units. The gaze degree calculation unit (D100) also extracts image features uncorrelated with changes in visibility as inter-image common features from camera images belonging to the same group.

[0024] The aggregation unit (E100) aggregates useful features, harmful features, and inter-image common features from multiple camera images belonging to the same group, and generates one attention map (H100) for each group. The aggregation unit (E100) generates the attention map so that useful features have a high attention level, and harmful features and inter-image common features have a low attention level. In this way, the attention map indicates the attention level of each image feature contained in the image in pixel units.

[0025] The image reconstruction unit (F100) incorporates a gaze map corresponding to the group to which each camera image belongs into each camera image to generate training images. Incorporating a gaze map into a camera image involves, for example, adding a channel indicating gaze map information (attention level information for each pixel of the image) to the three channels (channels representing color information for red, green, and blue (RGB)) that represent color information for each pixel of the image. This allows each pixel of the image to have color information and gaze level information. By incorporating the gaze map into the camera image in this way, the visibility estimation AI can learn each image feature of the training image, including the gaze level indicated by the gaze map, when learning the training images. In particular, the visibility estimation AI will focus on learning image features with high gaze levels.

[0026] By having the visibility estimation AI learn the training images in this way, the information from the gaze map is fed back to the visibility estimation AI as domain knowledge for learning the visibility estimation AI.

[0027] FIG. 2 is a diagram showing an example of a processing flow of the inference phase by a computing device in an embodiment. In the inference phase, a gaze map corresponding to a group to which each camera image to be inferred belongs is selected from the gaze maps (H100) generated in the learning phase, and the gaze map is incorporated into the camera image for each group to generate an inference image. When estimating visibility for the inference image into which the gaze map has been incorporated, the visibility estimation AI focuses on image features in the inference image that are highly gazed upon as the basis for estimation. By having the visibility estimation AI estimate visibility for the inference image in this way, information on the gaze map is fed back to the visibility estimation AI as domain knowledge for the inference of the visibility estimation AI. Note that the gaze map is incorporated into the camera image in the same manner as in the example described above.

[0028] FIG. 3 is a diagram illustrating the effect of an embodiment. With reference to FIG. 3, the effect of determining image features that serve as the basis for estimating visibility from multiple camera images with substantially the same shooting angle of view, shooting direction, and estimation target will be described. In the evaluation shown in FIG. 3, the estimation accuracy is evaluated for an AI that estimates the visibility of an arbitrary target. In FIG. 3, TN represents true negative, FP represents false positive, TP represents true positive, and FN represents false negative. In FIG. 3, TP / (TP+FN) on the horizontal axis represents the estimation accuracy rate (recall rate) for visible targets, and TN / (TN+FP) on the vertical axis represents the estimation accuracy rate (specificity rate) for invisible targets. High accuracy rates for both indicate that visibility can be more accurately estimated. As shown in Figure 3, it was confirmed that when multiple camera images with approximately the same shooting angle of view, shooting direction, and estimated object are used (embodiment), the accuracy of visual recognition is improved compared to when only one image is used (conventional example).

[0029] Next, the details of the embodiment will be described. First Embodiment Fig. 4 is a diagram showing an example of the overall configuration of the arithmetic device in the first embodiment. The configuration of the arithmetic device in the first embodiment will be described with reference to Fig. 4. Note that explanations that are the same as those already explained may be omitted.

[0030] The calculation device (1) includes a memory (10), an image grouping unit (B100), a data refining unit (C100), a gaze calculation unit (D100), an aggregation unit (E100), an image reconstruction unit (F100), a learning unit (G100), and an inference unit (I100).

[0031] 5 is a diagram showing an example of the configuration of the gaze degree calculation unit (D100) in the first embodiment. The gaze degree calculation unit (D100) includes an image extraction unit (D111), an image feature strength calculation unit (D112), and an image comparison unit within the same group (D113).

[0032] 6 is a diagram showing an example of the configuration of the aggregation unit (E100) in the first embodiment. The aggregation unit (E100) includes an accumulation unit (E111), a normalization unit (E112), a weak strength truncation unit (E113), and a gaze map generation unit (E114).

[0033] As shown in FIG. 4, the calculation device (1) is connected to a plurality of input devices (input device 1 (11), input device 2 (12), input device 3 (13), input device 4 (14), input device 5 (15)) and a plurality of output devices (output device 1 (21), output device 2 (22), output device 3 (23)).

[0034] The input device 1 (11) is, for example, a camera, and inputs a captured camera image (A100) into the memory 10. The input device 1 (11) may be a device other than a camera (for example, an information processing terminal such as a personal computer (PC), a tablet, or a smartphone, or a server device) that captures the camera image (A100) from the camera and inputs it into the memory 10.

[0035] The input device 2 (12) inputs metadata including shooting conditions (e.g., shooting direction, zoom magnification, etc.) and shooting information (e.g., shooting time, shooting position, etc.) for each camera image into the memory (10). The input device 2 (12) may be, for example, the same device as the input device 1 (11) (e.g., the camera itself that captured the image), or may be a device different from the input device 1 (11) (e.g., an information processing terminal such as a personal computer (PC), tablet, or smartphone, or a server device) that imports metadata from the camera or from a device that mounts or operates the camera and inputs it into the memory (10).

[0036] The input device 3 (13) is a device (for example, a personal computer (PC), tablet, smartphone, or other information processing terminal, or a server device) that inputs the trained AI model 1 into the memory (10). As will be described in detail below, the trained AI model 1 is used by the data refinement unit (C100) to quantify the impact of camera images on the visibility estimation AI and to extract useful and harmful images, and is also used by the gaze calculation unit (D100) to calculate the strength of useful and harmful features. The trained AI model 1 is an AI different from the visibility estimation AI (trained AI model 2) generated by the calculation device (1), and is prepared in advance by training it with various images.

[0037] As will be described in detail below, the output device 1 (21) is a device (for example, an information processing terminal such as a personal computer (PC), a tablet, or a smartphone, a server device, or a display) that reads out the gaze map (H100) generated by the aggregation unit (E100) and stored in the memory (10) from the memory (10) and displays it on a screen. The output device 1 (21) can also read out a gaze map that is currently being generated from the memory (10) and display it on a screen.

[0038] The input device 4 (14) is a device (for example, a personal computer (PC), a tablet, a smartphone, or other information processing terminal, or a server device) that inputs the gaze map (H100) to the memory (10). The gaze map input from the input device 4 (14) may be the gaze map itself read out by the output device 1 (21), or may be a gaze map that has been read out by the output device 1 (21) and then modified, for example, by a user. The input device 4 (14) may be the same device as the output device 1 (21), or may be a device different from the output device 1 (21).

[0039] As will be described in detail below, the output device 2 (22) is a device (e.g., an information processing terminal such as a personal computer (PC), a tablet, or a smartphone, a server device, a storage device, etc.) that reads out from the memory (10) and stores the trained AI model 2 (visibility estimation AI) that is generated by having the visibility estimation AI learn a training image (G101) and stored in the memory (10).

[0040] The input device 5 (15) is a device (for example, an information processing terminal such as a personal computer (PC), tablet, or smartphone, a server device, or a storage device) that inputs the trained AI model 2 into the memory (10). The input device 5 (15) may be the same device as the output device 2 (22), or may be a device different from the output device 2 (22).

[0041] The output device 3 (23) is, for example, a display device (such as a display), which receives the inference result output by the inference unit (I100) via the memory (10) and displays it on a screen, etc. The output device 3 (23) may be a device equipped with a display or the like other than a display device (for example, an information processing terminal such as a personal computer (PC), a tablet, or a smartphone, or a server device, etc.).

[0042] The input devices and output devices may all be the same device, or some of the input devices or output devices may be the same device, or all may be different devices as described above.

[0043] The contents and operations of each component of the arithmetic device (1) shown in FIGS. 4 to 6 will be described in detail with reference to the following drawings.

[0044] FIG. 7 is a diagram showing an example of a processing flow of the learning phase executed in the arithmetic device (1) in the first embodiment.

[0045] 7, one or more types of camera images captured by a camera with different shooting angles of view, shooting directions, and estimated objects are input in advance to memory 10 as camera images (A100). The camera images (A100) include multiple camera images of each type with approximately the same shooting angles of view, shooting directions, and estimated objects. Metadata related to each camera image is also input in advance to memory 10.

[0046] The image grouping unit (B100) reads the camera images (A100) and metadata from the memory (10). The image grouping unit (B100) divides (groups) the camera images (A100) into 1 to k groups according to the metadata and stores them in the memory (10) (B101). The image grouping unit (B100) performs grouping so that multiple camera images with approximately the same shooting angle of view, shooting direction, and estimated object belong to the same group.

[0047] The data refinement unit (C100) reads from the memory (10) the trained AI model 1 and grouped camera images that have been stored in advance in the memory (10). The data refinement unit (C100) quantifies the impact of each camera image on the learning convergence and estimation accuracy of the visibility estimation AI for each group (C101). In the quantification (C101), for example, an influence function, an accuracy evaluation index, or the like is appropriately used. The data refinement unit (C100) extracts one or more camera images that contribute to improving the learning convergence and estimation accuracy as beneficial images for each group according to the degree of the quantified impact, tags these images, and stores them in the memory (10) (C102). Similarly, the data refinement unit (C100) extracts one or more camera images that hinder the improvement of the learning convergence and estimation accuracy as harmful images, tags these images, and stores them in the memory (10) (C103). The data refinement unit (C100) performs each of the above processes using the trained AI model 1.

[0048] The gaze degree calculation unit (D100) extracts useful features and harmful features from each of the useful and harmful images for each group, calculates the strength of each of the useful and harmful features (D101, D102), and further extracts common features between images for each group using the useful images (C101) and harmful images (C102) (D103).

[0049] Specifically, in the gaze degree calculation unit (D100), the image extraction unit (D111) reads useful images and harmful images tagged by the data refinement unit (C100) from the memory (10) for each group. The image extraction unit (D111) extracts useful images from the read useful images and harmful images for each group and extracts useful features possessed by the useful images. Similarly, the image extraction unit (D111) extracts harmful images from the read useful images and harmful images for each group and extracts harmful features possessed by the harmful images. Useful features are specifically objects, structures, buildings, etc. that can serve as a basis for estimating visibility, such as distant buildings that appear to be on the border between visible and invisible to the naked eye (such as buildings that appear blurred or hazy). On the other hand, harmful features are specifically, for example, the above-mentioned occlusions and disturbances.

[0050] The image feature intensity calculation unit (D112) uses the trained AI model 1 loaded from the memory (10) to calculate the intensity of beneficial features for each pixel in each camera image belonging to the same group for each group, and stores the calculated values ​​in the memory (10) (D101). Similarly, the image feature intensity calculation unit (D112) calculates the intensity of harmful features for each pixel in each camera image belonging to the same group for each group, and stores the calculated values ​​in the memory (10) (D102). The intensity refers to the degree to which the beneficial features (D101) and harmful features (D102) affect the visibility estimation, and represents the degree of attention paid to the beneficial features (D101) and harmful features (D102) when estimating the visibility. In calculating the intensity of the beneficial features (D101) and harmful features (D102), a technique for quantifying the estimation basis on a pixel-by-pixel basis using XAI, such as Grad-CAM, is appropriately used.

[0051] Meanwhile, the same-group image comparison unit (D113) reads useful images and harmful images for each group from the memory (10). The same-group image comparison unit (D113) compares the image features of one or more useful images (C101) and harmful images (C102) belonging to the same group for each group, extracts image features uncorrelated with changes in visibility from those images as inter-image common features, and stores them in the memory (10) (D103). Specific examples of inter-image common features include buildings, the sky, and the ground near the shooting location, and dirt on the camera lens captured in the image. The comparison of image features may involve, for example, comparing haze values ​​on a pixel-by-pixel basis, or comparing pixel values ​​on a pixel-by-pixel basis, as appropriate.

[0052] The aggregation unit (E100) reads the intensities of the beneficial and harmful features calculated by the image feature intensity calculation unit (D112) for each group from the memory (10), aggregates the intensities of each image feature, and generates a gaze map (H100) for each group.

[0053] Specifically, in the aggregation unit (E100), the accumulation unit (E111) reads the intensities calculated by the image feature intensity calculation unit (D111) for the useful features from the memory (10), and adds up the intensities of the useful features at corresponding pixels in multiple camera images belonging to the same group.

[0054] The normalization unit (E112) calculates the average value of the intensity of the useful feature for each corresponding pixel of each camera image (performs intensity averaging) by dividing the sum result by the number of camera images belonging to the same group (E101).

[0055] The weak intensity truncation unit (E113) replaces (truncates) the value of the intensity of the useful feature with "0" for pixels whose averaged intensity (average value) is smaller than a predetermined (predetermined) threshold (E102).

[0056] Similarly, the accumulation unit (E111) reads the intensities calculated by the image feature intensity calculation unit (D111) for harmful features from the memory (10) and adds up the intensities of the harmful features at corresponding pixels in multiple camera images belonging to the same group.

[0057] The normalization unit (E112) averages the intensity of the harmful feature for each corresponding pixel of each camera image by dividing the sum by the number of camera images belonging to the same group (E103).

[0058] The weak intensity truncation unit (E113) truncates the intensity value of the harmful feature to "0" for pixels whose averaged intensity is smaller than a predetermined threshold (E104). By this process, the aggregation unit (E100) aggregates the intensity of each image feature.

[0059] Then, the attention map generating unit (E114) generates an attention map (H100) by setting the attention degree on a pixel-by-pixel basis for each group so that the attention degree of the beneficial features is higher than the attention degree of the harmful features and the inter-image common features (D103), and stores the attention map (H100) in the memory (10) (E105). Regarding the setting of the attention degree, for example, the attention degree is set by using the averaged intensity of the beneficial features as is, and overwriting the attention degree of the pixels corresponding to the harmful features and the inter-image common features with the value "0".

[0060] The image reconstruction unit (F100) reads the grouped camera images and the gaze maps (H100) corresponding to the groups to which each camera image belongs from the memory (10), and incorporates each gaze map (H100) into each camera image belonging to the corresponding group (F101). The image reconstruction unit (F100) stores the camera images into which the gaze maps have been incorporated for each group as learning images (G101) in the memory (10). Note that the gaze maps (H100) are incorporated into the camera images in the same manner as in the example described above.

[0061] Although not shown in FIG. 7, the learning unit (G100) reads training images (G101) from the memory (10) and has the visibility estimation AI learn these training images (G101). During learning, the visibility estimation AI learns each image feature of the training images according to the gaze level indicated by the gaze map (H100). This allows the visibility estimation AI to focus on learning image features with high gaze levels, i.e., useful features. For example, an attention branch network can be used as a method for generating a gaze map or for having the visibility estimation AI learn training images incorporating the gaze map. The learning unit (G100) stores the visibility estimation AI that has learned the training images (G101) in the memory (10) as a trained AI model 2.

[0062] FIG. 8 is a diagram showing an example of a processing flow of the inference phase executed in the arithmetic device (1) in the first embodiment.

[0063] In Fig. 8, the image reconstruction unit (F100) reads from the memory (10) the grouped camera images and the gaze maps corresponding to the groups to which each camera image belongs, and incorporates each gaze map into each camera image belonging to the corresponding group (F101). The image reconstruction unit (F100) stores the camera images into which the gaze maps have been incorporated for each group as inference images (G102) in the memory (10). Note that the gaze map (H100) is incorporated into the camera images in the same manner as in the example described above.

[0064] Although not shown in FIG. 8, the inference unit (I100) reads the inference image (G102) and the trained AI model 2 (visibility estimation AI) that learned the training image (G101) in the training phase from the memory (10). The inference unit (I100) causes the visibility estimation AI to estimate visibility for the inference image (G102). When estimating visibility, the visibility estimation AI determines whether to use each image feature of the inference image as the basis for estimation based on the degree of attention. This allows the visibility estimation AI to prioritize image features with high attention levels, i.e., useful features, as the basis for estimation. For example, an attention branch network can be used as a method for causing the visibility estimation AI to estimate visibility based on the information on the gaze map incorporated in the inference image. The inference unit (I100) stores the visibility estimation results obtained by the visibility estimation AI in the memory (10). The estimation result is output from the memory (10) to the output device 3 (23) as needed, and is displayed on a screen provided in the output device 3 (23).

[0065] A user (e.g., an observer) can determine the visibility by referring to the estimation result displayed on the output device 3 (23). Note that the display format on the output device 3 (23) is not particularly limited as long as the user can properly grasp the visibility of the target.

[0066] As described above, the computing device in the first embodiment extracts useful features, harmful features, and common features between images from multiple camera images with approximately the same shooting angle of view, shooting direction, and estimation target, calculates and averages the intensity of each image feature on a pixel-by-pixel basis, generates a gaze map based on the averaged intensity, and incorporates the gaze map into training images and inference images. This enables learning and inference of visibility estimation AI that focuses on useful features that are highly correlated with visibility. In other words, the visibility estimation AI focuses on useful features with high gaze levels according to the gaze map as the basis for learning or inference, improving robustness against occlusions and disturbances that are uncorrelated with visibility and achieving high accuracy in automatic visibility estimation. Second Embodiment In the first embodiment, an example of a computing device that estimates visibility for one or more types of camera images that have different shooting angles of view, shooting directions, and estimation targets has been described. In the second embodiment, an example of a computing device that estimates visibility for camera images that have approximately the same shooting direction will be described.

[0067] The configuration example of the entire arithmetic device in the second embodiment is the same as the configuration example of the arithmetic device (1) in the first embodiment shown in FIG.

[0068] Fig. 9 is a diagram showing an example of a processing flow of a learning phase executed in the arithmetic device in the second embodiment. Also, Fig. 10 is a diagram showing an example of a processing flow of an inference phase executed in the arithmetic device in the second embodiment. In Figs. 9 and 10, the same configurations and processing contents as those in the processing flows shown in Figs. 7 and 8 are denoted by the same reference numerals.

[0069] Using Fig. 9, a description will be given of the configuration and processing contents that are different from the example processing flow of the learning phase in the first embodiment shown in Fig. 7. In Fig. 9, the camera image (A100) input in advance to the memory 10 is a camera image captured in approximately the same direction.

[0070] The image grouping unit (B100) reads the camera images (A100) and metadata from the memory (10). The image grouping unit (B100) divides (groups) the camera images (A100) into 1 to k groups according to the metadata and stores them in the memory (10) (B101). At this time, the image grouping unit (B100) performs grouping so that multiple camera images with approximately the same shooting angle of view belong to the same group.

[0071] The gaze degree calculation unit (D100) reads useful images tagged by the data refinement unit (C100) from the memory (10) for each group and extracts useful features possessed by the useful images. The gaze degree calculation unit (D100) calculates the intensity of the useful features for each pixel in each camera image belonging to the same group for each group and stores the calculated values ​​in the memory (10) (D101). The gaze degree calculation unit (D100) also reads useful images and harmful images for each group from the memory (10). The gaze degree calculation unit (D100) compares the image features of one or more useful images (C101) and harmful images (C102) belonging to the same group for each group, extracts image features uncorrelated with changes in visibility in those images as inter-image common features, and stores the extracted features in the memory (10) (D103).

[0072] On the other hand, the gaze degree calculation unit (D100) does not read the harmful image, extract the harmful features that the harmful image has, or calculate the strength of the harmful features. This is because harmful features in camera images taken from approximately the same direction are often disturbances, etc., as described in the first embodiment, and are often uncorrelated with the position on the camera image.

[0073] Since the gaze degree calculation unit (D100) does not calculate the intensity of harmful features, the aggregation unit (E100) also does not average the intensity of harmful features or truncate the intensity after averaging. In addition, the aggregation unit (E100) generates a gaze map (H100) for useful features and common features between images.

[0074] The configuration and processing contents other than those described above are the same as the example of the processing flow of the learning phase in the first embodiment shown in Fig. 7. In addition, the example of the processing flow of the inference phase shown in Fig. 10 is the same as the example of the processing flow of the inference phase in the first embodiment shown in Fig. 8.

[0075] As described above, the computing device of the second embodiment extracts useful features and inter-image common features from multiple camera images captured from approximately the same direction, calculates and averages the intensity of the useful features on a pixel-by-pixel basis, generates a gaze map based on the averaged intensity, and incorporates the gaze map into training images and inference images. This enables learning and inference by a visibility estimation AI that focuses on useful features that are highly correlated with visibility. In other words, the visibility estimation AI focuses on useful features with high gaze levels according to the gaze map as the basis for learning or inference, improving robustness against occlusions and disturbances uncorrelated with visibility and achieving high accuracy in automatic visibility estimation. <Third embodiment> In the first embodiment, an example of a computing device that estimates visibility based on a camera image is described. In the third embodiment, an example of a computing device that estimates visibility of an arbitrary target from a camera image is described.

[0076] The overall configuration of the arithmetic device in the third embodiment is the same as the configuration of the arithmetic device (1) in the first embodiment shown in FIG.

[0077] Fig. 11 is a diagram showing an example of a processing flow of the learning phase executed in the arithmetic device in the third embodiment. Also, Fig. 12 is a diagram showing an example of a processing flow of the inference phase executed in the arithmetic device in the third embodiment. In Figs. 11 and 12, the same configurations and processing contents as those in the processing flows shown in Figs. 7 and 8 are denoted by the same reference numerals.

[0078] Using Fig. 11, a description will be given of the configuration and processing contents that are different from the example of the processing flow of the learning phase in the first embodiment shown in Fig. 7. In this example, it is assumed that one or more targets to be subject to estimation of visibility are determined in advance.

[0079] 11, the image grouping unit (B100) reads camera images (A100) and metadata from the memory (10). The image grouping unit (B100) divides (groups) the camera images (A100) into 1 to k groups according to the metadata and stores them in the memory (10) (B101). At this time, the image grouping unit (B100) performs grouping so that multiple camera images that have the same predetermined target are assigned to the same group.

[0080] Unlike the example of the processing flow of the learning phase in the first embodiment shown in FIG. 7, the data refining unit (C100) does not perform any processing.

[0081] Meanwhile, the gaze degree calculation unit (D100) reads the grouped camera images from the memory (10). The gaze degree calculation unit (D100) extracts the target object to be estimated in each group as a useful feature from the multiple camera images belonging to each group. The processing thereafter is the same as the example processing flow of the learning phase in the first embodiment shown in FIG. 7. However, the gaze degree calculation unit (D100) does not read the harmful images, extract the harmful features possessed by the harmful images, or calculate the strength of the harmful features. Furthermore, the gaze degree calculation unit (D100) does not perform the processing of extracting common features between images.

[0082] Since the attention degree calculation unit (D100) does not calculate the intensity of harmful features, the aggregation unit (E100) also does not average the intensity of harmful features or truncate the intensity after averaging. Also, since the attention degree calculation unit (D100) does not extract common features between images, the aggregation unit (E100) generates a attention map (H100) targeting useful features. In generating the attention map (H100), the attention degree of the image area where the useful features are mapped is set to "1" and the attention degree of the other image areas is set to "0".

[0083] The configuration and processing contents other than those described above are the same as the example of the processing flow of the learning phase in the first embodiment shown in Fig. 7. In addition, the example of the processing flow of the inference phase shown in Fig. 12 is the same as the example of the processing flow of the inference phase in the first embodiment shown in Fig. 8.

[0084] As described above, the computing device of the third embodiment extracts useful features from multiple camera images in which the target object is substantially the same, calculates and averages the intensity of the useful features on a pixel-by-pixel basis, generates a gaze map based on the averaged intensity, and incorporates the gaze map into training images and inference images. This enables learning and inference by a gaze recognition / non-recognition estimation AI that focuses on the target object. In other words, the gaze recognition / non-recognition estimation AI focuses on useful features with high gaze levels according to the gaze map, improving robustness against occlusions and disturbances uncorrelated with the target object and achieving high accuracy in automatic estimation of the target object's gaze recognition / non-recognition.

[0085] An example of the hardware configuration of the arithmetic unit in the first to third embodiments described above will be described below. Fig. 13 is a diagram showing an example of the hardware configuration of the arithmetic unit.

[0086] The arithmetic device 1 includes a processor 31, a storage unit 32, and an interface unit 35. The processor 31 may be, for example, a CPU, and is a main body that executes predetermined processing. However, the processor 31 may be configured using other appropriate semiconductor devices (for example, a GPU) as long as it is capable of executing the predetermined processing.

[0087] The storage unit 32 can be configured using a main storage device 33 and an auxiliary storage device 34. Here, the main storage device 33 corresponds to the memory 10 in FIG. 4, etc. The main storage device 33 can be, for example, a RAM (Random Access Memory), and the processor 31 temporarily stores data in the main storage device 33 during data processing.

[0088] The auxiliary storage device (34) may be, for example, a hard disk drive (HDD), and stores programs used for various data processing (such as arithmetic processing and display processing) executed by the arithmetic device (1). The auxiliary storage device (34) stores, for example, programs such as an image grouping unit (B100), a data refining unit (C100), a gaze degree calculation unit (D100), an aggregation unit (E100), an image reconstruction unit (F100), a learning unit (G100), and an inference unit (I100). The auxiliary storage device (34) also stores data such as generated gaze maps. Note that the configurations of the main storage device (33) and the auxiliary storage device (34) described here are merely examples, and the configuration of the storage unit (32) may be changed as appropriate.

[0089] The interface unit (35) is an interface for inputting and outputting data. The calculation device (1) acquires, for example, a camera image (A100) from the input device 1 (11) via the interface unit (35). The calculation device (1) also outputs data to be displayed on the output device 3 (23) via the interface unit (35). The calculation device (1) may also receive input of a user's operation via the interface unit (35). The user's operation is performed using an appropriate input device (not shown), and examples of the user's operation include selecting a target in the camera image for which visibility is estimated.

[0090] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to the above-described embodiments, and various design modifications can be made without departing from the spirit of the present invention as defined in the claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations. [Explanation of symbols]

[0091] 1...Arithmetic unit 10. Memory 11 Input device 1 12 Input device 2 13 Input device 3 14 Input device 4 15. Input device 5 21 Output device 1 22 Output device 2 23 Output device 3 31 Processor 32...Storage section 33...Main memory 34...Auxiliary storage device 35 Interface section A100...Camera image B100: Image grouping unit C100 Data Refinement Department D100...Attention level calculation unit D111: Image extraction unit D112: Image feature intensity calculation unit D113: Image comparison section within the same group E100.....Collecting unit E111... Accumulation section E112...Normalization section E113....weak intensity truncation E114: Gaze map generation unit F100...Image reconstruction unit G100...Learning section G101...Learning image G102...Inference image H100···Gaze Map I100...Inference section

Claims

1. A computing device that estimates visibility using an image, an attention level calculation unit that extracts beneficial features or harmful features from a plurality of images having at least substantially the same photographing angle of view, photographing direction, or estimated object, and calculates the intensity of the beneficial features or harmful features in the plurality of images; an aggregation unit that aggregates the intensities of the beneficial features or the harmful features and generates an attention map indicating the attention degree of the beneficial features or the harmful features based on the aggregated intensities; an image reconstruction unit that incorporates the gaze map into each of the plurality of images; A computing device comprising:

2. 2. The computing device according to claim 1, Further comprising a learning unit that generates a trained AI model, the aggregation unit generates the attention map by setting a higher attention level for the beneficial feature than for the harmful feature; the image reconstruction unit incorporates the gaze map into each of the plurality of images to generate a learning image; The learning unit generates the trained AI model by learning the training image, The trained AI model focuses on learning the useful features for which the attention level is set high according to the attention map incorporated in the training image. Computing device.

3. 3. The computing device according to claim 2, An inference unit that estimates visibility using the trained AI model, the image reconstruction unit incorporates the gaze map into each of the plurality of images to generate an inference image; The inference unit estimates visibility using the trained AI model according to the gaze map incorporated in the inference image, with a focus on the useful features for which the gaze level is set high as an estimation basis. Computing device.

4. 2. The computing device according to claim 1, Further comprising a data refinement unit that extracts one or more useful images or harmful images from the plurality of images, The gaze degree calculation unit extracts the useful features from the one or more useful images, or extracts the harmful features from the one or more harmful images. Computing device.

5. 5. The computing device according to claim 4, an image grouping unit that groups one or more images that are included in a plurality of images input from outside and have at least the same photographing angle of view, the same photographing direction, or the same estimated object into groups of images that have at least the same photographing angle of view, the same photographing direction, or the same estimated object; The data refining unit extracts the one or more useful images or harmful images from the plurality of images for each of the grouped image groups, the attention degree calculation unit extracts the beneficial feature or the harmful feature for each of the image groups; Computing device.

6. 5. The computing device according to claim 4, the gaze degree calculation unit compares image features of the one or more useful images or harmful images, and further extracts common features between images that are uncorrelated with changes in visibility in the one or more useful images or harmful images; the aggregation unit generates the attention map by setting a higher attention level for the beneficial features than for the harmful features and the inter-image common features. Computing device.

7. 5. The computing device according to claim 4, The gaze degree calculation unit an image extractor configured to extract the beneficial features from the one or more beneficial images or to extract the harmful features from the one or more harmful images; an image feature intensity calculation unit that calculates the intensity of the beneficial feature or the harmful feature for each pixel in each of the plurality of images; Computing device.

8. 2. The computing device according to claim 1, The collecting unit is an accumulation unit that adds up the intensities of the beneficial features or the harmful features in the plurality of images; a normalization unit that divides the added intensities of the beneficial features or the harmful features by the number of the plurality of images. Computing device.

9. 1. A method for estimating visibility using images, comprising: Extracting useful features or harmful features from a plurality of images having at least substantially the same photographing angle of view, photographing direction, or estimated object; calculating the intensities of the beneficial or harmful features in the plurality of images; aggregating the strength of the beneficial or harmful characteristics; generating an attention map indicating the degree of attention to the beneficial feature or the harmful feature based on the intensities after the aggregation; incorporating the gaze map for each of the plurality of images; Visibility estimation methods.

10. 10. The visibility estimation method according to claim 9, generating the attention map by setting the attention level of the beneficial feature to be higher than that of the harmful feature; Incorporating the gaze map into each of the plurality of images to generate a training image; Generate a trained AI model by training the training image; The trained AI model focuses on learning the useful features for which the attention level is set high according to the attention map incorporated in the training image. Visibility estimation methods.

11. The visibility estimation method according to claim 10, Incorporating the gaze map into each of the plurality of images to generate an inference image; Using the trained AI model, the visibility is estimated based on the gaze map incorporated in the inference image, with a focus on the useful features for which the gaze level is set high. Visibility estimation methods.

12. 10. The visibility estimation method according to claim 9, extracting one or more useful or harmful images from the plurality of images; extracting the beneficial features from the one or more beneficial images or extracting the harmful features from the one or more harmful images; Visibility estimation methods.

13. The visibility estimation method according to claim 12, Receive multiple images from outside, one or more images that are included in the plurality of images received from the outside and have at least the same photographing angle of view, the same photographing direction, or the same estimated object are grouped into groups of images that have at least the same photographing angle of view, the same photographing direction, or the same estimated object; extracting the one or more useful images or harmful images from the plurality of images for each of the grouped images; extracting the beneficial features or the harmful features for each of the image groups; Visibility estimation methods.

14. The visibility estimation method according to claim 12, comparing image characteristics in the one or more useful or detrimental images; Further extracting common features between images that are uncorrelated with changes in visibility in the one or more useful or harmful images; generating the attention map by setting a higher attention level for the beneficial features than for the harmful features and the common features between images; Visibility estimation methods.

Citation Information

Patent Citations

  • Visibility range evaluation device, visibility range evaluation system and visibility range evaluation method

    JP2022020911A